Evaluation of Exam Questions Using Bootstrapping: Practical Applications in R and SPSS with a Case Study
Saved in:
| Title: | Evaluation of Exam Questions Using Bootstrapping: Practical Applications in R and SPSS with a Case Study |
|---|---|
| Language: | English |
| Authors: | Changiz Mohiyeddini |
| Source: | Anatomical Sciences Education. 2025 18(8):858-880. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 23 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Test Items, Sampling, Statistical Inference, Nonparametric Statistics, Difficulty Level, Correlation, Test Construction |
| DOI: | 10.1002/ase.70082 |
| ISSN: | 1935-9772 1935-9780 |
| Abstract: | This article presents a step-by-step guide to using R and SPSS to bootstrap exam questions. Bootstrapping, a versatile nonparametric analytical technique, can help to improve the psychometric qualities of exam questions in the process of quality assurance. Bootstrapping is particularly useful in disciplines such as medical education, where student cohorts are normally too small to reliably use parametric analysis to evaluate the quality of exam questions. Traditional parametric approaches need large samples; otherwise, they can yield unreliable estimates of metrics such as item difficulty and point-biserial correlations with small cohorts, potentially misleading the evaluation of exam questions and consequently leading to flawed assessments. By employing bootstrapping, educators can resample data to obtain robust confidence intervals for key metrics. This allows for a more accurate evaluation of question quality. This guide provides a step-by-step approach using R and SPSS, along with explaining the necessary code to bootstrap exam question means, standard deviations, item difficulty, and point-biserial correlations. In addition, the code includes automated visualizations and the capability to export results in reader-friendly tables, enhancing time efficiency and streamlining both data analysis and presentation processes. Furthermore, this article includes a case study in which the code is applied and the results are discussed to showcase how bootstrapping can inform decisions regarding exam question revisions. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1478968 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGY27V0EhNozBFeIIoziZbYAAAA4TCB3gYJKoZIhvcNAQcGoIHQMIHNAgEAMIHHBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDNRobgjX0CwMZK8h6AIBEICBmQUN_9IStquRv21Vovlz7PfL_SsNnP0LdfysG49DL42lhFpUUy-ayTpPd10D774D53_0llyl2wqEZJrjoAfQ3dYZnPczWT6GG4h2Ggs_uXHxBYduPZzBZb-jRKsKXX_66TXF4x8hBbZDbk6O6JH9LUuD4Zj3cEuhqb2p7L5-knrCj9Lg-jn606p2n_oHJIBlR38TbJurV7fm4Q== Text: Availability: 1 Value: <anid>AN0187112559;[8z8k]01aug.25;2025Aug06.02:56;v2.2.500</anid> <title id="AN0187112559-1">Evaluation of exam questions using bootstrapping: Practical applications in R and SPSS with a case study </title> <p>This article presents a step‐by‐step guide to using R and SPSS to bootstrap exam questions. Bootstrapping, a versatile nonparametric analytical technique, can help to improve the psychometric qualities of exam questions in the process of quality assurance. Bootstrapping is particularly useful in disciplines such as medical education, where student cohorts are normally too small to reliably use parametric analysis to evaluate the quality of exam questions. Traditional parametric approaches need large samples; otherwise, they can yield unreliable estimates of metrics such as item difficulty and point‐biserial correlations with small cohorts, potentially misleading the evaluation of exam questions and consequently leading to flawed assessments. By employing bootstrapping, educators can resample data to obtain robust confidence intervals for key metrics. This allows for a more accurate evaluation of question quality. This guide provides a step‐by‐step approach using R and SPSS, along with explaining the necessary code to bootstrap exam question means, standard deviations, item difficulty, and point‐biserial correlations. In addition, the code includes automated visualizations and the capability to export results in reader‐friendly tables, enhancing time efficiency and streamlining both data analysis and presentation processes. Furthermore, this article includes a case study in which the code is applied and the results are discussed to showcase how bootstrapping can inform decisions regarding exam question revisions.</p> <p>Keywords: bootstrapping; confidence interval; examination; item difficulty; point‐biserial correlation; R; RStudio; SPSS</p> <hd id="AN0187112559-2">INTRODUCTION</hd> <p>In medical education, the reliability and validity of exam questions are vital to ensure that assessments accurately reflect student learning and competencies.[<reflink idref="bib1" id="ref1">1</reflink>] Reliable exam questions consistently differentiate between high‐ and low‐performing students and yield similar results under consistent conditions, meaning they can effectively distinguish between varying levels of ability within a cohort.[<reflink idref="bib2" id="ref2">2</reflink>] Valid questions, on the other hand, measure the intended knowledge or skills, accurately reflect the learning objectives, and assess the intended competencies.[<reflink idref="bib3" id="ref3">3</reflink>]</p> <p>The quality of exam questions, henceforth used interchangeably with item, is particularly important in high‐stakes exams, where flawed questions could lead to inaccurate evaluations of students' abilities, ultimately impacting their future clinical performance and, by extension, the health and safety of their patients.[<reflink idref="bib4" id="ref4">4</reflink>] However, creating fair, reliable, and content‐valid exam questions is one of the most challenging tasks for medical educators.</p> <p>Undoubtedly, the benchmark for evaluating the quality of exam questions is the exploration of their psychometric qualities.[<reflink idref="bib5" id="ref5">5</reflink>] However, in medical education, it is common to have relatively few students compared to other disciplines, which poses statistical challenges to using parametric indicators and analysis.[<reflink idref="bib6" id="ref6">6</reflink>] Because small cohorts often produce unreliable estimates of key metrics, such as item difficulty and point‐biserial correlations, they may lead to inaccurate, unstable estimates and conclusions, which can make it difficult to evaluate whether a question is too easy, too hard, or fails to discriminate well between students.[[<reflink idref="bib7" id="ref7">7</reflink>]]</p> <hd id="AN0187112559-3">Bootstrapping and its benefits for evaluation of exam questions in small samples of students</hd> <p>Traditional assessment frameworks, such as Classical Test Theory (CTT) and Item Response Theory (IRT), are widely used to evaluate the psychometric quality of exam questions.[[<reflink idref="bib9" id="ref8">9</reflink>], [<reflink idref="bib11" id="ref9">11</reflink>]] CTT provides straightforward metrics like item difficulty and discrimination, but it relies on large, representative samples for reliable estimates, making it less effective in small sample contexts.[[<reflink idref="bib10" id="ref10">10</reflink>], [<reflink idref="bib12" id="ref11">12</reflink>]] IRT, on the other hand, offers a more sophisticated analysis by modeling individual item characteristics and student abilities, providing precise insights into item performance.[[<reflink idref="bib12" id="ref12">12</reflink>], [<reflink idref="bib14" id="ref13">14</reflink>]] However, similar to CTT, IRT also requires extensive data, and its complexity makes it less accessible in educational settings such as medical education with a limited size of student cohorts.[[<reflink idref="bib14" id="ref14">14</reflink>], [<reflink idref="bib16" id="ref15">16</reflink>]]</p> <p>To address these limitations, bootstrapping[<reflink idref="bib17" id="ref16">17</reflink>] offers a flexible, nonparametric alternative that fills gaps left by CTT and IRT. Bootstrapping is particularly useful in small sample contexts, as it generates numerous simulated samples from the available data, allowing for more robust confidence intervals and reliable item metrics even when data are limited.[[<reflink idref="bib6" id="ref17">6</reflink>], [<reflink idref="bib18" id="ref18">18</reflink>], [<reflink idref="bib20" id="ref19">20</reflink>]] By resampling the data with replacement, bootstrapping allows educators to create thousands of new datasets, enabling the calculation of more robust confidence intervals for key metrics such as item means, standard deviations, item difficulty, and point‐biserial correlations. These resampled datasets provide a valid picture of the variability in the data, helping educators make more psychometrically informed decisions about the quality of their exam questions.[<reflink idref="bib20" id="ref20">20</reflink>] In this edition, Mohiyeddini[<reflink idref="bib6" id="ref21">6</reflink>] provides an overview of bootstrapping resampling and its applications in educational assessment.</p> <hd id="AN0187112559-4">Relevance of bootstrapping for medical education and the scope of the present article</hd> <p>Improving item discrimination through bootstrapping methods has valuable implications for medical education, where assessment accuracy is critical for evaluating students' readiness for being a physician.[[<reflink idref="bib1" id="ref22">1</reflink>], [<reflink idref="bib21" id="ref23">21</reflink>]] For example, refining questions to better differentiate between high‐ and low‐performing students can help identify specific areas where students may need further instruction, additional teaching or practice material, or peer support, thereby guiding curriculum adjustments. High‐quality exam questions that reliably indicate students' competencies ensure that assessments more accurately reflect the clinical skills and knowledge required in real‐world healthcare settings.[<reflink idref="bib22" id="ref24">22</reflink>] This alignment between assessment quality and educational outcomes ultimately supports the preparation of well‐qualified healthcare professionals, underscoring the practical value of robust psychometric analysis.[<reflink idref="bib23" id="ref25">23</reflink>] Therefore, the primary focus of the article is on providing a detailed, practical guide for using bootstrapping to evaluate exam questions, specifically aimed at educators dealing with smaller student cohorts where traditional parametric methods might be less effective. This guide offers practical steps for using R and SPSS to bootstrap key metrics for exam question evaluation, including item means, standard deviations, item difficulty, and point‐biserial correlations. In addition, the article will provide a step‐by‐step explanation of how to visualize bootstrapped results using confidence intervals and how to export the results into reader‐friendly tables that can be easily adjusted for publications. By providing detailed coding instructions and clear explanations of each step, the guide will equip medical educators with the tools they need to apply bootstrapping to improve the quality of their exam questions.</p> <hd id="AN0187112559-5">Item mean and standard deviation</hd> <p>Item means referring to the average score students achieved on a particular question. In multiple‐choice exams, each item is scored as either correct (<reflink idref="bib1" id="ref26">1</reflink>) or incorrect (0), and the mean represents the proportion of students who answered the item correctly. The standard deviation (SD) reflects how spread out the scores are, indicating the consistency of students' performance on the item. For data that are normally distributed, about 68% of values lie within 1 SD of the mean, 95% within 2 SDs, and 99.7% within 3 SDs. If a standard deviation is small, data points are tightly clustered around the mean. If it is large, the data points are more spread out. The specific interpretations of SD can vary across different fields of study. For instance, in psychological testing, a SD of around 1 on a 1–5 or 1–7 Likert scale is generally considered moderate, while values significantly above that might be viewed as large.[[<reflink idref="bib11" id="ref27">11</reflink>], [<reflink idref="bib24" id="ref28">24</reflink>]] Analyzing these metrics in the process of exam question evaluations and revisions is essential to ensure the quality of the review process.[<reflink idref="bib10" id="ref29">10</reflink>]</p> <hd id="AN0187112559-6">Item difficulty (ID)</hd> <p>ID is calculated as the proportion of students who answer an item correctly. In classical test theory, ID is often represented by the "<emph>p</emph>‐value" (P stands for proportion) which indicates the proportion of test‐takers who answered the item correctly. As Clauser and Hambelton[<reflink idref="bib25" id="ref30">25</reflink>] correctly state "It is unfortunate that the statistic is not called "item‐easiness" since the high <emph>p</emph> values describe easy items (<reflink idref="bib357" id="ref31">357</reflink>). The <emph>p</emph>‐value ranges from 0 to 1. While there is no universally accepted standard for interpreting ID, the following guidelines may be helpful[[<reflink idref="bib9" id="ref32">9</reflink>], [<reflink idref="bib26" id="ref33">26</reflink>]]:</p> <p></p> <ulist> <item> High difficulty (too difficult): <emph>p</emph> ≤ 0.30 (&lt;30% of test‐takers answer correctly).</item> <p></p> <item> Moderate difficulty (ideal range): <emph>p</emph> between 0.30 and 0.70 (30%–70% of test‐takers answer correctly).</item> <p></p> <item> Low difficulty (too easy): <emph>p</emph> ≥ 0.70 (more than 70% of test‐takers answer correctly).</item> </ulist> <p>Ideally, exams should contain a mix of items with varying levels of difficulty to challenge all students adequately. Extremely easy or difficult items, however, may not effectively differentiate between high‐ and low‐performing students, reducing the overall discriminatory power of the exam.[<reflink idref="bib27" id="ref34">27</reflink>] There are instances in assessment contexts where items with very high or very low difficulty are essential. For example, in IQ testing, extremely challenging items are often used to identify gifted individuals, while very simple items can help diagnose cognitive impairments.[<reflink idref="bib28" id="ref35">28</reflink>] Similarly, in clinical diagnostics, symptoms vary in their "difficulty" to diagnose due to their prevalence across conditions. For instance, a common symptom like pain might be considered an "easy" symptom because it occurs in many conditions, and hence, it cannot solely inform a clinical diagnosis. In contrast, rare symptoms, such as visual hallucinations in certain neurological disorders, might be underreported by patients, making them "difficult" yet crucial indicators for diagnosing rare health condition.[<reflink idref="bib29" id="ref36">29</reflink>] Therefore, incorporating these "difficult" diagnostic items is vital for accurate and early diagnosis of these rare diseases.</p> <p>However, for dichotomous items (such as multiple‐choice questions in medical education with a correct vs. incorrect response), the mean represents the ID. For example, if the mean score on a question is 0.45, it indicates that 45% of students answered that question correctly, which simultaneously represents the difficulty level of the question. In this context, bootstrapping the mean equates to bootstrapping the ID, and the confidence interval (CI) of the bootstrapped mean serves as the CI for the bootstrapped ID.</p> <hd id="AN0187112559-7">Point‐biserial correlation (PBC)</hd> <p>The PBC is a measure of how well an exam question discriminates between students who performed well on the exam overall and those who performed poorly. It compares the scores of students who answered a particular item correctly with their overall exam scores.</p> <p>It is important to highlight that PBC and the Pearson correlation share the same mathematical calculation. The key difference lies in what each item is correlated with: If we calculate the correlation between question 1 (Q1) and the overall test score including Q1, the result would be a standard Pearson correlation. In this case, the resulting correlation would be artificially inflated. This happens because including Q1 in the total score adds a shared variance between the item and the overall score, making it seem more correlated than it actually is in relation to the other items. To prevent this, we calculate the correlation between Q1 and the total test score excluding Q1. This correlation is known as PBC. In other words, excluding Q1 from the total score transforms the Pearson correlation into the PBC. This is also why analytic frameworks such as R or SPSS do not have a designated tab to calculate PBC. Instead, they require creating total scores excluding the item itself first and then calculating the Pearson correlation between the question and the total score.</p> <p>A high positive PBC indicates that students who scored high on the exam were more likely to answer the question correctly, showing that the item effectively discriminates between high‐ and low performers. A low or negative PBC suggests the item may be flawed, as lower performing students might be answering it correctly more often than higher performing ones.[<reflink idref="bib30" id="ref37">30</reflink>] This metric is critical for refining exam questions to improve their ability to assess students' true understanding.[<reflink idref="bib31" id="ref38">31</reflink>]</p> <p>While universally defined thresholds for what constitutes a very high, high, moderate, or low PBC are missing, the following table provides a synthesis of guidelines for interpreting PBC values in item analysis (Table 1). These guidelines are based on general best practices in educational and psychological testing, as drawn from the literature.[[<reflink idref="bib9" id="ref39">9</reflink>], [<reflink idref="bib26" id="ref40">26</reflink>], [<reflink idref="bib30" id="ref41">30</reflink>], [<reflink idref="bib32" id="ref42">32</reflink>], [<reflink idref="bib34" id="ref43">34</reflink>]]</p> <p>1 TABLE Interpretation guidelines and recommended actions for point‐biserial correlation (PBC) values in item analysis.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;PBC range&lt;/th&gt;&lt;th align="left"&gt;Interpretation&lt;/th&gt;&lt;th align="left"&gt;Recommended action&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#62;0.50&amp;#8211;0.60&lt;/td&gt;&lt;td align="left"&gt;Too high (Possibly redundant)&lt;/td&gt;&lt;td align="left"&gt;Revise or retain if critical for the exam&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#8805;0.30&lt;/td&gt;&lt;td align="left"&gt;High (Good discrimination)&lt;/td&gt;&lt;td align="left"&gt;Retain the question&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.20&amp;#8211;0.29&lt;/td&gt;&lt;td align="left"&gt;Moderate (Acceptable discrimination)&lt;/td&gt;&lt;td align="left"&gt;Retain, consider minor revisions&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;0.10&amp;#8211;0.19&lt;/td&gt;&lt;td align="left"&gt;Low (Limited discrimination)&lt;/td&gt;&lt;td align="left"&gt;Revise, consider removal&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#60;0.10 (or negative)&lt;/td&gt;&lt;td align="left"&gt;Critically low (Poor discrimination)&lt;/td&gt;&lt;td align="left"&gt;Remove&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0187112559-8">Do ID and PBC deliver the same information?</hd> <p>The answer is no. ID simply reflects how many students answered correctly, while PBC measures how well the item discriminates between high‐ and low‐performing students. Therefore, a low ID (an easy item) does not necessarily imply a low PBC. Consider a multiple‐choice question in an organ system exam, such as "What is the primary function of red blood cells?" Most medical students may answer this correctly as "to transport oxygen", making it an easy item (e.g., 75% of students get it right, so ID = 0.75). However, this item could still have a high PBC if the students who answer it incorrectly are predominantly lower performing students on the overall test. In contrast, the above example regarding pain as a common symptom is an example of an easy item with low PBC. Therefore, in the process of exam question evaluation, both ID and PBC must be considered.</p> <hd id="AN0187112559-9">R‐based bootstrapping for exam question analysis: A comprehensive step‐by‐step guide</hd> <p>R is a versatile open‐source analytic tool that does not require licensing and is freely available.[<reflink idref="bib35" id="ref44">35</reflink>]</p> <hd id="AN0187112559-10">Step 1: Preparing the dataset</hd> <p>First, save the dataset containing student responses to the exam as a "CSV" file (short for Comma‐Separated Values file). Among others, CSV files can be generated in Excel, Google Sheets, R, Python, and SPSS.</p> <p>The file should have each question's responses as columns (e.g., Q1 and Q2), and each row should represent a student's answers across all questions. Each column should have a clear header indicating the question number, which will simplify the analysis process later. For example:</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Q1, Q2, Q3, &amp;#x2026;, Q101, 0, 1, &amp;#x2026;, 10, 1, 0, &amp;#x2026;, 01, 1, 1, &amp;#x2026;, 0&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This setup makes it easy to import the data into R and access each question separately, which is essential for calculating metrics like the PBC for individual items. Ensure each question column is coded as 1 (correct) or 0 (incorrect), so that item metrics can be calculated accurately.</p> <hd id="AN0187112559-11">Step 2: Install R, RStudio, and required packages</hd> <p>To perform bootstrapping in R, you need both R and RStudio, along with specific packages.</p> <hd id="AN0187112559-12">Installing R and RStudio</hd> <p>Download R from the Comprehensive R Archive Network (CRAN) (https://cran.r‐project.org/).</p> <p>Download RStudio from "Posit" (https://posit.co/download/rstudio‐desktop/). RStudio is a user‐friendly interface for R that can help to organize code, output, and files in one layout, making it much easier to work through data analysis projects, especially for users new to R.</p> <hd id="AN0187112559-13">Installing required packages</hd> <p>In RStudio (and R in general), "packages" are collections of pre‐written code, functions, data, and documentation that extend R's capabilities. They are like "add‐ons" or "libraries" that allow you to perform specific tasks or analyses more efficiently. Each package is designed for a particular purpose, such as data manipulation, visualization, machine learning, or statistical analysis. For the purpose of this article, we need to install the following packages, which are needed for bootstrapping, visualization, and exporting results by running the following commands:</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;install.packages("boot") Bootstrappinginstall.packages("psych") Psychological and educational statisticsinstall.packages("ggplot2") Visualizationinstall.packages("officer") Export to Wordinstall.packages("flextable") Table formatting for Word&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Each package serves a unique function:</p> <p></p> <ulist> <item> 'boot': Performs bootstrapping and creates resampled datasets.</item> <p></p> <item> 'psych': Includes functions to compute PBCs.</item> <p></p> <item> 'ggplot2': A powerful package for data visualization in R.</item> <p></p> <item> 'officer': Allows for the creation and manipulation of Word documents</item> <p></p> <item> 'flextable': Used to format tables for exporting to Word documents. 'officer' and 'flextable' work together to create and format tables in Word documents.</item> </ulist> <hd id="AN0187112559-14">Step 3. Loading and preparing the dataset</hd> <p>To load the CSV file containing exam results, use the following code:</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt; # Load the datasetexam&amp;#95;data &amp;#60;- read.csv("D:/Examdata/exam.csv")# Check the data structurestr(exam&amp;#95;data)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Note:</p> <p></p> <ulist> <item> In R, the # symbol is used to add comments within the code. Anything written after # on the same line is ignored by R when running the code. This allows programmers to include explanations, notes, or descriptions alongside the code, helping both themselves and others understand the purpose of each section or line.</item> <p></p> <item> This code assumes the file is named "exam.csv" and is saved in the specified directory "Examdata" on the D drive. (Adjust the file path if the file has a different name or is located in a subfolder).</item> </ulist> <p>This step imports the data into R and lets you verify that each question is represented as a column and that data types are correctly assigned.</p> <p></p> <ulist> <item> read.csv: Loads the dataset from the specified file path.</item> <p></p> <item> str: Displays the structure of the dataset, including the number of observations and variables, as well as data types for each column. This helps verify that the data were loaded correctly and gives insight into its format (e.g., if the data contains other numbers than 0 and 1).</item> </ulist> <hd id="AN0187112559-15">Step 4: Calculating total scores</hd> <p>Before calculating PBC, compute a total score for each student by summing responses across all questions. This total score represents the overall exam performance for each student and will be used as a benchmark for calculating PBCs for each question.</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;# Calculate the total score for each studentexam&amp;#95;data$total&amp;#95;score &amp;#60;- rowSums(exam&amp;#95;data[, c("Q1", "Q2", "Q3", "Q4", "Q5", "Q6", "Q7", "Q8", "Q9", "Q10")])&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Step 4 calculates the total score for each student across all exam questions and creates a new column called "total_score" (in exam.csv), which contains the sum of each student's responses to all questions. The total score represents each student's overall performance on the exam by summing their scores across all questions, excluding any one specific question when calculating the PBC for that question. Calculating the total score is essential for determining metrics like the PBC, which require comparing individual item responses (e.g., responses to Q1) to the overall exam performance across other items (e.g., total score excluding Q1).</p> <p></p> <ulist> <item> exam_data[, paste0("Q", 1:10)]: This part selects the columns in exam_data that correspond to the questions (Q1 to Q10).</item> <p></p> <item> paste0("Q", 1:10) generates a vector of column names: Q1, Q2, ..., Q10. This assumes that each column represents a question and holds scores for that question.</item> <p></p> <item> rowSums() calculates the sum of specified columns across each row. In this case, it sums each student's responses from columns Q1 to Q10.</item> <p></p> <item> exam_data$total_score&lt;‐ ...: The calculated row sums are assigned to a new column total_score in the exam_data data frame, storing each student's total score.</item> </ulist> <hd id="AN0187112559-16">Step 5: Bootstrapping item metrics</hd> <p>Step 5 uses bootstrapping to calculate key metrics that evaluate the quality of each exam question, such as the item mean, SD, ID, and PBC. Each metric provides unique insights into the quality of exam questions. While detailed, the following steps involve complex coding, which may be overwhelming for readers new to R. To enhance clarity, brief explanations of the coding segments are provided, especially around looping structures and parameter settings to make this guide more accessible.</p> <p>Step 5.1: Initialize lists to store results</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;mean&amp;#95;results&amp;#95;list &amp;#60;- list()sd&amp;#95;results&amp;#95;list &amp;#60;- list()pb&amp;#95;results&amp;#95;list &amp;#60;- list()&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>These lines create three empty lists to store the results of each metric (mean, SD, and PBC) for each question. Each question's results will be saved in these lists after bootstrapping.</p> <p>Step 5.2: Setting up vectors to hold actual (non‐bootstrapped) values</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;actual&amp;#95;means &amp;#60;- numeric(10)actual&amp;#95;sds &amp;#60;- numeric(10)actual&amp;#95;point&amp;#95;biserials &amp;#60;- numeric(10)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>These lines create three empty vectors to store the actual, non‐bootstrapped values of each metric. Since we are dealing with 10 questions, these vectors are set to hold 10 values each.</p> <p>Step 5.3: Starting the loop for each question</p> <p>In this looping structure, each iteration dynamically processes each question in the dataset. Explanations for each line of code follow to aid those less familiar with R's loop syntax and parameter selection.</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;for (i in 1:10) { question &amp;#60;- paste0("Q", i) # Set the question column name dynamically&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Note: In R, the placement of a {on a new line (or using a} at the end of a line) is a stylistic choice often seen in code for readability and clarity, but it does not affect the functionality. In R, you can place the closing bracket} either on the same line (after the expression within the function or block) or on a new line (below the final line within the function or block, making each function more readable, especially in complex scripts).</p> <p>This loop goes through each question in exam_data, labeled Q1, Q2, etc. By using paste0("Q", i), we create a dynamic column name (Q1, Q2, etc.) that changes for each iteration.</p> <p>5.4: Function to calculate the means/item difficulties for bootstrapping</p> <p>Inside the loop, we define functions for each statistic. These functions will calculate the statistic for a particular question based on a resampled dataset.</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt; item&amp;#95;mean &amp;#60;- function(data, indices) { d &amp;#60;- data[indices, ] return(mean(d[[question]], na.rm = TRUE)) }&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This function defines how the mean for a specific question will be calculated when bootstrapped:</p> <p></p> <ulist> <item> "data[indices, ]" creates a resampled dataset using the given indices.</item> <p></p> <item> "mean(d[[question]])" calculates the mean score for the current question in the resampled data.</item> </ulist> <p>Step 5.5: Function to calculate the standard deviation for bootstrapping</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt; item&amp;#95;sd &amp;#60;- function(data, indices) { d &amp;#60;- data[indices, ] return(sd(d[[question]], na.rm = TRUE)) }&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This function defines how the SD of scores for the question is calculated for bootstrapping:</p> <p></p> <ulist> <item> Similar to 'item_mean', it resamples 'data' using the 'indices'.</item> <p></p> <item> 'sd(d[[question]])' calculates the SD for the specified question.</item> </ulist> <p>Step 5.6: Calculate total score excluding the current question</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt; exam&amp;#95;data$total&amp;#95;score&amp;#95;excl &amp;#60;- rowSums(exam&amp;#95;data[, paste0("Q", setdiff(1:10, i))])&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Here, 'total_score_excl' is calculated for each row (student) by summing scores for all questions except the current one. This allows calculation of the PBC, which requires the total score excluding the item being analyzed.</p> <p>Step 5.7: Function to calculate PBC excluding the item</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;point&amp;#95;biserial &amp;#60;- function(data, indices) { d &amp;#60;- data[indices, ] return(cor(d[[question]], d$total&amp;#95;score&amp;#95;excl, method = "pearson"))}&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This function calculates the PBC between the current question's score and 'total_score_excl' in the resampled data:</p> <p></p> <ulist> <item> 'cor(d[[question]], d$total_score_excl, method = "pearson")' computes the Pearson correlation between the question's score and the total score excluding that question.</item> </ulist> <p>Step 5.8: Bootstrapping mean, SD, and PBC with 1000 resamples</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;mean&amp;#95;results &amp;#60;- boot(data = exam&amp;#95;data, statistic = item&amp;#95;mean, R&amp;#x2009;=&amp;#x2009;1000)sd&amp;#95;results &amp;#60;- boot(data = exam&amp;#95;data, statistic = item&amp;#95;sd, R&amp;#x2009;=&amp;#x2009;1000)difficulty&amp;#95;results &amp;#60;- boot(data = exam&amp;#95;data, statistic = item&amp;#95;difficulty, R&amp;#x2009;=&amp;#x2009;1000)pb&amp;#95;results &amp;#60;- boot(data = exam&amp;#95;data, statistic = point&amp;#95;biserial, R&amp;#x2009;=&amp;#x2009;1000)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Using the 'boot' function, each metric is bootstrapped with 1000 resamples:</p> <p></p> <ulist> <item> 'data = exam_data' specifies the dataset.</item> <p></p> <item> 'statistic = ...' defines the function (e.g., 'item_mean', 'item_sd') to calculate the metric.</item> <p></p> <item> 'R = 1000' sets the number of resamples to 1000, which allows estimation of the metric's variability.</item> </ulist> <p>Step 5.9: Store results</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;mean&amp;#95;results&amp;#95;list[[question]] &amp;#60;- mean&amp;#95;resultssd&amp;#95;results&amp;#95;list[[question]] &amp;#60;- sd&amp;#95;resultsdifficulty&amp;#95;results&amp;#95;list[[question]] &amp;#60;- difficulty&amp;#95;resultspb&amp;#95;results&amp;#95;list[[question]] &amp;#60;- pb&amp;#95;results&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>The bootstrapped results for each metric are stored in their respective lists, with each question's result saved under its respective name (e.g., 'Q1', 'Q2', etc.).</p> <p>Step 5.10: Calculate actual values</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;actual&amp;#95;means[i] &amp;#60;- mean(exam&amp;#95;data[[question]])actual&amp;#95;sds[i] &amp;#60;- sd(exam&amp;#95;data[[question]])actual&amp;#95;difficulties[i] &amp;#60;- mean(exam&amp;#95;data[[question]])actual&amp;#95;point&amp;#95;biserials[i] &amp;#60;- cor(exam&amp;#95;data[[question]], exam&amp;#95;data$total&amp;#95;score&amp;#95;excl, method = "pearson")&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This final step calculates the actual (non‐bootstrapped) values of each metric for the current question:</p> <p></p> <ulist> <item> 'mean(exam_data[[question]])' and 'sd(exam_data[[question]])' calculate the mean and standard deviation.</item> <p></p> <item> The difficulty is again calculated as the mean of the question's scores.</item> <p></p> <item> 'cor(exam_data[[question]], exam_data$total_score_excl, method = "pearson")' calculates the actual PBC. Each of these values is stored in its respective vector ('actual_means', 'actual_sds', etc.) for later reference.</item> </ulist> <hd id="AN0187112559-17">Step 6: Visualize bootstrapped confidence intervals for ID and PBC</hd> <p>This step visualizes the bootstrapped results for ID and PBC to understand each metric's variability across resampled data. We use boxplots to display these results, and each plot is saved as an image file. Boxplots are particularly useful for visualizing bootstrapped results because they effectively display variability and distribution across resampled data, which is crucial for understanding the stability of metrics like ID and PBC.</p> <p>Step 6.1: Creating the loop</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;for (i in 1:10) { question &amp;#60;- paste0("Q", i)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>A "for" loop iterates through each question, Q1 to Q10. In each iteration, the variable question is set dynamically to the current question label (e.g., Q1, Q2).</p> <p>Step 6.2: Visualize PBC using a boxplot</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;boot&amp;#95;pb&amp;#95;df &amp;#60;- as.data.frame(pb&amp;#95;results&amp;#95;list[[question]]$t)colnames(boot&amp;#95;pb&amp;#95;df) &amp;#60;- "Bootstrapped&amp;#95;&amp;#x200B;PointBiserial"&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p></p> <ulist> <item> pb_results_list[[question]]$t extracts the bootstrapped PBC results for the current question.</item> <p></p> <item> as.data.frame(...) converts these results into a data frame.</item> <p></p> <item> colnames(...) &lt;‐ "Bootstrapped_PointBiserial" renames the column to Bootstrapped_PointBiserial.</item> <p></p> </ulist> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;pb&amp;#95;plot &amp;#60;- ggplot(boot&amp;#95;pb&amp;#95;df, aes(y = Bootstrapped&amp;#95;PointBiserial)) + geom&amp;#95;boxplot(fill = "green", alpha&amp;#x2009;=&amp;#x2009;0.6) + labs(title = paste("Bootstrapped Point-Biserial Correlations for", question), y = "Point-Biserial Correlation") + theme&amp;#95;minimal()&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This code creates a boxplot for the bootstrapped PBC:</p> <p></p> <ulist> <item> ggplot(boot_pb_df, aes(y = Bootstrapped_PointBiserial)) initializes the plot with Bootstrapped_PointBiserial on the <emph>y</emph> ‐axis.</item> <p></p> <item> geom_boxplot(fill = "green", alpha = 0.6) adds a green boxplot with 60% opacity.</item> <p></p> <item> labs(...) sets the title and <emph>y</emph> ‐axis label.</item> <p></p> <item> theme_minimal() applies a minimal theme to the plot.</item> <p></p> </ulist> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;ggsave(filename = paste0("D:/Examdata/bootstrapped&amp;#95;pointbiserial&amp;#95;", question, "&amp;#95;boxplot.png"), plot = pb&amp;#95;plot)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>The ggsave function saves the PBC plot as a PNG file, with a filename that incorporates the question identifier (e.g., bootstrapped_pointbiserial_Q1_boxplot.png).</p> <hd id="AN0187112559-18">Step 7: Export results to word</hd> <p>This step exports a summary of the actual and bootstrapped item metrics (mean, SD, difficulty, and PBC) to a Word document, including confidence intervals for each bootstrapped metric. By automating the export process, it helps save a significant amount of time in collecting and organizing results, as the comprehensive summary of bootstrapped estimates and confidence intervals is readily formatted and accessible for review. This enables educators and researchers to quickly analyze and interpret the data without manual data entry or formatting, thereby streamlining the process of evaluating item quality.</p> <p>Step 7.1: Create a data frame to summarize results</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;results&amp;#95;table &amp;#60;- data.frame(Question = paste0("Q", 1:10), Actual&amp;#95;Mean = round(actual&amp;#95;means, 3), Bootstrapped&amp;#95;Mean = round(sapply(mean&amp;#95;results&amp;#95;list, function(x) mean(x$t, na.rm = TRUE)), 3), Mean&amp;#95;CI&amp;#95;Lower = round(sapply(mean&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = "bca")$bca[4] } else { NA } }), 3), Mean&amp;#95;CI&amp;#95;Upper = round(sapply(mean&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = "bca")$bca[5] } else { NA }}), 3), Actual&amp;#95;SD = round(actual&amp;#95;sds, 3), Bootstrapped&amp;#95;SD = round(sapply(sd&amp;#95;results&amp;#95;list, function(x) mean(x$t, na.rm = TRUE)), 3), SD&amp;#95;CI&amp;#95;Lower = round(sapply(sd&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = "bca")$bca[4] } else { NA } }), 3), SD&amp;#95;CI&amp;#95;Upper = round(sapply(sd&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = "bca")$bca[5] } else { NA } }), 3), Actual&amp;#95;PointBiserial = round(actual&amp;#95;point&amp;#95;biserials, 3), Bootstrapped&amp;#95;PointBiserial = round(sapply(pb&amp;#95;results&amp;#95;list, function(x) mean(x$t, na.rm = TRUE)), 3),&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt; PB&amp;#95;CI&amp;#95;Lower = round(sapply(pb&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[4] } else { NA } }), 3), PB&amp;#95;CI&amp;#95;Upper = round(sapply(pb&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[5] } else { NA } }), 3))&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This block of code creates a data frame, results_table, that consolidates the actual and bootstrapped metrics for each question (Q1 to Q10). Here is what each part does:</p> <p></p> <ulist> <item> Question = paste0("Q", 1:10): Generates the question labels Q1 to Q10.</item> <p></p> <item> Actual_Mean through PB_CI_Upper: These columns contain actual values, bootstrapped means, and bootstrapped confidence intervals for each metric (mean, standard deviation, difficulty, and PBC).</item> <p></p> <item> Actual_Mean, Actual_SD, etc.: Directly use the actual values calculated earlier.</item> <p></p> <item> Bootstrapped_Mean, Bootstrapped_SD, etc.: Calculate the mean of the bootstrapped results.</item> <p></p> <item> Confidence Intervals (CI): Use boot.ci(x, type = "bca")$bca[<reflink idref="bib4" id="ref45">4</reflink>] and boot.ci(x, type = "bca")$bca[<reflink idref="bib5" id="ref46">5</reflink>] to get the lower and upper bounds of the bootstrapped confidence intervals.</item> </ulist> <p>Each value is rounded to three decimal places for readability.</p> <p>Step 7.2: Convert to flextable</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;ft &amp;#60;- flextable(results&amp;#95;table)ft &amp;#60;- set&amp;#95;header&amp;#95;labels(ft, Actual&amp;#95;Mean = "Actual Mean", Bootstrapped&amp;#95;Mean = "Bootstrapped Mean", Mean&amp;#95;CI&amp;#95;Lower = "Mean CI Lower", Mean&amp;#95;CI&amp;#95;Upper = "Mean CI Upper", Actual&amp;#95;SD = &amp;#x201C;Actual SD&amp;#x201D;, Bootstrapped&amp;#95;SD = &amp;#x201C;Bootstrapped SD&amp;#x201D;, SD&amp;#95;CI&amp;#95;Lower = &amp;#x201C;SD CI Lower&amp;#x201D;, SD&amp;#95;CI&amp;#95;Upper = &amp;#x201C;SD CI Upper&amp;#x201D;,&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt; Actual&amp;#95;PointBiserial = &amp;#x201C;Actual Point-Biserial&amp;#x201D;, Bootstrapped&amp;#95;PointBiserial = &amp;#x201C;Bootstrapped Point-Biserial&amp;#x201D;, PB&amp;#95;CI&amp;#95;Lower = &amp;#x201C;PB CI Lower&amp;#x201D;, PB&amp;#95;CI&amp;#95;Upper = &amp;#x201C;PB CI Upper&amp;#x201D;)ft &amp;#60;- bold(ft, part = &amp;#x201C;header&amp;#x201D;)ft &amp;#60;- fontsize(ft, size&amp;#x2009;=&amp;#x2009;12)ft &amp;#60;- theme&amp;#95;vanilla(ft)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This section formats results_table as a flextable for exporting to Word:</p> <p></p> <ulist> <item> Convert data frame to flextable: ft &lt;‐ flextable(results_table) converts the data frame results_table into a flextable object.</item> <p></p> <item> Set header labels: set_header_labels(...) customizes the column headers to make them more descriptive and easier to read.</item> <p></p> <item> Apply styling:</item> <p></p> <item> bold(ft, part = "header"): Makes the header row bold.</item> <p></p> <item> fontsize(ft, size = 12): Sets the font size for the entire table to 12.</item> <p></p> <item> theme_vanilla(ft): Applies a clean, minimal style to the table.</item> </ulist> <p>Step 7.3: Create word document and add the table</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;doc &amp;#60;- read&amp;#95;docx()doc &amp;#60;- body&amp;#95;add&amp;#95;flextable(doc, value = ft)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This part creates a new Word document and adds the formatted table:</p> <p></p> <ulist> <item> read_docx(): Initializes an empty Word document.</item> <p></p> <item> body_add_flextable(doc, value = ft): Adds the formatted flextable (ft) to the body of the Word document.</item> </ulist> <p>Step 7.4: Save the word document</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;print(doc, target = "D:/Examdata/bootstrapped&amp;#95;results.docx")&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This line saves the Word document containing the table to a specified location. The document is named bootstrapped_results.docx and is saved in the folder D:/Examdata/.</p> <p>Step 7.5: Confirm file creation</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;print("The bootstrapped results table has been saved as D:/Examdata/bootstrapped&amp;#95;results.docx")&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>This line prints a confirmation message in the console, letting the user know that the Word document has been successfully saved.</p> <p>The box below presents the code and covers the full workflow, from preparing data to exporting results, providing a robust way to evaluate exam question quality.</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;# Step 1: Prepare your dataset# Ensure "exam.csv" is saved in the "Examdata" folder on your D drive# Each column should represent a question (e.g., Q1 to Q10)# Step 2: Load necessary librariesinstall.packages("boot") # For bootstrappinginstall.packages("psych") # For educational statistics, including point-biserial correlationinstall.packages("ggplot2") # For data visualizationinstall.packages("officer") # For exporting results to Wordinstall.packages("flextable") # For formatting tables for Wordlibrary(boot)library(psych)library(ggplot2)library(officer)library(flextable)# Step 3: Load and inspect your dataexam&amp;#95;data &amp;#60;- read.csv("D:/Examdata/exam&amp;#95;data3.csv")str(exam&amp;#95;data) # Verify the structure of the data# Convert any character columns to numeric (specifically Q7 and Q8 in this case)exam&amp;#95;data$Q7 &amp;#60;- as.numeric(exam&amp;#95;data$Q7)exam&amp;#95;data$Q8 &amp;#60;- as.numeric(exam&amp;#95;data$Q8)# Step 4: Calculate total scores for each studentexam&amp;#95;data$total&amp;#95;score &amp;#60;- rowSums(exam&amp;#95;data[, paste0("Q", 1:10)])# Step 5: Bootstrapping Item Metrics (Mean, SD, Point-Biserial Correlation)#Step 5.1: Initialize lists to store resultsmean&amp;#95;results&amp;#95;list &amp;#60;- list()sd&amp;#95;results&amp;#95;list &amp;#60;- list()pb&amp;#95;results&amp;#95;list &amp;#60;- list()# Step 5.2: Initialize vectors to store actual valuesactual&amp;#95;means &amp;#60;- numeric(10)actual&amp;#95;sds &amp;#60;- numeric(10)actual&amp;#95;point&amp;#95;biserials &amp;#60;- numeric(10)# Step 5.3: Starting the Loop for Each Questionfor (i in 1:10) { question &amp;#60;- paste0("Q", i) # Set the question column name dynamically # Step 5.4: Function to calculate the mean for bootstrapping item&amp;#95;mean &amp;#60;- function(data, indices) { d &amp;#60;- data[indices, ] return(mean(d[[question]], na.rm = TRUE)) }&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt; # Step 5.5: Function to calculate the standard deviation for bootstrapping item&amp;#95;sd &amp;#60;- function(data, indices) { d &amp;#60;- data[indices, ] return(sd(d[[question]], na.rm = TRUE)) }# Step 5.6: Calculate total score excluding the current question exam&amp;#95;data$total&amp;#95;score&amp;#95;excl &amp;#60;- rowSums(exam&amp;#95;data[, paste0(&amp;#x201C;Q&amp;#x201D;, setdiff(1:10, i))], na.rm = TRUE) # Step 5.7: Function to calculate point-biserial correlation excluding the item point&amp;#95;biserial &amp;#60;- function(data, indices) { d &amp;#60;- data[indices, ] return(cor(d[[question]], d$total&amp;#95;score&amp;#95;excl, method = &amp;#x201C;pearson&amp;#x201D;, use = &amp;#x201C;complete.obs&amp;#x201D;)) } # Step 5.8: Bootstrapping mean, SD, and point-biserial correlation with 1000 resamples mean&amp;#95;results &amp;#60;- boot(data = exam&amp;#95;data, statistic = item&amp;#95;mean, R&amp;#x2009;=&amp;#x2009;1000) sd&amp;#95;results &amp;#60;- boot(data = exam&amp;#95;data, statistic = item&amp;#95;sd, R&amp;#x2009;=&amp;#x2009;1000) pb&amp;#95;results &amp;#60;- boot(data = exam&amp;#95;data, statistic = point&amp;#95;biserial, R&amp;#x2009;=&amp;#x2009;1000) # Step 5.9: Store results mean&amp;#95;results&amp;#95;list[[question]] &amp;#60;- mean&amp;#95;results sd&amp;#95;results&amp;#95;list[[question]] &amp;#60;- sd&amp;#95;results pb&amp;#95;results&amp;#95;list[[question]] &amp;#60;- pb&amp;#95;results # Step 5.10: Calculate actual values actual&amp;#95;means[i] &amp;#60;- mean(exam&amp;#95;data[[question]], na.rm = TRUE) actual&amp;#95;sds[i] &amp;#60;- sd(exam&amp;#95;data[[question]], na.rm = TRUE) actual&amp;#95;point&amp;#95;biserials[i] &amp;#60;- cor(exam&amp;#95;data[[question]], exam&amp;#95;data$total&amp;#95;score&amp;#95;excl, method = &amp;#x201C;pearson&amp;#x201D;, use = &amp;#x201C;complete.obs&amp;#x201D;)}# Step 6: Visualize Bootstrapped Confidence Intervals for Point-Biserial Correlation# Step 6.1: Creating the loop for (i in 1:10) { question &amp;#60;- paste0(&amp;#x201C;Q&amp;#x201D;, i) # Step 6.12: Visualize Point-Biserial Correlation using boxplot, removing any non-finite values boot&amp;#95;pb&amp;#95;df &amp;#60;- as.data.frame(pb&amp;#95;results&amp;#95;list[[question]]$t) boot&amp;#95;pb&amp;#95;df &amp;#60;- boot&amp;#95;pb&amp;#95;df[is.finite(boot&amp;#95;pb&amp;#95;df[,1]), drop = FALSE] # Filter non-finite values colnames(boot&amp;#95;pb&amp;#95;df) &amp;#60;- &amp;#x201C;Bootstrapped&amp;#95;PointBiserial&amp;#x201D; pb&amp;#95;plot &amp;#60;- ggplot(boot&amp;#95;pb&amp;#95;df, aes(y = Bootstrapped&amp;#95;PointBiserial)) + geom&amp;#95;boxplot(fill = &amp;#x201C;green&amp;#x201D;, alpha&amp;#x2009;=&amp;#x2009;0.6) + labs(title = paste(&amp;#x201C;Bootstrapped Point-Biserial Correlations for&amp;#x201D;, question), y = &amp;#x201C;Point-Biserial Correlation&amp;#x201D;) + theme&amp;#95;minimal() ggsave(filename = paste0(&amp;#x201C;D:/Examdata/bootstrapped&amp;#95;pointbiserial&amp;#95;&amp;#x201D;, question, &amp;#x201C;&amp;#95;boxplot.png&amp;#x201D;), plot = pb&amp;#95;plot)}&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;# Step 7: Export Results to Word# Step 7.1: Create a data frame to summarize resultsresults&amp;#95;table &amp;#60;- data.frame(Question = paste0(&amp;#x201C;Q&amp;#x201D;, 1:10), Actual&amp;#95;Mean = round(actual&amp;#95;means, 3), Bootstrapped&amp;#95;Mean = round(sapply(mean&amp;#95;results&amp;#95;list, function(x) mean(x$t, na.rm = TRUE)), 3), Mean&amp;#95;CI&amp;#95;Lower = round(sapply(mean&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[4] } else { NA } }), 3), Mean&amp;#95;CI&amp;#95;Upper = round(sapply(mean&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[5] } else { NA } }), 3), Actual&amp;#95;SD = round(actual&amp;#95;sds, 3), Bootstrapped&amp;#95;SD = round(sapply(sd&amp;#95;results&amp;#95;list, function(x) mean(x$t, na.rm = TRUE)), 3), SD&amp;#95;CI&amp;#95;Lower = round(sapply(sd&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[4] } else { NA } }), 3), SD&amp;#95;CI&amp;#95;Upper = round(sapply(sd&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[5] } else { NA } }), 3), Actual&amp;#95;PointBiserial = round(actual&amp;#95;point&amp;#95;biserials, 3), Bootstrapped&amp;#95;PointBiserial = round(sapply(pb&amp;#95;results&amp;#95;list, function(x) mean(x$t, na.rm = TRUE)), 3), PB&amp;#95;CI&amp;#95;Lower = round(sapply(pb&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) { boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[4] } else { NA } }), 3), PB&amp;#95;CI&amp;#95;Upper = round(sapply(pb&amp;#95;results&amp;#95;list, function(x) { if (!any(is.na(x$t))) {&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt; boot.ci(x, type = &amp;#x201C;bca&amp;#x201D;)$bca[5] } else { NA } }), 3))# Step 7.2: Convert to flextableft &amp;#60;- flextable(results&amp;#95;table)ft &amp;#60;- set&amp;#95;header&amp;#95;labels(ft, Actual&amp;#95;Mean = &amp;#x201C;Actual Mean&amp;#x201D;, Bootstrapped&amp;#95;Mean = &amp;#x201C;Bootstrapped Mean&amp;#x201D;, Mean&amp;#95;CI&amp;#95;Lower = &amp;#x201C;Mean CI Lower&amp;#x201D;, Mean&amp;#95;CI&amp;#95;Upper = &amp;#x201C;Mean CI Upper&amp;#x201D;, Actual&amp;#95;SD = &amp;#x201C;Actual SD&amp;#x201D;, Bootstrapped&amp;#95;SD = &amp;#x201C;Bootstrapped SD&amp;#x201D;, SD&amp;#95;CI&amp;#95;Lower = &amp;#x201C;SD CI Lower&amp;#x201D;, SD&amp;#95;CI&amp;#95;Upper = &amp;#x201C;SD CI Upper&amp;#x201D;, Actual&amp;#95;PointBiserial = &amp;#x201C;Actual Point-Biserial&amp;#x201D;, Bootstrapped&amp;#95;PointBiserial = &amp;#x201C;Bootstrapped Point-Biserial&amp;#x201D;, PB&amp;#95;CI&amp;#95;Lower = &amp;#x201C;PB CI Lower&amp;#x201D;, PB&amp;#95;CI&amp;#95;Upper = &amp;#x201C;PB CI Upper&amp;#x201D;)ft &amp;#60;- bold(ft, part = &amp;#x201C;header&amp;#x201D;)ft &amp;#60;- fontsize(ft, size&amp;#x2009;=&amp;#x2009;12)ft &amp;#60;- theme&amp;#95;vanilla(ft)# Step 7.3: Create Word document and add the tabledoc &amp;#60;- read&amp;#95;docx()doc &amp;#60;- body&amp;#95;add&amp;#95;flextable(doc, value = ft)# Step 7.4: Save the Word documentprint(doc, target = &amp;#x201C;D:/Examdata/bootstrapped&amp;#95;results.docx&amp;#x201D;)# Step 7.4: Confirm file creationprint(&amp;#x201C;The bootstrapped results table has been saved as D:/Examdata/bootstrapped&amp;#95;results.docx&amp;#x201D;)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Summary of the steps</p> <p></p> <ulist> <item> Step 1–4: Set up and load your dataset, ensuring it is correctly structured, and then calculate each student's total score.</item> <p></p> <item> Step 5: Perform bootstrapping for item mean (ID), SD, and PBC, with functions for each metric and confidence intervals to assess reliability.</item> <p></p> <item> Step 6: Visualize the PBC confidence intervals using ggplot2.</item> <p></p> <item> Step 7: Export all results in a formatted table to a Word document with officer and flextable, making the data presentation‐ready.</item> </ulist> <hd id="AN0187112559-19">How to tailor this code to your needs?</hd> <p>Assuming you have 15 questions instead of 10 (you can have as many questions as you wish, you need to change the code in the same way). Here is all that you need to change:</p> <p>Step 4: Calculate total scores for each student</p> <p></p> <ulist> <item> <bold> Change rowSums(exam_data[, paste0("Q", 1:10)]) to rowSums(exam_data[, paste0("Q", 1: _B_15</bold>)]).</item> </ulist> <p>Step 5: Bootstrapping item metrics</p> <p>Update all occurrences of 1:10 to 1:<bold>15</bold> in the following areas:</p> <p></p> <ulist> <item> <bold> Loop initialization: Modify for (i in 1:10) to for (i in 1: _B_15</bold>) so that the loop iterates over all 15 items.</item> <p></p> <item> <bold> Question exclusion in point‐biserial calculation: Modify setdiff(1:10, i) to setdiff(1: _B_15</bold> , i) so the total score calculation excludes the specific item from the new total of 15 items.</item> <p></p> <item> Initialize vectors for actual values: Update the vectors' initialization from 10 to 15:</item> <p></p> <item> actual_means &lt;‐ numeric(<reflink idref="bib15" id="ref47">15</reflink>)</item> <p></p> <item> actual_sds &lt;‐ numeric(<reflink idref="bib15" id="ref48">15</reflink>)</item> <p></p> <item> actual_point_biserials &lt;‐ numeric(<reflink idref="bib15" id="ref49">15</reflink>)</item> </ulist> <p>Step 6: Visualization and storage of results</p> <p>In the visualization section, replace 1:10 with 1:15 in the for loop.</p> <p>for (i in 1:15) {</p> <p>Step 7: Export results to word</p> <p></p> <ulist> <item> In the results_table data frame, replace paste0("Q", 1:10) with paste0("Q", 1:15) to accommodate all 15 items.</item> <p></p> <item> results_table &lt;‐ data.frame(</item> <p></p> <item> Question = paste0("Q", 1:15),</item> <p></p> <item> Actual_Mean = round(actual_means, 3),</item> </ulist> <p>With these five minor changes, you can modify the code to analyze your entire exam questions. Needless to mention that you need to create a new.CSV file that contains all the questions (coded as 0 or 1).</p> <hd id="AN0187112559-20">Step‐by‐step guide for bootstrapping item statistics (mean, SD, ID, and PBCs) in SPSS</hd> <p>SPSS (IBM)[<reflink idref="bib36" id="ref50">36</reflink>] stands for Statistical Package for the Social Sciences and is a powerful software suite for data analysis, data management, and statistical reporting. Widely used in fields like psychology, education, and social sciences, SPSS simplifies complex data analysis tasks through a user‐friendly interface and extensive support for statistical procedures. It provides tools for descriptive statistics, inferential statistics, regression, and advanced modeling, as well as tools for visualizing data.</p> <p>However, in contrast to R, which is open‐source and freely available, SPSS requires a costly annual subscription. SPSS licenses typically need to be renewed yearly, which can be a significant investment, especially for institutions or individuals needing access to its advanced features.</p> <p>For this analysis, we will</p> <p></p> <ulist> <item> Import the dataset (exam.csv) and load it into SPSS and save it as "Exam.sav."</item> <p></p> <item> Calculate mean (ID) and standard deviation for each item.</item> <p></p> <item> Compute PBC for each question.</item> <p></p> <item> Use bootstrapping to obtain confidence intervals for these statistics.</item> </ulist> <p>Step 1: Import the data into SPSS</p> <p></p> <ulist> <item> Open SPSS and go to File &gt; Open &gt; Data.</item> <p></p> <item> Navigate to the file location, for example, 'D:\examdata\exam.sav', and open the file.</item> <p></p> <item> Verify the data structure: each row represents a student, and each column (e.g., 'Q1' to 'Q10') represents a question scored as '0' (incorrect) or '1' (correct).</item> </ulist> <p>Step 2: Calculate descriptive statistics with bootstrap (mean and standard deviation)</p> <p></p> <ulist> <item> Go to Analyze &gt; Descriptive Statistics &gt; Descriptives.</item> <p></p> <item> Select all items ('Q1' to 'Q10') and move them to the Variable(s) box.</item> <p></p> <item> Click Bootstrap at the bottom of the Descriptives dialog box.</item> <p></p> <item> Check Perform bootstrapping and specify the number of bootstrap samples, typically 1000, for reliable estimates.</item> <p></p> <item> Ensure Bias Corrected Accelerated (BCa) Confidence Intervals is selected to obtain bootstrap confidence intervals.</item> <p></p> <item> Click OK.</item> </ulist> <p>SPSS will provide the mean and SD for each item, along with bootstrapped confidence intervals, giving insight into the stability of ID and spread across samples.</p> <p>Step 3: Calculate item difficulty with bootstrap</p> <p>In general, SPSS does not have a designated function or tab specifically for calculating ID or PBC; these parameters must be manually coded. However, as previously mentioned, for dichotomous items, the mean of a question represents the ID. Therefore, bootstrapping the mean in the previous section equates to bootstrapping the ID, and the CI of the bootstrapped mean serves as the CI for the bootstrapped ID.</p> <p>Step 4: Calculate PBCs with bootstrap</p> <p>The following section includes a macro with a sequence of commands, some of which may require familiarity with SPSS syntax. To assist users, explanations are added for parameter adjustments and setup, particularly for looping commands and key settings for bootstrapping.</p> <p>Suppose we have an exam with 10 questions (Q1 to Q10). To calculate the PBC for each question:</p> <p></p> <ulist> <item> The PBC for Q1 should be calculated between the scores on Q1 and the sum of scores from Q2 to Q10. This way, the correlation shows how well Q1 relates to the overall test performance excluding Q1 itself.</item> <p></p> <item> Similarly, for Q2, the PBC would require calculating the total score over Q1 and Q3 to Q10 (excluding Q2), and so on for each item.</item> </ulist> <p>SPSS does not have a dedicated function for calculating PBCs. However, there are two ways to work around this limitation in SPSS:</p> <hd id="AN0187112559-21">Approach 1: Using reliability analysis</hd> <p>SPSS's Reliability Analysis function provides an option that can serve as a workaround for calculating PBCs. Here's how:</p> <p></p> <ulist> <item> Go to Analyze &gt; Scale &gt; Reliability Analysis.</item> <p></p> <item> Move all items (e.g., Q1 to Q10) into the Items box.</item> <p></p> <item> Click Statistics and check Correlations under the Inter‐Item tab. Then, under the Descriptives for tab, check Scale if item deleted.</item> <p></p> <item> Click OK.</item> </ulist> <p>In the output, the "Corrected Item‐Total Correlation" under the "Item‐Total Statistics" table provides the PBC for each item, calculated by correlating each item with the total test score excluding that item. However, the main shortcoming of this user‐friendly approach to calculate PBC is that SPSS does not offer any option to bootstrap Corrected Item‐Total Correlation. Therefore, the PBC and bootstrap must be coded. The good news is that, comparing to R, the coding is simpler with another caveat: SPSS does not allow to extract the results automatically in Word tables.</p> <hd id="AN0187112559-22">Approach 2: Manual calculation of PBC with bootstrap in SPSS</hd> <p>Since SPSS does not have a dedicated analysis tab for PBC or an option to bootstrap the Corrected Item‐Total Correlation within reliability analysis, we need to follow these steps to accomplish this:</p> <p></p> <ulist> <item> Create a separate overall test score for each item that excludes that specific item. For example, for Q1, calculate a total score by summing the scores of Q2 to Q10 (for Q2, calculate a total score by summing the scores of Q1, Q3 to Q10), thereby creating a total score that excludes Q1.</item> <p></p> <item> Calculate the Pearson correlation between Q1 and this adjusted total score, and then apply bootstrapping to this Pearson correlation.</item> <p></p> <item> Repeat this process for each item in the test.</item> </ulist> <p>The following macro can accomplish this task. Here are the steps:</p> <p></p> <ulist> <item> Open a new workbook:</item> <p></p> <item> Go to File &gt; New &gt; Workbook (as shown in your menu).</item> <p></p> <item> Enter the macro syntax in the workbook:</item> <p></p> <item> In the Workbook, you can type or paste the syntax code directly into a cell, much like a spreadsheet or document editor.</item> <p></p> <item> SPSS will recognize the code you enter as syntax.</item> <p></p> <item> Run the Syntax from the Workbook:</item> <p></p> <item> Highlight the syntax lines or the specific cell containing the syntax you want to run.</item> <p></p> <item> Click on "Run Syntax" which appears automatically under the box.</item> </ulist> <p>Here are all the components of the macro, which will be provided in its entirety below.</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#42; Compute the total scores excluding each item manually.COMPUTE Total&amp;#95;Excl&amp;#95;Q1 = SUM(Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q2 = SUM(Q1, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q3 = SUM(Q1, Q2, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q4 = SUM(Q1, Q2, Q3, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q5 = SUM(Q1, Q2, Q3, Q4, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q6 = SUM(Q1, Q2, Q3, Q4, Q5, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q7 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q8 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q9 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q10 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q9).&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Explanation:</p> <p></p> <ulist> <item> These lines create new variables (Total_Excl_Q1, Total_Excl_Q2, etc.) representing total scores for the test, excluding each specific item.</item> <p></p> <item> For instance, Total_Excl_Q1 is calculated by summing all items except Q1. This is done so that the correlation for each item is calculated with the overall test score minus that item.</item> <p></p> <item> This helps avoid inflated correlations by excluding the item itself from the total.</item> <p></p> </ulist> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#42; Compute the total scores excluding each item manually.COMPUTE Total&amp;#95;Excl&amp;#95;Q1 = SUM(Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q2 = SUM(Q1, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q3 = SUM(Q1, Q2, Q4, Q5, Q6, Q7, Q8, Q9, Q10).EXECUTE.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p></p> <ulist> <item> This command forces SPSS to immediately process all pending transformations, ensuring that the computed total scores are added to the dataset before proceeding.</item> <p></p> <item> Without EXECUTE, SPSS may wait to apply the transformations until it encounters a procedure command.</item> <p></p> </ulist> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#42; Display the variable list to confirm they were created.DISPLAY VARIABLES.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p></p> <ulist> <item> This command lists all variables in the dataset, allowing you to confirm that the new variables (Total_Excl_Q1, Total_Excl_Q2, etc.) were successfully created.</item> <p></p> <item> It' is useful for verifying that each total score excluding the relevant item has been added to the active dataset.</item> <p></p> </ulist> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;BOOTSTRAP /SAMPLING METHOD=SIMPLE /VARIABLES INPUT=Q1 Total&amp;#95;Excl&amp;#95;Q1 /CRITERIA CILEVEL=95 CITYPE=PERCENTILE NSAMPLES=1000 /MISSING USERMISSING=EXCLUDE.CORRELATIONS /VARIABLES=Q1 Total&amp;#95;Excl&amp;#95;Q1 /PRINT=TWOTAIL NOSIG FULL /MISSING=PAIRWISE.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p> <bold>BOOTSTRAP</bold>: This command sets up bootstrapping for estimating the correlation between Q1 and Total_Excl_Q1.</p> <p></p> <ulist> <item> /SAMPLING METHOD=SIMPLE specifies simple random sampling for bootstrapping.</item> <p></p> <item> /VARIABLES INPUT=Q1 Total_Excl_Q1 identifies the variables to bootstrap (i.e., Q1 and Total_Excl_Q1).</item> <p></p> <item> /CRITERIA CILEVEL=95 CITYPE=PERCENTILE NSAMPLES=1000 sets the confidence interval level to 95%, the type to percentile, and the number of bootstrap samples to 1000.</item> <p></p> <item> /MISSING USERMISSING=EXCLUDE specifies that user‐defined missing values are to be excluded.</item> </ulist> <p>Correlations:</p> <p></p> <ulist> <item> This command calculates the Pearson correlation between Q1 and Total_Excl_Q1 (which is in case of dichotomous response options identical with PBC) using pairwise exclusion for missing data.</item> <p></p> <item> /VARIABLES=Q1 Total_Excl_Q1 specifies the variables to correlate.</item> <p></p> <item> /PRINT=TWOTAIL NOSIG FULL requests the two‐tailed significance, suppresses asterisks for significance, and provides the full output of the correlation.</item> <p></p> <item> /MISSING=PAIRWISE handles missing data by using pairwise exclusion.</item> </ulist> <p>This box provides the entire syntax</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#42; Compute the total scores excluding each item manually.COMPUTE Total&amp;#95;Excl&amp;#95;Q1 = SUM(Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q2 = SUM(Q1, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q3 = SUM(Q1, Q2, Q4, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q4 = SUM(Q1, Q2, Q3, Q5, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q5 = SUM(Q1, Q2, Q3, Q4, Q6, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q6 = SUM(Q1, Q2, Q3, Q4, Q5, Q7, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q7 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q8, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q8 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q9, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q9 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q10).COMPUTE Total&amp;#95;Excl&amp;#95;Q10 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q9).&amp;#42; Ensure the new variables are saved to the active dataset.EXECUTE.&amp;#42; Display the variable list to confirm they were created.DISPLAY VARIABLES.BOOTSTRAP /SAMPLING METHOD=SIMPLE /VARIABLES INPUT=Q1 Total&amp;#95;Excl&amp;#95;Q1 /CRITERIA CILEVEL=95 CITYPE=PERCENTILE NSAMPLES=1000 /MISSING USERMISSING=EXCLUDE.CORRELATIONS /VARIABLES=Q1 Total&amp;#95;Excl&amp;#95;Q1 /PRINT=TWOTAIL NOSIG FULL /MISSING=PAIRWISE.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>By copying and pasting the BOOTSTRAP and CORRELATIONS sections of the syntax and replacing Q1 with Q2 (and subsequently with Q3, Q4, etc.), you can calculate the PBC for each remaining question and apply bootstrapping.</p> <hd id="AN0187112559-23">How to tailor the SPSS code</hd> <p>Including additional items in the descriptive analysis (with bootstrapping) is straightforward as highlighted in Step2.</p> <p>To extend the PBC calculation, for instance, adding Q11, two steps are needed.</p> <p>1. Adding a line to the "Compute Total" code for Q11: e.g.</p> <p>COMPUTE Total_Excl_Q11 = SUM(Q1, Q2, Q3, Q4, Q5, Q6, Q7, Q8, Q9, Q10).</p> <p>Note: This command can be shortened using a "to" statement, such as SUM(Q1 to Q10). However, using "to" will include any variables in exam.sav that are positioned between Q1 and Q10. Therefore, we need to ensure that only the intended items are located within this range. Otherwise, we must separate each item with a comma.</p> <p>2. Paste the bootstrap code and adjust it as shown below.</p> <p></p> <p> <ephtml> &lt;table&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;BOOTSTRAP /SAMPLING METHOD=SIMPLE /VARIABLES INPUT=Q11 Total&amp;#95;Excl&amp;#95;Q11 /CRITERIA CILEVEL=95 CITYPE=PERCENTILE NSAMPLES=1000 /MISSING USERMISSING=EXCLUDE.CORRELATIONS /VARIABLES=Q11 Total&amp;#95;Excl&amp;#95;Q11 /PRINT=TWOTAIL NOSIG FULL /MISSING=PAIRWISE.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0187112559-24">Practical application of the codes: A case study with sample student data</hd> <p>In this case study, <emph>N</emph> = 100 students participated in an exam with 10 multiple‐choice questions (Q1 to Q10), each with only one correct response option coded as 1 = correct and 0 = incorrect.</p> <hd id="AN0187112559-25">Application of the R code</hd> <p>For the R application, the student performance data was entered in an excel sheet and saved as "exam.csv" and stored in the "Examdata" folder on the D drive.</p> <p>Rstudio has four main windows (or panes). Each pane is customizable, so you can adjust its layout to suit your workflow.</p> <p>1. Source/Script Editor: Located in the top‐left (by default), this window is where you write and edit scripts, functions, and code. 2. Console: Located in the bottom‐left, the Console is where code is executed directly. You can type commands here to get immediate results, which is great for testing snippets or running quick calculations. 3. Environment/History: This window, located in the top‐right, shows your Environment tab, listing all active objects (data frames, variables, functions, etc.) in your current R session. 4. Files/Plots/Packages/Help/Viewer: This window, located in the bottom‐right, has multiple tabs such as Files, Plots, or Packages.</p> <p>The R code was entered in the script editor pane, the code was highlighted, and then executed by pressing Ctrl + Enter (or Command + Enter on Mac).</p> <p>Table 2 presents the outcome of the analysis that was generated automatically by this code.</p> <p>2 TABLE Summary of actual and bootstrapped item metrics: Mean (difficulty), standard deviation, PBC.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;A&lt;/th&gt;&lt;th align="left"&gt;B&lt;/th&gt;&lt;th align="left"&gt;C&lt;/th&gt;&lt;th align="left"&gt;D&lt;/th&gt;&lt;th align="left"&gt;E&lt;/th&gt;&lt;th align="left"&gt;F&lt;/th&gt;&lt;th align="left"&gt;G&lt;/th&gt;&lt;th align="left"&gt;H&lt;/th&gt;&lt;th align="left"&gt;I&lt;/th&gt;&lt;th align="left"&gt;G&lt;/th&gt;&lt;th align="left"&gt;K&lt;/th&gt;&lt;th align="left"&gt;L&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Q1&lt;/td&gt;&lt;td align="char" char="."&gt;0.61&lt;/td&gt;&lt;td align="char" char="."&gt;0.611&lt;/td&gt;&lt;td align="char" char="."&gt;0.510&lt;/td&gt;&lt;td align="char" char="."&gt;0.69&lt;/td&gt;&lt;td align="char" char="."&gt;0.490&lt;/td&gt;&lt;td align="char" char="."&gt;0.488&lt;/td&gt;&lt;td align="char" char="."&gt;0.451&lt;/td&gt;&lt;td align="char" char="."&gt;0.502&lt;/td&gt;&lt;td align="char" char="."&gt;0.268&lt;/td&gt;&lt;td align="char" char="."&gt;0.266&lt;/td&gt;&lt;td align="char" char="."&gt;0.081&lt;/td&gt;&lt;td align="char" char="."&gt;0.429&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q2&lt;/td&gt;&lt;td align="char" char="."&gt;0.50&lt;/td&gt;&lt;td align="char" char="."&gt;0.498&lt;/td&gt;&lt;td align="char" char="."&gt;0.390&lt;/td&gt;&lt;td align="char" char="."&gt;0.59&lt;/td&gt;&lt;td align="char" char="."&gt;0.503&lt;/td&gt;&lt;td align="char" char="."&gt;0.500&lt;/td&gt;&lt;td align="char" char="."&gt;0.502&lt;/td&gt;&lt;td align="char" char="."&gt;0.503&lt;/td&gt;&lt;td align="char" char="."&gt;0.333&lt;/td&gt;&lt;td align="char" char="."&gt;0.331&lt;/td&gt;&lt;td align="char" char="."&gt;0.141&lt;/td&gt;&lt;td align="char" char="."&gt;0.482&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q3&lt;/td&gt;&lt;td align="char" char="."&gt;0.50&lt;/td&gt;&lt;td align="char" char="."&gt;0.502&lt;/td&gt;&lt;td align="char" char="."&gt;0.392&lt;/td&gt;&lt;td align="char" char="."&gt;0.59&lt;/td&gt;&lt;td align="char" char="."&gt;0.503&lt;/td&gt;&lt;td align="char" char="."&gt;0.500&lt;/td&gt;&lt;td align="char" char="."&gt;0.502&lt;/td&gt;&lt;td align="char" char="."&gt;0.503&lt;/td&gt;&lt;td align="char" char="."&gt;0.405&lt;/td&gt;&lt;td align="char" char="."&gt;0.404&lt;/td&gt;&lt;td align="char" char="."&gt;0.237&lt;/td&gt;&lt;td align="char" char="."&gt;0.528&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q4&lt;/td&gt;&lt;td align="char" char="."&gt;0.19&lt;/td&gt;&lt;td align="char" char="."&gt;0.188&lt;/td&gt;&lt;td align="char" char="."&gt;0.110&lt;/td&gt;&lt;td align="char" char="."&gt;0.27&lt;/td&gt;&lt;td align="char" char="."&gt;0.394&lt;/td&gt;&lt;td align="char" char="."&gt;0.390&lt;/td&gt;&lt;td align="char" char="."&gt;0.314&lt;/td&gt;&lt;td align="char" char="."&gt;0.446&lt;/td&gt;&lt;td align="char" char="."&gt;0.155&lt;/td&gt;&lt;td align="char" char="."&gt;0.154&lt;/td&gt;&lt;td align="char" char="."&gt;&amp;#8722;0.034&lt;/td&gt;&lt;td align="char" char="."&gt;0.342&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q5&lt;/td&gt;&lt;td align="char" char="."&gt;0.17&lt;/td&gt;&lt;td align="char" char="."&gt;0.170&lt;/td&gt;&lt;td align="char" char="."&gt;0.100&lt;/td&gt;&lt;td align="char" char="."&gt;0.24&lt;/td&gt;&lt;td align="char" char="."&gt;0.378&lt;/td&gt;&lt;td align="char" char="."&gt;0.372&lt;/td&gt;&lt;td align="char" char="."&gt;0.288&lt;/td&gt;&lt;td align="char" char="."&gt;0.429&lt;/td&gt;&lt;td align="char" char="."&gt;0.249&lt;/td&gt;&lt;td align="char" char="."&gt;0.244&lt;/td&gt;&lt;td align="char" char="."&gt;0.064&lt;/td&gt;&lt;td align="char" char="."&gt;0.423&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q6&lt;/td&gt;&lt;td align="char" char="."&gt;0.75&lt;/td&gt;&lt;td align="char" char="."&gt;0.748&lt;/td&gt;&lt;td align="char" char="."&gt;0.640&lt;/td&gt;&lt;td align="char" char="."&gt;0.82&lt;/td&gt;&lt;td align="char" char="."&gt;0.435&lt;/td&gt;&lt;td align="char" char="."&gt;0.432&lt;/td&gt;&lt;td align="char" char="."&gt;0.368&lt;/td&gt;&lt;td align="char" char="."&gt;0.473&lt;/td&gt;&lt;td align="char" char="."&gt;0.327&lt;/td&gt;&lt;td align="char" char="."&gt;0.323&lt;/td&gt;&lt;td align="char" char="."&gt;0.127&lt;/td&gt;&lt;td align="char" char="."&gt;0.512&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q7&lt;/td&gt;&lt;td align="char" char="."&gt;0.83&lt;/td&gt;&lt;td align="char" char="."&gt;0.831&lt;/td&gt;&lt;td align="char" char="."&gt;0.740&lt;/td&gt;&lt;td align="char" char="."&gt;0.89&lt;/td&gt;&lt;td align="char" char="."&gt;0.378&lt;/td&gt;&lt;td align="char" char="."&gt;0.373&lt;/td&gt;&lt;td align="char" char="."&gt;0.302&lt;/td&gt;&lt;td align="char" char="."&gt;0.435&lt;/td&gt;&lt;td align="char" char="."&gt;0.302&lt;/td&gt;&lt;td align="char" char="."&gt;0.301&lt;/td&gt;&lt;td align="char" char="."&gt;0.072&lt;/td&gt;&lt;td align="char" char="."&gt;0.483&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q8&lt;/td&gt;&lt;td align="char" char="."&gt;0.76&lt;/td&gt;&lt;td align="char" char="."&gt;0.762&lt;/td&gt;&lt;td align="char" char="."&gt;0.650&lt;/td&gt;&lt;td align="char" char="."&gt;0.82&lt;/td&gt;&lt;td align="char" char="."&gt;0.429&lt;/td&gt;&lt;td align="char" char="."&gt;0.426&lt;/td&gt;&lt;td align="char" char="."&gt;0.359&lt;/td&gt;&lt;td align="char" char="."&gt;0.469&lt;/td&gt;&lt;td align="char" char="."&gt;0.325&lt;/td&gt;&lt;td align="char" char="."&gt;0.325&lt;/td&gt;&lt;td align="char" char="."&gt;0.102&lt;/td&gt;&lt;td align="char" char="."&gt;0.509&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q9&lt;/td&gt;&lt;td align="char" char="."&gt;0.78&lt;/td&gt;&lt;td align="char" char="."&gt;0.777&lt;/td&gt;&lt;td align="char" char="."&gt;0.680&lt;/td&gt;&lt;td align="char" char="."&gt;0.85&lt;/td&gt;&lt;td align="char" char="."&gt;0.416&lt;/td&gt;&lt;td align="char" char="."&gt;0.414&lt;/td&gt;&lt;td align="char" char="."&gt;0.349&lt;/td&gt;&lt;td align="char" char="."&gt;0.461&lt;/td&gt;&lt;td align="char" char="."&gt;0.338&lt;/td&gt;&lt;td align="char" char="."&gt;0.337&lt;/td&gt;&lt;td align="char" char="."&gt;0.125&lt;/td&gt;&lt;td align="char" char="."&gt;0.525&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q10&lt;/td&gt;&lt;td align="char" char="."&gt;0.84&lt;/td&gt;&lt;td align="char" char="."&gt;0.839&lt;/td&gt;&lt;td align="char" char="."&gt;0.743&lt;/td&gt;&lt;td align="char" char="."&gt;0.90&lt;/td&gt;&lt;td align="char" char="."&gt;0.368&lt;/td&gt;&lt;td align="char" char="."&gt;0.363&lt;/td&gt;&lt;td align="char" char="."&gt;0.288&lt;/td&gt;&lt;td align="char" char="."&gt;0.423&lt;/td&gt;&lt;td align="char" char="."&gt;0.217&lt;/td&gt;&lt;td align="char" char="."&gt;0.208&lt;/td&gt;&lt;td align="char" char="."&gt;0.005&lt;/td&gt;&lt;td align="char" char="."&gt;0.432&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 <emph>Note</emph>: Q1–Q10: Exam questions 1 to 10; A. Actual mean: The mean (average) of scores for each question across all students. B. Bootstrapped mean: The mean value calculated from the bootstrapped samples, providing an estimate that includes variability. C. Mean CI lower: The lower bound of the 95% CI for the bootstrapped mean, indicating the range in which the true mean is likely to fall. D. Mean CI upper: The upper bound of the 95% CI for the bootstrapped mean. E. Actual SD: The SD of scores for each question, representing the spread or variability of the data. F. Bootstrapped SD: The SD calculated from bootstrapped samples, providing an estimate that includes sampling variability. G. SD CI lower: The lower bound of the 95% CI for the bootstrapped standard deviation. H. SD CI upper: The upper bound of the 95% CI for the bootstrapped standard deviation. I. Actual PBC: The PBC between each question's score and the total score, indicating how well the item discriminates between high and low‐performing students. J. Bootstrapped PBC: The bootstrapped estimate of the PBC. K. PB CI lower: The lower bound of the 95% CI for the bootstrapped PBC. L. PB CI upper: The upper bound of the 95% CI for the bootstrapped PBC.</p> <hd id="AN0187112559-26">DISCUSSION</hd> <p></p> <hd id="AN0187112559-27">Item means/difficulty</hd> <p>The actual means of the items range from 0.17 (Q5) to 0.84 (Q10), indicating a broad spectrum of difficulty across the questions. In the context of multiple‐choice format with only one correct response option per item, the mean also represents the ID.</p> <p>Items with means around 0.5, such as Q2 and Q3, show average difficulty, indicating a balanced distribution in student performance on these questions. In contrast, items with lower means, like Q5 (mean = 0.17), are more challenging: Only 17% of students provided the correct answer. Conversely, Q10 has a higher mean (0.84), making it a relatively easier question, with 84% of students answering correctly. Bootstrapped means are closely aligned with actual means, confirming stability in ID across 1000 generated samples. The narrow 95% CI for the bootstrapped means indicate high precision, suggesting with 95% certainty that future applications of these questions will yield similar item means, underscoring the reliability of ID across different test administrations.</p> <hd id="AN0187112559-28">Standard deviations (SDs)</hd> <p>The actual SDs range from approximately 0.37 (Q10) to 0.50 (Q2), indicating moderate variability in student performance for each item. This range reflects a reasonable spread in responses, with no item showing extreme variability. The bootstrapped SDs closely mirror the actual SDs, supporting consistent variability across bootstrap sampling. The narrow confidence intervals around these bootstrapped SDs suggest that, with 95% certainty, future test administrations will exhibit similar variability to the actual SDs. This consistency reinforces that each item reliably captures a range of student responses without excessive clustering toward correct or incorrect responses.</p> <hd id="AN0187112559-29">PBCs</hd> <p>The actual PBCs vary, showing positive values for most items, indicating that these items generally differentiate well between higher and lower performing students. Items with higher PBCs, such as Q3 (0.405), demonstrate good discrimination, meaning they effectively distinguish student performance levels. The bootstrapped PBCs closely align with actual values for most items, suggesting stable discrimination power across samples. However, Q4 and Q7 warrant closer examination, as they highlight why relying on a single parameter (ID or PBC) is insufficient. This demonstrates the importance of bootstrapping in evaluating exam items to avoid misleading conclusions about item quality.</p> <p>The PBC for Q4 is 0.155, indicating a positive relationship between item performance and overall test performance and does not necessarily suggest Q4 requires revision. With an ID of 0.19 (indicating that only 19% of students answered this question correctly), Q4 is challenging; however, it may still be considered a valuable question for covering a broad range of difficulties within the exam. However, the lower bound of the CI for the bootstrapped PBC is −0.034, suggesting that Q4 may exhibit inconsistent discrimination power in future administrations. This variability implies that, although Q4 currently distinguishes between performance levels, it may not be as effective in the future. This important diagnostic insight would have been missed without the application of bootstrapping. To ensure the psychometric quality of Q4, this finding suggests that Q4 would benefit from revision.</p> <p>The ID for Q7 is 0.83, with 83% of students answering correctly, placing it within a typical range for easier items and its PBC of 0.302, which is normally considered very good. However, the 95% CI from bootstrapping (0.072 to 0.483) suggests potential inconsistencies. The lower bound of CI suggests it may not consistently differentiate between higher and lower performing students. Accordingly, like Q4, Q7 can benefit from a careful revision. Q2 and Q3 show ideal difficulty and PBCs. Exploring the wording and format of these two questions can guide the revision of Q4 and Q7 to align them better with the intended concepts and skills of the exam, which could improve its clarity and effectiveness in assessing student understanding accurately.</p> <hd id="AN0187112559-30">Analysis of item metrics through box plots</hd> <p>To illustrate how visualizations can enhance understanding of item performance, the visualizations for PBC of Q4 (as a poor performing question) and Q2 (as a well performing question) that were generated by the code are presented and explained.</p> <p>The boxplot of the bootstrapped PBC for Q4 provides several insights into the performance of this exam item in distinguishing between higher and lower performing students (Figure 1).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/8Z8K/01aug25/ase70082-fig-0001.jpg?ephost1=dGJyMNLe80Sepq84zdnyOLCmsE6epq5Srqa4SK6WxWXS" alt="ase70082-fig-0001.jpg" title="1 The boxplot for bootstrapped PBC values for Q4 generated by R." /> </p> <p></p> <p>Central tendency and interquartile range: The boxplot shows that the median PBC value for Q4 lies around 0.1 to 0.15, indicating a weak positive discrimination ability. Although this is not close to zero, it still reflects a limited capacity to effectively differentiate between high‐ and low‐performing students. The interquartile range (the length of the box), which contains the middle 50% of the bootstrapped PBC values, is relatively narrow and remains mostly in the low‐positive range, suggesting only weak discrimination.</p> <p>Range of values: The PBC values range from slightly below zero to around 0.25. The presence of values below zero, along with the narrow spread, indicates that Q4 may not consistently separate students based on overall performance. This suggests that the item's discriminative power is variable, which could be a concern for its effectiveness in the exam.</p> <p>Outliers: There are no extreme outliers in the boxplot, but the inclusion of values close to zero (and even slightly negative) emphasizes that high‐performing students did not consistently score better on this item than low‐performing students. This inconsistency in discrimination suggests that Q4 may not be an optimal question for assessing student understanding across different performance levels.</p> <p>Implications for item revision: The distribution of PBC values around 0.1 to 0.15, combined with a narrow range, underscores the limited utility of Q4 as a reliable discriminator in the test. The visual analysis suggests that Q4 may benefit from revision to improve its clarity, alignment with learning objectives, or content relevance, ensuring it more effectively differentiates between students of varying abilities.</p> <p>The boxplot of the bootstrapped PBC for Q2 also provides several insights into the performance of this exam item in distinguishing between higher and lower performing students (Figure 2).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/8Z8K/01aug25/ase70082-fig-0002.jpg?ephost1=dGJyMNLe80Sepq84zdnyOLCmsE6epq5Srqa4SK6WxWXS" alt="ase70082-fig-0002.jpg" title="2 The boxplot for bootstrapped PBC values for Q2 generated by R." /> </p> <p></p> <p>Central tendency and interquartile range: The boxplot shows that the median PBC value for this item is around 0.35 to 0.4, which indicates a strong positive discrimination ability. This median suggests that, on average, this item effectively differentiates between high‐ and low‐performing students. The interquartile range (the length of the box), which contains the middle 50% of the bootstrapped PBC values, is relatively narrow and lies entirely within the positive range, reinforcing the item's consistent performance as a discriminative question.</p> <p>Range of values: The PBC values range from slightly above zero to around 0.4. The absence of values near zero or below indicates that this item consistently shows a positive relationship between performance on this item and overall test performance. This consistent range further supports its strength as a well‐performing question that reliably separates students based on their overall abilities.</p> <p>Outliers: There are no extreme outliers in the boxplot, indicating a lack of extreme variability. This suggests that high‐performing students consistently scored better on this item than low‐performing students, without any significant deviations. The consistent positive PBC values imply that this item is effective in assessing student understanding at varying levels of performance.</p> <p>Implications for item quality: The distribution of PBC values around a median of approximately 0.35–0.4, along with a narrow range, highlights the utility of this item as a reliable discriminator in the test. The visual analysis confirms that this item is well designed for assessing student performance and does not require any immediate revision. It demonstrates stable, positive discrimination, making it a strong and effective item on the exam.</p> <hd id="AN0187112559-33">Bootstrapping in SPSS</hd> <p>The data (Exam.csv) that was analyzed with R application was imported into SPSS 30 and saved as "exam.sav." In the first step, the Means and SDs were calculated and bootstrapped, followed by the calculation of PBCs.</p> <hd id="AN0187112559-34">RESULTS</hd> <p></p> <hd id="AN0187112559-35">Means (ID) and SDs</hd> <p>The calculation of means and SDs using the descriptive statistics tab revealed the following results:</p> <p>Table 3 shows the output. For the purpose of brevity, only the first three items are depicted:</p> <p>3 TABLE Bootstrapped means (IDs) and SDs.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;Descriptive statistics&lt;/th&gt;&lt;th align="left"&gt;Bootstrap&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;95% confidence interval&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;Statistic&lt;/th&gt;&lt;th align="left"&gt;Bias&lt;/th&gt;&lt;th align="left"&gt;SE&lt;/th&gt;&lt;th align="left"&gt;Lower&lt;/th&gt;&lt;th align="left"&gt;Upper&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Q1&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;N&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Mean&lt;/td&gt;&lt;td align="left"&gt;0.6100&lt;/td&gt;&lt;td align="left"&gt;0.0013&lt;/td&gt;&lt;td align="left"&gt;0.0503&lt;/td&gt;&lt;td align="left"&gt;0.5100&lt;/td&gt;&lt;td align="left"&gt;0.7100&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SD&lt;/td&gt;&lt;td align="left"&gt;0.49021&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.00307&lt;/td&gt;&lt;td align="left"&gt;0.01241&lt;/td&gt;&lt;td align="left"&gt;0.45605&lt;/td&gt;&lt;td align="left"&gt;0.50212&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;N&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Mean&lt;/td&gt;&lt;td align="left"&gt;0.5000&lt;/td&gt;&lt;td align="left"&gt;0.0011&lt;/td&gt;&lt;td align="left"&gt;0.0511&lt;/td&gt;&lt;td align="left"&gt;0.4003&lt;/td&gt;&lt;td align="left"&gt;0.6097&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SD&lt;/td&gt;&lt;td align="left"&gt;0.50252&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.00265&lt;/td&gt;&lt;td align="left"&gt;0.00409&lt;/td&gt;&lt;td align="left"&gt;0.48783&lt;/td&gt;&lt;td align="left"&gt;0.50252&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;N&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Mean&lt;/td&gt;&lt;td align="left"&gt;0.5000&lt;/td&gt;&lt;td align="left"&gt;0.0007&lt;/td&gt;&lt;td align="left"&gt;0.0508&lt;/td&gt;&lt;td align="left"&gt;0.4000&lt;/td&gt;&lt;td align="left"&gt;0.6000&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SD&lt;/td&gt;&lt;td align="left"&gt;0.50252&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.00261&lt;/td&gt;&lt;td align="left"&gt;0.00358&lt;/td&gt;&lt;td align="left"&gt;0.49021&lt;/td&gt;&lt;td align="left"&gt;0.50252&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Valid N (listwise)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;N&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>2 a Unless otherwise noted, bootstrap results are based on 1000 bootstrap samples.</p> <p>For each exam question, SPSS provides the number of valid cases (<emph>N</emph>), mean, and SD. SPSS also automatically includes bias and standard error for each parameter.</p> <p>The Bias column represents the difference between the original sample estimate (e.g., the mean or SD of the item responses) and the mean of the statistics obtained from the bootstrap samples. This bias value indicates whether the bootstrapped estimates systematically differ from the original sample statistic. For example, a bias of 0.00 suggests that the bootstrap estimates closely match the original sample estimates, meaning the resampling process did not introduce significant deviation. If the bias value were larger, it would indicate a tendency for the bootstrapped estimates to be either consistently higher or lower than the original statistic, potentially revealing systematic differences introduced by resampling, which would hint that the initial sample is too small and not representative.</p> <p>The standard error column, on the other hand, measures the variability of the bootstrapped statistic across the 1000 bootstrap samples. Standard error provides insight into the precision of the statistic: a lower standard error indicates that the estimate is relatively stable across different bootstrap samples, whereas a higher standard error suggests more variability. In this table, the standard error reflects the reliability of the mean and SD estimates across multiple bootstrap samples. This is particularly useful when interpreting confidence intervals, as a smaller standard error would result in narrower confidence intervals, indicating greater precision in the estimates.</p> <p>Together, Bias and Standard Error help evaluate the stability and accuracy of the bootstrap estimates, providing a clearer understanding of the reliability of the calculated statistics. Additionally, the 95% CI (Lower and Upper) columns give the range within which the true population parameter likely falls with 95% CI, based on the bootstrap distribution. These values, combined with bias and standard error, are valuable for assessing the robustness of the results.</p> <hd id="AN0187112559-36">Calculating PBCs with bootstrap in SPSS</hd> <p>The code that was presented earlier was inserted in a "Workbook" in SPSS and executed. Tables 4–6 show the results in their original SPSS format to maintain the authenticity and accuracy of the data presentation.</p> <p>4 TABLE Inserting new variables to calculate PBCs in SPSS.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Variable information&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;Variables&lt;/th&gt;&lt;th align="left"&gt;Position&lt;/th&gt;&lt;th align="left"&gt;Label&lt;/th&gt;&lt;th align="left"&gt;Measurement level&lt;/th&gt;&lt;th align="left"&gt;Role&lt;/th&gt;&lt;th align="left"&gt;Print format&lt;/th&gt;&lt;th align="left"&gt;Write format&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Q1&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q2&lt;/td&gt;&lt;td align="left"&gt;2&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q3&lt;/td&gt;&lt;td align="left"&gt;3&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q4&lt;/td&gt;&lt;td align="left"&gt;4&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q5&lt;/td&gt;&lt;td align="left"&gt;5&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q6&lt;/td&gt;&lt;td align="left"&gt;6&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q7&lt;/td&gt;&lt;td align="left"&gt;7&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q8&lt;/td&gt;&lt;td align="left"&gt;8&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q9&lt;/td&gt;&lt;td align="left"&gt;9&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Q10&lt;/td&gt;&lt;td align="left"&gt;10&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q1&lt;/td&gt;&lt;td align="left"&gt;11&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q2&lt;/td&gt;&lt;td align="left"&gt;12&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q3&lt;/td&gt;&lt;td align="left"&gt;13&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q4&lt;/td&gt;&lt;td align="left"&gt;14&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q5&lt;/td&gt;&lt;td align="left"&gt;15&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q6&lt;/td&gt;&lt;td align="left"&gt;16&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q7&lt;/td&gt;&lt;td align="left"&gt;17&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q8&lt;/td&gt;&lt;td align="left"&gt;18&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q9&lt;/td&gt;&lt;td align="left"&gt;19&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q10&lt;/td&gt;&lt;td align="left"&gt;20&lt;/td&gt;&lt;td align="left"&gt;&amp;#60;none&amp;#62;&lt;/td&gt;&lt;td align="left"&gt;Nominal&lt;/td&gt;&lt;td align="left"&gt;Input&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;td align="left"&gt;F8.2&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ulist> <item>3 <emph>Note</emph>: Variables in the working file.</item> <item>5 TABLE Bootstrap specifications in SPSS.</item> </ulist> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Bootstrap specifications&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;Sampling method&lt;/th&gt;&lt;th align="left"&gt;Simple&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Number of samples&lt;/td&gt;&lt;td align="left"&gt;1000&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Confidence interval level&lt;/td&gt;&lt;td align="left"&gt;95.0%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Confidence interval type&lt;/td&gt;&lt;td align="left"&gt;Percentile&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>6 TABLE Bootstrapped PBC for Q1 in SPSS.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Correlations&lt;/th&gt;&lt;th align="left"&gt;Q1&lt;/th&gt;&lt;th align="left"&gt;Total&amp;#95;Excl&amp;#95;Q1&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Q1&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Pearson correlation&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;0.268&lt;xref ref-type="fn" rid="tfn6" /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Sig. (two&amp;#8208;tailed)&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.007&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;N&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Bootstrap&lt;xref ref-type="fn" rid="tfn5" /&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Bias&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.001&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SE&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;td align="left"&gt;0.092&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;95% confidence interval&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Lower&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;0.087&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Upper&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;0.441&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&amp;#95;Excl&amp;#95;Q1&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Pearson correlation&lt;/td&gt;&lt;td align="left"&gt;0.268&lt;xref ref-type="fn" rid="tfn6" /&gt;&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Sig. (two&amp;#8208;tailed)&lt;/td&gt;&lt;td align="left"&gt;0.007&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;N&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;td align="left"&gt;100&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Bootstrap&lt;xref ref-type="fn" rid="tfn5" /&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Bias&lt;/td&gt;&lt;td align="left"&gt;&amp;#8722;0.001&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SE&lt;/td&gt;&lt;td align="left"&gt;0.092&lt;/td&gt;&lt;td align="left"&gt;0&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;95% confidence interval&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Lower&lt;/td&gt;&lt;td align="left"&gt;0.087&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Upper&lt;/td&gt;&lt;td align="left"&gt;0.441&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ulist> <item>4 <emph>Note</emph>: In SPSS, the correlation results are often shown twice in the output when using bootstrapping because SPSS displays both the original correlation estimate and the bootstrapped estimates, including the bias, standard error, and confidence intervals derived from the bootstrap samples. This double reporting is normal and helps in comparing the original statistics with the bootstrapped results.</item> <item>5 a Unless otherwise noted, bootstrap results are based on 1000 bootstrap samples.</item> <item>6 ** Correlation is significant at the 0.01 level (two‐tailed).</item> </ulist> <p>This part of the output shows that "Compute Total" command was executed successfully. For each exam question, a total score—excluding the exam question—was calculated and added to the data sheet (e.g., Total_Excl_Q1 displays the overall exam performance of students excluding Q1).</p> <p>This part of the output shows that we have activated bootstrap, and we are calculating the 95% CI based on 1000 samples.</p> <p>It is important to note that the slight differences in the CIs for the PBC between SPSS (0.203–0.541) and R (0.185–0.528) likely stem from variations in the bootstrapping methods and CI calculation algorithms used by each software. SPSS and R may apply different bias corrections, bootstrap methodologies, or CI types (e.g., percentile‐based vs bias‐corrected and accelerated). Additionally, differences in random seed settings and rounding conventions can contribute to minor discrepancies. These variations are expected and generally acceptable, as both intervals provide a reasonable estimate of the CI for the PBC. However, these differences are normally minor, and relying on either process will not lead to incorrect decisions. These variations are expected and generally acceptable, as both intervals provide a reasonable estimate of the CI for the PBC.</p> <p>This part of the output shows the PBC and the outcome of bootstrapping. For the purpose of brevity, we only discuss Q1 here.</p> <p>The first row under Q1 shows the original Pearson correlation between Q1 and Total_Excl_Q1 without bootstrapping. This is the raw, non‐bootstrapped PBC.</p> <p>The second row under Total_Excl_Q1 displays the bootstrapped results, including:</p> <p></p> <ulist> <item> The bootstrapped correlation (the same as the original correlation if there is no bias).</item> <p></p> <item> Bias: The difference between the original and bootstrapped correlation values.</item> <p></p> <item> Standard error: An estimate of the variability in the correlation across bootstrap samples.</item> <p></p> <item> 95% CI: The percentile‐based interval derived from the 1000 bootstrap samples.</item> </ulist> <p>The Pearson correlation/PBC between Q1 (the individual item score) and Total_Excl_Q1 (the total test score excluding Q1) is 0.268, which is statistically significant at the 0.01 level (as indicated by the "**" next to the value and the <emph>p</emph> &lt; 0.001). This positive correlation suggests that higher scores on Q1 are associated with higher overall test performance (excluding Q1 itself), indicating that Q1 aligns well with the test's overall construct.</p> <p>The bootstrap results provide additional insights into the stability and precision of the PBC. The bias for the correlation estimate is −0.001. This small bias indicates that the bootstrapped estimate is very close to the original correlation value, suggesting minimal distortion due to sampling variability. The standard error of the bootstrapped correlation is 0.084, providing a measure of the variability of the correlation estimate across the 1000 bootstrap samples. The lower bound of the 95% CI for the bootstrapped correlation is 0.087 and the upper bound of the 95% CI is 0.441. Accordingly, this CI suggests that the true PBC for Q1 is likely to fall within this range. Since the interval does not include zero, it supports the statistical significance of the correlation. These results allow us to conclude that Q1 has a moderate positive PBC with the rest of the test (Total_Excl_Q1), suggesting that it is a consistent and meaningful question in measuring student exam performance. However, the lower bound also suggests that while the correlation is positive, it is relatively weak at the lower end of the confidence interval. This implies that there is some uncertainty about the strength of the relationship. Consequently, while Q1 is generally a meaningful item in assessing student performance, its alignment with the overall test may not be as robust as items with higher correlation ranges.</p> <hd id="AN0187112559-37">DISCUSSION</hd> <p>The bootstrapping approach provided critical insights into item quality that would have likely remained unnoticed with traditional item metrics alone. Bootstrapping revealed confidence intervals for each question's PBC, highlighting the inconsistency in Q4 and Q7's discrimination power across bootstrapped samples. For instance, Q4's bootstrapped PBC intervals included negative values, signaling potential issues in its discriminatory function that a single PBC value might not have conveyed. Similarly, the bootstrapped PBC of Q7, which showed a negative median, underscored its poor alignment with test performance. Without bootstrapping, these subtleties might have been overlooked, potentially leading to the retention of poorly performing items.</p> <hd id="AN0187112559-38">Strengths of the bootstrapping approach</hd> <p>The bootstrapping method used in this analysis offers several strengths. It provides robust estimates of item performance by generating confidence intervals, capturing variability across different samples. This approach allows for a more nuanced evaluation of each question's discrimination ability, revealing underlying inconsistencies that could affect future test reliability. By exposing items like Q4 and Q7 to further scrutiny, bootstrapping facilitated targeted revisions that enhanced the test's overall validity. This iterative process not only improves individual item quality but also contributes to a more reliable and valid assessment tool, ultimately benefiting future cohorts through a more refined evaluation framework.</p> <p>Overall, this article demonstrates that bootstrapping can offer a significant advantage in assessing item metrics when dealing with small sample sizes. By providing CIs for ID and PBCs, bootstrapping adds a layer of robustness to the evaluation process, enabling educators to make informed adjustments to exam questions.</p> <p>The article is valuable for researchers and medical educators, particularly those involved in assessment design and analysis. It provides practical insights into handling small sample sizes, a common issue in medical education, where cohorts are often limited. The inclusion of confidence intervals for item metrics enables a more nuanced understanding of question performance, helping educators make evidence‐based revisions. Narrow confidence intervals indicate a high level of precision and consistency in the metric across samples, suggesting that the item performs reliably in distinguishing student performance or in representing a specific level of difficulty. This precision gives educators confidence that any observed metric (e.g., a specific difficulty level) will likely generalize to new cohorts, making the item dependable for future exams. Conversely, wide confidence intervals indicate greater variability, which may arise from small sample sizes or inconsistencies in how students respond to the item. A wide interval for ID might suggest that students' performance fluctuates significantly, perhaps due to ambiguous question phrasing or varying levels of prior knowledge. For PBCs, a wide interval may imply that the item inconsistently distinguishes between high‐ and low‐performing students, potentially undermining the exam's reliability. When confidence intervals are wide, educators might consider revising or replacing the item to enhance its stability or using the item with caution. Thus, examining confidence interval widths enables educators to make informed judgments about which items provide reliable assessments and which may need refinement for improved educational outcomes.</p> <hd id="AN0187112559-39">Relevance of bootstrapping for medical education</hd> <p>Improving item discrimination through bootstrapping methods has valuable implications for medical education, where assessment accuracy is critical for evaluating students' readiness for being a physician. For example, refining questions to better differentiate between high‐ and low‐performing students can help identify specific areas where students may need further instruction, thereby guiding curriculum adjustments. High‐quality exam questions that reliably indicate students' competencies ensure that assessments more accurately reflect the clinical skills and knowledge required in real‐world healthcare settings. This alignment between assessment quality and educational outcomes ultimately supports the preparation of well‐qualified healthcare professionals, underscoring the practical value of robust psychometric analysis.</p> <hd id="AN0187112559-40">Limitations of bootstrapping</hd> <p>Bootstrapping offers a valuable alternative to parametric methods, especially in small sample contexts. However, it is not without its limitations. Bootstrapping assumes that the sample data is representative of the population from which it is drawn. This means that in cases of highly skewed or sparse data, bootstrapped estimates may become unreliable, particularly for metrics like confidence intervals. For instance, if a small sample contains an unrepresentative proportion of high or low scorers, the bootstrapped confidence intervals for ID or discrimination could be skewed, limiting their applicability to other student groups. Furthermore, bootstrapping relies on random resampling, which can lead to increased variability in estimates if sample sizes are extremely limited. Educators and researchers should be mindful of these limitations and consider augmenting bootstrapping with additional validation techniques, such as cross‐validation or alternative resampling methods, to confirm the robustness of findings.</p> <p>Furthermore, while bootstrapping offers insights into the variability of metrics, it does not address inherent biases in item phrasing or structure, which require further qualitative analysis of the content of exam questions (including response options). In other words, while bootstrapping can enhance the quantitative assessment of item metrics, it does not address the equally important qualitative aspects of item quality, such as question clarity, alignment with learning objectives, and relevance to curriculum goals. For example, an item may statistically perform well but still be confusing or misleading due to ambiguous wording, thus impacting student comprehension and engagement. Additionally, items should be crafted to assess skills and knowledge directly aligned with curriculum objectives; otherwise, even a statistically reliable item may fail to measure what is educationally intended. By integrating qualitative item review—such as peer reviews, expert feedback, and alignment checks—with quantitative analyses, educators can ensure that questions are both technically sound and pedagogically meaningful.</p> <p>Therefore, future studies should consider combining bootstrapping with other psychometric methods (e.g., Growth Curve Analysis) to yield a more holistic assessment of item quality. Moreover, the reliance on R and SPSS may limit accessibility for educators unfamiliar with these tools, suggesting a need for user‐friendly guidelines, codes, and automated systems to facilitate broader implementation.</p> <p>In conclusion, bootstrapping is a valuable approach for enhancing assessment quality, particularly when implemented within a comprehensive evaluation system[<reflink idref="bib37" id="ref51">37</reflink>] that is mindful of educational objectives and goals, aligned with course content, and incorporates content analysis, expert ratings, and student feedback, and is complemented and supported by alternative analytic approaches such as Bayesian methods,[<reflink idref="bib38" id="ref52">38</reflink>] or Rasch analysis.[<reflink idref="bib39" id="ref53">39</reflink>] By ensuring that assessments are reflective of both the curriculum and intended competencies, and by supplementing bootstrapped metrics with expert insights, educators can create robust evaluations that accurately measure student performance and support meaningful learning outcomes.</p> <hd id="AN0187112559-41">AUTHOR CONTRIBUTIONS</hd> <p> <bold>Changiz Mohiyeddini:</bold> Conceptualization; investigation; writing – original draft; methodology; validation; visualization; writing – review and editing; software; formal analysis; project administration; supervision; resources.</p> <ref id="AN0187112559-42"> <title> REFERENCES </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> Perleth M, Picker C. High‐stakes exams in medical education: a systematic review. J Med Educ. 2011 ; 8 (2): 94 – 99. https://doi.org/10.5116/ijme.5b2e.aa44</bibtext> </blist> <blist> <bibl id="bib2" idref="ref2" type="bt">2</bibl> <bibtext> Nitko AJ, Brookhart SM. Educational assessment of students. 8th ed. Upper Saddle River, NJ : Pearson ; 2020.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref3" type="bt">3</bibl> <bibtext> Downing SM. Validity: on meaningful interpretation of assessment data. Med Educ. 2003 ; 37 (9): 830 – 837. https://doi.org/10.1046/j.1365‐2923.2003.01594.x</bibtext> </blist> <blist> <bibl id="bib4" idref="ref4" type="bt">4</bibl> <bibtext> Liu O, Frankel L, Crotts Rohr K. Assessing critical thinking in higher education: current state and directions for next‐generation assessment. ETS Res Rep Ser. 2014 ; 2014 (1): 1 – 23. https://doi.org/10.1002/ets2.12009</bibtext> </blist> <blist> <bibl id="bib5" idref="ref5" type="bt">5</bibl> <bibtext> Mody S, Johnson J, Bostrom A. The effect of exam question clarity on student performance. Med Educ. 2017 ; 51 (5): 516 – 524. https://doi.org/10.1111/medu.13225</bibtext> </blist> <blist> <bibl id="bib6" idref="ref6" type="bt">6</bibl> <bibtext> Mohiyeddini C. Enhancing exam question quality in medical education through bootstrapping. Anat Sci Educ. 2024 ; 1 – 6. https://doi.org/10.1002/ase.2522</bibtext> </blist> <blist> <bibl id="bib7" idref="ref7" type="bt">7</bibl> <bibtext> Bonett DG. Sample size requirements for testing and estimating coefficient alpha. J Educ Behav Stat. 2002 ; 27 (4): 335 – 340. https://doi.org/10.3102/10769986027004335</bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Muthén LK, Muthén BO. How to use a Monte Carlo study to decide on sample size and determine power. Struct Equ Model. 2002 ; 9 (4): 599 – 620. https://doi.org/10.1207/S15328007SEM0904_8</bibtext> </blist> <blist> <bibl id="bib9" idref="ref8" type="bt">9</bibl> <bibtext> Allen MJ, Yen WM. Introduction to measurement theory. Long Grove, IL : Waveland Press ; 2001.</bibtext> </blist> <blist> <bibtext> Crocker L, Algina J. Introduction to classical and modern test theory. Belmont, CA : Wadsworth Publishing ; 2006.</bibtext> </blist> <blist> <bibtext> Kline TJB. Psychological testing: a practical approach to design and evaluation. Thousand Oaks, CA : SAGE Publications ; 2005.</bibtext> </blist> <blist> <bibtext> Hambleton RK, Swaminathan H, Rogers HJ. Fundamentals of item response theory. Thousand Oaks, CA : SAGE Publications ; 1991.</bibtext> </blist> <blist> <bibtext> McDonald RP. Test theory: a unified treatment. Mahwah, NJ : Lawrence Erlbaum Associates ; 1999.</bibtext> </blist> <blist> <bibtext> Embretson SE, Reise SP. Item response theory for psychologists. Mahwah, NJ : Lawrence Erlbaum Associates ; 2000.</bibtext> </blist> <blist> <bibtext> Van der Linden WJ, Hambleton RK, editors. Handbook of modern item response theory. New York : Springer ; 1997.</bibtext> </blist> <blist> <bibtext> DeMars C. Item response theory ; online edn. Oxford Academic ; 2010. https://doi.org/10.1093/acprof:oso/9780195377033.001.0001</bibtext> </blist> <blist> <bibtext> Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979 ; 7 (1): 1 – 26. https://doi.org/10.1214/aos/1176344552</bibtext> </blist> <blist> <bibtext> Chernick MR. Bootstrap methods: a guide for practitioners and researchers. Hoboken, NJ : John Wiley &amp; Sons ; 2008. https://doi.org/10.1002/9780470192573</bibtext> </blist> <blist> <bibtext> Davison AC, Hinkley DV. Bootstrap methods and their application. Cambridge : Cambridge University Press ; 1997. https://doi.org/10.1017/CBO9780511802843</bibtext> </blist> <blist> <bibtext> Efron B, Tibshirani RJ. An introduction to the bootstrap. Boca Raton, FL : Chapman &amp; Hall/CRC ; 1993. https://doi.org/10.1201/9780429246593</bibtext> </blist> <blist> <bibtext> Hassan S. Standard setting in medical education: standards, methods, and psychometrics. In: Zaidi SH, Hassan S, Bigdeli S, Zehra T, editors. Global medical education in normal and challenging times. Advances in science, technology &amp; innovation. Cham : Springer ; 2024. https://doi.org/10.1007/978‐3‐031‐51244‐5_16</bibtext> </blist> <blist> <bibtext> Levine MB, Redick RJ. Examining the examination in medical education. Med Educ. 2011 ; 45 (12): 1138 – 1144. https://doi.org/10.1111/j.1365‐2923.2011.03951.x</bibtext> </blist> <blist> <bibtext> Brown A, Nidumolu A, McConnell M, Hecker K, Grierson L. Development and psychometric evaluation of an instrument to measure knowledge, skills, and attitudes towards quality improvement in health professions education: the beliefs, attitudes, skills, and confidence in quality improvement (BASiC‐QI) scale. Perspect Med Educ. 2019 ; 8 (3): 167 – 176. https://doi.org/10.1007/s40037‐019‐0511‐8</bibtext> </blist> <blist> <bibtext> Furr RM, Bacharach VR. Psychometrics: An introduction. Los Angeles, CA : Sage Publications, Inc ; 2008.</bibtext> </blist> <blist> <bibtext> Clauser BE, Hambleton RK. Item analysis. In: Kubiszyn T, editor. Educational testing and measurement. 12th ed. New York, NY : Wiley ; 2018. p. 349 – 372.</bibtext> </blist> <blist> <bibtext> Downing SM, Haladyna TM, editors. Handbook of test development. Mahwah, NJ : Lawrence Erlbaum Associates ; 2006.</bibtext> </blist> <blist> <bibtext> Downing SM, Haladyna TM. Validity and reliability of assessment in medical education. Assessment in health professions education. New York : Routledge ; 2004. p. 41 – 55.</bibtext> </blist> <blist> <bibtext> Sattler JM. Assessment of children: Cognitive foundations and applications. 6th ed. San Diego, CA : Jerome M. Sattler, Publisher, Inc ; 2018.</bibtext> </blist> <blist> <bibtext> Loscalzo J, Fauci A, Kasper D, Hauser S, Longo D, Jameson JL, editors. Harrison's principles of internal medicine. 21st ed. New York : McGraw Hill ; 2022.</bibtext> </blist> <blist> <bibtext> Ebel RL, Frisbie DA. Essentials of educational measurement. 4th ed. Englewood Cliffs, NJ : Prentice‐Hall ; 1986.</bibtext> </blist> <blist> <bibtext> Smith SR, Levin M, Everhart DE. Evaluating the effectiveness of test items and examinations. In: Herman JL, Haertel E, editors. Assessment in medical education and health professions. New York, NY : Routledge ; 2012. p. 145 – 162.</bibtext> </blist> <blist> <bibtext> Anastasi A, Urbina S. Psychological testing. 7th ed. Upper Saddle River, NJ : Prentice Hall ; 1997.</bibtext> </blist> <blist> <bibtext> Ebel RL. Essentials of educational measurement. Englewood Cliffs, NJ : Prentice‐Hall ; 1972.</bibtext> </blist> <blist> <bibtext> Kelley TL. The selection of upper and lower groups for the validation of test items. J Educ Psychol. 1939 ; 30 (1): 17 – 24. https://doi.org/10.1037/h0057123</bibtext> </blist> <blist> <bibtext> R Core Team. R: A language and environment for statistical computing. Vienna : R Foundation for Statistical Computing ; 2023.</bibtext> </blist> <blist> <bibtext> IBM Corp. IBM SPSS statistics for windows, version 30.0. Armonk, NY : IBM Corp ; 2024.</bibtext> </blist> <blist> <bibtext> Mukurunge E, Nyoni CN, Hugo L. Assessment approaches in undergraduate health professions education: towards the development of feasible assessment approaches for low‐resource settings. BMC Med Educ. 2024 ; 24 : 318. https://doi.org/10.1186/s12909‐024‐05264‐x</bibtext> </blist> <blist> <bibtext> Kreiter CD. A Bayesian perspective on constructing a written assessment of probabilistic clinical reasoning in experienced clinicians. J Eval Clin Pract. 2017 ; 23 (1) 44 – 48. https://doi.org/10.1111/jep.12469</bibtext> </blist> <blist> <bibtext> Farlie MK, Johnson C, Wilkinson T, Keating J. Refining assessment: Rasch analysis in health professional education. Focus Health Prof Educ. 2021 ; 22 (2): 88 – 104. https://doi.org/10.11157/fohpe.v22i2.569</bibtext> </blist> </ref> <aug> <p>By Changiz Mohiyeddini</p> <p>Reported by Author</p> <p></p> <p>Changiz Mohiyeddini, Ph.D. is a professor of Behavioral Medicine and Psychopathology in the Department of Foundational Medical Studies at Oakland University William Beaumont School of Medicine. His research focuses on human resiliency, emotion regulation, medical education, faculty development, quality assurance, and student engagement, success, and well‐being. Recently, he developed the theory of self‐directed teaching and launched a research program on cross‐cultural medical education. He is also interested in the application of advanced quantitative methods, evaluation, and assessment. Dr. Mohiyeddini currently serves as the Editor‐in‐Chief of Frontiers in Health Psychology.</p> </aug> <nolink nlid="nl1" bibid="bib11" firstref="ref9"></nolink> <nolink nlid="nl2" bibid="bib10" firstref="ref10"></nolink> <nolink nlid="nl3" bibid="bib12" firstref="ref11"></nolink> <nolink nlid="nl4" bibid="bib14" firstref="ref13"></nolink> <nolink nlid="nl5" bibid="bib16" firstref="ref15"></nolink> <nolink nlid="nl6" bibid="bib17" firstref="ref16"></nolink> <nolink nlid="nl7" bibid="bib18" firstref="ref18"></nolink> <nolink nlid="nl8" bibid="bib20" firstref="ref19"></nolink> <nolink nlid="nl9" bibid="bib21" firstref="ref23"></nolink> <nolink nlid="nl10" bibid="bib22" firstref="ref24"></nolink> <nolink nlid="nl11" bibid="bib23" firstref="ref25"></nolink> <nolink nlid="nl12" bibid="bib24" firstref="ref28"></nolink> <nolink nlid="nl13" bibid="bib25" firstref="ref30"></nolink> <nolink nlid="nl14" bibid="bib357" firstref="ref31"></nolink> <nolink nlid="nl15" bibid="bib26" firstref="ref33"></nolink> <nolink nlid="nl16" bibid="bib27" firstref="ref34"></nolink> <nolink nlid="nl17" bibid="bib28" firstref="ref35"></nolink> <nolink nlid="nl18" bibid="bib29" firstref="ref36"></nolink> <nolink nlid="nl19" bibid="bib30" firstref="ref37"></nolink> <nolink nlid="nl20" bibid="bib31" firstref="ref38"></nolink> <nolink nlid="nl21" bibid="bib32" firstref="ref42"></nolink> <nolink nlid="nl22" bibid="bib34" firstref="ref43"></nolink> <nolink nlid="nl23" bibid="bib35" firstref="ref44"></nolink> <nolink nlid="nl24" bibid="bib15" firstref="ref47"></nolink> <nolink nlid="nl25" bibid="bib36" firstref="ref50"></nolink> <nolink nlid="nl26" bibid="bib37" firstref="ref51"></nolink> <nolink nlid="nl27" bibid="bib38" firstref="ref52"></nolink> <nolink nlid="nl28" bibid="bib39" firstref="ref53"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1478968 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Evaluation of Exam Questions Using Bootstrapping: Practical Applications in R and SPSS with a Case Study – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Changiz+Mohiyeddini%22">Changiz Mohiyeddini</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Anatomical+Sciences+Education%22"><i>Anatomical Sciences Education</i></searchLink>. 2025 18(8):858-880. – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 23 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Sampling%22">Sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Inference%22">Statistical Inference</searchLink><br /><searchLink fieldCode="DE" term="%22Nonparametric+Statistics%22">Nonparametric Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Correlation%22">Correlation</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Construction%22">Test Construction</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1002/ase.70082 – Name: ISSN Label: ISSN Group: ISSN Data: 1935-9772<br />1935-9780 – Name: Abstract Label: Abstract Group: Ab Data: This article presents a step-by-step guide to using R and SPSS to bootstrap exam questions. Bootstrapping, a versatile nonparametric analytical technique, can help to improve the psychometric qualities of exam questions in the process of quality assurance. Bootstrapping is particularly useful in disciplines such as medical education, where student cohorts are normally too small to reliably use parametric analysis to evaluate the quality of exam questions. Traditional parametric approaches need large samples; otherwise, they can yield unreliable estimates of metrics such as item difficulty and point-biserial correlations with small cohorts, potentially misleading the evaluation of exam questions and consequently leading to flawed assessments. By employing bootstrapping, educators can resample data to obtain robust confidence intervals for key metrics. This allows for a more accurate evaluation of question quality. This guide provides a step-by-step approach using R and SPSS, along with explaining the necessary code to bootstrap exam question means, standard deviations, item difficulty, and point-biserial correlations. In addition, the code includes automated visualizations and the capability to export results in reader-friendly tables, enhancing time efficiency and streamlining both data analysis and presentation processes. Furthermore, this article includes a case study in which the code is applied and the results are discussed to showcase how bootstrapping can inform decisions regarding exam question revisions. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1478968 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1478968 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/ase.70082 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 23 StartPage: 858 Subjects: – SubjectFull: Test Items Type: general – SubjectFull: Sampling Type: general – SubjectFull: Statistical Inference Type: general – SubjectFull: Nonparametric Statistics Type: general – SubjectFull: Difficulty Level Type: general – SubjectFull: Correlation Type: general – SubjectFull: Test Construction Type: general Titles: – TitleFull: Evaluation of Exam Questions Using Bootstrapping: Practical Applications in R and SPSS with a Case Study Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Changiz Mohiyeddini IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 08 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1935-9772 – Type: issn-electronic Value: 1935-9780 Numbering: – Type: volume Value: 18 – Type: issue Value: 8 Titles: – TitleFull: Anatomical Sciences Education Type: main |
| ResultId | 1 |