When and How Biases Seep In: Enhancing Debiasing Approaches for Fair Educational Predictive Analytics
Saved in:
| Title: | When and How Biases Seep In: Enhancing Debiasing Approaches for Fair Educational Predictive Analytics |
|---|---|
| Language: | English |
| Authors: | Lin Li (ORCID |
| Source: | British Journal of Educational Technology. 2025 56(6):2478-2501. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 24 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Bias, Attitude Change, Prediction, Learning Analytics, Social Bias, Computer Mediated Communication, Social Media, Stereotypes, Labeling (of Persons) |
| DOI: | 10.1111/bjet.13575 |
| ISSN: | 0007-1013 1467-8535 |
| Abstract: | The use of predictive analytics powered by machine learning (ML) to model educational data has increasingly been identified to exhibit bias towards marginalized populations, prompting the need for more equitable applications of these techniques. To tackle bias that emerges in training data or models at different stages of the ML modelling pipeline, numerous debiasing approaches have been proposed. Yet, research into state-of-the-art techniques for effectively employing these approaches to enhance fairness in educational predictive scenarios remains limited. Prior studies often focused on mitigating bias from a single source at a specific stage of model construction within narrowly defined scenarios, overlooking the complexities of bias originating from multiple sources across various stages. Moreover, these approaches were often evaluated using typical threshold-dependent fairness metrics, which fail to account for real-world educational scenarios where thresholds are typically unknown before evaluation. To bridge these gaps, this study systematically examined a total of 28 representative debiasing approaches, categorized by the sources of bias and the stage they targeted, for two critical educational predictive tasks, namely forum post classification and student career prediction. Both tasks involve a two-phase modelling process where features learned from upstream models in the first phase are fed into classical ML models for final predictions, which is a common yet under-explored setting for educational data modelling. The study observed that addressing local stereotypical bias, label bias or proxy discrimination in training data, as well as imposing fairness constraints on models, can effectively enhance predictive fairness. But their efficacy was often compromised when features from upstream models were inherently biased. Beyond that, this study proposes two novel strategies, namely Multi-Stage and Multi-Source debiasing to integrate existing approaches. These strategies demonstrated substantial improvements in mitigating unfairness, underscoring the importance of unified approaches capable of addressing biases from various sources across multiple stages. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1486314 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFeOIXEKlY_jvq2A0CYgw0-AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDDUhquONM0WNGjUBQQIBEICBm_MXhopZTVRAP3k_lLBdeKuqA03B6BZaQSqCQonokrA1i2HK4BvtALwnsEtbRkRxsiX6_LsO0VoszW4Rd0HQKFA4uIMAQy3IaWIVnzhbVRjfc8GXdC8ku42T4gX4MdwwI3g3b0PIu7BOVNRkVapLak-bJkl1t8IKAqQLKrUTS64o9WKsBkFewPs-7oLUnUmD_IdWJgtYhvTgo6Ws Text: Availability: 1 Value: <anid>AN0188606177;58i01nov.25;2025Oct14.06:24;v2.2.500</anid> <title id="AN0188606177-1">When and how biases seep in: Enhancing debiasing approaches for fair educational predictive analytics </title> <p>The use of predictive analytics powered by machine learning (ML) to model educational data has increasingly been identified to exhibit bias towards marginalized populations, prompting the need for more equitable applications of these techniques. To tackle bias that emerges in training data or models at different stages of the ML modelling pipeline, numerous debiasing approaches have been proposed. Yet, research into state‐of‐the‐art techniques for effectively employing these approaches to enhance fairness in educational predictive scenarios remains limited. Prior studies often focused on mitigating bias from a single source at a specific stage of model construction within narrowly defined scenarios, overlooking the complexities of bias originating from multiple sources across various stages. Moreover, these approaches were often evaluated using typical threshold‐dependent fairness metrics, which fail to account for real‐world educational scenarios where thresholds are typically unknown before evaluation. To bridge these gaps, this study systematically examined a total of 28 representative debiasing approaches, categorized by the sources of bias and the stage they targeted, for two critical educational predictive tasks, namely forum post classification and student career prediction. Both tasks involve a two‐phase modelling process where features learned from upstream models in the first phase are fed into classical ML models for final predictions, which is a common yet under‐explored setting for educational data modelling. The study observed that addressing local stereotypical bias, label bias or proxy discrimination in training data, as well as imposing fairness constraints on models, can effectively enhance predictive fairness. But their efficacy was often compromised when features from upstream models were inherently biased. Beyond that, this study proposes two novel strategies, namely Multi‐Stage and Multi‐Source debiasing to integrate existing approaches. These strategies demonstrated substantial improvements in mitigating unfairness, underscoring the importance of unified approaches capable of addressing biases from various sources across multiple stages. Practitioner notesWhat is already known about this topic Predictive analytics for educational data modelling often exhibit bias against students from certain demographic groups based on sensitive attributes.Bias can emerge in training data or models at different time points of the ML modelling pipeline, resulting in unfair final predictions.Numerous debiasing approaches have been developed to tackle bias at different stages, including pre‐processing training data, in‐processing models, and post‐processing predicted outcomes or trained models.What this paper adds A systematic evaluation of 28 state‐of‐the‐art debiasing approaches covering multiple sources of biases and multiple stages across two different educational predictive scenarios, identifying leading sources of data biases contributing to predictive unfairness.Further enhancing predictive fairness with proposed debiasing strategies considering the multi‐source and multi‐stage characteristics of biases.Revealing potential risks of debiasing focused on a single sensitive attribute.Implications for practitioners Pre‐processing approaches, particularly those addressing stereotypical bias, label bias and proxy discrimination, are generally effective for improving fairness in educational predictions. Re‐weighing methods are especially useful for smaller datasets to tackle stereotypical bias.When dealing with two‐phase modelling, biases inherently encoded in the features generated from upstream models might not be effectively addressed by debiasing approaches applied to downstream models.Combining debiasing approaches to tackle multiple sources of biases across multiple stages significantly enhances predictive fairness.</p> <p>Keywords: debiasing approaches; ethical learning analytics; fairness; predictive analytics</p> <hd id="AN0188606177-2">INTRODUCTION</hd> <p>The widespread adoption of ML techniques across various real‐world decision‐making settings to support students has increasingly raised concerns about their potential to exhibit bias against certain demographic groups based on sensitive attributes, such as gender (Baker &amp; Hawn, [<reflink idref="bib7" id="ref1">7</reflink>]). Numerous studies have highlighted instances of such bias in predictions (Li et al., [<reflink idref="bib28" id="ref2">28</reflink>]; [<reflink idref="bib29" id="ref3">29</reflink>]), including higher error rates when predicting course success for female students compared with male students (Litman et al., [<reflink idref="bib30" id="ref4">30</reflink>]). These biased predictions carry particularly far‐reaching consequences in education (Kizilcec &amp; Lee, [<reflink idref="bib26" id="ref5">26</reflink>]), as decisions informed by them may directly influence students' access to resources or opportunities (e.g., determining whether a student requires targeted learning support based on at‐risk predictions). Such decisions can ultimately shape students' educational outcomes, including course grades and future career trajectories. Therefore, to ensure that students from any demographic group are not disadvantaged by predictive models, it is crucial to develop and apply fairness‐aware approaches in education.</p> <p>To enhance fairness in predictions, i.e., predictive fairness, various debiasing approaches (Caton &amp; Haas, [<reflink idref="bib14" id="ref6">14</reflink>]) have been proposed. Unfairness in predictions can stem from diverse sources of bias that may emerge at distinct stages of building a ML model. For instance, during the data preparation stage, distribution‐related biases, including class imbalance, representation bias (uneven group populations (Dominguez‐Catena et al., [<reflink idref="bib18" id="ref7">18</reflink>])), local stereotypical bias (i.e., over‐ or under‐representation of specific demographic groups within a class label) and global stereotypical bias (i.e., uneven joint distributions of sensitive attributes and class labels), label bias (i.e., mismatch between the true outcome of interest and the available data (Jiang &amp; Nachum, [<reflink idref="bib24" id="ref8">24</reflink>])) and proxy discrimination (e.g., when seemingly neutral attributes are strongly correlated with sensitive attributes, resulting in bias in predictions (Tschantz, [<reflink idref="bib41" id="ref9">41</reflink>])), have been cited as primary sources of data biases. Particularly in two‐phase tasks like text classification, features learned from pre‐trained large language models (LLMs) may encode bias, known as <emph>embedding bias</emph> (BehnamGhader &amp; Milios, [<reflink idref="bib10" id="ref10">10</reflink>]). In education, data collected from students is particularly prone to these biases due to factors, such as subjective judgements (Artelt, [<reflink idref="bib5" id="ref11">5</reflink>]; Campbell, [<reflink idref="bib13" id="ref12">13</reflink>]), socio‐cultural influences (García &amp; Weiss, [<reflink idref="bib20" id="ref13">20</reflink>]) and historical inequalities (Sadker and Sadker, [<reflink idref="bib38" id="ref14">38</reflink>]). For instance, female students have historically been under‐represented in STEM fields (Beede et al., [<reflink idref="bib9" id="ref15">9</reflink>]). Teachers' perceptions or grading practices might unintentionally favour certain groups (Campbell, [<reflink idref="bib13" id="ref16">13</reflink>]; Malouff &amp; Thorsteinsson, [<reflink idref="bib32" id="ref17">32</reflink>]), leading to biased labels. Additionally, student interactions in educational settings (Ashong &amp; Commander, [<reflink idref="bib6" id="ref18">6</reflink>]; Zhuhadar et al., [<reflink idref="bib48" id="ref19">48</reflink>]), such as classrooms, online forums or assessments, may exhibit significantly different patterns related to gender or race, potentially posing fairness risks when these patterns are linked to certain outcomes. For example, research (Aguillon et al., [<reflink idref="bib2" id="ref20">2</reflink>]; Moss‐Racusin et al., [<reflink idref="bib34" id="ref21">34</reflink>]) has shown that female students participate less than expected in college STEM classrooms as compared to male students, which can disadvantage them if classroom participation is considered a factor for grading. These different sources of data bias may be further encoded during model training, resulting in biased models that generate unfair predictions. Consequently, a diverse range of approaches that pre‐process the data, in‐process the model or post‐process the model predictions have been developed to address these biases (Li et al., [<reflink idref="bib28" id="ref22">28</reflink>]).</p> <p>Although numerous studies (Deho et al., [<reflink idref="bib17" id="ref23">17</reflink>]; Han, Shen, et al., [<reflink idref="bib23" id="ref24">23</reflink>]; Lamba et al., [<reflink idref="bib27" id="ref25">27</reflink>]) have systematically investigated these approaches, understanding their effective use to address real‐world educational problems still remains limited for several reasons: (i) Existing studies often overlook the complexities of data bias originating from diverse sources, covering only a limited range of biases. For example, <emph>label bias</emph> (Jiang &amp; Nachum, [<reflink idref="bib24" id="ref26">24</reflink>]), as a critical source of bias, has been largely ignored in existing evaluation studies. (ii) Previous studies predominantly measure fairness using metrics focused on binary class labels, which require a pre‐determined fixed threshold for evaluation. This approach may be impractical for real‐world scenarios and is often unsuitable for educational tasks where the threshold is often unknown before evaluation (Gardner et al., [<reflink idref="bib21" id="ref27">21</reflink>]). (iii) Educational predictive tasks increasingly involve multiple models in combination, where upstream models first extract features that are then used for downstream predictions. For example, in text modelling, large language models (LLMs) serve as upstream models to generate features for downstream tasks (Clavié &amp; Gal, [<reflink idref="bib16" id="ref28">16</reflink>]). Similarly, in student modelling tasks, features characterizing students are learned via one model and then used for task‐specific predictions (Yeung &amp; Yeung, [<reflink idref="bib46" id="ref29">46</reflink>]). However, fairness remains rarely explored in these contexts, limiting insights into the generalizability of existing debiasing approaches. (iv) Debiasing approaches typically address bias related to a single sensitive attribute, yet students often belong to multiple demographic groups. Ensuring fairness for one attribute (e.g., gender) does not necessarily ensure fairness for another (e.g., race). The interactions and trade‐offs between addressing fairness for different sensitive attributes have been largely overlooked in prior studies.</p> <p>To fill these gaps, this study systematically investigated 28 state‐of‐the‐art debiasing approaches across two critical educational tasks, that is, forum post classification (i.e., predicting whether a forum post is relevant to course content) and STEM career prediction (i.e., predicting whether a student will pursue a STEM career after college based on their middle school math learning records). This study aimed to address the following research questions: <emph>RQ1</emph>: <emph>When and how</emph> can predictive fairness be effectively enhanced using existing debiasing approaches? <emph>RQ2</emph>: How does debiasing focused on one sensitive attribute impact fairness regarding another sensitive attribute? and <emph>RQ3</emph>: To what extent can fairness be further enhanced by combining multiple debiasing approaches (considering multiple sources of bias across multiple stages)? The contributions of this study to fair educational predictive analytics included as follows: (i) evaluating pre‐processing approaches targeting six distinct sources of data bias, as well as in‐processing and post‐processing methods, (ii) focusing on two‐phase predictive tasks involving textual data and tabular data, (iii) assessing predictive fairness with a threshold‐independent metric tailored for the education domain, (iv) exploring the effectiveness of combined debiasing strategies for tackling multi‐source bias across multiple stages and (v) examining how debiasing focused on one sensitive attribute affects students categorized by another sensitive attribute.</p> <hd id="AN0188606177-3">RELATED WORK</hd> <p>Predictive fairness concerns the impartiality of ML model decisions for users with different sensitive attributes like gender (Li et al., [<reflink idref="bib28" id="ref30">28</reflink>]). In education, it applies to decisions, such as determining whether a student should receive additional support (Bayer et al., [<reflink idref="bib8" id="ref31">8</reflink>]), which can significantly affect students' opportunities and outcomes. Researchers address fairness through two key aspects (Baker &amp; Hawn, [<reflink idref="bib7" id="ref32">7</reflink>]; Li et al., [<reflink idref="bib28" id="ref33">28</reflink>]; Mehrabi et al., [<reflink idref="bib33" id="ref34">33</reflink>]), that is, <emph>measurement</emph>, which involves assessing the fairness of predicted outcomes, and <emph>enhancement</emph>, which focuses on improving predictive fairness by mitigating any identified biases in the predictive process.</p> <hd id="AN0188606177-4">Measuring predictive fairness</hd> <p>In terms of fairness <emph>measurement</emph>, two key frequently discussed concepts are group fairness and individual fairness (Li et al., [<reflink idref="bib28" id="ref35">28</reflink>]). This study focused on group fairness, which has been commonly assessed in classification tasks using well‐established metrics, such as <emph>Demographic Parity (DP)</emph>, <emph>Equal Opportunity (EqOpp)</emph> and <emph>Equalized Odds (EqOdds)</emph>. These metrics are based on class labels and often require a pre‐defined threshold, which limits their applicability since thresholds depend on specific deployment contexts (Gardner et al., [<reflink idref="bib21" id="ref36">21</reflink>]). On this premise, Gardner et al. proposed a metric named Absolute Between‐ROC Area (ABROCA) to measure the difference in overall error rate across groups using the area between ROC curves of different groups. In a different vein, a more recent metric called <emph>Model Absolute Density Distance (MADD)</emph> (Verger et al., [<reflink idref="bib43" id="ref37">43</reflink>]) measures the difference in the distribution of predicted probabilities between groups, regardless of the actual class labels. It examines differences in probabilities across the entire range of possible values, which could be considered a generalized version of Demographic Parity that typically compares the number of instances with a predicted probability above a certain threshold (e.g., 0.5).</p> <p>It is important to recognize that some fairness metrics may be more suitable than others, depending on the specific scenario. For situations involving decisions related to provisions of necessary learning support or interventions, such as identifying students at risk of failing a course and requiring tutoring, metrics prioritizing equal predictive performance (i.e., identifying at‐risk students with equally accurate rates across different groups) may be more appropriate (Caton &amp; Haas, [<reflink idref="bib14" id="ref38">14</reflink>]; Kizilcec &amp; Lee, [<reflink idref="bib26" id="ref39">26</reflink>]). Metrics, such as EqOdds and ABROCA, which target equal error rates (e.g., false positive and false negative rates) between groups, are particularly suited to these situations (Gardner et al., [<reflink idref="bib21" id="ref40">21</reflink>]). In contrast, for scenarios emphasizing balanced representation of different demographic groups in receiving favourable outcomes, such as access to extracurricular programmes or activities, metrics like DP and MADD are more relevant (Mehrabi et al., [<reflink idref="bib33" id="ref41">33</reflink>]). These metrics evaluate distributional differences to ensure that favourable predicted outcomes are fairly distributed across groups. This study focuses on scenarios where necessary learning supports are critical. Therefore, metrics that emphasize equal predictive performance were prioritized, and ABROCA was chosen for fairness evaluation in this study given its threshold‐independent nature.</p> <hd id="AN0188606177-5">Enhancing predictive fairness</hd> <p>Predictive unfairness can be attributed to various sources of biases emerging at different stages of constructing an ML model (Suresh &amp; Guttag, [<reflink idref="bib40" id="ref42">40</reflink>]). When preparing the data, representation bias (Dominguez‐Catena et al., [<reflink idref="bib18" id="ref43">18</reflink>]) emerges when the student population varies across different demographic groups (e.g., more males than females). Class labels might be also imbalanced. <emph>Local stereotypical bias</emph> (Dominguez‐Catena et al., [<reflink idref="bib18" id="ref44">18</reflink>]) arises when one group is more associated with either positive or negative class labels, while <emph>global stereotypical bias</emph> refers to uneven joint distributions of class labels and demographic labels. Additionally, class labels themselves might be incorrect, for example, samples from a certain group might be more frequently mislabelled, which is referred to as <emph>label bias</emph> (Jiang &amp; Nachum, [<reflink idref="bib24" id="ref45">24</reflink>]). Furthermore, some seemingly neutral features fed into a model might also encode bias by being highly correlated with a sensitive attribute as a proxy to result in unfair predictions, often called <emph>proxy discrimination</emph> (Tschantz, [<reflink idref="bib41" id="ref46">41</reflink>]). These data biases might perpetuate through model training, resulting in biased and unfair ML predictions.</p> <p>To address these biases, a wide range of methods have been proposed to take effect at different stages of the ML pipeline, categorized broadly into pre‐processing, in‐processing and post‐processing techniques. <emph>Pre‐processing</emph> approaches manipulate training data in various ways to tackle different sources of bias. These include balancing data via re‐sampling or re‐weighing to tackle distribution‐related biases (Kamiran &amp; Calders, [<reflink idref="bib25" id="ref47">25</reflink>]), correcting labels or removing mislabled samples (Jiang &amp; Nachum, [<reflink idref="bib24" id="ref48">24</reflink>]) and transforming features by excluding sensitive attributes or learning new ones via optimization (Vasquez Verdugo et al., [<reflink idref="bib42" id="ref49">42</reflink>]). <emph>In‐processing</emph> approaches either constrain a model with a desired fairness metric (Agarwal et al., [<reflink idref="bib1" id="ref50">1</reflink>]) or alter the optimization procedure itself (Zhang et al., [<reflink idref="bib47" id="ref51">47</reflink>]). <emph>Post‐processing</emph> approaches primarily focus on adjusting the classification threshold or decision boundaries (Caton &amp; Haas, [<reflink idref="bib14" id="ref52">14</reflink>]) to modify the predicted binary labels. More recent approaches have shifted focus to modifying probabilistic classifiers or their outputs, which are more suitable for educational scenarios.</p> <p>Despite several studies comparing different debiasing approaches (Friedler et al., [<reflink idref="bib19" id="ref53">19</reflink>]; Han, Shen, et al., [<reflink idref="bib23" id="ref54">23</reflink>]; Lamba et al., [<reflink idref="bib27" id="ref55">27</reflink>]), these efforts were not targeted at the education domain and overlooked key sources of bias, such as label bias and proxy discrimination. Moreover, all tasks analysed in these studies were single‐phase and focused exclusively on tabular data. Only two studies (Deho et al., [<reflink idref="bib17" id="ref56">17</reflink>]; Vasquez Verdugo et al., [<reflink idref="bib42" id="ref57">42</reflink>]) specifically addressed educational predictive tasks. However, both relied on fairness metrics based on binary class labels that require a pre‐defined threshold and failed to consider label bias and stereotypical bias. In summary, the characteristics of multi‐stage and multi‐source biases have been largely neglected in existing research, and debiasing approaches remain under‐explored regarding their effectiveness in fair educational predictive analytics. This gap motivated us to systematically evaluate existing debiasing approaches and to propose strategies for addressing multi‐source and multi‐stage biases to further enhance predictive fairness. A detailed comparison between this study and prior work is provided in Table 1.</p> <p>1 TABLE Comparison of the current study and previous studies in terms of data bias coverage, bias‐handling methods during training, types of post‐processed outputs, use of two‐phase modelling, focus on the educational domain, data types, use of threshold‐independent fairness metrics and use of combined debiasing strategies.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Characteristics&lt;/th&gt;&lt;th align="left"&gt;The current study&lt;/th&gt;&lt;th align="left"&gt;Deho et&amp;#160;al.&amp;#160;(&lt;xref ref-type="bibr" rid="bibr17"&gt;2022&lt;/xref&gt;)&lt;/th&gt;&lt;th align="left"&gt;Vasquez Verdugo et&amp;#160;al.&amp;#160;(&lt;xref ref-type="bibr" rid="bibr42"&gt;2022&lt;/xref&gt;)&lt;/th&gt;&lt;th align="left"&gt;Han, Shen, et&amp;#160;al.&amp;#160;(&lt;xref ref-type="bibr" rid="bibr23"&gt;2022&lt;/xref&gt;)&lt;/th&gt;&lt;th align="left"&gt;Lamba et&amp;#160;al.&amp;#160;(&lt;xref ref-type="bibr" rid="bibr27"&gt;2021&lt;/xref&gt;)&lt;/th&gt;&lt;th align="left"&gt;Friedler et&amp;#160;al.&amp;#160;(&lt;xref ref-type="bibr" rid="bibr19"&gt;2019&lt;/xref&gt;)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Tackling data bias&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Feature/embedding bias&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Class imbalance&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Representation bias&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Stereotypical bias&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Proxy discrimination&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Label bias&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Handling bias in model training&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Regularization&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Adversarial learning&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Optimization&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Modifying predicted outputs&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Predicted probabilistic scores&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Trained probabilistic classifier&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Education domain&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Two&amp;#8208;phase modelling process&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Data types&lt;/td&gt;&lt;td align="left"&gt;Textual &amp; tabular&lt;/td&gt;&lt;td align="left"&gt;Tabular&lt;/td&gt;&lt;td align="left"&gt;Tabular&lt;/td&gt;&lt;td align="left"&gt;Textual&lt;/td&gt;&lt;td align="left"&gt;Tabular&lt;/td&gt;&lt;td align="left"&gt;Tabular&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Threshold&amp;#8208;independent fairness metric&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Combined Debiasing&lt;/td&gt;&lt;td align="left"&gt;&amp;#10003;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;td align="left"&gt;&amp;#10005;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0188606177-6">METHODOLOGY</hd> <p>This section covers: (i) the two representative educational predictive tasks considered in this study, along with the datasets and models used; (ii) the debiasing approaches selected for enhancing predictive fairness, (iii) the metrics for evaluating the predictive accuracy and fairness, and (iv) the study setup details to ensure reproducibility.</p> <hd id="AN0188606177-7">Tasks, datasets and models</hd> <p>Two tasks, that is, <emph>Education Forum Post Classification</emph> and <emph>STEM Career Prediction</emph>, were selected as the main focus and introduced below.</p> <hd id="AN0188606177-8">Education forum post classification</hd> <p>Educational discussion forums serve as primary platforms for students to communicate their comprehension of course content, and productive forum discussions have been found to positively impact learning outcomes (Almatrafi et al., [<reflink idref="bib4" id="ref58">4</reflink>]). Therefore, it is crucial for instructors to identify educational forum messages relevant to the course content to facilitate meaningful discussions. For this purpose, to avoid demographic imbalances associated with certain courses, a dataset encompassing students' posts from diverse courses of Information Technology, Engineering, Education, Business, Economics and other disciplines was collected from the learning management system at the authors' university, along with students' gender and first‐language backgrounds. A random sample of 3693 posts was selected and manually labelled as either <emph>content‐relevant</emph> or <emph>content‐irrelevant</emph>, indicating whether a post is relevant to the course content or not. Each post was initially labelled by a junior teaching staff and subsequently reviewed by two independent senior teaching staff. Approximately 10% of the data were corrected by the two senior teaching staff. To classify these forum posts, a logistic regression coupled with BERT embeddings was adopted, which has demonstrated notable performance in text classification tasks (Clavié &amp; Gal, [<reflink idref="bib16" id="ref59">16</reflink>]).</p> <hd id="AN0188606177-9">STEM career prediction</hd> <p>To support early identification of students at risk of losing interest in math, which may affect their future career choices in STEM fields, the 2017 ASSISTments Data Mining competition aimed to predict students' future STEM career choices based on their interactions with the ASSISTments math learning system during middle school. The competition provided a publicly available dataset[<reflink idref="bib1" id="ref60">1</reflink>] (Patikorn et al., [<reflink idref="bib35" id="ref61">35</reflink>]) consisting of 942,816 interactions from 1709 students. Among these, 484 students indicated their career choice as either STEM or non‐STEM, along with their gender. For this task, this study replicate the model architecture proposed by Yeung and Yeung using their open‐source code,[<reflink idref="bib2" id="ref62">2</reflink>] which won the first place in the competition. This architecture utilizes the 'Deep Knowledge Tracing' model to extract the knowledge state of each student from their clickstream data. Subsequently, logistic regression is employed to make predictions, using both knowledge state features and student profile features.</p> <p>It is worth emphasizing that the selection of these two tasks was driven by several considerations, as detailed below: (i) These tasks represent important yet distinct categories of educational predictive analytics that focus on different levels of educational outcomes. <emph>Education Forum Post Classification</emph> aims to enhance immediate learning experiences and improve short‐term course‐related learning outcomes, while <emph>STEM Career Prediction</emph> focuses on students' behavioural patterns in acquiring knowledge and skills, aiming to improve long‐term educational outcomes, such as career trajectories in STEM fields. (ii) Additionally, it has been widely acknowledged and documented that these two task settings are vulnerable to predicted bias, which could be attributed to various factors related to either data or models, as highlighted by previous studies (Beede et al., [<reflink idref="bib9" id="ref63">9</reflink>]; Li et al., [<reflink idref="bib28" id="ref64">28</reflink>]). (iii) Moreover, including these two tasks, which deal with different types of educational data (i.e., textual data and tabular data), allows for an exploration of the generalisability of state‐of‐the‐art debiasing approaches, as well as an investigation into potential fairness challenges unique to each data type. More detailed information of the two datasets are provided in Table 2.</p> <p>2 TABLE Distribution w.r.t. sensitive attributes and class labels of training data for forum classification and career prediction.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Task&lt;/th&gt;&lt;th align="left"&gt;Class labels&lt;/th&gt;&lt;th align="left"&gt;Gender&lt;/th&gt;&lt;th align="left"&gt;Language&lt;/th&gt;&lt;th align="left"&gt;Total&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;Male&lt;/th&gt;&lt;th align="left"&gt;Female&lt;/th&gt;&lt;th align="left"&gt;English&lt;/th&gt;&lt;th align="left"&gt;Non&amp;#8208;English&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Forum post classification&lt;/td&gt;&lt;td align="left"&gt;Context&amp;#8208;relevant&lt;/td&gt;&lt;td align="char" char="("&gt;550 (74.83%)&lt;/td&gt;&lt;td align="char" char="("&gt;1046 (70.49%)&lt;/td&gt;&lt;td align="char" char="("&gt;616 (74.94%)&lt;/td&gt;&lt;td align="char" char="("&gt;980 (70.15%)&lt;/td&gt;&lt;td align="char" char="."&gt;1596&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Content&amp;#8208;irrelevant&lt;/td&gt;&lt;td align="char" char="("&gt;185 (25.17%)&lt;/td&gt;&lt;td align="char" char="("&gt;438 (29.51%)&lt;/td&gt;&lt;td align="char" char="("&gt;206 (25.06%)&lt;/td&gt;&lt;td align="char" char="("&gt;417 (29.85%)&lt;/td&gt;&lt;td align="char" char="."&gt;623&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&lt;/td&gt;&lt;td align="char" char="("&gt;735 (33.12%)&lt;/td&gt;&lt;td align="char" char="("&gt;1484 (66.88%)&lt;/td&gt;&lt;td align="char" char="("&gt;822 (37.04%)&lt;/td&gt;&lt;td align="char" char="("&gt;1397 (62.96%)&lt;/td&gt;&lt;td align="char" char="."&gt;2219&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Career prediction&lt;/td&gt;&lt;td align="left"&gt;STEM&lt;/td&gt;&lt;td align="char" char="("&gt;24 (14.04%)&lt;/td&gt;&lt;td align="char" char="("&gt;14 (7.73%)&lt;/td&gt;&lt;td align="char" char="("&gt;&amp;#8211;&lt;/td&gt;&lt;td align="char" char="("&gt;&amp;#8211;&lt;/td&gt;&lt;td align="char" char="."&gt;38&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Non&amp;#8208;STEM&lt;/td&gt;&lt;td align="char" char="("&gt;147 (85.96%)&lt;/td&gt;&lt;td align="char" char="("&gt;167 (92.27%)&lt;/td&gt;&lt;td align="char" char="("&gt;&amp;#8211;&lt;/td&gt;&lt;td align="char" char="("&gt;&amp;#8211;&lt;/td&gt;&lt;td align="char" char="."&gt;314&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Total&lt;/td&gt;&lt;td align="char" char="("&gt;171 (48.58%)&lt;/td&gt;&lt;td align="char" char="("&gt;181 (51.42%)&lt;/td&gt;&lt;td align="char" char="("&gt;&amp;#8211;&lt;/td&gt;&lt;td align="char" char="("&gt;&amp;#8211;&lt;/td&gt;&lt;td align="char" char="."&gt;352&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0188606177-10">Debiasing approaches</hd> <p>This study selected a diverse list of 28 representative debiasing approaches that operate on the training data (<emph>pre‐processing</emph>), model (<emph>in‐processing</emph>) and predicted outcomes (<emph>post‐processing</emph>).</p> <hd id="AN0188606177-11">Pre‐processing approaches</hd> <p>Educational scenarios are widely acknowledged as being prone to a range of factors that can affect the data collected from students, resulting in various types of biases (Baker &amp; Hawn, [<reflink idref="bib7" id="ref65">7</reflink>]). This study considers pre‐processing approaches targeting three categories of representative data biases that have been widely discussed yet remain under‐explored in education. These include (i) data balancing approaches to tackle distribution‐related biases, (ii) approaches that correct labels to tackle label bias, and (iii) approaches that <emph>transform</emph> features to mitigate proxy discrimination.</p> <p>Distribution‐related biases, including class imbalance, demographic population imbalance and the imbalance in the combination of these two, are common in educational contexts. For example, in the context of passing a STEM course, there may be an imbalance between the number of students who pass and fail (class imbalance), under‐representation of females in STEM courses (demographic population imbalance) and a lower failure rate among male students compared with females (interaction of both biases). The assumption for tackling these biases is that, such imbalances might lead to certain groups being over‐ or under‐represented during training, resulting in unequal attention given to different groups and, consequently, impacting the fairness of final predictions (Caton &amp; Haas, [<reflink idref="bib14" id="ref66">14</reflink>]). To address these imbalances, two main categories of data balancing approaches, namely re‐sampling and re‐weighting, have been proposed. Re‐sampling methods either over‐sample or under‐sample instances to balance the distribution across demographic groups and/or classes, while re‐weighting methods assign weights to instances in order to amplify or reduce the influence of specific groups or classes without altering the sample size. For each type of imbalance, this study selected the most representative approach from each category.</p> <p></p> <ulist> <item> For <emph>class imbalance</emph> , three representative approaches were selected. These included the following: (i) SMOTE‐Class, derived from the state‐of‐the‐art oversampling technique SMOTE (Chawla et al., [<reflink idref="bib15" id="ref67">15</reflink>]), which over‐samples the minority class labels from generated synthetic instances, (ii) Balanced‐Class (Han, Baldwin, &amp; Cohn, [<reflink idref="bib22" id="ref68">22</reflink>]), which re‐weighs each instance based on the proportion of its class label relative to half of the total sample size, and (iii) CB‐Class (Wang et al., [<reflink idref="bib44" id="ref69">44</reflink>]), which re‐weighs instances based on the proportion of their class label relative to the group size. Note that SMOTE‐Class and Balanced‐Class share the assumption that unfairness may primarily result from insufficient training for the minority class (Han, Shen, et al., [<reflink idref="bib23" id="ref70">23</reflink>]). To address this, a balanced representation of different class labels during training is necessary, which can be achieved through either re‐sampling (as in SMOTE‐Class) or re‐weighing (as in Balanced‐Class). While CB‐Class assumes that unfairness is more likely caused by class imbalance within each demographic group.</item> <p></p> <item> For imbalances in demographic group populations, i.e., <emph>representation bias</emph> , only one approach could be found to strictly tackle this bias via re‐weighing, that is, Balanced‐Demo (Han, Baldwin, &amp; Cohn, [<reflink idref="bib22" id="ref71">22</reflink>]). This approach re‐weighs each instance based on its group size, with the weight calculated as the ratio of half of the total sample size to the group size. The underlying assumption is the under‐represented group might be more likely to receive insufficient training, which could lead to unfairness.</item> <p></p> <item> Imbalance in the representation of different class labels within each demographic group, referred to as <emph>local stereotypical bias</emph> in Dominguez‐Catena et al. ([<reflink idref="bib18" id="ref72">18</reflink>]), has been widely assumed to be one of the primary factors contributing to unfairness. This study selected five methods, encompassing both classical and state‐of‐the‐art approaches, to address this bias through either re‐sampling or re‐weighing. Based on the assumption that the distribution of demographic groups should be independent of class labels to avoid potential unfairness, (i) Uniform Sampling (Kamiran &amp; Calders, [<reflink idref="bib25" id="ref73">25</reflink>]) uniformly re‐samples instances with replacement, while (ii) Independence (Kamiran &amp; Calders, [<reflink idref="bib25" id="ref74">25</reflink>]) re‐weighs each instance to satisfy the independence assumption. Both methods aim to adjust the distribution of instances such that the proportions of class labels within each demographic group match the product of the marginal probabilities of the class label and the demographic group. Another straightforward idea is to ensure the equal representation of different groups for each class. For this, one of the most representative re‐sampling approaches, i.e., (iii) SMOTE‐Demo (Chawla et al., [<reflink idref="bib15" id="ref75">15</reflink>]) was selected, which over‐samples the under‐represented demographic group from generated synthetic instances. While (iv) CB‐Demo (Han, Baldwin, &amp; Cohn, [<reflink idref="bib22" id="ref76">22</reflink>]) is the representative for re‐weighing approaches, with weights calculated as the ratio of half of the group size to the proportion of the class within the group. In addition, (v) FairBalance (Yan et al., [<reflink idref="bib45" id="ref77">45</reflink>]) was included due to its popularity and unique characteristic of not relying on demographic attributes. Instead, it assumes that instances within the same cluster share similar features, and thus, class labels should be balanced within each cluster, for which SMOTE is used to over‐sample the under‐represented class.</item> <p></p> <item> When considering the under‐ or over‐representation of combinations of class labels and group memberships (e.g., male students who pass a course) across the dataset, often referred to as <emph>global stereotypical bias</emph> , the idea is that any under‐represented combination of class labels and group memberships may result in insufficient training for instances of that combination, potentially leading to unfairness. Hence, the goal of addressing this bias is to ensure balanced representation of all demographic and class label pairs, which can be achieved through re‐sampling or re‐weighing. (i) SMOTE‐Joint (Chawla et al., [<reflink idref="bib15" id="ref78">15</reflink>]) as one of the most representative re‐sampling methods was selected. It uses SMOTE to over‐sample instances, ensuring an equal number of samples for each (class label and group membership) pair. (ii) CB‐Joint (Wang et al., [<reflink idref="bib44" id="ref79">44</reflink>]), a re‐weighting approach, adjusts instance weights based on the prevalence of each (class label, group membership) pair in the whole training data.</item> </ulist> <p>In addition to the distribution‐related biases mentioned above, certain groups of students might be more likely to be mischaracterized and associated with incorrect labels, leading to <emph>label bias</emph>. This has been also observed in education (Moss‐Racusin et al., [<reflink idref="bib34" id="ref80">34</reflink>]), for example, female in STEM may be unfairly labelled as less capable compared with their male counterparts. This bias is also considered an important factor contributing to unfairness. The main idea to address this bias involves identifying instances that are prone to being mislabelled (i.e., borderline instances), and then relabelling, re‐weighing or removing (via re‐sampling) them. Five approaches were included, with at least one representative method for each type of the three strategies. (i) Massaging (Kamiran &amp; Calders, [<reflink idref="bib25" id="ref81">25</reflink>]) is one of the most classical relabelling methods, which directly relabels borderline instances across different demographic groups, i.e., instances close to the decision boundary. The idea is that unequal representation of mislabelled instances, whether as positive or negative, can introduce bias into the model. (ii) Preferential Sampling (Kamiran &amp; Calders, [<reflink idref="bib25" id="ref82">25</reflink>]) is one of the most well‐known re‐sampling approaches to tackle label bias, which preferentially duplicates or removes borderline instances to ensure that each group reaches the desired size. (iii) For re‐weighing, methods proposed in a recent influential work (Jiang &amp; Nachum, [<reflink idref="bib24" id="ref83">24</reflink>]), namely Label Debias‐Dp, Label Debias‐EqOpp and Label Debias‐EqOdds, were considered due to their significant impact, as evidenced by its high citation. These methods represent three variants, each corresponding to a different fairness constraint: demographic parity, equal opportunity and equalized odds. The underlying assumption is that the observed labels are biased approximations of the true labels, and the approach involves a re‐weighting function that incorporates fairness constraints over the training examples.</p> <p>In contrast to the biases discussed above that focus on class labels and demographic group population, <emph>proxy discrimination</emph> refers to the bias that occurs when seemingly neutral or irrelevant features serves as a proxy for sensitive attributes. These features may correlate with the sensitive attribute in a way that reinforces stereotypes or disadvantages a particular group, potentially leading biased outcomes. In educational settings, this bias might be prevalent but hard to directly identify. For example, as indicated by a recent study (Zhuhadar et al., [<reflink idref="bib48" id="ref84">48</reflink>]), students of different gender have significantly different interaction patterns with intelligent tutoring systems. If these interaction patterns are used to predict student performance, the model may inadvertently introduce gender bias, since different gender groups might perform differently for reasons unrelated to their actual capabilities—such as societal expectations, prior experiences or the design of the system itself (Tschantz, [<reflink idref="bib41" id="ref85">41</reflink>]). To tackle this bias, this study selected one state‐of‐the‐art method that has been widely included in fairness evaluation studies, namely Correlation Remover (Bird et al., [<reflink idref="bib11" id="ref86">11</reflink>]). This method learns new features by minimizing the distances between the new and original features while enforcing orthogonality between them.</p> <hd id="AN0188606177-12">In‐processing approaches</hd> <p>A common assumption underlying in‐processing approaches is that models can inadvertently encode biases present in the training data (Caton &amp; Haas, [<reflink idref="bib14" id="ref87">14</reflink>]). To address this, in‐processing methods adjust the training process to balance prediction accuracy and fairness. These methods typically fall into three main strategies (Caton &amp; Haas, [<reflink idref="bib14" id="ref88">14</reflink>]): imposing fairness constraints through regularization, modifying the batch optimization process and employing adversarial training. This study considered six representative methods, selecting at least one classical approach from each of these strategies. (i) Reduction‐Dp &amp; Reduction‐EqOdds are two variants derived from <emph>Reduction</emph> (Agarwal et al., [<reflink idref="bib1" id="ref89">1</reflink>]), which constrained a classifier with demographic parity and equalized odds respectively. (ii) FairBatch‐Dp, FairBatch‐EqOpp and FairBatch‐EqOdds are three variants derived from <emph>FairBatch</emph> (Roh et al., [<reflink idref="bib37" id="ref90">37</reflink>]) constrained with demographic parity, equal opportunity and equalized odds, which dynamically samples varied number of instances from different groups to construct each batch during training to optimize the fairness constraint. (iii) Adversarial Debiasing (Zhang et al., [<reflink idref="bib47" id="ref91">47</reflink>]) tries to maximize the predictive accuracy for predicting class labels while minimizing its ability to predict the sensitive attribute from predictions via adversarial training.</p> <hd id="AN0188606177-13">Post‐processing approaches</hd> <p>Recognizing the predicted outputs can be unfair, post‐processing approaches often transform the outputs to achieve fairness while preserving accuracy. Compared with pre‐processing and in‐processing, post‐processing approaches are relatively less investigated in the literature. To provide a balanced perspective, this study included one well‐established method and one recent state‐of‐the‐art method, representing both foundational and cutting‐edge advancements in the field. (i) CalibratedEqOdds (Pleiss et al., [<reflink idref="bib36" id="ref92">36</reflink>]), a representative post‐processing approach, aims to maintain equal error rates (such as false negatives and false positives) by setting the scores of a specified proportion of instances to the class mean. (ii) FairProjection (Alghamdi et al., [<reflink idref="bib3" id="ref93">3</reflink>]) is a more recent method trying to transform the output of the trained classifier by solving a meta‐optimization problem using information projection theory.</p> <hd id="AN0188606177-14">First‐phase model debiasing</hd> <p>All methods described above address downstream tasks in the second phase, where models are trained using features extracted during the first phase. However, bias can also be inherently encoded in the extracted features from the first phase, referred to as <emph>embedding bias</emph>. Addressing this issue, particularly for textual data, aligns with the growing recognition in the literature that large language models often reflect and amplify biases present in their training corpus (Caliskan et al., [<reflink idref="bib12" id="ref94">12</reflink>]; Liu et al., [<reflink idref="bib31" id="ref95">31</reflink>]). To tackle that for forum classification, this study leveraged a recent model DebiasedBERT (Sha et al., [<reflink idref="bib39" id="ref96">39</reflink>]), which tries to reduce the awareness of the BERT on sensitive attributes by continuing training it with additional balanced domain‐specific data selected via active learning. For career prediction, task similar debiasing is not applicable as gender information are only available for students with career labels.</p> <hd id="AN0188606177-15">Combined approaches</hd> <p>Considering the characteristics of bias that may arise from multiple sources and emerge at multiple stages during ML modelling, this study proposed two strategies to further combine existing approaches, that is, Multi‐Stage Debiasing and Multi‐Source Debiasing. Multi‐Stage Debiasing subsequently apply the best approach at each stage during ML modelling, while Multi‐Source Debiasing focuses on pre‐processing stage, combining the best approaches to tackle different sources of bias, such as distribution‐related bias, label bias and proxy discrimination.</p> <hd id="AN0188606177-16">Evaluation metrics</hd> <p>This study use AUC (Area‐Under‐ROC) to evaluate predictive accuracy of the educational tasks, as it has been widely adopted in classification tasks and noted for its robustness to imbalanced data. For fairness evaluation, ABROCA (Gardner et al., [<reflink idref="bib21" id="ref97">21</reflink>]) was used, which calculates the area between the ROC curves of the two groups. The focus on ABROCA is multi‐faceted: (i) it specifically addresses predictive unfairness related to unequal error rates across different groups, which aligns with the goals of this study, as unequal error rates indicate a greater misalignment of support or resources for students from certain backgrounds, potentially exacerbating educational inequalities, (ii) it is the first fairness metric specifically proposed for educational classification tasks, uniquely not relying on any single pre‐defined threshold. This characteristic is particularly important for educational scenarios as the actual threshold may depend on the specific context after obtaining predictions.</p> <hd id="AN0188606177-17">Study setup</hd> <p>For forum post classification, the data were split into training, validation and test sets in an 6:2:2 ratio. While for career prediction, 20% of the data were sampled as test data, and 10% of the remaining data were used as validation data considering its smaller size. The validation and test sets were randomly sampled to ensure unbiased distribution in terms of both class labels and group membership. All base models were implemented using TensorFlow and optimized with Adam for a maximum of 1000 epochs, with early stopping if there were no performance improvements for 50 epochs. The best hyperparameters for each debiasing approach were determined based on the averaged optimization loss on the validation set over 10 runs. The final results reported were the average performance on test data over 10 runs. More details can be found in the Appendix S1[<reflink idref="bib3" id="ref98">3</reflink>] and the authors' open‐source Git repository.[<reflink idref="bib4" id="ref99">4</reflink>]</p> <hd id="AN0188606177-18">RESULTS</hd> <p>The results for RQ1, regarding the effectiveness of various debiasing approaches across different stages for the two tasks, are presented in Section "RQ1: Effectiveness of various debiasing approaches at different stages". For RQ2, which examines the impact of debiasing focused on one attribute on another, the results are discussed in Section "RQ2: Impact of debiasing focused on one attribute on another". Finally, the results for RQ3, evaluating the effectiveness of the proposed combined debiasing strategies—namely, Multi‐Stage and Multi‐Source debiasing—are provided in Section "RQ3: Effectiveness of debiasing approaches across multiple stages and multiple sources".</p> <hd id="AN0188606177-19">RQ1: Effectiveness of various debiasing approaches at different stages</hd> <p>Table 3 shows the predictive accuracy and fairness of all debiasing approaches w.r.t. gender and first‐language backgrounds for content relevancy classification for forum posts, and w.r.t. gender for STEM career prediction. The following key observations can be drawn from the results.</p> <p>3 TABLE Predictive accuracy and fairness of various debiasing approaches, where values in bold represent improved fairness and accuracy.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Stage&lt;/th&gt;&lt;th align="left"&gt;Category&lt;/th&gt;&lt;th align="left"&gt;RowID&lt;/th&gt;&lt;th align="left"&gt;Approach&lt;/th&gt;&lt;th align="left"&gt;Forum post classification&lt;/th&gt;&lt;th align="left"&gt;STEM career prediction&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;Gender&lt;/th&gt;&lt;th align="left"&gt;Language&lt;/th&gt;&lt;th align="left"&gt;Gender&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;AUC &lt;p&gt;&lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0001" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8593;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;ABROCA &lt;p&gt;&lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0002" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8595;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;AUC &lt;p&gt;&lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0003" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8593;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;ABROCA &lt;p&gt;&lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0004" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8595;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;AUC &lt;p&gt;&lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0005" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8593;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;ABROCA &lt;p&gt;&lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0006" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics xmlns=""&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8595;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#8211;&lt;/td&gt;&lt;td align="left"&gt;&amp;#8211;&lt;/td&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;Undebiased&lt;/td&gt;&lt;td align="left"&gt;0.9555&lt;/td&gt;&lt;td align="left"&gt;0.0133&lt;/td&gt;&lt;td align="left"&gt;0.9555&lt;/td&gt;&lt;td align="left"&gt;0.0230&lt;/td&gt;&lt;td align="left"&gt;0.5675&lt;/td&gt;&lt;td align="left"&gt;0.1316&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;First&amp;#8208;phase modelling&lt;/td&gt;&lt;td align="left"&gt;Tackle Embedding Bias&lt;/td&gt;&lt;td align="left"&gt;2&lt;/td&gt;&lt;td align="left"&gt;DebiasedBERT&lt;/td&gt;&lt;td align="left"&gt;0.9526 (&amp;#8722;0.30%)&lt;/td&gt;&lt;td align="left"&gt;0.0156 (&amp;#8722;17.29%)&lt;/td&gt;&lt;td align="left"&gt;0.9519 (&amp;#8722;0.38%)&lt;/td&gt;&lt;td align="left"&gt;0.0200 (13.04%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;&amp;#8211;&lt;/td&gt;&lt;td align="left"&gt;&amp;#8211;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Pre&amp;#8208;processing&lt;/td&gt;&lt;td align="left"&gt;Balancing data to tackle Class Imbalance&lt;/td&gt;&lt;td align="left"&gt;3&lt;/td&gt;&lt;td align="left"&gt;SMOTE&amp;#8208;Class&lt;/td&gt;&lt;td align="left"&gt;0.9561 (0.06%)&lt;/td&gt;&lt;td align="left"&gt;0.0136 (&amp;#8722;2.26%)&lt;/td&gt;&lt;td align="left"&gt;0.9567 (0.13%)&lt;/td&gt;&lt;td align="left"&gt;0.0246 (&amp;#8722;6.96%)&lt;/td&gt;&lt;td align="left"&gt;0.4745 (&amp;#8722;16.39%)&lt;/td&gt;&lt;td align="left"&gt;0.1490 (&amp;#8722;13.22%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;4&lt;/td&gt;&lt;td align="left"&gt;Balanced&amp;#8208;Class&lt;/td&gt;&lt;td align="left"&gt;0.9558 (0.03%)&lt;/td&gt;&lt;td align="left"&gt;0.0133 (0.00%)&lt;/td&gt;&lt;td align="left"&gt;0.9563 (0.08%)&lt;/td&gt;&lt;td align="left"&gt;0.0235 (&amp;#8722;2.17%)&lt;/td&gt;&lt;td align="left"&gt;0.5685 (0.18%)&lt;/td&gt;&lt;td align="left"&gt;0.1304 (0.91%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;5&lt;/td&gt;&lt;td align="left"&gt;CB&amp;#8208;Class&lt;/td&gt;&lt;td align="left"&gt;0.9564 (0.09%)&lt;/td&gt;&lt;td align="left"&gt;0.0131 (1.50%)&lt;/td&gt;&lt;td align="left"&gt;0.9560 (0.05%)&lt;/td&gt;&lt;td align="left"&gt;0.0237 (&amp;#8722;3.04%)&lt;/td&gt;&lt;td align="left"&gt;0.5703 (0.49%)&lt;/td&gt;&lt;td align="left"&gt;0.1306 (0.76%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Balancing data to tackle Representation Bias&lt;/td&gt;&lt;td align="left"&gt;6&lt;/td&gt;&lt;td align="left"&gt;Balanced&amp;#8208;Demo&lt;/td&gt;&lt;td align="left"&gt;0.9558 (0.03%)&lt;/td&gt;&lt;td align="left"&gt;0.0135 (&amp;#8722;1.50%)&lt;/td&gt;&lt;td align="left"&gt;0.9562 (0.07%)&lt;/td&gt;&lt;td align="left"&gt;0.0229 (0.43%)&lt;/td&gt;&lt;td align="left"&gt;0.5217 (&amp;#8722;8.07%)&lt;/td&gt;&lt;td align="left"&gt;0.1323 (&amp;#8722;0.53%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Balancing data to tackle Local Stereotypical Bias&lt;/td&gt;&lt;td align="left"&gt;7&lt;/td&gt;&lt;td align="left"&gt;Uniform Sampling&lt;/td&gt;&lt;td align="left"&gt;0.9554 (&amp;#8722;0.01%)&lt;/td&gt;&lt;td align="left"&gt;0.0151 (&amp;#8722;13.53%)&lt;/td&gt;&lt;td align="left"&gt;0.9528 (&amp;#8722;0.28%)&lt;/td&gt;&lt;td align="left"&gt;0.0258 (&amp;#8722;12.17%)&lt;/td&gt;&lt;td align="left"&gt;0.5557 (&amp;#8722;2.08%)&lt;/td&gt;&lt;td align="left"&gt;0.1323 (&amp;#8722;0.53%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;8&lt;/td&gt;&lt;td align="left"&gt;Independence&lt;/td&gt;&lt;td align="left"&gt;0.9561 (0.06%)&lt;/td&gt;&lt;td align="left"&gt;0.0131 (1.50%)&lt;/td&gt;&lt;td align="left"&gt;0.9565 (0.10%)&lt;/td&gt;&lt;td align="left"&gt;0.0232 (&amp;#8722;0.87%)&lt;/td&gt;&lt;td align="left"&gt;0.5694 (0.33%)&lt;/td&gt;&lt;td align="left"&gt;0.1302 (1.06%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;9&lt;/td&gt;&lt;td align="left"&gt;SMOTE&amp;#8208;Demo&lt;/td&gt;&lt;td align="left"&gt;0.9543 (&amp;#8722;0.13%)&lt;/td&gt;&lt;td align="left"&gt;0.0132 (0.75%)&lt;/td&gt;&lt;td align="left"&gt;0.9554 (&amp;#8722;0.01%)&lt;/td&gt;&lt;td align="left"&gt;0.0224 (2.61%)&lt;/td&gt;&lt;td align="left"&gt;0.5737 (1.09%)&lt;/td&gt;&lt;td align="left"&gt;0.1300 (1.22%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;10&lt;/td&gt;&lt;td align="left"&gt;CB&amp;#8208;Demo&lt;/td&gt;&lt;td align="left"&gt;0.9553 (&amp;#8722;0.02%)&lt;/td&gt;&lt;td align="left"&gt;0.0129 (3.01%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.9557 (0.02%)&lt;/td&gt;&lt;td align="left"&gt;0.0222 (3.48%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.5706 (0.55%)&lt;/td&gt;&lt;td align="left"&gt;0.1286 (2.28%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;11&lt;/td&gt;&lt;td align="left"&gt;FairBalance&lt;/td&gt;&lt;td align="left"&gt;0.9569 (0.15%)&lt;/td&gt;&lt;td align="left"&gt;0.0127 (4.51%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.9561 (0.06%)&lt;/td&gt;&lt;td align="left"&gt;0.0237 (&amp;#8722;3.04%)&lt;/td&gt;&lt;td align="left"&gt;0.4722 (&amp;#8722;16.79%)&lt;/td&gt;&lt;td align="left"&gt;0.1450 (&amp;#8722;10.18%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Balancing data to tackle Global Stereotypical Bias&lt;/td&gt;&lt;td align="left"&gt;12&lt;/td&gt;&lt;td align="left"&gt;SMOTE&amp;#8208;Joint&lt;/td&gt;&lt;td align="left"&gt;0.9535 (&amp;#8722;0.21%)&lt;/td&gt;&lt;td align="left"&gt;0.0141 (&amp;#8722;6.02%)&lt;/td&gt;&lt;td align="left"&gt;0.9564 (0.09%)&lt;/td&gt;&lt;td align="left"&gt;0.0228 (0.87%)&lt;/td&gt;&lt;td align="left"&gt;0.4719 (&amp;#8722;16.85%)&lt;/td&gt;&lt;td align="left"&gt;0.1462 (&amp;#8722;11.09%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;13&lt;/td&gt;&lt;td align="left"&gt;CB&amp;#8208;Joint&lt;/td&gt;&lt;td align="left"&gt;0.9558 (0.03%)&lt;/td&gt;&lt;td align="left"&gt;0.0132 (0.75%)&lt;/td&gt;&lt;td align="left"&gt;0.9562 (0.07%)&lt;/td&gt;&lt;td align="left"&gt;0.0235 (&amp;#8722;2.17%)&lt;/td&gt;&lt;td align="left"&gt;0.5666 (&amp;#8722;0.16%)&lt;/td&gt;&lt;td align="left"&gt;0.1328 (&amp;#8722;0.91%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Correcting labels to tackle Label Debias&lt;/td&gt;&lt;td align="left"&gt;14&lt;/td&gt;&lt;td align="left"&gt;Massaging&lt;/td&gt;&lt;td align="left"&gt;0.9549 (&amp;#8722;0.06%)&lt;/td&gt;&lt;td align="left"&gt;0.0109 (18.05%)&lt;/td&gt;&lt;td align="left"&gt;0.9555 (0.00%)&lt;/td&gt;&lt;td align="left"&gt;0.0246 (&amp;#8722;6.96%)&lt;/td&gt;&lt;td align="left"&gt;0.5747 (1.27%)&lt;/td&gt;&lt;td align="left"&gt;0.1208 (8.21%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;15&lt;/td&gt;&lt;td align="left"&gt;Preferential Sampling&lt;/td&gt;&lt;td align="left"&gt;0.9550 (&amp;#8722;0.05%)&lt;/td&gt;&lt;td align="left"&gt;0.0122 (8.27%)&lt;/td&gt;&lt;td align="left"&gt;0.9560 (0.05%)&lt;/td&gt;&lt;td align="left"&gt;0.0255 (&amp;#8722;10.87%)&lt;/td&gt;&lt;td align="left"&gt;0.5486 (&amp;#8722;3.33%)&lt;/td&gt;&lt;td align="left"&gt;0.1332 (&amp;#8722;1.22%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;16&lt;/td&gt;&lt;td align="left"&gt;Label Debias&amp;#8208;Dp&lt;/td&gt;&lt;td align="left"&gt;0.9565 (0.10%)&lt;/td&gt;&lt;td align="left"&gt;0.0134 (&amp;#8722;0.75%)&lt;/td&gt;&lt;td align="left"&gt;0.9556 (0.01%)&lt;/td&gt;&lt;td align="left"&gt;0.0239 (&amp;#8722;3.91%)&lt;/td&gt;&lt;td align="left"&gt;0.5520 (&amp;#8722;2.73%)&lt;/td&gt;&lt;td align="left"&gt;0.1328 (&amp;#8722;0.91%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;17&lt;/td&gt;&lt;td align="left"&gt;Label Debias&amp;#8208;EqOpp&lt;/td&gt;&lt;td align="left"&gt;0.9558 (0.03%)&lt;/td&gt;&lt;td align="left"&gt;0.0133 (0.00%)&lt;/td&gt;&lt;td align="left"&gt;0.9564 (0.09%)&lt;/td&gt;&lt;td align="left"&gt;0.0234 (&amp;#8722;1.74%)&lt;/td&gt;&lt;td align="left"&gt;0.5700 (0.44%)&lt;/td&gt;&lt;td align="left"&gt;0.1318 (&amp;#8722;0.15%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;18&lt;/td&gt;&lt;td align="left"&gt;Label&amp;#8208;Debias&amp;#8208;EqOdds&lt;/td&gt;&lt;td align="left"&gt;0.9553 (&amp;#8722;0.02%)&lt;/td&gt;&lt;td align="left"&gt;0.0128 (3.76%)&lt;/td&gt;&lt;td align="left"&gt;0.9560 (0.05%)&lt;/td&gt;&lt;td align="left"&gt;0.0240 (&amp;#8722;4.35%)&lt;/td&gt;&lt;td align="left"&gt;0.5667 (&amp;#8722;0.14%)&lt;/td&gt;&lt;td align="left"&gt;0.1295 (1.60%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Transforming features to tackle Proxy Discrimination&lt;/td&gt;&lt;td align="left"&gt;19&lt;/td&gt;&lt;td align="left"&gt;Correlation Remover&lt;/td&gt;&lt;td align="left"&gt;0.9562 (0.07%)&lt;/td&gt;&lt;td align="left"&gt;0.0121 (9.02%)&lt;/td&gt;&lt;td align="left"&gt;0.9568 (0.14%)&lt;/td&gt;&lt;td align="left"&gt;0.0244 (&amp;#8722;6.09%)&lt;/td&gt;&lt;td align="left"&gt;0.5704 (0.51%)&lt;/td&gt;&lt;td align="left"&gt;0.1290 (1.98%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;In&amp;#8208;processing&lt;/td&gt;&lt;td align="left"&gt;Constraints&lt;/td&gt;&lt;td align="left"&gt;20&lt;/td&gt;&lt;td align="left"&gt;Reduction&amp;#8208;Dp&lt;/td&gt;&lt;td align="left"&gt;0.9568 (0.14%)&lt;/td&gt;&lt;td align="left"&gt;0.0117 (12.03%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.9561 (0.06%)&lt;/td&gt;&lt;td align="left"&gt;0.0244 (&amp;#8722;6.09%)&lt;/td&gt;&lt;td align="left"&gt;0.5303 (&amp;#8722;6.56%)&lt;/td&gt;&lt;td align="left"&gt;0.1333 (&amp;#8722;1.29%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;21&lt;/td&gt;&lt;td align="left"&gt;Reduction&amp;#8208;EqOdds&lt;/td&gt;&lt;td align="left"&gt;0.9563 (0.08%)&lt;/td&gt;&lt;td align="left"&gt;0.0126 (5.26%)&lt;/td&gt;&lt;td align="left"&gt;0.9559 (0.04%)&lt;/td&gt;&lt;td align="left"&gt;0.0243 (&amp;#8722;5.65%)&lt;/td&gt;&lt;td align="left"&gt;0.5298 (&amp;#8722;6.64%)&lt;/td&gt;&lt;td align="left"&gt;0.1101 (16.34%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Optimization&lt;/td&gt;&lt;td align="left"&gt;22&lt;/td&gt;&lt;td align="left"&gt;FairBatch&amp;#8208;Dp&lt;/td&gt;&lt;td align="left"&gt;0.9563 (0.08%)&lt;/td&gt;&lt;td align="left"&gt;0.0136 (&amp;#8722;2.26%)&lt;/td&gt;&lt;td align="left"&gt;0.9513 (&amp;#8722;0.44%)&lt;/td&gt;&lt;td align="left"&gt;0.0241 (&amp;#8722;4.78%)&lt;/td&gt;&lt;td align="left"&gt;0.5687 (0.21%)&lt;/td&gt;&lt;td align="left"&gt;0.1333 (&amp;#8722;1.29%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;23&lt;/td&gt;&lt;td align="left"&gt;FairBatch&amp;#8208;EqOpp&lt;/td&gt;&lt;td align="left"&gt;0.9551 (&amp;#8722;0.04%)&lt;/td&gt;&lt;td align="left"&gt;0.0125 (6.02%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.9564 (0.09%)&lt;/td&gt;&lt;td align="left"&gt;0.0241 (&amp;#8722;4.78%)&lt;/td&gt;&lt;td align="left"&gt;0.5682 (0.12%)&lt;/td&gt;&lt;td align="left"&gt;0.1347 (&amp;#8722;2.36%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;24&lt;/td&gt;&lt;td align="left"&gt;FairBatch&amp;#8208;EqOdds&lt;/td&gt;&lt;td align="left"&gt;0.9537 (&amp;#8722;0.19%)&lt;/td&gt;&lt;td align="left"&gt;0.0124 (6.77%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.9556 (0.01%)&lt;/td&gt;&lt;td align="left"&gt;0.0240 (&amp;#8722;4.35%)&lt;/td&gt;&lt;td align="left"&gt;0.5809 (2.36%)&lt;/td&gt;&lt;td align="left"&gt;0.1257 (4.48%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Adversarial leaning&lt;/td&gt;&lt;td align="left"&gt;25&lt;/td&gt;&lt;td align="left"&gt;Adversarial Debiasing&lt;/td&gt;&lt;td align="left"&gt;0.9561 (0.06%)&lt;/td&gt;&lt;td align="left"&gt;0.0136 (&amp;#8722;2.26%)&lt;/td&gt;&lt;td align="left"&gt;0.9547 (&amp;#8722;0.08%)&lt;/td&gt;&lt;td align="left"&gt;0.0252 (&amp;#8722;9.57%)&lt;/td&gt;&lt;td align="left"&gt;0.4957 (&amp;#8722;12.65%)&lt;/td&gt;&lt;td align="left"&gt;0.1241 (5.70%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Post&amp;#8208;processing&lt;/td&gt;&lt;td align="left"&gt;Score transformation&lt;/td&gt;&lt;td align="left"&gt;26&lt;/td&gt;&lt;td align="left"&gt;CalibratedEqOdds&amp;#8208;fpr&lt;/td&gt;&lt;td align="left"&gt;0.9489 (0.02%)&lt;/td&gt;&lt;td align="left"&gt;0.0240 (&amp;#8722;80.45%)&lt;/td&gt;&lt;td align="left"&gt;0.9538 (&amp;#8722;0.18%)&lt;/td&gt;&lt;td align="left"&gt;0.0202 (12.17%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.5578 (&amp;#8722;1.71%)&lt;/td&gt;&lt;td align="left"&gt;0.1462 (&amp;#8722;11.09%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;27&lt;/td&gt;&lt;td align="left"&gt;CalibratedEqOdds&amp;#8208;fnr&lt;/td&gt;&lt;td align="left"&gt;0.9557 (0.01%)&lt;/td&gt;&lt;td align="left"&gt;0.0117 (12.03%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.9561 (0.06%)&lt;/td&gt;&lt;td align="left"&gt;0.0227 (1.30%)&lt;/td&gt;&lt;td align="left"&gt;0.5675 (0.00%)&lt;/td&gt;&lt;td align="left"&gt;0.1330 (&amp;#8722;1.06%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;28&lt;/td&gt;&lt;td align="left"&gt;CalibratedEqOdds&amp;#8208;weighted&lt;/td&gt;&lt;td align="left"&gt;0.9556 (0.12%)&lt;/td&gt;&lt;td align="left"&gt;0.0148 (&amp;#8722;11.28%)&lt;/td&gt;&lt;td align="left"&gt;0.9547 (&amp;#8722;0.08%)&lt;/td&gt;&lt;td align="left"&gt;0.0219 (4.78%)&lt;xref ref-type="fn" rid="tfn2" /&gt;&lt;/td&gt;&lt;td align="left"&gt;0.5422 (&amp;#8722;4.46%)&lt;/td&gt;&lt;td align="left"&gt;0.1403 (&amp;#8722;6.61%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Meta&amp;#8208;optimization&lt;/td&gt;&lt;td align="left"&gt;29&lt;/td&gt;&lt;td align="left"&gt;FairProjection&lt;/td&gt;&lt;td align="left"&gt;0.9566 (0.12%)&lt;/td&gt;&lt;td align="left"&gt;0.0130 (2.26%)&lt;/td&gt;&lt;td align="left"&gt;0.9546 (&amp;#8722;0.09%)&lt;/td&gt;&lt;td align="left"&gt;0.0251 (&amp;#8722;9.13%)&lt;/td&gt;&lt;td align="left"&gt;0.5477 (&amp;#8722;3.49%)&lt;/td&gt;&lt;td align="left"&gt;0.1299 (1.29%)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 <emph>Note</emph>: The signs <ephtml> &lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0007" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8593;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0008" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8595;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> indicate whether a higher ( <ephtml> &lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0009" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8593;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ) or lower ( <ephtml> &lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0010" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8595;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ) value is favourable.</p> <p>2 * Significant improvements in mitigating bias.</p> <p>For pre‐processing approaches, several observations can be made from Table 3 as follows. (i) Predictive fairness can be enhanced in most cases with pre‐processing approaches that tackle distribution‐related biases, label bias and proxy discrimination. For example, at least one approach targeting different biases significantly mitigated unfairness, with ABROCA reduced by 4.51%, 18.05% and 9.02% with FairBalance, Massaging and Correlation Remover for forum post classification (Rows 11, 14 and 19), and 2.28%, 8.21% and 1.98% with CB‐demo, Massaging and Correlation Remover for STEM career prediction (Rows 10, 14 and 19). (ii) Among these approaches, SMOTE‐Demo and CB‐Demo consistently reduced unfairness in all three examined cases (see Rows 9 and 10). Both tackle local stereotypical bias but in different ways: SMOTE‐Demo through re‐sampling and CB‐Demo through re‐weighing. (iii) Notably, Massaging despite its simplicity, achieved largest improvements in fairness w.r.t. gender for both tasks. Similar to pre‐processing, in‐processing approaches also showed overall comparable performance in improving fairness w.r.t. the gender attribute for both tasks. In particular, Reduction achieved the largest improvements, reducing ABRACA by 12.03% and 16.34% with Reduction‐Dp and Reduction‐EqOdds for predicting forum post relevancy and STEM career choices respectively (see Rows 20 and 21). For post‐processing approaches, it is observed that, CalibratedEqOdds greatly improved fairness for both gender and language attributes in forum post classification but increased unfairness in STEM career prediction, while FairProjection achieved enhanced fairness over gender for both predictive tasks, it increased unfairness over language for forum post classification. In terms of first stage debiasing, it can be noticed that in the case of (forum post classification, language), DebiasedBERT achieved the best performance across all approaches (Row 2), while a majority of other approaches focused on second‐phase debiasing failed to enhance fairness. This indicates that text embeddings from BERT encoded significant bias regarding language, which is rather difficult to be removed through debiasing approaches at later stages.</p> <p>To summarize, pre‐processing and in‐processing approaches generally appear more reliable for effectively mitigating bias across different cases. Specifically, <emph>local stereotypical bias</emph>, <emph>label bias</emph> and <emph>proxy discrimination</emph> are identified as most important sources of data bias leading to predictive unfairness. These biases can be effectively addressed by employing state‐of‐the‐art debiasing approaches, such as re‐sampling or re‐weighing training samples, relabelling borderline instances and transforming features. Additionally, regularizing the model or altering the optimization procedure to satisfy equalized odds can also effectively prevent the model from amplifying bias. However, these debiasing approaches might not be useful when features generated in the initial phase are inherently biased. This is especially challenging when modelling textual data, as bias encoded in text embeddings generated from pre‐trained language models can be a significant factor leading to unfairness, which might be very hard to tackle in subsequent phases.</p> <hd id="AN0188606177-20">RQ2: Impact of debiasing focused on one attribute on another</hd> <p>The fairness with respect to first‐language background was also assessed when debiasing focused on gender (and vice versa) to examine the impact of targeting a single sensitive attribute (RQ2). As shown in Figure 1, it can be noticed that when debiasing focused on gender, most approaches appeared to slightly exacerbate unfairness w.r.t. first‐language, with over 30% of the approaches that showed improved fairness in gender exhibiting increased unfairness on language by 0.6% ~ 5%. Only four approaches, i.e., SMOTE‐Demo, FairBalance, Reduction‐dp and Fairbatch‐EqOdds showed simultaneous fairness improvements on both attributes. Conversely, when considering the impact on gender while debiasing along first‐language, it can be observed that in most cases bias regarding the gender attribute are barely changed or just slightly reduced in most cases. This discrepancy in two cases suggests that unfairness mitigation in one sensitive attribute does not necessarily guarantee the reduced unfairness w.r.t. other sensitive attributes. In particular, it is likely to further amplify the bias w.r.t. the other attribute, which may potentially lead to more severe discrimination towards users with other specific attribute.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/58I/01nov25/bjet13575-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="bjet13575-fig-0001.jpg" title="1 Upper: Impact of various debiasing approaches on fairness w.r.t. first‐language background when addressing bias regarding gender. Bottom: Impact of various debiasing approaches on fairness w.r.t. gender when addressing bias regarding first‐language. Points in different colours and shapes correspond to the approaches from different stages, namely pre‐processing, in‐processing and post‐processing." /> </p> <p></p> <hd id="AN0188606177-22">RQ3: Effectiveness of debiasing approaches across multiple stages and multiple sources</hd> <p>Table 4 presents the results of combined approaches, indicating that applying debiasing with both multi‐stage and multi‐source strategies substantially improves predictive accuracy and fairness. Unfairness was further reduced by 21.03% and 16.27% for (forum post classification, gender) and (STEM career prediction, gender) respectively. Notably, only combining approaches from multiple stages or approaches tacking multiple sources of biases cannot guarantee the improved performance. For example, when considering gender bias in STEM career prediction, fairness dropped by 9.18% compared with the best performance achieved by Reduction after combining best approaches from each stage. It should be noted that, for STEM career prediction, when using both multi‐stage and multi‐source strategies, CB‐Demo was replaced with SMOTE‐Demo rather than being combined with Reduction. This choice was made because both Reduction and CB‐Demo rely on re‐weighing instances, and combining them could lead to ineffective debiasing by potentially undermining each other's effect. Therefore, it is crucial to consider how different approaches interact with data and model to ensure they complement each other effectively.</p> <p>4 TABLE Predictive accuracy and fairness of combined debiasing.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Task&amp;#8208;sensitive attribute&lt;/th&gt;&lt;th align="left"&gt;Strategy&lt;/th&gt;&lt;th align="left"&gt;First&amp;#8208;phase&lt;/th&gt;&lt;th align="left"&gt;Pre&amp;#8208;processing&lt;/th&gt;&lt;th align="left"&gt;In&amp;#8208;processing&lt;/th&gt;&lt;th align="left"&gt;Post&amp;#8208;processing&lt;/th&gt;&lt;th align="left"&gt;Combined&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;Embedding bias&lt;/th&gt;&lt;th align="left"&gt;Local stereotypical bias&lt;/th&gt;&lt;th align="left"&gt;Label bias&lt;/th&gt;&lt;th align="left"&gt;Proxy discrimination&lt;/th&gt;&lt;th align="left"&gt;&amp;#8211;&lt;/th&gt;&lt;th align="left"&gt;&amp;#8211;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Forum post classification gender&lt;/td&gt;&lt;td align="left"&gt;Best Approaches&lt;/td&gt;&lt;td align="left"&gt;DebiasedBERT&lt;/td&gt;&lt;td align="left"&gt;FairBalance&lt;/td&gt;&lt;td align="left"&gt;Massaging&lt;/td&gt;&lt;td align="left"&gt;Correlation Remover&lt;/td&gt;&lt;td align="left"&gt;Reduction&lt;/td&gt;&lt;td align="left"&gt;CalibratedEqOdds&lt;/td&gt;&lt;td align="left"&gt;AUC&lt;/td&gt;&lt;td align="left"&gt;ABROCA&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Top Single&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.9549&lt;/td&gt;&lt;td align="left"&gt;0.0109&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.9539 (&amp;#8722;0.10%)&lt;/td&gt;&lt;td align="left"&gt;0.0118 (&amp;#8722;8.25%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.9532 (&amp;#8722;0.18%)&lt;/td&gt;&lt;td align="left"&gt;0.0123 (&amp;#8722;12.73%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.9549 (0.00%)&lt;/td&gt;&lt;td align="left"&gt;0.0096 (11.92%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage &amp; Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.9567 (0.19%)&lt;/td&gt;&lt;td align="left"&gt;0.0086 (21.03%)&amp;#42;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage &amp; Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.9569 (0.21%)&amp;#42;&lt;/td&gt;&lt;td align="left"&gt;0.0095 (12.96%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Forum post classification language&lt;/td&gt;&lt;td align="left"&gt;Best Approaches&lt;/td&gt;&lt;td align="left"&gt;DebiasedBERT&lt;/td&gt;&lt;td align="left"&gt;CB&amp;#8208;Demo&lt;/td&gt;&lt;td align="left"&gt;Massaging&lt;/td&gt;&lt;td align="left"&gt;Correlation Remover&lt;/td&gt;&lt;td align="left"&gt;Reduction&lt;/td&gt;&lt;td align="left"&gt;CalibratedEqOdds&lt;/td&gt;&lt;td align="left"&gt;AUC&lt;/td&gt;&lt;td align="left"&gt;ABROCA&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Top Single&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.9519&lt;/td&gt;&lt;td align="left"&gt;0.0200&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.9528 (0.09%)&amp;#42;&lt;/td&gt;&lt;td align="left"&gt;0.0198 (0.75%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.9516 (&amp;#8722;0.03%)&lt;/td&gt;&lt;td align="left"&gt;0.0195 (2.65%)&amp;#42;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage &amp; Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.9524 (0.05%)&lt;/td&gt;&lt;td align="left"&gt;0.0195 (2.65%)&amp;#42;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;STEM Career Prediction Gender&lt;/td&gt;&lt;td align="left"&gt;Best Approaches&lt;/td&gt;&lt;td align="left"&gt;&amp;#8211;&lt;/td&gt;&lt;td align="left"&gt;CB&amp;#8208;Demo/SMOTE&amp;#8208;Demo&lt;/td&gt;&lt;td align="left"&gt;Massaging&lt;/td&gt;&lt;td align="left"&gt;Correlation Remover&lt;/td&gt;&lt;td align="left"&gt;Reduction&lt;/td&gt;&lt;td align="left"&gt;FairProj&lt;/td&gt;&lt;td align="left"&gt;AUC&lt;/td&gt;&lt;td align="left"&gt;ABROCA&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Top Single&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.5298&lt;/td&gt;&lt;td align="left"&gt;0.1101&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.5863 (10.66%)&lt;/td&gt;&lt;td align="left"&gt;0.1200 (&amp;#8722;8.96%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;0.5809 (9.65%)&lt;/td&gt;&lt;td align="left"&gt;0.1240 (&amp;#8722;12.59%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.5899 (11.34%)&lt;/td&gt;&lt;td align="left"&gt;0.1202 (&amp;#8722;9.18%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage &amp; Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.6104 (15.21%)&amp;#42;&lt;/td&gt;&lt;td align="left"&gt;0.0974 (11.54%)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Multi&amp;#8208;Stage &amp; Multi&amp;#8208;Source&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;&amp;#9733;&lt;/td&gt;&lt;td align="left"&gt;0.6030 (13.82%)&lt;/td&gt;&lt;td align="left"&gt;0.0922 (16.27%)&amp;#42;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>3 <emph>Note</emph>: For each task‐sensitive attribute, the best‐performing approaches to tackle bias of different sources at each stage were listed in <emph>Best Approaches</emph> row. In <emph>Top Single</emph> row, <ephtml> &lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0011" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8902;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> marks the best single debiasing approach, with AUC and ABROCA provided in the last two columns. For other rows, <ephtml> &lt;math altimg="urn:x-wiley:00071013:media:bjet13575:bjet13575-math-0012" display="inline" overflow="scroll" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;&amp;#8902;&lt;/mo&gt;&lt;/mrow&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> indicates the debiasing approaches involved in combined debiasing with <emph>Multi‐Source</emph> and/or <emph>Multi‐Stage</emph> strategies. Note that the percentage in parentheses represents the performance improvements compared with the best‐achieving result in the <emph>Top Single</emph> row.</p> <hd id="AN0188606177-23">DISCUSSION AND CONCLUSION</hd> <p>The study examined a comprehensive set of state‐of‐the‐art debiasing approaches, focusing on their application in addressing various biases in two‐phase modelling for both textual and tabular data. It also explored the impact of debiasing with respect to one sensitive attribute on students grouped by a different sensitive attribute. Additionally, two simple yet effective strategies for combining these debiasing approaches—Multi‐Source and Multi‐Stage—were proposed. The study concludes by highlighting key findings and offering practical insights to inform future efforts.</p> <p>Based on results from RQ1, this study highlights several key findings that provide general insights for improving fairness across both scenarios, including (i) among the three stages (i.e., pre‐processing, in‐post‐processing and post‐processing), pre‐processing and in‐processing approaches tend to provide more consistent and reliable debiasing outcomes across various contexts. This is particularly evident in the significant fairness improvements achieved by most methods, especially with regard to the gender attribute. These improvements were observed in both the forum relevancy classification task, which involves a larger textual dataset, and the career prediction task, which operates on a significantly smaller tabular dataset. (ii) Notably, in both scenarios fairness was significantly enhanced by addressing <emph>local stereotypical bias</emph>, <emph>label bias</emph> and <emph>proxy discrimination</emph>, underscoring the importance of identifying and mitigating these biases across different types of data, including textual and tabular. Looking into data distribution, note that over‐represented demographic groups consistently suffer from significant stereotypical bias in both datasets. For instance, female students and students whose first language is not English contributed 66.88% and 62.96% of forum posts, respectively, but their posts were less frequently labelled as <emph>content‐relevant</emph>. Similarly, female students pursuing STEM careers were under‐represented. This alignment between the observed <emph>local stereotypical bias</emph> in the data and the fairness improvements achieved by targeted approaches suggests that, unfairness can be effectively mitigated using existing approaches addressing <emph>local stereotypical bias</emph> at the early stage of data preparation either via re‐sampling or re‐weighing. In particular, when the training data are relatively small (e.g., STEM career prediction data), re‐weighing approaches are more favourable compared with re‐sampling that involves synthetic instance generation, as the quality of generated instances might not be guaranteed for small datasets. (iii) Different from distribution‐related bias, <emph>label bias</emph> and <emph>proxy discrimination</emph> are hard to be directly identified but appear to be contributing factors for predictive unfairness. This highlights the need to effectively measure the two types of bias for effectively auditing the data. (iv) Among approaches associated with a fairness constraint, i.e., Label Debias, Reduction and FairBatch, the variant imposed with <emph>equalized odds</emph> tend to be more effective, which is reasonable as <emph>equalized odds</emph> is a specific case of <emph>ABROCA</emph> with threshold as 0.5. This implies the possibility of approaching desired fairness targets by breaking it down to more specific cases easier to tackle, when the targets are difficult to be directly integrated into models. (v) Notably, it is worth pointing out that when dealing with two‐phase modelling, particularly with textual data, most debiasing approaches focused on the downstream task (Rows 3–29 of Table 3), i.e., second‐phase modelling, have failed to enhance fairness w.r.t. first‐language attribute as compared to the gender attribute. The best fairness outcomes were achieved by DebiasedBERT, which adjusted the upstream language model in the first‐phase modelling to modify the text embeddings. This suggests that posts generated by students from different first‐language backgrounds may vary significantly, with these differences often being inherently captured in the textual embeddings produced by large language models. Such inherent disparities are difficult to mitigate through downstream debiasing techniques, highlighting the need for more debiasing research at the first phase of the modelling pipeline.</p> <p>Along with the results from RQ3, the significant fairness improvements achieved by the combined debiasing strategy using Multi‐Source and Multi‐Stage underscore the complexity of biases originating from multiple sources across various stages. Interestingly, Multi‐Source debiasing alone did not outperform the best pre‐processing approach targeting a single source of bias. However, consistent improvements were observed when this strategy was combined with the Multi‐Stage approach. One hypothesis is that this is due to the interplay between biases from different sources; for example, <emph>local stereotypical bias</emph> may contribute to <emph>proxy discrimination</emph>. Based on these findings, the study advocate for further research to better understand the interactions between data biases and to develop more advanced approaches for addressing biases from multiple sources.</p> <p>Beyond the effectiveness of various debiasing approaches, results from RQ2 reveal that while debiasing focused on a single protected attribute can enhance fairness for groups associated with that attribute, it may unfairly impact groups with other unexamined demographic attributes, potentially leading to increased discrimination towards groups with different demographic characteristics. This underscores the need to consider multiple protected attributes in designing future debiasing approaches, specifically approaches aimed at addressing intersectional bias.</p> <hd id="AN0188606177-24">LIMITATIONS</hd> <p>The study acknowledges certain limitations. Firstly, similar to previous research, the evaluation and discussion were confined to binary protected attributes. Most of the debiasing approaches evaluated were originally designed for binary attributes, limiting the scope of the discussion on non‐binary attributes. However, recognizing their importance, the study intends to extend the evaluation to consider non‐binary attributes, including continuous ones. Secondly, the study focuses on only two educational tasks, providing limited insights into more complex scenarios, such as those involving multi‐modal data or tasks requiring deeper contextual understanding. To address this limitation, future work will incorporate a wider range of diverse and complex educational tasks. Thirdly, as a preliminary investigation, the study focused on simple strategies for combining the most effective debiasing approaches. Given the significant improvements observed with such simple combinations, future research should explore more complex, multi‐stage and multi‐source debiasing methods. Lastly, the study specifically focused on fairness metrics that prioritize equal error rates, as these directly affect students' opportunities to receive the necessary learning support. However, it also emphasizes the need for future efforts to measure and improve fairness from another critical perspective: ensuring equal representation of different groups in relation to specific outcomes, such as equal rates of receiving positive predictions, by employing metrics like MADD.</p> <hd id="AN0188606177-25">ACKNOWLEDGEMENT</hd> <p>Open access publishing facilitated by Monash University, as part of the Wiley ‐ Monash University agreement via the Council of Australian University Librarians.</p> <hd id="AN0188606177-26">CONFLICT OF INTEREST STATEMENT</hd> <p>The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.</p> <hd id="AN0188606177-27">DATA AVAILABILITY STATEMENT</hd> <p>We used two datasets in our study, where the first forum post dataset was collected from university, which will not be publicly available due to the regulations related to privacy and ethical guidelines. The second one is an open‐sourced dataset available on https://sites.google.com/view/assistmentsdatamining/dataset.</p> <hd id="AN0188606177-28">ETHICS STATEMENT</hd> <p>The data we used has been passed the critical ethical inspection by University ethic committee and obtained ethical approval.</p> <p>GRAPH: Data S1.</p> <p>GRAPH: Data S2.</p> <ref id="AN0188606177-29"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref50" type="bt">1</bibl> <bibtext> https://sites.google.com/view/assistmentsdatamining/dataset.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref20" type="bt">2</bibl> <bibtext> https://github.com/ckyeungac/ADM2017.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref93" type="bt">3</bibl> <bibtext> https://bit.ly/fairLA.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref58" type="bt">4</bibl> <bibtext> https://github.com/lilylikeart/fairness‐evaluation‐edu</bibtext> </blist> </ref> <ref id="AN0188606177-30"> <title> REFERENCES </title> <blist> <bibtext> Agarwal, A., Beygelzimer, A., Dudík, M., Langford, J., &amp; Wallach, H. (2018). A reductions approach to fair classification. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmassan, Stockholm, Sweden (pp. 60 – 69). PMLR.</bibtext> </blist> <blist> <bibtext> Aguillon, S. M., Siegmund, G.‐F., Petipas, R. H., Drake, A. G., Cotner, S., &amp; Ballen, C. J. (2020). Gender differences in student participation in an active‐learning classroom. CBE–Life Sciences Education, 19, ar12.</bibtext> </blist> <blist> <bibtext> Alghamdi, W., Hsu, H., Jeong, H., Wang, H., Michalak, P., Asoodeh, S., &amp; Calmon, F. (2022). Beyond adult and compas: Fair multi‐class prediction via information projection. Advances in Neural Information Processing Systems, 35, 38747 – 38760.</bibtext> </blist> <blist> <bibtext> Almatrafi, O., Johri, A., &amp; Rangwala, H. (2018). Needle in a haystack: Identifying learner posts that require urgent response in MOOC discussion forums. Computers &amp; Education, 118, 1 – 9.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref11" type="bt">5</bibl> <bibtext> Artelt, C. (2015). Teacher judgments and their role in the educational process. In Emerging trends in the social and behavioral sciences: An interdisciplinary, searchable, and linkable resource (pp. 1 – 16). Wiley.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref18" type="bt">6</bibl> <bibtext> Ashong, C. Y., &amp; Commander, N. E. (2012). Ethnicity, gender, and perceptions of online learning in higher education. MERLOT Journal of Online Learning and Teaching, 8, 98 – 100.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref1" type="bt">7</bibl> <bibtext> Baker, R. S., &amp; Hawn, A. (2021). Algorithmic bias in education. International Journal of Artificial Intelligence in Education, 32, 1052 – 1092.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref31" type="bt">8</bibl> <bibtext> Bayer, V., Hlosta, M., &amp; Fernandez, M. (2021). Learning analytics and fairness: Do existing algorithms serve everyone equally? In International Conference on Artificial Intelligence in Education, Utrecht, The Netherlands, pp. 71 – 75. Springer.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref15" type="bt">9</bibl> <bibtext> Beede, D. N., Julian, T. A., Langdon, D., McKittrick, G., Khan, B., &amp; Doms, M. E. (2011). Women in stem: A gender gap to innovation. Economics and Statistics Administration Issue Brief.</bibtext> </blist> <blist> <bibtext> BehnamGhader, P., &amp; Milios, A. (2022). An analysis of social biases present in BERT variants across multiple languages. In Workshop on Trustworthy and Socially Responsible Machine Learning (p. 2022). NeurIPS.</bibtext> </blist> <blist> <bibtext> Bird, S., Dudík, M., Edgar, R., Horn, B., Lutz, R., Milan, V., Sameki, M., Wallach, H., &amp; Walker, K. (2020). Fairlearn: A toolkit for assessing and improving fairness in AI. Microsoft, Tech. Rep. MSR‐TR‐2020‐32.</bibtext> </blist> <blist> <bibtext> Caliskan, A., Bryson, J. J., &amp; Narayanan, A. (2017). Semantics derived automatically from language corpora contain human‐like biases. Science, 356, 183 – 186.</bibtext> </blist> <blist> <bibtext> Campbell, T. (2015). Stereotyped at seven? Biases in teacher judgement of pupils' ability and attainment. Journal of Social Policy, 44, 517 – 547.</bibtext> </blist> <blist> <bibtext> Caton, S., &amp; Haas, C. (2020). Fairness in machine learning: A survey. arXiv preprint arXiv:2010.04053.</bibtext> </blist> <blist> <bibtext> Chawla, N. V., Bowyer, K. W., Hall, L. O., &amp; Kegelmeyer, W. P. (2002). Smote: Synthetic minority over‐sampling technique. Journal of Artificial Intelligence Research, 16, 321 – 357.</bibtext> </blist> <blist> <bibtext> Clavié, B., &amp; Gal, K. (2019). Edubert: Pretrained deep language models for learning analytics. arXiv preprint arXiv:1912.00690.</bibtext> </blist> <blist> <bibtext> Deho, O. B., Zhan, C., Li, J., Liu, J., Liu, L., &amp; Duy Le, T. (2022). How do the existing fairness metrics and unfairness mitigation algorithms contribute to ethical learning analytics? British Journal of Educational Technology, 53, 822 – 843.</bibtext> </blist> <blist> <bibtext> Dominguez‐Catena, I., Paternain, D., &amp; Galar, M. (2024). Metrics for dataset demographic bias: A case study on facial expression recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46, 5209 – 5226.</bibtext> </blist> <blist> <bibtext> Friedler, S. A., Scheidegger, C., Venkatasubramanian, S., Choudhary, S., Hamilton, E. P., &amp; Roth, D. (2019). A comparative study of fairness‐enhancing interventions in machine learning. In Proceedings of the Conference on Fairness, Accountability, and Transparency, Atlanta, Georgia, USA, pp. 329 – 338. ACM.</bibtext> </blist> <blist> <bibtext> García, E., &amp; Weiss, E. (2017). Education inequalities at the school starting gate: Gaps, trends, and strategies to address them. Economic Policy Institute.</bibtext> </blist> <blist> <bibtext> Gardner, J., Brooks, C., &amp; Baker, R. (2019). Evaluating the fairness of predictive student models through slicing analysis. In Proceedings of the 9th International Conference on Learning Analytics &amp; Knowledge, Tempe, AZ, USA, pp. 225 – 234. ACM.</bibtext> </blist> <blist> <bibtext> Han, X., Baldwin, T., &amp; Cohn, T. (2022). Balancing out bias: Achieving fairness through balanced training. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 11335 – 11350). Association for Computational Linguistics.</bibtext> </blist> <blist> <bibtext> Han, X., Shen, A., Cohn, T., Baldwin, T., &amp; Frermann, L. (2022). Systematic evaluation of predictive fairness. In Proceedings of the 2nd Conference of the Asia‐Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 68 – 81, Online only. Association for Computational Linguistics.</bibtext> </blist> <blist> <bibtext> Jiang, H., &amp; Nachum, O. (2020). Identifying and correcting label bias in machine learning. In International Conference on Artificial Intelligence and Statistics, Palermo, Italy, pp. 702 – 712. PMLR.</bibtext> </blist> <blist> <bibtext> Kamiran, F., &amp; Calders, T. (2012). Data preprocessing techniques for classification without discrimination. Knowledge and Information Systems, 33, 1 – 33.</bibtext> </blist> <blist> <bibtext> Kizilcec, R. F., &amp; Lee, H. (2022). Algorithmic fairness in education. In The ethics of artificial intelligence in education (pp. 174 – 202). Routledge.</bibtext> </blist> <blist> <bibtext> Lamba, H., Rodolfa, K. T., &amp; Ghani, R. (2021). An empirical comparison of bias reduction methods on real‐world problems in high‐stakes policy settings. ACM SIGKDD Explorations Newsletter, 23, 69 – 85.</bibtext> </blist> <blist> <bibtext> Li, L., Sha, L., Li, Y., Raković, M., Rong, J., Joksimovic, S., Selwyn, N., Gašević, D., &amp; Chen, G. (2023). Moral machines or tyranny of the majority? A systematic review on predictive bias in education. In LAK23: 13th International Learning Analytics and Knowledge Conference, Arlinton, TX, USA, pp. 499 – 508. ACM.</bibtext> </blist> <blist> <bibtext> Li, L., Srivastava, N., Rong, J., Pianta, G., Varanasi, R., Gašević, D., &amp; Chen, G. (2024). Unveiling Goods and Bads: A Critical Analysis of Machine Learning Predictions of Standardized Test Performance in Early Childhood Education. Proceedings of the 14th Learning Analytics and Knowledge Conference, Kyoto, Japan, pp. 608 – 619. ACM.</bibtext> </blist> <blist> <bibtext> Litman, D., Zhang, H., Correnti, R., Matsumura, L. C., &amp; Wang, E. (2021). A fairness evaluation of automated methods for scoring text evidence usage in writing. In International Conference on Artificial Intelligence in Education, Utrecht, The Netherlands, pp. 255 – 267. Springer.</bibtext> </blist> <blist> <bibtext> Liu, H., Jin, W., Karimi, H., Liu, Z., &amp; Tang, J. (2021). The authors matter: Understanding and mitigating implicit bias in deep text classification. arXiv preprint arXiv:2105.02778.</bibtext> </blist> <blist> <bibtext> Malouff, J. M., &amp; Thorsteinsson, E. B. (2016). Bias in grading: A meta‐analysis of experimental research findings. Australian Journal of Education, 60, 245 – 256.</bibtext> </blist> <blist> <bibtext> Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., &amp; Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR), 54, 1 – 35.</bibtext> </blist> <blist> <bibtext> Moss‐Racusin, C. A., Dovidio, J. F., Brescoll, V. L., Graham, M. J., &amp; Handelsman, J. (2012). Science faculty's subtle gender biases favor male students. Proceedings of the National Academy of Sciences, 109, 16474 – 16479.</bibtext> </blist> <blist> <bibtext> Patikorn, T., Baker, R. S., &amp; Heffernan, N. T. (2020). Assistments longitudinal data mining competition special issue: A preface. Journal of Educational Data Mining, 12, i – xi.</bibtext> </blist> <blist> <bibtext> Pleiss, G., Raghavan, M., Wu, F., Kleinberg, J., &amp; Weinberger, K. Q. (2017). On fairness and calibration. Advances in Neural Information Processing Systems, 30, 5690 – 5689.</bibtext> </blist> <blist> <bibtext> Roh, Y., Lee, K., Whang, S. E., &amp; Suh, C. (2021). Fairbatch: Batch selection for model fairness. In 9th International Conference on Learning Representations, Online only.</bibtext> </blist> <blist> <bibtext> Sadker, M., &amp; Sadker, D. (2010). Failing at fairness: How America's schools cheat girls. Simon and Schuster.</bibtext> </blist> <blist> <bibtext> Sha, L., Li, Y., Gasevic, D., &amp; Chen, G. (2022). Bigger data or fairer data? Augmenting BERT via active sampling for educational text classification. In Proceedings of the 29th International Conference on Computational Linguistics, Gyeongju, Republic of Korea, pp. 1275 – 1285.</bibtext> </blist> <blist> <bibtext> Suresh, H., &amp; Guttag, J. (2021). A framework for understanding sources of harm throughout the machine learning life cycle. In Equity and access in algorithms, mechanisms, and optimization (pp. 1 – 9). ACM.</bibtext> </blist> <blist> <bibtext> Tschantz, M. C. (2022). What is proxy discrimination? In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea, pp. 1993 – 2003. ACM.</bibtext> </blist> <blist> <bibtext> Vasquez Verdugo, J., Gitiaux, X., Ortega, C., &amp; Rangwala, H. (2022). Faired: A systematic fairness analysis approach applied in a higher educational context. In LAK22: 12th International Learning Analytics and Knowledge Conference, Online, USA, pp. 271 – 281. ACM.</bibtext> </blist> <blist> <bibtext> Verger, M., Lallé, S., Bouchet, F., &amp; Luengo, V. (2023). Is your model "MADD"? a novel metric to evaluate algorithmic fairness for predictive student models. arXiv preprint arXiv:2305.15342.</bibtext> </blist> <blist> <bibtext> Wang, T., Zhao, J., Yatskar, M., Chang, K.‐W., &amp; Ordonez, V. (2019). Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Los Alamitos, CA, USA, pp. 5310 – 5319. IEEE Computer Society.</bibtext> </blist> <blist> <bibtext> Yan, S., Kao, H.‐t., &amp; Ferrara, E. (2020). Fair class balancing: Enhancing model fairness without observing sensitive attributes. In Proceedings of the 29th ACM International Conference on Information &amp; Knowledge Management, Virtual Event, Ireland, pp. 1715 – 1724. ACM.</bibtext> </blist> <blist> <bibtext> Yeung, C.‐K., &amp; Yeung, D.‐Y. (2019). Incorporating features learned by an enhanced deep knowledge tracing model for stem/non‐stem job prediction. International Journal of Artificial Intelligence in Education, 29, 317 – 341.</bibtext> </blist> <blist> <bibtext> Zhang, B. H., Lemoine, B., &amp; Mitchell, M. (2018). Mitigating unwanted biases with adversarial learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, New Orleans, LA, USA, pp. 335 – 340. ACM.</bibtext> </blist> <blist> <bibtext> Zhuhadar, L., Marklin, S., Thrasher, E., &amp; Lytras, M. D. (2016). Is there a gender difference in interacting with intelligent tutoring system? Can bayesian knowledge tracing and learning curve analysis models answer this question? Computers in Human Behavior, 61, 198 – 204.</bibtext> </blist> </ref> <aug> <p>By Lin Li; Namrata Srivastava; Jia Rong; Quanlong Guan; Dragan Gašević and Guanliang Chen</p> <p>Reported by Author; Author; Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib28" firstref="ref2"></nolink> <nolink nlid="nl2" bibid="bib29" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib30" firstref="ref4"></nolink> <nolink nlid="nl4" bibid="bib26" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib14" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib18" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib24" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib41" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib10" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib13" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib20" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib38" firstref="ref14"></nolink> <nolink nlid="nl13" bibid="bib32" firstref="ref17"></nolink> <nolink nlid="nl14" bibid="bib48" firstref="ref19"></nolink> <nolink nlid="nl15" bibid="bib34" firstref="ref21"></nolink> <nolink nlid="nl16" bibid="bib17" firstref="ref23"></nolink> <nolink nlid="nl17" bibid="bib23" firstref="ref24"></nolink> <nolink nlid="nl18" bibid="bib27" firstref="ref25"></nolink> <nolink nlid="nl19" bibid="bib21" firstref="ref27"></nolink> <nolink nlid="nl20" bibid="bib16" firstref="ref28"></nolink> <nolink nlid="nl21" bibid="bib46" firstref="ref29"></nolink> <nolink nlid="nl22" bibid="bib33" firstref="ref34"></nolink> <nolink nlid="nl23" bibid="bib43" firstref="ref37"></nolink> <nolink nlid="nl24" bibid="bib40" firstref="ref42"></nolink> <nolink nlid="nl25" bibid="bib25" firstref="ref47"></nolink> <nolink nlid="nl26" bibid="bib42" firstref="ref49"></nolink> <nolink nlid="nl27" bibid="bib47" firstref="ref51"></nolink> <nolink nlid="nl28" bibid="bib19" firstref="ref53"></nolink> <nolink nlid="nl29" bibid="bib35" firstref="ref61"></nolink> <nolink nlid="nl30" bibid="bib15" firstref="ref67"></nolink> <nolink nlid="nl31" bibid="bib22" firstref="ref68"></nolink> <nolink nlid="nl32" bibid="bib44" firstref="ref69"></nolink> <nolink nlid="nl33" bibid="bib45" firstref="ref77"></nolink> <nolink nlid="nl34" bibid="bib11" firstref="ref86"></nolink> <nolink nlid="nl35" bibid="bib37" firstref="ref90"></nolink> <nolink nlid="nl36" bibid="bib36" firstref="ref92"></nolink> <nolink nlid="nl37" bibid="bib12" firstref="ref94"></nolink> <nolink nlid="nl38" bibid="bib31" firstref="ref95"></nolink> <nolink nlid="nl39" bibid="bib39" firstref="ref96"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1486314 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: When and How Biases Seep In: Enhancing Debiasing Approaches for Fair Educational Predictive Analytics – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Lin+Li%22">Lin Li</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4205-7975">0000-0002-4205-7975</externalLink>)<br /><searchLink fieldCode="AR" term="%22Namrata+Srivastava%22">Namrata Srivastava</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-4194-318X">0000-0003-4194-318X</externalLink>)<br /><searchLink fieldCode="AR" term="%22Jia+Rong%22">Jia Rong</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-9462-3924">0000-0002-9462-3924</externalLink>)<br /><searchLink fieldCode="AR" term="%22Quanlong+Guan%22">Quanlong Guan</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6911-3853">0000-0001-6911-3853</externalLink>)<br /><searchLink fieldCode="AR" term="%22Dragan+Gaševic%22">Dragan Gaševic</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9265-1908">0000-0001-9265-1908</externalLink>)<br /><searchLink fieldCode="AR" term="%22Guanliang+Chen%22">Guanliang Chen</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-8236-3133">0000-0002-8236-3133</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22British+Journal+of+Educational+Technology%22"><i>British Journal of Educational Technology</i></searchLink>. 2025 56(6):2478-2501. – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 24 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Bias%22">Bias</searchLink><br /><searchLink fieldCode="DE" term="%22Attitude+Change%22">Attitude Change</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Analytics%22">Learning Analytics</searchLink><br /><searchLink fieldCode="DE" term="%22Social+Bias%22">Social Bias</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Mediated+Communication%22">Computer Mediated Communication</searchLink><br /><searchLink fieldCode="DE" term="%22Social+Media%22">Social Media</searchLink><br /><searchLink fieldCode="DE" term="%22Stereotypes%22">Stereotypes</searchLink><br /><searchLink fieldCode="DE" term="%22Labeling+%28of+Persons%29%22">Labeling (of Persons)</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/bjet.13575 – Name: ISSN Label: ISSN Group: ISSN Data: 0007-1013<br />1467-8535 – Name: Abstract Label: Abstract Group: Ab Data: The use of predictive analytics powered by machine learning (ML) to model educational data has increasingly been identified to exhibit bias towards marginalized populations, prompting the need for more equitable applications of these techniques. To tackle bias that emerges in training data or models at different stages of the ML modelling pipeline, numerous debiasing approaches have been proposed. Yet, research into state-of-the-art techniques for effectively employing these approaches to enhance fairness in educational predictive scenarios remains limited. Prior studies often focused on mitigating bias from a single source at a specific stage of model construction within narrowly defined scenarios, overlooking the complexities of bias originating from multiple sources across various stages. Moreover, these approaches were often evaluated using typical threshold-dependent fairness metrics, which fail to account for real-world educational scenarios where thresholds are typically unknown before evaluation. To bridge these gaps, this study systematically examined a total of 28 representative debiasing approaches, categorized by the sources of bias and the stage they targeted, for two critical educational predictive tasks, namely forum post classification and student career prediction. Both tasks involve a two-phase modelling process where features learned from upstream models in the first phase are fed into classical ML models for final predictions, which is a common yet under-explored setting for educational data modelling. The study observed that addressing local stereotypical bias, label bias or proxy discrimination in training data, as well as imposing fairness constraints on models, can effectively enhance predictive fairness. But their efficacy was often compromised when features from upstream models were inherently biased. Beyond that, this study proposes two novel strategies, namely Multi-Stage and Multi-Source debiasing to integrate existing approaches. These strategies demonstrated substantial improvements in mitigating unfairness, underscoring the importance of unified approaches capable of addressing biases from various sources across multiple stages. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1486314 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1486314 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/bjet.13575 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 24 StartPage: 2478 Subjects: – SubjectFull: Bias Type: general – SubjectFull: Attitude Change Type: general – SubjectFull: Prediction Type: general – SubjectFull: Learning Analytics Type: general – SubjectFull: Social Bias Type: general – SubjectFull: Computer Mediated Communication Type: general – SubjectFull: Social Media Type: general – SubjectFull: Stereotypes Type: general – SubjectFull: Labeling (of Persons) Type: general Titles: – TitleFull: When and How Biases Seep In: Enhancing Debiasing Approaches for Fair Educational Predictive Analytics Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Lin Li – PersonEntity: Name: NameFull: Namrata Srivastava – PersonEntity: Name: NameFull: Jia Rong – PersonEntity: Name: NameFull: Quanlong Guan – PersonEntity: Name: NameFull: Dragan Gaševic – PersonEntity: Name: NameFull: Guanliang Chen IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 11 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 0007-1013 – Type: issn-electronic Value: 1467-8535 Numbering: – Type: volume Value: 56 – Type: issue Value: 6 Titles: – TitleFull: British Journal of Educational Technology Type: main |
| ResultId | 1 |