A multi-class imbalanced data stream classification algorithm based on sample weighting and adaptive oversampling.

Saved in:
Bibliographic Details
Title: A multi-class imbalanced data stream classification algorithm based on sample weighting and adaptive oversampling.
Authors: Han, Meng1 (AUTHOR), Zhu, Shineng1 (AUTHOR) 2003051@nmu.edu.cn, Yang, Shurong1 (AUTHOR), Dai, Zhenlong1 (AUTHOR), Yang, Wenyan1 (AUTHOR), Ding, Jian1 (AUTHOR), Li, Juan1 (AUTHOR)
Source: Data Mining & Knowledge Discovery. Jun2026, Vol. 40 Issue 3, p1-38. 38p.
Subjects: Ensemble learning, Loss functions (Statistics)
Abstract: Data stream learning is typically affected by class distribution imbalance, which causes learning algorithms to be severely biased toward the majority class due to the scarcity of minority samples, thereby significantly reducing their ability to identify minority instances. At the same time, data streams commonly exhibit concept drift, defined as the phenomenon where the input–output relationship learned by the model changes over time, requiring models to maintain ongoing adaptability. The challenges of multi-class imbalance and concept drift greatly weaken the generalization ability and classification accuracy of models. To address these issues, an ensemble classification algorithm tailored for multi-class imbalanced data streams is proposed. First, an expert group validation strategy based on feature-importance updates is designed to effectively capture the dynamic characteristics of multi-class imbalanced data streams under concept drift, thereby enhancing the adaptability of the ensemble model. Second, an adaptive oversampling strategy is introduced, which flexibly adjusts the sample generation process to produce instances consistent with the current data distribution, thus improving the recognition ability of the model for minority classes in drifting environments. Finally, a sample weighting strategy based on Focal Loss and boundary proximity is developed to dynamically adjust the training weights of samples, thereby strengthening the model's focus on minority or highly uncertain instances. Experimental results demonstrate that the proposed algorithm achieves superior performance on various multi-class imbalanced data streams with concept drift, outperforming existing methods. [ABSTRACT FROM AUTHOR]
Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 192346966
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A multi-class imbalanced data stream classification algorithm based on sample weighting and adaptive oversampling.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Han%2C+Meng%22">Han, Meng</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhu%2C+Shineng%22">Zhu, Shineng</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> 2003051@nmu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Yang%2C+Shurong%22">Yang, Shurong</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Dai%2C+Zhenlong%22">Dai, Zhenlong</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Yang%2C+Wenyan%22">Yang, Wenyan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Ding%2C+Jian%22">Ding, Jian</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Li%2C+Juan%22">Li, Juan</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Data+Mining+%26+Knowledge+Discovery%22">Data Mining & Knowledge Discovery</searchLink>. Jun2026, Vol. 40 Issue 3, p1-38. 38p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Ensemble+learning%22">Ensemble learning</searchLink><br /><searchLink fieldCode="DE" term="%22Loss+functions+%28Statistics%29%22">Loss functions (Statistics)</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Data stream learning is typically affected by class distribution imbalance, which causes learning algorithms to be severely biased toward the majority class due to the scarcity of minority samples, thereby significantly reducing their ability to identify minority instances. At the same time, data streams commonly exhibit concept drift, defined as the phenomenon where the input–output relationship learned by the model changes over time, requiring models to maintain ongoing adaptability. The challenges of multi-class imbalance and concept drift greatly weaken the generalization ability and classification accuracy of models. To address these issues, an ensemble classification algorithm tailored for multi-class imbalanced data streams is proposed. First, an expert group validation strategy based on feature-importance updates is designed to effectively capture the dynamic characteristics of multi-class imbalanced data streams under concept drift, thereby enhancing the adaptability of the ensemble model. Second, an adaptive oversampling strategy is introduced, which flexibly adjusts the sample generation process to produce instances consistent with the current data distribution, thus improving the recognition ability of the model for minority classes in drifting environments. Finally, a sample weighting strategy based on Focal Loss and boundary proximity is developed to dynamically adjust the training weights of samples, thereby strengthening the model's focus on minority or highly uncertain instances. Experimental results demonstrate that the proposed algorithm achieves superior performance on various multi-class imbalanced data streams with concept drift, outperforming existing methods. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=192346966
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10618-026-01197-9
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 38
        StartPage: 1
    Subjects:
      – SubjectFull: Ensemble learning
        Type: general
      – SubjectFull: Loss functions (Statistics)
        Type: general
    Titles:
      – TitleFull: A multi-class imbalanced data stream classification algorithm based on sample weighting and adaptive oversampling.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Han, Meng
      – PersonEntity:
          Name:
            NameFull: Zhu, Shineng
      – PersonEntity:
          Name:
            NameFull: Yang, Shurong
      – PersonEntity:
          Name:
            NameFull: Dai, Zhenlong
      – PersonEntity:
          Name:
            NameFull: Yang, Wenyan
      – PersonEntity:
          Name:
            NameFull: Ding, Jian
      – PersonEntity:
          Name:
            NameFull: Li, Juan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 13845810
          Numbering:
            – Type: volume
              Value: 40
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Data Mining & Knowledge Discovery
              Type: main
ResultId 1