Don't push the button! Exploring data leakage risks in machine learning and transfer learning.

Saved in:
Bibliographic Details
Title: Don't push the button! Exploring data leakage risks in machine learning and transfer learning.
Authors: Apicella, Andrea1 (AUTHOR) andapicella@unisa.it, Isgrò, Francesco2 (AUTHOR) francesco.isgro@unina.it, Prevete, Roberto2 (AUTHOR) roberto.prevete@unina.it
Source: Artificial Intelligence Review. Nov2025, Vol. 58 Issue 11, p1-58. 58p.
Subjects: Machine learning, Transfer of training, Evaluation methodology, Supervised learning, Data security failures
Abstract: Machine Learning (ML) has revolutionized various domains, offering predictive capabilities in several areas. However, there is growing evidence in the literature that ML approaches are not always used appropriately, leading to incorrect and sometimes overly optimistic results. One reason for this inappropriate use of ML may be the increasing availability of machine learning tools, leading to what we call the "push the button" approach. While this approach provides convenience, it raises concerns about the reliability of outcomes, leading to challenges such as incorrect performance evaluation. In particular, this paper addresses a critical issue in ML, known as data leakage, where unintended information contaminates the training data, impacting model performance evaluation. Indeed, crucial steps in ML pipeline can be inadvertently overlooked, leading to optimistic performance estimates that may not hold in real-world scenarios. The discrepancy between evaluated and actual performance on new data is a significant concern. In particular, this paper categorizes data leakage in ML, discussing how certain conditions can propagate through the ML approach workflow. Furthermore, it explores the connection between data leakage and the specific task being addressed, investigates its occurrence in Transfer Learning framework, and compares standard inductive ML with transductive ML paradigms. The conclusion summarizes key findings, emphasizing the importance of addressing data leakage for robust and reliable ML applications considering tasks and generalization goals. [ABSTRACT FROM AUTHOR]
Copyright of Artificial Intelligence Review is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 187436280
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Don't push the button! Exploring data leakage risks in machine learning and transfer learning.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Apicella%2C+Andrea%22">Apicella, Andrea</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> andapicella@unisa.it</i><br /><searchLink fieldCode="AR" term="%22Isgrò%2C+Francesco%22">Isgrò, Francesco</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> francesco.isgro@unina.it</i><br /><searchLink fieldCode="AR" term="%22Prevete%2C+Roberto%22">Prevete, Roberto</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> roberto.prevete@unina.it</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Artificial+Intelligence+Review%22">Artificial Intelligence Review</searchLink>. Nov2025, Vol. 58 Issue 11, p1-58. 58p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Transfer+of+training%22">Transfer of training</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+methodology%22">Evaluation methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Supervised+learning%22">Supervised learning</searchLink><br /><searchLink fieldCode="DE" term="%22Data+security+failures%22">Data security failures</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Machine Learning (ML) has revolutionized various domains, offering predictive capabilities in several areas. However, there is growing evidence in the literature that ML approaches are not always used appropriately, leading to incorrect and sometimes overly optimistic results. One reason for this inappropriate use of ML may be the increasing availability of machine learning tools, leading to what we call the "push the button" approach. While this approach provides convenience, it raises concerns about the reliability of outcomes, leading to challenges such as incorrect performance evaluation. In particular, this paper addresses a critical issue in ML, known as data leakage, where unintended information contaminates the training data, impacting model performance evaluation. Indeed, crucial steps in ML pipeline can be inadvertently overlooked, leading to optimistic performance estimates that may not hold in real-world scenarios. The discrepancy between evaluated and actual performance on new data is a significant concern. In particular, this paper categorizes data leakage in ML, discussing how certain conditions can propagate through the ML approach workflow. Furthermore, it explores the connection between data leakage and the specific task being addressed, investigates its occurrence in Transfer Learning framework, and compares standard inductive ML with transductive ML paradigms. The conclusion summarizes key findings, emphasizing the importance of addressing data leakage for robust and reliable ML applications considering tasks and generalization goals. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Artificial Intelligence Review is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=187436280
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10462-025-11326-3
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 58
        StartPage: 1
    Subjects:
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Transfer of training
        Type: general
      – SubjectFull: Evaluation methodology
        Type: general
      – SubjectFull: Supervised learning
        Type: general
      – SubjectFull: Data security failures
        Type: general
    Titles:
      – TitleFull: Don't push the button! Exploring data leakage risks in machine learning and transfer learning.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Apicella, Andrea
      – PersonEntity:
          Name:
            NameFull: Isgrò, Francesco
      – PersonEntity:
          Name:
            NameFull: Prevete, Roberto
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 11
              Text: Nov2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 02692821
          Numbering:
            – Type: volume
              Value: 58
            – Type: issue
              Value: 11
          Titles:
            – TitleFull: Artificial Intelligence Review
              Type: main
ResultId 1