Training and Evaluating with Human Label Variation: An Empirical Study.

Saved in:
Bibliographic Details
Title: Training and Evaluating with Human Label Variation: An Empirical Study.
Authors: Kurniawan, Kemal1 (AUTHOR) kurniawan.k@unimelb.edu.au, Mistica, Meladel1 (AUTHOR) kurniawan.k@unimelb.edu.au, Baldwin, Timothy1,2 (AUTHOR) kurniawan.k@unimelb.edu.au, Lau, Jey Han1 (AUTHOR) kurniawan.k@unimelb.edu.au
Source: Computational Linguistics. Mar2026, Vol. 52 Issue 1, p85-111. 27p.
Subjects: Annotations, Evaluation methodology, Fuzzy sets, Machine learning, Empirical research
Abstract: Human label variation (HLV) challenges the standard assumption that a labeled instance has a single ground truth, instead embracing the natural variation in human annotation to train and evaluate models. While various training methods and metrics for HLV have been proposed, it is still unclear which methods and metrics perform best in what settings. We propose new evaluation metrics for HLV leveraging fuzzy set theory. Because these new proposed metrics are differentiable, we then in turn experiment with using these metrics as training objectives. We conduct an extensive study over 6 HLV datasets testing 14 training methods and 6 evaluation metrics. We find that training on either disaggregated annotations or soft labels performs best across metrics, outperforming training using the proposed training objectives with differentiable metrics. We also show that our proposed soft micro F1 score is one of the best metrics for HLV data. [ABSTRACT FROM AUTHOR]
Copyright of Computational Linguistics is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 192910418
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Training and Evaluating with Human Label Variation: An Empirical Study.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kurniawan%2C+Kemal%22">Kurniawan, Kemal</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> kurniawan.k@unimelb.edu.au</i><br /><searchLink fieldCode="AR" term="%22Mistica%2C+Meladel%22">Mistica, Meladel</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> kurniawan.k@unimelb.edu.au</i><br /><searchLink fieldCode="AR" term="%22Baldwin%2C+Timothy%22">Baldwin, Timothy</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> kurniawan.k@unimelb.edu.au</i><br /><searchLink fieldCode="AR" term="%22Lau%2C+Jey+Han%22">Lau, Jey Han</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> kurniawan.k@unimelb.edu.au</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Computational+Linguistics%22">Computational Linguistics</searchLink>. Mar2026, Vol. 52 Issue 1, p85-111. 27p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Annotations%22">Annotations</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+methodology%22">Evaluation methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Fuzzy+sets%22">Fuzzy sets</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Empirical+research%22">Empirical research</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Human label variation (HLV) challenges the standard assumption that a labeled instance has a single ground truth, instead embracing the natural variation in human annotation to train and evaluate models. While various training methods and metrics for HLV have been proposed, it is still unclear which methods and metrics perform best in what settings. We propose new evaluation metrics for HLV leveraging fuzzy set theory. Because these new proposed metrics are differentiable, we then in turn experiment with using these metrics as training objectives. We conduct an extensive study over 6 HLV datasets testing 14 training methods and 6 evaluation metrics. We find that training on either disaggregated annotations or soft labels performs best across metrics, outperforming training using the proposed training objectives with differentiable metrics. We also show that our proposed soft micro F1 score is one of the best metrics for HLV data. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Computational Linguistics is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=192910418
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1162/COLI.a.578
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 27
        StartPage: 85
    Subjects:
      – SubjectFull: Annotations
        Type: general
      – SubjectFull: Evaluation methodology
        Type: general
      – SubjectFull: Fuzzy sets
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Empirical research
        Type: general
    Titles:
      – TitleFull: Training and Evaluating with Human Label Variation: An Empirical Study.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kurniawan, Kemal
      – PersonEntity:
          Name:
            NameFull: Mistica, Meladel
      – PersonEntity:
          Name:
            NameFull: Baldwin, Timothy
      – PersonEntity:
          Name:
            NameFull: Lau, Jey Han
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: Mar2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 08912017
          Numbering:
            – Type: volume
              Value: 52
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Computational Linguistics
              Type: main
ResultId 1