Training and Evaluating with Human Label Variation: An Empirical Study.
Saved in:
| Title: | Training and Evaluating with Human Label Variation: An Empirical Study. |
|---|---|
| Authors: | Kurniawan, Kemal1 (AUTHOR) kurniawan.k@unimelb.edu.au, Mistica, Meladel1 (AUTHOR) kurniawan.k@unimelb.edu.au, Baldwin, Timothy1,2 (AUTHOR) kurniawan.k@unimelb.edu.au, Lau, Jey Han1 (AUTHOR) kurniawan.k@unimelb.edu.au |
| Source: | Computational Linguistics. Mar2026, Vol. 52 Issue 1, p85-111. 27p. |
| Subjects: | Annotations, Evaluation methodology, Fuzzy sets, Machine learning, Empirical research |
| Abstract: | Human label variation (HLV) challenges the standard assumption that a labeled instance has a single ground truth, instead embracing the natural variation in human annotation to train and evaluate models. While various training methods and metrics for HLV have been proposed, it is still unclear which methods and metrics perform best in what settings. We propose new evaluation metrics for HLV leveraging fuzzy set theory. Because these new proposed metrics are differentiable, we then in turn experiment with using these metrics as training objectives. We conduct an extensive study over 6 HLV datasets testing 14 training methods and 6 evaluation metrics. We find that training on either disaggregated annotations or soft labels performs best across metrics, outperforming training using the proposed training objectives with differentiable metrics. We also show that our proposed soft micro F1 score is one of the best metrics for HLV data. [ABSTRACT FROM AUTHOR] |
| Copyright of Computational Linguistics is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Human label variation (HLV) challenges the standard assumption that a labeled instance has a single ground truth, instead embracing the natural variation in human annotation to train and evaluate models. While various training methods and metrics for HLV have been proposed, it is still unclear which methods and metrics perform best in what settings. We propose new evaluation metrics for HLV leveraging fuzzy set theory. Because these new proposed metrics are differentiable, we then in turn experiment with using these metrics as training objectives. We conduct an extensive study over 6 HLV datasets testing 14 training methods and 6 evaluation metrics. We find that training on either disaggregated annotations or soft labels performs best across metrics, outperforming training using the proposed training objectives with differentiable metrics. We also show that our proposed soft micro F1 score is one of the best metrics for HLV data. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 08912017 |
| DOI: | 10.1162/COLI.a.578 |