Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties.

Saved in:
Bibliographic Details
Title: Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties.
Authors: Liu, Suyuan1 (AUTHOR) suyuan97@student.ubc.ca, Sóskuthy, Márton1 (AUTHOR), Zhang, Sijia1 (AUTHOR)
Source: Journal of the Acoustical Society of America. Apr2026, Vol. 159 Issue 4, p3164-3180. 17p.
Subjects: Mandarin dialects, Best practices, Phonetics, Speech processing systems, Statistical accuracy, Hierarchical Bayes model, Speech
Abstract: Forced alignment is widely used in phonetic research to align transcripts with acoustic signals. Yet there exists a lack of agreement on conventions for evaluating forced alignment, and our understanding of the reliability of forced aligners rests primarily on results from English. This study aims to fill these gaps by examining the concrete issue of forced aligning different Mandarin varieties. It evaluates machine-generated alignments from Montreal Forced Aligner (MFA); McAuliffe, Socolof, Mihuc, Wagner and Sonderegger [Proc. Interspeech 2017, 498–502 (2017a)] against two sets of independent human baselines using a Bayesian hierarchical multivariate regression model. Our findings suggest closer agreement between human aligners than between humans and MFA; large differences in alignment accuracy across different sequence types, with somewhat divergent patterns of errors across humans and MFA; some effects of speech rate and speaker-specific variation; and essentially no variation in robustness across varieties. These results serve (i) to reinforce previous results on the robustness of forced alignment across different varieties of the same language; and (ii) to provide a set of important methodological recommenations for evaluating forced alignment accuracy. [ABSTRACT FROM AUTHOR]
Copyright of Journal of the Acoustical Society of America is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 193402363
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Liu%2C+Suyuan%22">Liu, Suyuan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> suyuan97@student.ubc.ca</i><br /><searchLink fieldCode="AR" term="%22Sóskuthy%2C+Márton%22">Sóskuthy, Márton</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhang%2C+Sijia%22">Zhang, Sijia</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+the+Acoustical+Society+of+America%22">Journal of the Acoustical Society of America</searchLink>. Apr2026, Vol. 159 Issue 4, p3164-3180. 17p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Mandarin+dialects%22">Mandarin dialects</searchLink><br /><searchLink fieldCode="DE" term="%22Best+practices%22">Best practices</searchLink><br /><searchLink fieldCode="DE" term="%22Phonetics%22">Phonetics</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+processing+systems%22">Speech processing systems</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+accuracy%22">Statistical accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Hierarchical+Bayes+model%22">Hierarchical Bayes model</searchLink><br /><searchLink fieldCode="DE" term="%22Speech%22">Speech</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Forced alignment is widely used in phonetic research to align transcripts with acoustic signals. Yet there exists a lack of agreement on conventions for evaluating forced alignment, and our understanding of the reliability of forced aligners rests primarily on results from English. This study aims to fill these gaps by examining the concrete issue of forced aligning different Mandarin varieties. It evaluates machine-generated alignments from Montreal Forced Aligner (MFA); McAuliffe, Socolof, Mihuc, Wagner and Sonderegger [Proc. Interspeech 2017, 498–502 (2017a)] against two sets of independent human baselines using a Bayesian hierarchical multivariate regression model. Our findings suggest closer agreement between human aligners than between humans and MFA; large differences in alignment accuracy across different sequence types, with somewhat divergent patterns of errors across humans and MFA; some effects of speech rate and speaker-specific variation; and essentially no variation in robustness across varieties. These results serve (i) to reinforce previous results on the robustness of forced alignment across different varieties of the same language; and (ii) to provide a set of important methodological recommenations for evaluating forced alignment accuracy. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of the Acoustical Society of America is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193402363
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1121/10.0043323
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 17
        StartPage: 3164
    Subjects:
      – SubjectFull: Mandarin dialects
        Type: general
      – SubjectFull: Best practices
        Type: general
      – SubjectFull: Phonetics
        Type: general
      – SubjectFull: Speech processing systems
        Type: general
      – SubjectFull: Statistical accuracy
        Type: general
      – SubjectFull: Hierarchical Bayes model
        Type: general
      – SubjectFull: Speech
        Type: general
    Titles:
      – TitleFull: Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Liu, Suyuan
      – PersonEntity:
          Name:
            NameFull: Sóskuthy, Márton
      – PersonEntity:
          Name:
            NameFull: Zhang, Sijia
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 04
              Text: Apr2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 00014966
          Numbering:
            – Type: volume
              Value: 159
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Journal of the Acoustical Society of America
              Type: main
ResultId 1