Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties.
Saved in:
| Title: | Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties. |
|---|---|
| Authors: | Liu, Suyuan1 (AUTHOR) suyuan97@student.ubc.ca, Sóskuthy, Márton1 (AUTHOR), Zhang, Sijia1 (AUTHOR) |
| Source: | Journal of the Acoustical Society of America. Apr2026, Vol. 159 Issue 4, p3164-3180. 17p. |
| Subjects: | Mandarin dialects, Best practices, Phonetics, Speech processing systems, Statistical accuracy, Hierarchical Bayes model, Speech |
| Abstract: | Forced alignment is widely used in phonetic research to align transcripts with acoustic signals. Yet there exists a lack of agreement on conventions for evaluating forced alignment, and our understanding of the reliability of forced aligners rests primarily on results from English. This study aims to fill these gaps by examining the concrete issue of forced aligning different Mandarin varieties. It evaluates machine-generated alignments from Montreal Forced Aligner (MFA); McAuliffe, Socolof, Mihuc, Wagner and Sonderegger [Proc. Interspeech 2017, 498–502 (2017a)] against two sets of independent human baselines using a Bayesian hierarchical multivariate regression model. Our findings suggest closer agreement between human aligners than between humans and MFA; large differences in alignment accuracy across different sequence types, with somewhat divergent patterns of errors across humans and MFA; some effects of speech rate and speaker-specific variation; and essentially no variation in robustness across varieties. These results serve (i) to reinforce previous results on the robustness of forced alignment across different varieties of the same language; and (ii) to provide a set of important methodological recommenations for evaluating forced alignment accuracy. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of the Acoustical Society of America is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 193402363 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Liu%2C+Suyuan%22">Liu, Suyuan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> suyuan97@student.ubc.ca</i><br /><searchLink fieldCode="AR" term="%22Sóskuthy%2C+Márton%22">Sóskuthy, Márton</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhang%2C+Sijia%22">Zhang, Sijia</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+the+Acoustical+Society+of+America%22">Journal of the Acoustical Society of America</searchLink>. Apr2026, Vol. 159 Issue 4, p3164-3180. 17p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Mandarin+dialects%22">Mandarin dialects</searchLink><br /><searchLink fieldCode="DE" term="%22Best+practices%22">Best practices</searchLink><br /><searchLink fieldCode="DE" term="%22Phonetics%22">Phonetics</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+processing+systems%22">Speech processing systems</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+accuracy%22">Statistical accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Hierarchical+Bayes+model%22">Hierarchical Bayes model</searchLink><br /><searchLink fieldCode="DE" term="%22Speech%22">Speech</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Forced alignment is widely used in phonetic research to align transcripts with acoustic signals. Yet there exists a lack of agreement on conventions for evaluating forced alignment, and our understanding of the reliability of forced aligners rests primarily on results from English. This study aims to fill these gaps by examining the concrete issue of forced aligning different Mandarin varieties. It evaluates machine-generated alignments from Montreal Forced Aligner (MFA); McAuliffe, Socolof, Mihuc, Wagner and Sonderegger [Proc. Interspeech 2017, 498–502 (2017a)] against two sets of independent human baselines using a Bayesian hierarchical multivariate regression model. Our findings suggest closer agreement between human aligners than between humans and MFA; large differences in alignment accuracy across different sequence types, with somewhat divergent patterns of errors across humans and MFA; some effects of speech rate and speaker-specific variation; and essentially no variation in robustness across varieties. These results serve (i) to reinforce previous results on the robustness of forced alignment across different varieties of the same language; and (ii) to provide a set of important methodological recommenations for evaluating forced alignment accuracy. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of the Acoustical Society of America is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193402363 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1121/10.0043323 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 17 StartPage: 3164 Subjects: – SubjectFull: Mandarin dialects Type: general – SubjectFull: Best practices Type: general – SubjectFull: Phonetics Type: general – SubjectFull: Speech processing systems Type: general – SubjectFull: Statistical accuracy Type: general – SubjectFull: Hierarchical Bayes model Type: general – SubjectFull: Speech Type: general Titles: – TitleFull: Best practices in evaluating forced-alignment accuracy: The case of Mandarin varieties. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Liu, Suyuan – PersonEntity: Name: NameFull: Sóskuthy, Márton – PersonEntity: Name: NameFull: Zhang, Sijia IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 04 Text: Apr2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 00014966 Numbering: – Type: volume Value: 159 – Type: issue Value: 4 Titles: – TitleFull: Journal of the Acoustical Society of America Type: main |
| ResultId | 1 |