Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling.
Saved in:
| Title: | Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling. |
|---|---|
| Authors: | Poncelet, Jakob1 (AUTHOR) jakob.poncelet@kuleuven.be, Van hamme, Hugo1 (AUTHOR) hugo.vanhamme@kuleuven.be |
| Source: | EURASIP Journal on Audio Speech & Music Processing. 3/4/2026, Vol. 2026 Issue 1, p1-20. 20p. |
| Subjects: | Automatic speech recognition, Closed captioning, Broadcasting industry, Dutch language, Machine learning, Low-resource languages, Supervised learning |
| Abstract: | The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for low-resource languages and dialects. This paper explores the integration of weakly supervised transcripts from TV subtitles into automatic speech recognition (ASR) systems, aiming to improve both verbatim transcriptions and automatically generated subtitles. To this end, verbatim data and subtitles are regarded as different domains or languages, due to their distinct characteristics. We propose and compare several end-to-end architectures that are designed to jointly model both modalities with separate or shared encoders and decoders. The proposed methods are able to jointly generate a verbatim transcription and a subtitle. Evaluation on Flemish (Belgian Dutch) demonstrates that a model with cascaded encoders and separate decoders allows to represent the differences between the two data types most efficiently while improving on both domains. Despite differences in domain and linguistic variations, combining verbatim transcripts with subtitle data leads to notable ASR improvements. The proposed methods circumvent the need for extensive preprocessing, filtering and cleaning of the subtitled data, which is commonly used to leverage such data for ASR training. Additionally, experiments with a large-scale subtitle dataset show the scalability of the proposed approach. The methods not only improve ASR accuracy but also generate subtitles that closely match standard written text, offering several potential applications. [ABSTRACT FROM AUTHOR] |
| Copyright of EURASIP Journal on Audio Speech & Music Processing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 193005327 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Poncelet%2C+Jakob%22">Poncelet, Jakob</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> jakob.poncelet@kuleuven.be</i><br /><searchLink fieldCode="AR" term="%22Van+hamme%2C+Hugo%22">Van hamme, Hugo</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> hugo.vanhamme@kuleuven.be</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22EURASIP+Journal+on+Audio+Speech+%26+Music+Processing%22">EURASIP Journal on Audio Speech & Music Processing</searchLink>. 3/4/2026, Vol. 2026 Issue 1, p1-20. 20p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Closed+captioning%22">Closed captioning</searchLink><br /><searchLink fieldCode="DE" term="%22Broadcasting+industry%22">Broadcasting industry</searchLink><br /><searchLink fieldCode="DE" term="%22Dutch+language%22">Dutch language</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Low-resource+languages%22">Low-resource languages</searchLink><br /><searchLink fieldCode="DE" term="%22Supervised+learning%22">Supervised learning</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for low-resource languages and dialects. This paper explores the integration of weakly supervised transcripts from TV subtitles into automatic speech recognition (ASR) systems, aiming to improve both verbatim transcriptions and automatically generated subtitles. To this end, verbatim data and subtitles are regarded as different domains or languages, due to their distinct characteristics. We propose and compare several end-to-end architectures that are designed to jointly model both modalities with separate or shared encoders and decoders. The proposed methods are able to jointly generate a verbatim transcription and a subtitle. Evaluation on Flemish (Belgian Dutch) demonstrates that a model with cascaded encoders and separate decoders allows to represent the differences between the two data types most efficiently while improving on both domains. Despite differences in domain and linguistic variations, combining verbatim transcripts with subtitle data leads to notable ASR improvements. The proposed methods circumvent the need for extensive preprocessing, filtering and cleaning of the subtitled data, which is commonly used to leverage such data for ASR training. Additionally, experiments with a large-scale subtitle dataset show the scalability of the proposed approach. The methods not only improve ASR accuracy but also generate subtitles that closely match standard written text, offering several potential applications. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of EURASIP Journal on Audio Speech & Music Processing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193005327 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1186/s13636-026-00450-9 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 20 StartPage: 1 Subjects: – SubjectFull: Automatic speech recognition Type: general – SubjectFull: Closed captioning Type: general – SubjectFull: Broadcasting industry Type: general – SubjectFull: Dutch language Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Low-resource languages Type: general – SubjectFull: Supervised learning Type: general Titles: – TitleFull: Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Poncelet, Jakob – PersonEntity: Name: NameFull: Van hamme, Hugo IsPartOfRelationships: – BibEntity: Dates: – D: 04 M: 03 Text: 3/4/2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 16874714 Numbering: – Type: volume Value: 2026 – Type: issue Value: 1 Titles: – TitleFull: EURASIP Journal on Audio Speech & Music Processing Type: main |
| ResultId | 1 |