Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling.

Saved in:
Bibliographic Details
Title: Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling.
Authors: Poncelet, Jakob1 (AUTHOR) jakob.poncelet@kuleuven.be, Van hamme, Hugo1 (AUTHOR) hugo.vanhamme@kuleuven.be
Source: EURASIP Journal on Audio Speech & Music Processing. 3/4/2026, Vol. 2026 Issue 1, p1-20. 20p.
Subjects: Automatic speech recognition, Closed captioning, Broadcasting industry, Dutch language, Machine learning, Low-resource languages, Supervised learning
Abstract: The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for low-resource languages and dialects. This paper explores the integration of weakly supervised transcripts from TV subtitles into automatic speech recognition (ASR) systems, aiming to improve both verbatim transcriptions and automatically generated subtitles. To this end, verbatim data and subtitles are regarded as different domains or languages, due to their distinct characteristics. We propose and compare several end-to-end architectures that are designed to jointly model both modalities with separate or shared encoders and decoders. The proposed methods are able to jointly generate a verbatim transcription and a subtitle. Evaluation on Flemish (Belgian Dutch) demonstrates that a model with cascaded encoders and separate decoders allows to represent the differences between the two data types most efficiently while improving on both domains. Despite differences in domain and linguistic variations, combining verbatim transcripts with subtitle data leads to notable ASR improvements. The proposed methods circumvent the need for extensive preprocessing, filtering and cleaning of the subtitled data, which is commonly used to leverage such data for ASR training. Additionally, experiments with a large-scale subtitle dataset show the scalability of the proposed approach. The methods not only improve ASR accuracy but also generate subtitles that closely match standard written text, offering several potential applications. [ABSTRACT FROM AUTHOR]
Copyright of EURASIP Journal on Audio Speech & Music Processing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 193005327
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Poncelet%2C+Jakob%22">Poncelet, Jakob</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> jakob.poncelet@kuleuven.be</i><br /><searchLink fieldCode="AR" term="%22Van+hamme%2C+Hugo%22">Van hamme, Hugo</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> hugo.vanhamme@kuleuven.be</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22EURASIP+Journal+on+Audio+Speech+%26+Music+Processing%22">EURASIP Journal on Audio Speech & Music Processing</searchLink>. 3/4/2026, Vol. 2026 Issue 1, p1-20. 20p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Closed+captioning%22">Closed captioning</searchLink><br /><searchLink fieldCode="DE" term="%22Broadcasting+industry%22">Broadcasting industry</searchLink><br /><searchLink fieldCode="DE" term="%22Dutch+language%22">Dutch language</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Low-resource+languages%22">Low-resource languages</searchLink><br /><searchLink fieldCode="DE" term="%22Supervised+learning%22">Supervised learning</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for low-resource languages and dialects. This paper explores the integration of weakly supervised transcripts from TV subtitles into automatic speech recognition (ASR) systems, aiming to improve both verbatim transcriptions and automatically generated subtitles. To this end, verbatim data and subtitles are regarded as different domains or languages, due to their distinct characteristics. We propose and compare several end-to-end architectures that are designed to jointly model both modalities with separate or shared encoders and decoders. The proposed methods are able to jointly generate a verbatim transcription and a subtitle. Evaluation on Flemish (Belgian Dutch) demonstrates that a model with cascaded encoders and separate decoders allows to represent the differences between the two data types most efficiently while improving on both domains. Despite differences in domain and linguistic variations, combining verbatim transcripts with subtitle data leads to notable ASR improvements. The proposed methods circumvent the need for extensive preprocessing, filtering and cleaning of the subtitled data, which is commonly used to leverage such data for ASR training. Additionally, experiments with a large-scale subtitle dataset show the scalability of the proposed approach. The methods not only improve ASR accuracy but also generate subtitles that closely match standard written text, offering several potential applications. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of EURASIP Journal on Audio Speech & Music Processing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193005327
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1186/s13636-026-00450-9
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 1
    Subjects:
      – SubjectFull: Automatic speech recognition
        Type: general
      – SubjectFull: Closed captioning
        Type: general
      – SubjectFull: Broadcasting industry
        Type: general
      – SubjectFull: Dutch language
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Low-resource languages
        Type: general
      – SubjectFull: Supervised learning
        Type: general
    Titles:
      – TitleFull: Leveraging broadcast media subtitle transcripts for automatic speech recognition and subtitling.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Poncelet, Jakob
      – PersonEntity:
          Name:
            NameFull: Van hamme, Hugo
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 04
              M: 03
              Text: 3/4/2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 16874714
          Numbering:
            – Type: volume
              Value: 2026
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: EURASIP Journal on Audio Speech & Music Processing
              Type: main
ResultId 1