Unlocking the power of transfer learning with Ad-Dabit-Al-Lughawi: A token classification approach for enhanced Arabic Text Diacritization.

Saved in:
Bibliographic Details
Title: Unlocking the power of transfer learning with Ad-Dabit-Al-Lughawi: A token classification approach for enhanced Arabic Text Diacritization.
Authors: Skiredj, Abderrahman1,2 (AUTHOR) abderrahman.skiredj@um6p.ma, Berrada, Ismail2 (AUTHOR) ismail.berrada@um6p.ma
Source: Expert Systems with Applications. Apr2025, Vol. 269, pN.PAG-N.PAG. 1p.
Subjects: Language models, Machine translating, Linguistic analysis, Data reduction, Error rates
Abstract: Automatic diacritization of Arabic text is a pivotal process that involves the addition of diacritical marks to enhance the clarity and comprehension of the text. This task is crucial for several reasons: it resolves ambiguity in undiacritized text by providing necessary vowel indications, thereby improving readability and understanding. Additionally, diacritization significantly benefits various Natural Language Processing (NLP) applications, including text-to-speech systems, machine translation, and more, by enhancing the accuracy of linguistic analysis and processing. Recognizing these impacts, this paper introduces Ad-dabit-Al-Lughawi 1 1 In Arabic, Ad-dabt refers to the process of adding diacritics to Arabic texts, implying the function of a diacritizer. The prefix "ad" evokes the acronym "AD," standing for "Arabic Diacritization," while Al-Lughawi pertains to the adherence to linguistic principles. , or Ad-dabit for brevity, a pioneering two-phase approach for the Arabic Text Diacritization (ATD) task. Initially, Ad-dabit undergoes simultaneous pre-finetuning on linguistically relevant tasks such as finetuning on CA texts, POS tagging, segmentation, and text diacritization, all framed as Masked Language Modeling (MLM) tasks. This enriches the model's contextual understanding by integrating knowledge from these tasks. Following this, Ad-dabit progresses into a finetuning phase, where ATD is treated as a token classification task. The effectiveness of Ad-dabit is demonstrated through evaluations on two benchmark datasets derived from the Tashkeela dataset, where it achieves state-of-the-art results, including a 30% reduction in Word Error Rate (WER) compared to the previous state-of-the-art (SOTA) models on these benchmarks. The enhancements in WER achieved by Ad-dabit not only set new standards in ATD but also enhance the performance of related NLP applications, contributing to more accurate and efficient processing of Arabic digital communications. [Display omitted] • Introducing "Ad-dabit" for improved Arabic Text Diacritization. • Surpasses previous SOTA with up to 61% DER and 52% WER reduction on augmented data. • Ablation confirms pre-finetuning and multi-task learning boost performance. • Utilizes BERT models for efficiency in diverse computing environments. • Ad-dabit's open-source materials foster community innovation. [ABSTRACT FROM AUTHOR]
Copyright of Expert Systems with Applications is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 182873297
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Unlocking the power of transfer learning with Ad-Dabit-Al-Lughawi: A token classification approach for enhanced Arabic Text Diacritization.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Skiredj%2C+Abderrahman%22">Skiredj, Abderrahman</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> abderrahman.skiredj@um6p.ma</i><br /><searchLink fieldCode="AR" term="%22Berrada%2C+Ismail%22">Berrada, Ismail</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> ismail.berrada@um6p.ma</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Expert+Systems+with+Applications%22">Expert Systems with Applications</searchLink>. Apr2025, Vol. 269, pN.PAG-N.PAG. 1p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+translating%22">Machine translating</searchLink><br /><searchLink fieldCode="DE" term="%22Linguistic+analysis%22">Linguistic analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Data+reduction%22">Data reduction</searchLink><br /><searchLink fieldCode="DE" term="%22Error+rates%22">Error rates</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Automatic diacritization of Arabic text is a pivotal process that involves the addition of diacritical marks to enhance the clarity and comprehension of the text. This task is crucial for several reasons: it resolves ambiguity in undiacritized text by providing necessary vowel indications, thereby improving readability and understanding. Additionally, diacritization significantly benefits various Natural Language Processing (NLP) applications, including text-to-speech systems, machine translation, and more, by enhancing the accuracy of linguistic analysis and processing. Recognizing these impacts, this paper introduces Ad-dabit-Al-Lughawi 1 1 In Arabic, Ad-dabt refers to the process of adding diacritics to Arabic texts, implying the function of a diacritizer. The prefix "ad" evokes the acronym "AD," standing for "Arabic Diacritization," while Al-Lughawi pertains to the adherence to linguistic principles. , or Ad-dabit for brevity, a pioneering two-phase approach for the Arabic Text Diacritization (ATD) task. Initially, Ad-dabit undergoes simultaneous pre-finetuning on linguistically relevant tasks such as finetuning on CA texts, POS tagging, segmentation, and text diacritization, all framed as Masked Language Modeling (MLM) tasks. This enriches the model's contextual understanding by integrating knowledge from these tasks. Following this, Ad-dabit progresses into a finetuning phase, where ATD is treated as a token classification task. The effectiveness of Ad-dabit is demonstrated through evaluations on two benchmark datasets derived from the Tashkeela dataset, where it achieves state-of-the-art results, including a 30% reduction in Word Error Rate (WER) compared to the previous state-of-the-art (SOTA) models on these benchmarks. The enhancements in WER achieved by Ad-dabit not only set new standards in ATD but also enhance the performance of related NLP applications, contributing to more accurate and efficient processing of Arabic digital communications. [Display omitted] • Introducing "Ad-dabit" for improved Arabic Text Diacritization. • Surpasses previous SOTA with up to 61% DER and 52% WER reduction on augmented data. • Ablation confirms pre-finetuning and multi-task learning boost performance. • Utilizes BERT models for efficiency in diverse computing environments. • Ad-dabit's open-source materials foster community innovation. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Expert Systems with Applications is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=182873297
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.eswa.2024.126166
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1
        StartPage: N.PAG
    Subjects:
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Machine translating
        Type: general
      – SubjectFull: Linguistic analysis
        Type: general
      – SubjectFull: Data reduction
        Type: general
      – SubjectFull: Error rates
        Type: general
    Titles:
      – TitleFull: Unlocking the power of transfer learning with Ad-Dabit-Al-Lughawi: A token classification approach for enhanced Arabic Text Diacritization.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Skiredj, Abderrahman
      – PersonEntity:
          Name:
            NameFull: Berrada, Ismail
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 15
              M: 04
              Text: Apr2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 09574174
          Numbering:
            – Type: volume
              Value: 269
          Titles:
            – TitleFull: Expert Systems with Applications
              Type: main
ResultId 1