Unlocking the power of transfer learning with Ad-Dabit-Al-Lughawi: A token classification approach for enhanced Arabic Text Diacritization.
Saved in:
| Title: | Unlocking the power of transfer learning with Ad-Dabit-Al-Lughawi: A token classification approach for enhanced Arabic Text Diacritization. |
|---|---|
| Authors: | Skiredj, Abderrahman1,2 (AUTHOR) abderrahman.skiredj@um6p.ma, Berrada, Ismail2 (AUTHOR) ismail.berrada@um6p.ma |
| Source: | Expert Systems with Applications. Apr2025, Vol. 269, pN.PAG-N.PAG. 1p. |
| Subjects: | Language models, Machine translating, Linguistic analysis, Data reduction, Error rates |
| Abstract: | Automatic diacritization of Arabic text is a pivotal process that involves the addition of diacritical marks to enhance the clarity and comprehension of the text. This task is crucial for several reasons: it resolves ambiguity in undiacritized text by providing necessary vowel indications, thereby improving readability and understanding. Additionally, diacritization significantly benefits various Natural Language Processing (NLP) applications, including text-to-speech systems, machine translation, and more, by enhancing the accuracy of linguistic analysis and processing. Recognizing these impacts, this paper introduces Ad-dabit-Al-Lughawi 1 1 In Arabic, Ad-dabt refers to the process of adding diacritics to Arabic texts, implying the function of a diacritizer. The prefix "ad" evokes the acronym "AD," standing for "Arabic Diacritization," while Al-Lughawi pertains to the adherence to linguistic principles. , or Ad-dabit for brevity, a pioneering two-phase approach for the Arabic Text Diacritization (ATD) task. Initially, Ad-dabit undergoes simultaneous pre-finetuning on linguistically relevant tasks such as finetuning on CA texts, POS tagging, segmentation, and text diacritization, all framed as Masked Language Modeling (MLM) tasks. This enriches the model's contextual understanding by integrating knowledge from these tasks. Following this, Ad-dabit progresses into a finetuning phase, where ATD is treated as a token classification task. The effectiveness of Ad-dabit is demonstrated through evaluations on two benchmark datasets derived from the Tashkeela dataset, where it achieves state-of-the-art results, including a 30% reduction in Word Error Rate (WER) compared to the previous state-of-the-art (SOTA) models on these benchmarks. The enhancements in WER achieved by Ad-dabit not only set new standards in ATD but also enhance the performance of related NLP applications, contributing to more accurate and efficient processing of Arabic digital communications. [Display omitted] • Introducing "Ad-dabit" for improved Arabic Text Diacritization. • Surpasses previous SOTA with up to 61% DER and 52% WER reduction on augmented data. • Ablation confirms pre-finetuning and multi-task learning boost performance. • Utilizes BERT models for efficiency in diverse computing environments. • Ad-dabit's open-source materials foster community innovation. [ABSTRACT FROM AUTHOR] |
| Copyright of Expert Systems with Applications is the property of Pergamon Press - An Imprint of Elsevier Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
Be the first to leave a comment!