Urdu paraphrased text reuse and plagiarism detection using pre-trained large language models and deep hybrid neural networks.

Saved in:
Bibliographic Details
Title: Urdu paraphrased text reuse and plagiarism detection using pre-trained large language models and deep hybrid neural networks.
Authors: Rizwan Iqbal, Hafiz1 (AUTHOR) rizwan.iqbal@itu.edu.pk, Sharjeel, Muhammad2 (AUTHOR) muhammadsharjeel@cuilahore.edu.pk, Shafi, Jawad2 (AUTHOR) jawadshafi@cuilahore.edu.pk, Mehmood, Usama1 (AUTHOR) usamamehmood@itu.edu.pk, Hassan, Saeed Ul1 (AUTHOR) saeed-ul-hassan@itu.edu.pk, Raza, Agha Ali3 (AUTHOR) agha.ali.raza@lums.edu.pk
Source: Multimedia Tools & Applications. Oct2025, Vol. 84 Issue 35, p43475-43497. 23p.
Subjects: Urdu language, Paraphrase, Long short-term memory, Plagiarism prevention, Language models, Artificial neural networks
Abstract: The growing prevalence of text reuse and plagiarism in various fields has led to an urgent need for reliable computational methods for detection. However, current commercial plagiarism detection systems are ineffective in identifying paraphrased cases of text reuse, highlighting the need for improvement. Previous research on paraphrased text reuse and plagiarism detection has mainly focused on English, European, Persian, and Arabic languages, and very few studies have been reported on the under-resourced Urdu language. This study aims to overcome this research gap by using a Deep Neural Network (DNN) based architecture and pre-trained Large Language Models (LLMs) for the task of Urdu paraphrased text reuse and plagiarism detection. The architecture called Deep Text Reuse and Paraphrased Plagiarism Detection (D-TRaPPD), relies on LLMs for input and utilizes CNN and LSTM to extract essential textual features. Moreover, we have proposed and evaluated two D-TRaPPD variants, Word Embeddings-D-TRaPPD (WE-D-TRaPPD) and Sentence Embeddings-D-TRaPPD (SE-D-TRaPPD), using two gold standard document-level corpora containing both real and simulated cases of Urdu paraphrased text reuse and plagiarism. The results demonstrate the effectiveness of the D-TRaPPD architecture, with SE-D-TRaPPD achieving the highest F 1 scores of 91.77 for real cases and 95.15 for simulated cases. Furthermore, the results highlight the superiority of our approaches over the state-of-the-art methods for Urdu paraphrased text reuse and plagiarism detection. [ABSTRACT FROM AUTHOR]
Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 188906123
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Urdu paraphrased text reuse and plagiarism detection using pre-trained large language models and deep hybrid neural networks.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Rizwan+Iqbal%2C+Hafiz%22">Rizwan Iqbal, Hafiz</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> rizwan.iqbal@itu.edu.pk</i><br /><searchLink fieldCode="AR" term="%22Sharjeel%2C+Muhammad%22">Sharjeel, Muhammad</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> muhammadsharjeel@cuilahore.edu.pk</i><br /><searchLink fieldCode="AR" term="%22Shafi%2C+Jawad%22">Shafi, Jawad</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> jawadshafi@cuilahore.edu.pk</i><br /><searchLink fieldCode="AR" term="%22Mehmood%2C+Usama%22">Mehmood, Usama</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> usamamehmood@itu.edu.pk</i><br /><searchLink fieldCode="AR" term="%22Hassan%2C+Saeed+Ul%22">Hassan, Saeed Ul</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> saeed-ul-hassan@itu.edu.pk</i><br /><searchLink fieldCode="AR" term="%22Raza%2C+Agha+Ali%22">Raza, Agha Ali</searchLink><relatesTo>3</relatesTo> (AUTHOR)<i> agha.ali.raza@lums.edu.pk</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Multimedia+Tools+%26+Applications%22">Multimedia Tools & Applications</searchLink>. Oct2025, Vol. 84 Issue 35, p43475-43497. 23p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Urdu+language%22">Urdu language</searchLink><br /><searchLink fieldCode="DE" term="%22Paraphrase%22">Paraphrase</searchLink><br /><searchLink fieldCode="DE" term="%22Long+short-term+memory%22">Long short-term memory</searchLink><br /><searchLink fieldCode="DE" term="%22Plagiarism+prevention%22">Plagiarism prevention</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The growing prevalence of text reuse and plagiarism in various fields has led to an urgent need for reliable computational methods for detection. However, current commercial plagiarism detection systems are ineffective in identifying paraphrased cases of text reuse, highlighting the need for improvement. Previous research on paraphrased text reuse and plagiarism detection has mainly focused on English, European, Persian, and Arabic languages, and very few studies have been reported on the under-resourced Urdu language. This study aims to overcome this research gap by using a Deep Neural Network (DNN) based architecture and pre-trained Large Language Models (LLMs) for the task of Urdu paraphrased text reuse and plagiarism detection. The architecture called Deep Text Reuse and Paraphrased Plagiarism Detection (D-TRaPPD), relies on LLMs for input and utilizes CNN and LSTM to extract essential textual features. Moreover, we have proposed and evaluated two D-TRaPPD variants, Word Embeddings-D-TRaPPD (WE-D-TRaPPD) and Sentence Embeddings-D-TRaPPD (SE-D-TRaPPD), using two gold standard document-level corpora containing both real and simulated cases of Urdu paraphrased text reuse and plagiarism. The results demonstrate the effectiveness of the D-TRaPPD architecture, with SE-D-TRaPPD achieving the highest F 1 scores of 91.77 for real cases and 95.15 for simulated cases. Furthermore, the results highlight the superiority of our approaches over the state-of-the-art methods for Urdu paraphrased text reuse and plagiarism detection. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=188906123
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s11042-025-20862-7
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 23
        StartPage: 43475
    Subjects:
      – SubjectFull: Urdu language
        Type: general
      – SubjectFull: Paraphrase
        Type: general
      – SubjectFull: Long short-term memory
        Type: general
      – SubjectFull: Plagiarism prevention
        Type: general
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Artificial neural networks
        Type: general
    Titles:
      – TitleFull: Urdu paraphrased text reuse and plagiarism detection using pre-trained large language models and deep hybrid neural networks.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Rizwan Iqbal, Hafiz
      – PersonEntity:
          Name:
            NameFull: Sharjeel, Muhammad
      – PersonEntity:
          Name:
            NameFull: Shafi, Jawad
      – PersonEntity:
          Name:
            NameFull: Mehmood, Usama
      – PersonEntity:
          Name:
            NameFull: Hassan, Saeed Ul
      – PersonEntity:
          Name:
            NameFull: Raza, Agha Ali
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 28
              M: 10
              Text: Oct2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 13807501
          Numbering:
            – Type: volume
              Value: 84
            – Type: issue
              Value: 35
          Titles:
            – TitleFull: Multimedia Tools & Applications
              Type: main
ResultId 1