Can ChatGPT Translate Like a Pro? A Pilot Benchmarking Study of English-Malay Translation Quality.

Saved in:
Bibliographic Details
Title: Can ChatGPT Translate Like a Pro? A Pilot Benchmarking Study of English-Malay Translation Quality.
Authors: SULAIMAN, M. ZAIN1 zain@ukm.edu.my, ZAINUDIN, INTAN SAFINAZ1, HAROON, HASLINA2
Source: 3L: Southeast Asian Journal of English Language Studies. Dec2025, Vol. 31 Issue 4, p259-278. 20p.
Subject Terms: *Translating & interpreting, *Certification, ChatGPT, Professional standards, Machine translating, Malay language
Abstract: Artificial intelligence (AI) tools such as ChatGPT have significantly advanced machine translation, yet their performance in low-resource language pairs, particularly English-Malay, lags behind. While existing studies have compared AI and human translation quality, most have relied on academic assessment frameworks, leaving a gap in evaluating AI translation through professional certification standards. From a professional standpoint, translation competence is most reliably assessed through formal certification frameworks that combine analytic rubrics, performance descriptors, and expert judgment. To determine whether AI systems can perform at a professional standard, they must be evaluated using the same criteria applied to human translators. This pilot study addresses that gap by benchmarking ChatGPT's English-Malay translation performance against a novice and a professional translator using the National Accreditation Authority for Translators and Interpreters (NAATI) Certified Translator examination framework. Thirteen professional raters from the Malaysian Translators Association assessed the translations based on Meaning Transfer, Textual Norms and Conventions, and Language Proficiency. Findings revealed a clear performance hierarchy--Professional Translator > ChatGPT > Novice Translator--indicating that while ChatGPT achieved near-professional competence in fluency and meaning accuracy, it remained limited in idiomatic precision and cultural adaptation. The study highlights ChatGPT's potential as an assistive tool for translation and training, while reaffirming the need for human oversight. It also validates the NAATI framework as a robust benchmark for evaluating AI translation quality. As AI models continue to evolve, future research involving larger translator samples and a wider range of language pairs is essential to evaluate ongoing progress and ensure the responsible integration of AI translation into professional practice. [ABSTRACT FROM AUTHOR]
Copyright of 3L: Southeast Asian Journal of English Language Studies is the property of 3L: Language, Linguistics, Literature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Description
Abstract:Artificial intelligence (AI) tools such as ChatGPT have significantly advanced machine translation, yet their performance in low-resource language pairs, particularly English-Malay, lags behind. While existing studies have compared AI and human translation quality, most have relied on academic assessment frameworks, leaving a gap in evaluating AI translation through professional certification standards. From a professional standpoint, translation competence is most reliably assessed through formal certification frameworks that combine analytic rubrics, performance descriptors, and expert judgment. To determine whether AI systems can perform at a professional standard, they must be evaluated using the same criteria applied to human translators. This pilot study addresses that gap by benchmarking ChatGPT's English-Malay translation performance against a novice and a professional translator using the National Accreditation Authority for Translators and Interpreters (NAATI) Certified Translator examination framework. Thirteen professional raters from the Malaysian Translators Association assessed the translations based on Meaning Transfer, Textual Norms and Conventions, and Language Proficiency. Findings revealed a clear performance hierarchy--Professional Translator > ChatGPT > Novice Translator--indicating that while ChatGPT achieved near-professional competence in fluency and meaning accuracy, it remained limited in idiomatic precision and cultural adaptation. The study highlights ChatGPT's potential as an assistive tool for translation and training, while reaffirming the need for human oversight. It also validates the NAATI framework as a robust benchmark for evaluating AI translation quality. As AI models continue to evolve, future research involving larger translator samples and a wider range of language pairs is essential to evaluate ongoing progress and ensure the responsible integration of AI translation into professional practice. [ABSTRACT FROM AUTHOR]
ISSN:01285157
DOI:10.17576/3L-2025-3104-17