The Great Detectives: Humans versus AI Detectors in Catching Large Language Model-Generated Medical Writing

Saved in:
Bibliographic Details
Title: The Great Detectives: Humans versus AI Detectors in Catching Large Language Model-Generated Medical Writing
Language: English
Authors: Jae Q. J. Liu, Kelvin T. K. Hui, Fadi Al Zoubi, Zing Z. X. Zhou, Dino Samartzis, Curtis C. H. Yu, Jeremy R. Chang, Arnold Y. L. Wong (ORCID 0000-0002-5911-5756)
Source: International Journal for Educational Integrity. 2024 20.
Availability: BioMed Central, Ltd. Available from: Springer Nature. 233 Spring Street, New York, NY 10013. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-348-4505; e-mail: customerservice@springernature.com; Web site: https://www.springer.com/gp/biomedical-sciences
Peer Reviewed: Y
Page Count: 14
Publication Date: 2024
Document Type: Journal Articles
Reports - Evaluative
Descriptors: Artificial Intelligence, Investigations, Identification, Human Factors Engineering, Academic Language, Natural Language Processing, Man Machine Systems, Writing (Composition), Ethics, Accuracy, Technology Uses in Education, Difficulty Level, Educational Quality, Writing Evaluation, Evaluation Methods
DOI: 10.1007/s40979-024-00155-6
ISSN: 1833-2595
Abstract: The application of artificial intelligence (AI) in academic writing has raised concerns regarding accuracy, ethics, and scientific rigour. Some AI content detectors may not accurately identify AI-generated texts, especially those that have undergone paraphrasing. Therefore, there is a pressing need for efficacious approaches or guidelines to govern AI usage in specific disciplines. Our study aims to compare the accuracy of mainstream AI content detectors and human reviewers in detecting AI-generated rehabilitation-related articles with or without paraphrasing. This cross-sectional study purposively chose 50 rehabilitation-related articles from four peer-reviewed journals, and then fabricated another 50 articles using ChatGPT. Specifically, ChatGPT was used to generate the introduction, discussion, and conclusion sections based on the original titles, methods, and results. Wordtune was then used to rephrase the ChatGPT-generated articles. Six common AI content detectors (Originality.ai, Turnitin, ZeroGPT, GPTZero, Content at Scale, and GPT-2 Output Detector) were employed to identify AI content for the original, ChatGPT-generated and AI-rephrased articles. Four human reviewers (two student reviewers and two professorial reviewers) were recruited to differentiate between the original articles and AI-rephrased articles, which were expected to be more difficult to detect. They were instructed to give reasons for their judgements.Originality.ai correctly detected 100% of ChatGPT-generated and AI-rephrased texts. ZeroGPT accurately detected 96% of ChatGPT-generated and 88% of AI-rephrased articles. The areas under the receiver operating characteristic curve (AUROC) of ZeroGPT were 0.98 for identifying human-written and AI articles. Turnitin showed a 0% misclassification rate for human-written articles, although it only identified 30% of AI-rephrased articles. Professorial reviewers accurately discriminated at least 96% of AI-rephrased articles, but they misclassified 12% of human-written articles as AI-generated. On average, students only identified 76% of AI-rephrased articles. Reviewers identified AI-rephrased articles based on 'incoherent content' (34.36%), followed by 'grammatical errors' (20.26%), and 'insufficient evidence' (16.15%).This study directly compared the accuracy of advanced AI detectors and human reviewers in detecting AI-generated medical writing after paraphrasing. Our findings demonstrate that specific detectors and experienced reviewers can accurately identify articles generated by Large Language Models, even after paraphrasing. The rationale employed by our reviewers in their assessments can inform future evaluation strategies for monitoring AI usage in medical education or publications. AI content detectors may be incorporated as an additional screening tool in the peer-review process of academic journals.
Abstractor: As Provided
Entry Date: 2024
Accession Number: EJ1424973
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: EJ1424973
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: The Great Detectives: Humans versus AI Detectors in Catching Large Language Model-Generated Medical Writing
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Jae+Q%2E+J%2E+Liu%22">Jae Q. J. Liu</searchLink><br /><searchLink fieldCode="AR" term="%22Kelvin+T%2E+K%2E+Hui%22">Kelvin T. K. Hui</searchLink><br /><searchLink fieldCode="AR" term="%22Fadi+Al+Zoubi%22">Fadi Al Zoubi</searchLink><br /><searchLink fieldCode="AR" term="%22Zing+Z%2E+X%2E+Zhou%22">Zing Z. X. Zhou</searchLink><br /><searchLink fieldCode="AR" term="%22Dino+Samartzis%22">Dino Samartzis</searchLink><br /><searchLink fieldCode="AR" term="%22Curtis+C%2E+H%2E+Yu%22">Curtis C. H. Yu</searchLink><br /><searchLink fieldCode="AR" term="%22Jeremy+R%2E+Chang%22">Jeremy R. Chang</searchLink><br /><searchLink fieldCode="AR" term="%22Arnold+Y%2E+L%2E+Wong%22">Arnold Y. L. Wong</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-5911-5756">0000-0002-5911-5756</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22International+Journal+for+Educational+Integrity%22"><i>International Journal for Educational Integrity</i></searchLink>. 2024 20.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: BioMed Central, Ltd. Available from: Springer Nature. 233 Spring Street, New York, NY 10013. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-348-4505; e-mail: customerservice@springernature.com; Web site: https://www.springer.com/gp/biomedical-sciences
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 14
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Evaluative
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Investigations%22">Investigations</searchLink><br /><searchLink fieldCode="DE" term="%22Identification%22">Identification</searchLink><br /><searchLink fieldCode="DE" term="%22Human+Factors+Engineering%22">Human Factors Engineering</searchLink><br /><searchLink fieldCode="DE" term="%22Academic+Language%22">Academic Language</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Man+Machine+Systems%22">Man Machine Systems</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+%28Composition%29%22">Writing (Composition)</searchLink><br /><searchLink fieldCode="DE" term="%22Ethics%22">Ethics</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Quality%22">Educational Quality</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Evaluation%22">Writing Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1007/s40979-024-00155-6
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1833-2595
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The application of artificial intelligence (AI) in academic writing has raised concerns regarding accuracy, ethics, and scientific rigour. Some AI content detectors may not accurately identify AI-generated texts, especially those that have undergone paraphrasing. Therefore, there is a pressing need for efficacious approaches or guidelines to govern AI usage in specific disciplines. Our study aims to compare the accuracy of mainstream AI content detectors and human reviewers in detecting AI-generated rehabilitation-related articles with or without paraphrasing. This cross-sectional study purposively chose 50 rehabilitation-related articles from four peer-reviewed journals, and then fabricated another 50 articles using ChatGPT. Specifically, ChatGPT was used to generate the introduction, discussion, and conclusion sections based on the original titles, methods, and results. Wordtune was then used to rephrase the ChatGPT-generated articles. Six common AI content detectors (Originality.ai, Turnitin, ZeroGPT, GPTZero, Content at Scale, and GPT-2 Output Detector) were employed to identify AI content for the original, ChatGPT-generated and AI-rephrased articles. Four human reviewers (two student reviewers and two professorial reviewers) were recruited to differentiate between the original articles and AI-rephrased articles, which were expected to be more difficult to detect. They were instructed to give reasons for their judgements.Originality.ai correctly detected 100% of ChatGPT-generated and AI-rephrased texts. ZeroGPT accurately detected 96% of ChatGPT-generated and 88% of AI-rephrased articles. The areas under the receiver operating characteristic curve (AUROC) of ZeroGPT were 0.98 for identifying human-written and AI articles. Turnitin showed a 0% misclassification rate for human-written articles, although it only identified 30% of AI-rephrased articles. Professorial reviewers accurately discriminated at least 96% of AI-rephrased articles, but they misclassified 12% of human-written articles as AI-generated. On average, students only identified 76% of AI-rephrased articles. Reviewers identified AI-rephrased articles based on 'incoherent content' (34.36%), followed by 'grammatical errors' (20.26%), and 'insufficient evidence' (16.15%).This study directly compared the accuracy of advanced AI detectors and human reviewers in detecting AI-generated medical writing after paraphrasing. Our findings demonstrate that specific detectors and experienced reviewers can accurately identify articles generated by Large Language Models, even after paraphrasing. The rationale employed by our reviewers in their assessments can inform future evaluation strategies for monitoring AI usage in medical education or publications. AI content detectors may be incorporated as an additional screening tool in the peer-review process of academic journals.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1424973
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1424973
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s40979-024-00155-6
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 14
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Investigations
        Type: general
      – SubjectFull: Identification
        Type: general
      – SubjectFull: Human Factors Engineering
        Type: general
      – SubjectFull: Academic Language
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Man Machine Systems
        Type: general
      – SubjectFull: Writing (Composition)
        Type: general
      – SubjectFull: Ethics
        Type: general
      – SubjectFull: Accuracy
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Difficulty Level
        Type: general
      – SubjectFull: Educational Quality
        Type: general
      – SubjectFull: Writing Evaluation
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
    Titles:
      – TitleFull: The Great Detectives: Humans versus AI Detectors in Catching Large Language Model-Generated Medical Writing
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Jae Q. J. Liu
      – PersonEntity:
          Name:
            NameFull: Kelvin T. K. Hui
      – PersonEntity:
          Name:
            NameFull: Fadi Al Zoubi
      – PersonEntity:
          Name:
            NameFull: Zing Z. X. Zhou
      – PersonEntity:
          Name:
            NameFull: Dino Samartzis
      – PersonEntity:
          Name:
            NameFull: Curtis C. H. Yu
      – PersonEntity:
          Name:
            NameFull: Jeremy R. Chang
      – PersonEntity:
          Name:
            NameFull: Arnold Y. L. Wong
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-electronic
              Value: 1833-2595
          Numbering:
            – Type: volume
              Value: 20
          Titles:
            – TitleFull: International Journal for Educational Integrity
              Type: main
ResultId 1