Automated Paraphrase Quality Assessment Using Recurrent Neural Networks and Language Models

Saved in:
Bibliographic Details
Title: Automated Paraphrase Quality Assessment Using Recurrent Neural Networks and Language Models
Language: English
Authors: Nicula, Bogdan, Dascalu, Mihai, Newton, Natalie, Orcutt, Ellen, McNamara, Danielle S.
Source: Grantee Submission. 2021.
Peer Reviewed: Y
Page Count: 9
Publication Date: 2021
Sponsoring Agency: Institute of Education Sciences (ED)
Office of Naval Research (ONR) (DOD)
Contract Number: R305A190063
R305A190050
N000141712300
N000141912424
Document Type: Speeches/Meeting Papers
Reports - Research
Education Level: Elementary Education
Descriptors: Phrase Structure, Networks, Semantics, Feedback (Response), Syntax, Computational Linguistics, Language Usage, Models, Teaching Methods, Classification, Artificial Intelligence, Linguistic Input, Intelligent Tutoring Systems, Natural Language Processing, Literacy Education, Elementary School Students
DOI: 10.1007/978-3-030-80421-3_36
Abstract: The ability to automatically assess the quality of paraphrases can be very useful for facilitating literacy skills and providing timely feedback to learners. Our aim is twofold: a) to automatically evaluate the quality of paraphrases across four dimensions: lexical similarity, syntactic similarity, semantic similarity and paraphrase quality, and b) to assess how well models trained for this task generalize. The task is modeled as a classification problem and three different methods are explored: (a) manual feature extraction combined with an Extra Trees model, (b) GloVe embeddings and a Siamese neural network, and (c) using a pre-trained BERT model fine-tuned on our task. Starting from a dataset of 1998 paraphrases from the User Language Paraphrase Corpus (ULPC), we explore how the three models trained on the ULPC dataset generalize when applied on a separate, small paraphrase corpus based on children inputs. The best out-of-the-box generalization performance is obtained by the Extra Trees model with at least 75% average F1-scores for the three similarity dimensions. We also show that the Siamese neural network and BERT models can obtain an improvement of at least 5% after fine-tuning across all dimensions.
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2023
Accession Number: ED628430
Database: ERIC
Be the first to leave a comment!
You must be logged in first