Age of Exposure 2.0: Estimating Word Complexity Using Iterative Models of Word Embeddings

Saved in:
Bibliographic Details
Title: Age of Exposure 2.0: Estimating Word Complexity Using Iterative Models of Word Embeddings
Language: English
Authors: Botarleanu, Robert-Mihai, Dascalu, Mihai (ORCID 0000-0002-4815-9227), Watanabe, Micah, Crossley, Scott Andrew, McNamara, Danielle S.
Source: Grantee Submission. 2022.
Peer Reviewed: Y
Page Count: 28
Publication Date: 2022
Sponsoring Agency: Institute of Education Sciences (ED)
Office of Naval Research (ONR) (DOD)
Contract Number: R305A180144
R305A180261
N000141712300
N000142012623
Document Type: Reports - Research
Descriptors: Age Differences, Vocabulary Development, Correlation, Reading Comprehension, Word Lists, Scores, Adults, Decision Making, Writing Skills, Children, Prediction, Error Patterns, Computational Linguistics, Reliability, Accuracy, Comparative Analysis, Language Acquisition, Linguistic Input, Models, Computer Software, Speech Communication, Word Frequency, Interpersonal Relationship, Readability
DOI: 10.3758/s13428-022-01797-5
Abstract: Age of acquisition (AoA) is a measure of word complexity which refers to the age at which a word is typically learned. AoA measures have shown strong correlations with reading comprehension, lexical decision times, and writing quality. AoA scores based on both adult and child data have limitations that allow for error in measurement, and increase the cost and effort to produce. In this paper, we introduce Age of Exposure (AoE) version 2, a proxy for human exposure to new vocabulary terms that expands AoA word lists through training regressors to predict AoA scores. Word2vec word embeddings are trained on cumulatively increasing corpora of texts, word exposure trajectories are generated by aligning the word2vec vector spaces, and features of words are derived for modeling AoA scores. Our prediction models achieve low errors (from 13% with a corresponding R[superscript 2] of 0.35 up to 7% with an R[superscript 2] of 0.74), can be uniformly applied to different AoA word lists, and generalize to the entire vocabulary of a language. Our method benefits from using existing readability indices to define the order of texts in the corpora, while the performed analyses confirm that the generated AoA scores accurately predicted the difficulty of texts (R[superscript 2] of 0.84, surpassing related previous work). Further, we provide evidence of the internal reliability of our word trajectory features, demonstrate the effectiveness of the word trajectory features when contrasted with simple lexical features, and show that the exclusion of features that rely on external resources does not significantly impact performance. [This is the online first version of an article published in "Behavior Research Methods."]
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2022
Accession Number: ED620060
Database: ERIC
Full text is not displayed to guests.
Description
Abstract:Age of acquisition (AoA) is a measure of word complexity which refers to the age at which a word is typically learned. AoA measures have shown strong correlations with reading comprehension, lexical decision times, and writing quality. AoA scores based on both adult and child data have limitations that allow for error in measurement, and increase the cost and effort to produce. In this paper, we introduce Age of Exposure (AoE) version 2, a proxy for human exposure to new vocabulary terms that expands AoA word lists through training regressors to predict AoA scores. Word2vec word embeddings are trained on cumulatively increasing corpora of texts, word exposure trajectories are generated by aligning the word2vec vector spaces, and features of words are derived for modeling AoA scores. Our prediction models achieve low errors (from 13% with a corresponding R[superscript 2] of 0.35 up to 7% with an R[superscript 2] of 0.74), can be uniformly applied to different AoA word lists, and generalize to the entire vocabulary of a language. Our method benefits from using existing readability indices to define the order of texts in the corpora, while the performed analyses confirm that the generated AoA scores accurately predicted the difficulty of texts (R[superscript 2] of 0.84, surpassing related previous work). Further, we provide evidence of the internal reliability of our word trajectory features, demonstrate the effectiveness of the word trajectory features when contrasted with simple lexical features, and show that the exclusion of features that rely on external resources does not significantly impact performance. [This is the online first version of an article published in "Behavior Research Methods."]
DOI:10.3758/s13428-022-01797-5