Age of Exposure 2.0: Estimating Word Complexity Using Iterative Models of Word Embeddings
Saved in:
| Title: | Age of Exposure 2.0: Estimating Word Complexity Using Iterative Models of Word Embeddings |
|---|---|
| Language: | English |
| Authors: | Botarleanu, Robert-Mihai, Dascalu, Mihai (ORCID |
| Source: | Grantee Submission. 2022. |
| Peer Reviewed: | Y |
| Page Count: | 28 |
| Publication Date: | 2022 |
| Sponsoring Agency: | Institute of Education Sciences (ED) Office of Naval Research (ONR) (DOD) |
| Contract Number: | R305A180144 R305A180261 N000141712300 N000142012623 |
| Document Type: | Reports - Research |
| Descriptors: | Age Differences, Vocabulary Development, Correlation, Reading Comprehension, Word Lists, Scores, Adults, Decision Making, Writing Skills, Children, Prediction, Error Patterns, Computational Linguistics, Reliability, Accuracy, Comparative Analysis, Language Acquisition, Linguistic Input, Models, Computer Software, Speech Communication, Word Frequency, Interpersonal Relationship, Readability |
| DOI: | 10.3758/s13428-022-01797-5 |
| Abstract: | Age of acquisition (AoA) is a measure of word complexity which refers to the age at which a word is typically learned. AoA measures have shown strong correlations with reading comprehension, lexical decision times, and writing quality. AoA scores based on both adult and child data have limitations that allow for error in measurement, and increase the cost and effort to produce. In this paper, we introduce Age of Exposure (AoE) version 2, a proxy for human exposure to new vocabulary terms that expands AoA word lists through training regressors to predict AoA scores. Word2vec word embeddings are trained on cumulatively increasing corpora of texts, word exposure trajectories are generated by aligning the word2vec vector spaces, and features of words are derived for modeling AoA scores. Our prediction models achieve low errors (from 13% with a corresponding R[superscript 2] of 0.35 up to 7% with an R[superscript 2] of 0.74), can be uniformly applied to different AoA word lists, and generalize to the entire vocabulary of a language. Our method benefits from using existing readability indices to define the order of texts in the corpora, while the performed analyses confirm that the generated AoA scores accurately predicted the difficulty of texts (R[superscript 2] of 0.84, surpassing related previous work). Further, we provide evidence of the internal reliability of our word trajectory features, demonstrate the effectiveness of the word trajectory features when contrasted with simple lexical features, and show that the exclusion of features that rely on external resources does not significantly impact performance. [This is the online first version of an article published in "Behavior Research Methods."] |
| Abstractor: | As Provided |
| IES Funded: | Yes |
| Entry Date: | 2022 |
| Accession Number: | ED620060 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHzVMAYe2OTiKU5Rp15Km7sAAAA4jCB3wYJKoZIhvcNAQcGoIHRMIHOAgEAMIHIBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDG7PpYHwi1bED-EVBAIBEICBmtoSTGDlmUytEX846xhUia1rmEOzIFNwvcBirggzW1P1J5jADSsyTDft2tN8BJ7wkrZ2-qX5Rjdqwo4SP-wdwjf6kEEUj9V0qJJfdhhcRexjJlteJATGBQA5nkIZx76w6REohg6ksCgJ5RqIuGGll-Ojap6B1pDrhlTUhPA9CxZqkca2L7h26UU4VRbTLhaHV1o1sgLES8jLRtk= Text: Availability: 1 Value: <anid>AN0160647162;[1wie]01dec.22;2022Dec09.02:42;v2.2.500</anid> <title id="AN0160647162-1">Age of Exposure 2.0: Estimating word complexity using iterative models of word embeddings </title> <p>Age of acquisition (AoA) is a measure of word complexity which refers to the age at which a word is typically learned. AoA measures have shown strong correlations with reading comprehension, lexical decision times, and writing quality. AoA scores based on both adult and child data have limitations that allow for error in measurement, and increase the cost and effort to produce. In this paper, we introduce Age of Exposure (AoE) version 2, a proxy for human exposure to new vocabulary terms that expands AoA word lists through training regressors to predict AoA scores. Word2vec word embeddings are trained on cumulatively increasing corpora of texts, word exposure trajectories are generated by aligning the word2vec vector spaces, and features of words are derived for modeling AoA scores. Our prediction models achieve low errors (from 13% with a corresponding R&lt;sup&gt;2&lt;/sup&gt; of.35 up to 7% with an R&lt;sup&gt;2&lt;/sup&gt; of.74), can be uniformly applied to different AoA word lists, and generalize to the entire vocabulary of a language. Our method benefits from using existing readability indices to define the order of texts in the corpora, while the performed analyses confirm that the generated AoA scores accurately predicted the difficulty of texts (R&lt;sup&gt;2&lt;/sup&gt; of.84, surpassing related previous work). Further, we provide evidence of the internal reliability of our word trajectory features, demonstrate the effectiveness of the word trajectory features when contrasted with simple lexical features, and show that the exclusion of features that rely on external resources does not significantly impact performance.</p> <p>Keywords: Age of acquisition; Age of exposure; Word embeddings; Word exposure</p> <hd id="AN0160647162-2">Introduction</hd> <p>Age of acquisition (AoA) is a measure of word complexity that attempts to account for the age at which a child learns a word. Previous studies have demonstrated that AoA scores predict readers' text processing and comprehension (Crossley et al., [<reflink idref="bib19" id="ref1">19</reflink>]), lexical decision times (Kuperman et al., [<reflink idref="bib44" id="ref2">44</reflink>]), and measures of writing quality (Crossley &amp; McNamara, [<reflink idref="bib17" id="ref3">17</reflink>]) above and beyond other measures of word complexity such as word frequency. AoA norms have been collected from adult populations (see Alonso et al., [<reflink idref="bib1" id="ref4">1</reflink>]; Cortese &amp; Khanna, [<reflink idref="bib15" id="ref5">15</reflink>]; Kuperman et al., [<reflink idref="bib44" id="ref6">44</reflink>]; Montefinese et al., [<reflink idref="bib61" id="ref7">61</reflink>]; Moors et al., [<reflink idref="bib62" id="ref8">62</reflink>]; and Stadthagen-Gonzales &amp; Davis, [<reflink idref="bib76" id="ref9">76</reflink>]) and child populations (see Álvarez &amp; Cuetos, [<reflink idref="bib2" id="ref10">2</reflink>]; Brysbaert &amp; Biemiller, [<reflink idref="bib12" id="ref11">12</reflink>]; Chalard et al., [<reflink idref="bib14" id="ref12">14</reflink>]; Frank et al., [<reflink idref="bib27" id="ref13">27</reflink>]; Grigoriev &amp; Oshhepkov, [<reflink idref="bib33" id="ref14">33</reflink>]; Morrison et al., [<reflink idref="bib63" id="ref15">63</reflink>]). AoA norms based on adult data allow for errors in the estimation because they are based on adults' perceptions and memory of word learning. On the other hand, AoA norms from child data may result in lists with specific types of words (Morrison et al., [<reflink idref="bib63" id="ref16">63</reflink>]), introduce error from parent reporting (Frank et al., [<reflink idref="bib27" id="ref17">27</reflink>]), or rely on older datasets that require updating (Brysbaert &amp; Biemiller, [<reflink idref="bib12" id="ref18">12</reflink>]). Finally, experiments to compile AoA scores are time-consuming, expensive, and require periodic updates to account for language change.</p> <p>An automated method for estimating AoA scores has the potential to alleviate the problems of current AoA collection methods. First, an automated method can simulate the word learning process, potentially removing the limitation of relying on adults' memory. Second, automating the process drastically increases the number of words for which an AoA score can be estimated. Finally, automating the process can significantly reduce the time and cost of collecting human word ratings. Thus, this paper presents an automated method capable of simulating human AoA ratings through a combination of features extracted from word embeddings trained on corpora of increasing sizes together with lexical features (such as hyponym and hypernym tree characteristics, synonym set cardinalities, and statistical attributes of words).</p> <p>Our work expands on the Word Maturity (WM) approach introduced by Landauer et al. ([<reflink idref="bib46" id="ref19">46</reflink>]) and on the Age of Exposure (AoE) model described by Dascalu et al. ([<reflink idref="bib20" id="ref20">20</reflink>]). Our objective is to examine the accuracy of a new version of the AoE algorithm (AoE 2.0) that not only exposes semantic models to an increasing number of texts, but also considers the ordering of text by readability. We assess the strength of our approach by comparing the results of AoE 2.0 to AoE 1.0, WM and to various AoA measures, such as those reported by Kuperman et al. ([<reflink idref="bib44" id="ref21">44</reflink>]), Bird (Bird et al., [<reflink idref="bib7" id="ref22">7</reflink>]), Bristol (Stadthagen-Gonzalez &amp; Davis, [<reflink idref="bib76" id="ref23">76</reflink>]), Cortese (Cortese &amp; Khanna, [<reflink idref="bib15" id="ref24">15</reflink>]), Shock (Shock et al., [<reflink idref="bib74" id="ref25">74</reflink>]), and Morrison (Morrison et al., [<reflink idref="bib63" id="ref26">63</reflink>]). Besides features generated from AoE trajectories, we also consider non-AoE lexical features to more effectively simulate human ratings. Our approach is inspired by Crossley et al. ([<reflink idref="bib18" id="ref27">18</reflink>]), wherein traditional word features such as word length, frequency, hypernymy, and polysemy were used in conjunction with features extracted from WordNet and LSA dimensions in order to predict human word ratings. The addition of these features to the AoE model increases the predictive power of our method, enabling us to better model human AoA ratings.</p> <hd id="AN0160647162-3">Age of acquisition</hd> <p>AoA is an estimate of the average age at which a word's meaning is acquired by an average speaker of the language. AoA scores are derived from human ratings and provide a linear scale for comparing words, with terms acquired in later age groups having a higher complexity compared to those acquired earlier. AoA scores correlate with other measures of word complexity such as familiarity and concreteness (Gilhooly &amp; Logie, [<reflink idref="bib30" id="ref28">30</reflink>]), word frequency, word length, and ease of production (Kuperman et al., [<reflink idref="bib44" id="ref29">44</reflink>]) with higher AoA scores indicating the words are more complex. AoA scores have been demonstrated to significantly predict readers' text processing and comprehension (Crossley et al., [<reflink idref="bib19" id="ref30">19</reflink>]), lexical processing speed (Johnston &amp; Barry, [<reflink idref="bib39" id="ref31">39</reflink>]), and have been incorporated within models of writing quality (Crossley &amp; McNamara, [<reflink idref="bib17" id="ref32">17</reflink>]).</p> <p>While AoA scores are valuable tools, there are limitations. The first challenge of AoA is the significant human effort required to generate AoA scores, since the lists are usually crowdsourced. For example, Kuperman et al. ([<reflink idref="bib44" id="ref33">44</reflink>]) generated an English AoA list of 30,000 words through crowdsourcing experiments, with 1960 participants tasked with estimating the age at which they learned a set of words. Other AoA word lists include: Bird (Bird et al., [<reflink idref="bib7" id="ref34">7</reflink>]) with 2522 terms collected from 45 British English speakers with a diverse age distribution (M = 60.7, SD = 15.5); Bristol (Stadthagen-Gonzalez &amp; Davis, [<reflink idref="bib76" id="ref35">76</reflink>]) with 3353 terms that constitute norms collected from 20 undergraduate subjects that were combined with previous G&amp;L norms (Gilhooly &amp; Logie, [<reflink idref="bib30" id="ref36">30</reflink>]) which provided 1944 AoA scores estimated by 36 student volunteers; Cortese (Cortese &amp; Khanna, [<reflink idref="bib15" id="ref37">15</reflink>]) with 2999 terms estimated by 32 undergraduate participants enrolled in a psychology course; and Shock (Shock et al., [<reflink idref="bib74" id="ref38">74</reflink>]) with 3000 terms collected from a study with 32 participants enrolled in undergraduate psychology courses.</p> <p>The crowdsourced process illuminates a second limitation: AoA lists are derived from adults' estimates of when a word was learned. Previous studies have demonstrated adults' ratings correlate with children's ratings and estimates of when words are known (see Ghyselinck et al., [<reflink idref="bib29" id="ref39">29</reflink>]). However, these correlations are run on small subsets of words, and often only on nouns. For instance, Morrison et al. ([<reflink idref="bib63" id="ref40">63</reflink>]) and Chalard et al. (2001) presented approximately 300 object pictures to 14 groups of children and an adult group. The researchers used the children's object naming accuracy to calculate an objective AoA norm for each word. The objective AoA norm, which was derived from the data from children, was highly correlated with traditional AoA scores (Chalard et al., [<reflink idref="bib14" id="ref41">14</reflink>]; Morrison et al., [<reflink idref="bib63" id="ref42">63</reflink>]). Furthermore, Chalard et al. ([<reflink idref="bib14" id="ref43">14</reflink>]) reported evidence that objective AoA norms derived from children accounted for more variance in a lexical decision response-time task than adult AoA ratings. These studies provide evidence that AoA norms derived from ratings by children who are learning the language may be more accurate than AoA norms from adults who are mature language users and recalling their exposure to words. However, collecting ratings from children is infeasible at large scales, and previous attempts to collect AoA ratings from children have used picture and object naming tasks to estimate when children learn a word, which limits the targeted AoA scores to nouns (see Álvarez &amp; Cuetos, [<reflink idref="bib2" id="ref44">2</reflink>]; Chalard et al., [<reflink idref="bib14" id="ref45">14</reflink>]; Grigoriev &amp; Oshhepkov, [<reflink idref="bib33" id="ref46">33</reflink>]; Morrison et al., [<reflink idref="bib63" id="ref47">63</reflink>]).</p> <p>There have been previous attempts to estimate AoA norms for large quantities of words from child-derived data. Frank et al. ([<reflink idref="bib27" id="ref48">27</reflink>]) utilized parent-reports of children's word acquisition, which allows researchers to estimate AoA scores (Braginsky et al., [<reflink idref="bib10" id="ref49">10</reflink>]). Brysbaert &amp; Biemiller, [<reflink idref="bib12" id="ref50">12</reflink>] created a test-based AoA estimate from previous studies on children's word acquisition using a regression to convert from grades to AoA scores. They compared the AoA estimates to the Kuperman adult AoA ratings, reporting a correlation of.757. While both databases are valuable resources, there is the possibility of error in estimation either due to parent-reporting (Frank et al., [<reflink idref="bib27" id="ref51">27</reflink>]), or the use of measures not specifically designed to estimate AoA (Brysbaert &amp; Biemiller, [<reflink idref="bib12" id="ref52">12</reflink>]).</p> <p>These limitations demonstrate the need for AoA scores that are easier to collect and based on proxies for the word learning processes. An automated method for estimating AoA scores can potentially alleviate problems inherent in the representation of AoA, as well as the time, cost, and subjectivity of collecting human word ratings. The concept of expanding AoA word lists using Machine Learning models has been explored in the past. Mandera et al. ([<reflink idref="bib54" id="ref53">54</reflink>]) investigated the usage of various algorithms for constructing semantic spaces, namely LSA, LDA, HAL (Lund &amp; Burgess, [<reflink idref="bib50" id="ref54">50</reflink>]), and a skip-gram based method (Mikolov, Le, &amp; Sutskever, [<reflink idref="bib59" id="ref55">59</reflink>]), in combination with either <emph>k</emph>-nearest neighbors or random forest models for extrapolating various subjective ratings. Authors report correlations of up to.737 with Kuperman AoA ratings using two random splits of the data, using a corpus compiled through downloading 204,408 documents containing English film and television subtitles from the Open Subtitles database (<ulink href="http://opensubtitles.org">http://opensubtitles.org</ulink>). However, the authors also reported that some of the extrapolation methods introduced artifacts to the data, potentially producing different conclusions from human ratings in the context of two lexical decision tasks. While the method presented by Mandera et al. ([<reflink idref="bib54" id="ref56">54</reflink>]) has many similarities with the approach adopted in the current study, they relied solely on word embeddings generated through different algorithms using a single training corpus, whereas the current study models AoA by leveraging exposure trajectories which aim to simulate the way language learners acquire words. In the following section, we describe word learning processes, specifically in relation to the impact of increasing exposure to words. We then describe two prior machine learning approaches to simulate increasing word exposure, the Word Maturity model (Biemiller et al., [<reflink idref="bib6" id="ref57">6</reflink>]; Landauer et al., [<reflink idref="bib46" id="ref58">46</reflink>]) and the AoE 1.0 model (Dascalu et al., [<reflink idref="bib20" id="ref59">20</reflink>]).</p> <hd id="AN0160647162-4">Word learning and exposure</hd> <p>Children's ability to learn a particular word depends largely on exposure to that word in their language environment (Hills et al., [<reflink idref="bib35" id="ref60">35</reflink>]; Hoff &amp; Naigles, [<reflink idref="bib37" id="ref61">37</reflink>]; Roy et al., [<reflink idref="bib70" id="ref62">70</reflink>]). As the primary source of word input, caregivers essentially manifest the language environment during the child's infancy (Weisleder &amp; Fernald, [<reflink idref="bib80" id="ref63">80</reflink>]). Caregiver names and concrete objects are typically among the first learned by children because of the constant exposure infants have to those words in their early language environment (Roy et al., [<reflink idref="bib70" id="ref64">70</reflink>]). Furthermore, a greater number of words in infants' language environments predicts larger and more expressive vocabularies several months later, indicating that mere exposure to a greater number of words affords infants opportunities to learn more words (Hoff, [<reflink idref="bib36" id="ref65">36</reflink>]; Pan et al., [<reflink idref="bib66" id="ref66">66</reflink>]). These studies demonstrate that word learning is dependent on word input from outside sources, which for children is primarily adults (Hills et al., [<reflink idref="bib35" id="ref67">35</reflink>]; Hoff &amp; Naigles, [<reflink idref="bib37" id="ref68">37</reflink>]).</p> <p>As children grow, the relationship between word exposure and word learning can be observed both in speech and in written language. Children whose peers are more linguistically skilled are more likely to learn new words compared to children whose peers are less linguistically skilled (Justice et al., [<reflink idref="bib40" id="ref69">40</reflink>]; Webb, [<reflink idref="bib79" id="ref70">79</reflink>]). In addition to speech, literate children will begin working with written language, and reading will provide children the opportunity to learn new words from the context of written passages (Nagy et al., [<reflink idref="bib64" id="ref71">64</reflink>]; Teng, [<reflink idref="bib77" id="ref72">77</reflink>]). Both studies demonstrate that exposure to new sources of words, either through peers or written language, affords students the opportunity for greater word learning. Even for adults, the ability to learn words in a second language depends on exposure to new words (Eckerth &amp; Tavakoli, [<reflink idref="bib24" id="ref73">24</reflink>]).</p> <p>Overall, the literature on word learning indicates that increased exposure to words over time is a foundational aspect of word learning; and yet estimating the rate of word exposure throughout childhood is fraught with barriers. Thus, computational simulations of increasing word exposure over time have strong potential to contribute to research and practice in areas related to literacy. Previous efforts to estimate increasing word exposure include the Word Maturity model and the AoE 1.0 model, each described in the following sections.</p> <hd id="AN0160647162-5">Word Maturity</hd> <p>The AoE 2.0 model builds on the Word Maturity model (Biemiller et al., [<reflink idref="bib6" id="ref74">6</reflink>]; Landauer et al., [<reflink idref="bib46" id="ref75">46</reflink>]), an automated model constructed to reflect the word exposure process by deriving word-occurrence patterns from incremental subcorpora of texts. The incremental subcorpora approximate the growing language environment experienced by children. As the subcorpora increase in size, the word-occurrence patterns change based on the repeated exposures to words, and associations of trained words to novel words.</p> <p>Word Maturity uses Latent Semantic Analysis (LSA; Landauer &amp; Dumais, [<reflink idref="bib45" id="ref76">45</reflink>]) to transform words into vector representations within a semantic vector space. LSA uses a term-document occurrence matrix whose dimensionality is reduced using Singular Value Decomposition. The newly formed vector representations hold latent information on the relationships between words such that words which co-occur in similar texts have similar meanings. LSA does not directly represent the probability of a word appearing in similar documents, but rather it deconstructs the meaning of a paragraph as a sum of the meanings of its component words. Thus, each word is represented using a high-dimensional vector with numerical values, such that the decomposition optimizes the least squares criterion.</p> <p>Word Maturity quantifies the evolution of any word's complexity throughout a speaker's process of language acquisition and approximates the manner in which language readers are exposed to new texts during their learning process. The method described by Landauer et al. ([<reflink idref="bib46" id="ref77">46</reflink>]) consists of splitting a text corpus into cumulative subcorpora. For each of these incremental subcorpora, LSA semantic spaces are trained, and then the generated intermediary vector representations are aligned to the mature semantic space via Procrustes rotation. Next, the Word Maturity model computes the averaged vector for all paragraphs containing a word for each intermediate semantic space; these averaged vectors are then used to measure the cosine distance to the mature vector representation. These cosine values for the incremental subcorpora model each word's trajectory such that subcorpora gather the word's vector representations across intermediary models, up until the mature space (for which we inherently have a perfect overlap, thus a cosine of 1). These trajectories can be either compared between words or viewed for singular words.</p> <p>These trajectories show equivalent sensitivity to word frequency as do conventional psychometric tests, as indicated by the strong correlations that were reported. This was measured through evaluating the relationship between human vocabulary knowledge and the various word metrics on a number of vocabulary tests, resulting in rank order correlations that ranged from 0.73 for the Kaufman Brief Intelligence Test-II (Kaufman &amp; Kaufman, [<reflink idref="bib42" id="ref78">42</reflink>]) and 0.76 for the Peabody Picture Vocabulary Test-III (Maddux, [<reflink idref="bib53" id="ref79">53</reflink>]), to 0.81 for the Kaufman Assessment Battery for Children-Verbal Knowledge (Kaufman &amp; Kaufman, [<reflink idref="bib41" id="ref80">41</reflink>]) and 0.83 for the Kaufman Assessment Battery for Children–Expressive Vocabulary (Kaufman &amp; Kaufman, [<reflink idref="bib41" id="ref81">41</reflink>]).</p> <p>The results indicate that the relation between the trajectories taken by individual words in the Word Maturity model can be viewed as an indicator of the word's complexity—i.e., a consequence of the amount of <emph>exposure</emph> needed before the latent representation of a model matures. Words with low complexity, acquired early in a learner's vocabulary, have latent representations in earlier intermediate models which are highly similar to those found in the mature model. Thus, the model requires fewer texts to adequately represent the word. In contrast, the meaning of complex words is acquired only after a certain degree of exposure to a language, which is not necessarily measured by grade level or age.</p> <hd id="AN0160647162-6">Age of Exposure version 1.0</hd> <p>AoE 1.0 (AoE; Dascalu et al., [<reflink idref="bib20" id="ref82">20</reflink>]) was developed to provide an alternative measurement of a word's complexity. One limitation of the Word Maturity model is that it is proprietary and thus neither the code nor the word complexity estimates were available to the public. The overarching objectives driving AoE 1.0 were to provide an open-source model and at the same time improve on the Word Maturity model. AoE 1.0 did so by utilizing latent topic probability distributions (Blei et al., [<reflink idref="bib8" id="ref83">8</reflink>]) in place of LSA. In contrast to LSA, Latent Dirichlet Allocation (LDA) establishes latent topics in which words have corresponding probabilities; higher probabilities of a word in multiple different topics are indicative of potential different word senses (i.e., LDA accounts to some extent for word polysemy). AoE provides measures of a word's relative complexity at various points during the iterative training by matching topics across intermediate models, with more difficult words having topic distributions further away from those of the mature semantic space.</p> <p>AoE 1.0 was computed by generating sequentially increasing corpora of documents that simulate a human's exposure to language. For each of the generated subcorpora, LDA models were trained. Each LDA model learned topics for both documents and words, which were then aligned to the mature model. The topic distributions were quantified by applying a flow algorithm over a bipartite graph consisting of the intermediate topics and the mature topics, with edges having weights computed as the Jensen-Shannon divergence between word probabilities corresponding to the two topics. After this matching, various features were extracted from the aligned intermediate and mature topic spaces by measuring the cosine similarities between the word's representation in the intermediate model and its topic distribution in the fully-trained model. Some of these extracted features included (a) the inverse average similarity between intermediate and mature models; (b) the inverse linear regression slope of the similarities; (c) the index of the intermediate model that first exceeds a similarity threshold of.3 to the mature model (i.e., the "index above threshold"); (d) the index of the first intermediate model, for which a grade 3 polynomial fit gives a score above a.4 threshold; and (e) the inflection point of the polynomial (see Dascalu et al., [<reflink idref="bib20" id="ref84">20</reflink>], for further explanations). All features extracted in the AoE model were intended to capture the way in which a word's representations evolve as more documents are used to train incremental LDA models. Dascalu et al. ([<reflink idref="bib20" id="ref85">20</reflink>]) reported high correlations between AoE indices and various word features (e.g., Kuperman AoA,.716 to.893; word frequencies, −.599 to −.774; word entropies, −.615 to −.780; word naming latencies,.611 to.779; lexical decision latencies,.616 to.766) as evidence for the method's adequacy in building word features that capture the way a word's representation evolves as a greater portion of the dataset is used.</p> <hd id="AN0160647162-7">Age of Exposure Version 2.0: The current study</hd> <p>Our objective in the current study is to enhance the accuracy of AoE estimates of human ratings by incorporating alternative computational approaches and assessing the advantages of two alternative methodological approaches. Computationally, first, AoE 2.0 incorporates the use of regression models to predict word features, enabling the generalization of AoA scores to words that are not included in the original AoA word lists. Second, AoE 2.0 leverages word2vec, which affords estimates of syntactic and semantic relations between words using vector geometry. Methodologically, we examine the impact of randomly introducing the texts to the model compared to introducing text sequentially using an automated readability score, Flesch Reading Ease (Flesch, [<reflink idref="bib26" id="ref86">26</reflink>]).</p> <p>AoE 2.0 leverages the generalization capabilities of regression models to expand human-collected AoA word lists by first training on the existing words in the list and subsequently generating predictions for other words from the vocabulary that are not available within an AoA list. This can be especially beneficial for AoA word lists that have a limited numbers of words (see e.g., Łuniewska et al., [<reflink idref="bib51" id="ref87">51</reflink>]). To simulate incremental word exposure, AoE 2.0 predicts words' AoA scores using features generated by word2vec models that are incrementally exposed to increasing corpora of text.</p> <p>Word2vec is an algorithm for training neural networks to estimate word vector representations through self-supervised learning. Context windows of words with fixed sizes are considered across the training corpus and word embeddings are learned. The principal difference between word2vec and LDA (used in AoE 1.0) is that LDA is a topic modeling technique, whereas word2vec generates word embeddings. The word vectors that word2vec generates reflect syntactic and semantic relations between words through vector geometry, which have been shown to be superior to LSA (Baroni et al., [<reflink idref="bib4" id="ref88">4</reflink>]; Lenci et al., [<reflink idref="bib47" id="ref89">47</reflink>]; Levy et al., [<reflink idref="bib48" id="ref90">48</reflink>]; Mikolov, Chen, et al., [<reflink idref="bib57" id="ref91">57</reflink>]; Mikolov, Sutskever, et al., [<reflink idref="bib58" id="ref92">58</reflink>]). Our method requires the alignment of vector spaces generated by models trained at different corpus exposure levels as well as the measurement of word similarities. Word2vec is well suited for this task because of the inherent arithmetic properties of the word representations. Additionally, LDA is computationally expensive on very large corpora, whereas word2vec learns embeddings in a distributed fashion through stochastic gradient descent.</p> <p>In addition to incorporating computational advantages, we examine the impact that the <emph>order</emph> in which texts are used during training has for the performance of the models. In essence, we compare an arbitrarily random order to an ascending order in which more readable texts are introduced first to the model. This sorting of paragraphs was performed using their Flesch Reading Ease (Flesch, [<reflink idref="bib26" id="ref93">26</reflink>]) readability score which provides an indication of the readability of a given text as a score in the range of 0–100.</p> <p>The accuracy of the models is measured using a 10-fold cross-validation with a random forest regressor (Breiman, [<reflink idref="bib11" id="ref94">11</reflink>]), a support vector regression, linear regression, and lasso regression. We selected these regressors because we found that other models, such as multilayered perceptron, produce weaker results on the datasets analyzed in our work. Moreover, random forest and linear regression models are more interpretable.</p> <p>To assess the generalizability of the models, we compare AoE 2.0, AoE 1.0, and Word Maturity estimates of AoA using the word list built by Kuperman et al. ([<reflink idref="bib44" id="ref95">44</reflink>]) in terms of mean absolute error (MAE), correlations, and <emph>R</emph><sups>2</sups>. Second, we assess the accuracy of AoE 2.0 predictions of AoA scores using various other word lists, namely Bird (Bird et al., [<reflink idref="bib7" id="ref96">7</reflink>]), Bristol (Stadthagen-Gonzalez &amp; Davis, [<reflink idref="bib76" id="ref97">76</reflink>]), Cortese (Cortese &amp; Khanna, [<reflink idref="bib15" id="ref98">15</reflink>]), Shock (Shock et al., [<reflink idref="bib74" id="ref99">74</reflink>]), and the AoA scores derived from children by Morrison et al. ([<reflink idref="bib63" id="ref100">63</reflink>]). Third, we analyze the effectiveness of the generated scores in the context of estimating textual difficulty and compare the results to Word Maturity and AoE 1.0. Finally, we explore the impact of the different types of features used for training the regressors by performing an ablation study in order to measure the potential benefits of adding WordNet and word trajectory features to the lexical feature set.</p> <p>Notably, the AoE 2.0 model is fully reproducible and has been released as an open-source project available at: https://github.com/readerbench/Age-of-Exposure. The generated AoE scores for the entire English vocabulary using our most predictive model are also available within the previously mentioned repository.</p> <hd id="AN0160647162-8">Method</hd> <p></p> <hd id="AN0160647162-9">Text corpus</hd> <p>A large text corpus was used to simulate the manner in which people are exposed to new words during reading and listening by splitting the full corpus into incrementally increasing subcorpora. Our collection of texts consists of a combination of <emph>TASA</emph> (Touchstone Applied Science Associates, Inc.), <emph>COCA</emph> (Corpus of Contemporary American English), and the <emph>Child Directed Speech</emph> (CDS) corpora selected from the Child Language Data Exchange System (CHILDES) dataset (MacWhinney, [<reflink idref="bib52" id="ref101">52</reflink>]). <emph>TASA</emph> contains short fragments from text documents that aim to provide a representative sample of educational English for various topics such as health, industrial arts, home economics, social studies, language arts, and others (Ivens &amp; Koslin, [<reflink idref="bib38" id="ref102">38</reflink>]). <emph>COCA</emph> contains short- and medium-length documents on a variety of topics—academic journals, fiction (i.e., short stories and plays), magazine articles taken from a variety of popular magazines, newspapers articles, as well as transcripts from TV and radio programs (Davies, [<reflink idref="bib21" id="ref103">21</reflink>]-). The CDS corpora consist of transcripts of conversations between young children and mature language speakers. The CDS datasets comprise dialogues between young children, typically before the beginning of formal education, and various interlocutors such as their parents. We include the CDS corpus because it includes words with low AoA scores (i.e., words acquired at an early age). A large-scale database for CDS is the CHILDES (MacWhinney, [<reflink idref="bib52" id="ref104">52</reflink>]) system which offers transcripts of various CDS experiments in different languages. In total, we selected the texts from 56 datasets found under the "North America" and "United Kingdom" English sections (please see the index found at https://childes.talkbank.org/access/ for the individual corpora and their citations).</p> <p>The texts from TASA and COCA were split into their constituent paragraphs, whereas the CDS datasets were split at the transcript level, to generate individual samples roughly equal in length. This was required since certain COCA texts were considerably larger by an order of magnitude in comparison to typical TASA texts. For CDS, we elected to use the entire transcript instead of individual utterances because the child utterances tend to be only a few words long. This approach rendered the transcripts, as a collection of child utterances, comparable in length to the TASA and COCA paragraphs. In total, 9391 CDS transcripts were added together with the combined TASA/COCA dataset to form our final training corpus of 233,060 documents (see Fig. 1 with the number of paragraphs for each category displayed in a logarithmic scale). We opted to use the logarithmic scale so that the largest categories would not dominate the plot, obscuring the least frequent collections on a linear scale.</p> <p>Graph: Fig. 1 Histogram showing the number of paragraphs in the TASA, COCA, and CDS corpora as a function of text type on a logarithmic scale</p> <p>Figure 2 depicts the readability scores using Flesch Reading Ease for the selected transcripts, as computed on transcripts as a whole. The density plots from Fig. 2 show that the CDS corpora have texts of lower complexity than the ones present in the COCA and TASA datasets. Notably, however, because they are transcripts of dialogue, there are significant differences between the structure of these documents and the paragraphs extracted from TASA and COCA.</p> <p>Graph: Fig. 2 Readability of CDS transcripts. Scores outside of the 0–100 range correspond to errors in the way sentences are separated during parsing caused by formatting errors</p> <p>Our aggregated dataset comprising of three corpora (TASA, COCA, and CDS) contains documents written at various grade levels and on various topics. The purpose of combining the three corpora is to provide an adequate representation of common texts in English that are representative of language that might be encountered by children in their home and school environments. These texts are representative of potential exposure to language from various sources such as school, television, internet, and interactions with caregivers, peers, and teachers. Overall, we aimed to simulate a variety of situations in which the usual language learner is exposed to new words. Nonetheless, the AoE 2.0 computational methodology can be applied to virtually any corpus if the aim is to build corpus-specific models.</p> <hd id="AN0160647162-10">Building iterative word embedding models</hd> <p>Multiple word2vec models were trained on an incrementally increasing datasets of documents from the aforementioned corpus, with a total of 465,682,595 tokens of which only lemmatized verbs, nouns, adjectives, and adverbs were retained. Lemmatization and the selection of only content words was performed to reduce noise from stop words, reduce all word forms into their dictionary form, and ensure alignment with words from the AoA word lists. The spaCy (https://spacy.io) framework was used for tokenization, part-of-speech tagging, and lemmatization with the largest available English model being selected. Of note is that we did not check the accuracy of this model's tagging and lemmatization. The use of part-of-speech (POS) taggers for CHILDES or other CDS datasets, or the confirmation of spaCy's accuracy on child transcripts may be of interest for future work that uses CDS data as a larger portion of the total corpus. Additionally, we did not check the accuracy of spaCy's taggers on the COCA and TASA documents either.</p> <p>To perform the iterative word embedding modeling, two components are required: a function describing the dataset <emph>growth</emph>, and a relative <emph>ordering</emph> function of the documents. The growth function simulates the manner in which people are exposed to an increasing number of texts during their lifetime. Because exposure to the content of texts and the number of texts themselves accumulate and have some potential to relate to previous knowledge, our model is also exposed to all texts from the previous stages. Our assumption is that a model trained at step <emph>t</emph> will also contain most of the semantic and syntactic information on the vocabulary words from models at previous steps (<emph>t</emph> ′ &lt; <emph>t</emph>).</p> <p>In the experiments that follow, we utilize a simple linear growth scheme, where each iteration of training adds an equal-sized portion of the total corpus. In Appendix A, <emph>Impact of the Growth Scheme</emph>, we describe in greater detail how different growth schemes can potentially impact the performance of the model. The order in which documents are introduced to the learner also plays an important role. The first approach involves random sampling of documents, whereas the second method orders documents by their Flesch Reading Ease score. The equation for this scoring method is given below:</p> <olist> <item> <ephtml> &lt;math display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;206.835&lt;/mn&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mn&gt;1.015&lt;/mn&gt;&lt;mspace width="4pt" /&gt;&lt;mfenced close=")" open="("&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mfenced&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mn&gt;84.6&lt;/mn&gt;&lt;mfenced close=")" open="("&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </item> </olist> <p>Graph</p> <p>The ordering makes simpler documents appear in earlier stages, while more complex texts appear in later stages. The purpose is to simulate the notion that people usually start with more readable texts or simpler language, and then progress to more difficult texts and language as their language proficiency increases. Moreover, less difficult words should be seen early in model training by placing texts with higher readability at the beginning.</p> <p>The Flesch Reading Ease readability score provides a score from 0 (very difficult to read) to 100 (very easy to read[<reflink idref="bib1" id="ref105">1</reflink>]) that is a function of the length of the sentences and the number of syllables per word, which is often used as a proxy for the readability level of a text. The Flesch Reading Ease score it does not consider multiple aspects that may impact text readability (e.g., see McNamara et al., [<reflink idref="bib55" id="ref106">55</reflink>]), but it is easily computed for any given text and is commonly used as a measure of a text's readability.</p> <hd id="AN0160647162-11">Training self-supervised word embedding models</hd> <p>The word2vec (Mikolov, Chen, et al., [<reflink idref="bib57" id="ref107">57</reflink>]; Mikolov, Sutskever, et al., [<reflink idref="bib58" id="ref108">58</reflink>]) model uses a shallow neural network with a single hidden layer that provides the embeddings and an output layer consisting of a softmax activation over the vocabulary. There are two main variants of training a word2vec model: continuous-bag-of-words (CBOW) and Skip-gram. These are illustrated in Fig. 3. In both cases, the context window of a word is composed of the neighboring terms from the text. For CBOW, the model is trained to predict the target word given an input consisting of the window words, for which the order is disregarded. However, for the Skip-gram model, the target word is given as input and used to predict which words are most probable to appear in the same context as the target word. Thus, the word2vec neural network learns embeddings that capture the local co-occurrence of words.</p> <p>Graph: Fig. 3 The two variants of the word2vec algorithm (Mikolov, Chen, et al., [<reflink idref="bib57" id="ref109">57</reflink>])</p> <p>Both the CBOW and Skip-gram algorithms generate models that reflect a probability distribution over the entire vocabulary by considering an output layer having a number of neurons equal to the size of the vocabulary, combined with a softmax activation, described in Eq. 3:</p> <p>2 <ephtml> &lt;math display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;msup&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;/msup&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msubsup&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;z&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mi mathvariant="double-struck"&gt;R&lt;/mi&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Graph</p> <p>The softmax activation transforms an array of numbers into a probability distribution, such that the resulting values have a sum of 1 and higher initial values correspond to higher values in the new distribution.</p> <p>In general, CBOW tends to have better performance on more common words, while Skip-gram is better at representing rare words (Mikolov, [<reflink idref="bib56" id="ref110">56</reflink>]). The word2vec model has numerous computational advantages because it does not require the computation of the global occurrence matrix; however, its main benefit consists in the vector representations themselves. As presented in the initial studies conducted by Mikolov, Chen, et al. ([<reflink idref="bib57" id="ref111">57</reflink>]), word embeddings capture both syntactic and semantic properties of words. Of particular interest are the highlighted mathematical properties, such as the ability to measure the cosine similarity between two words to determine if they have similar meanings, or the ability to perform algebraic operations on the vectors in order to determine semantic relationships between words (e.g., by adding the vector for "king" and subtracting the representation for "man" to the vector of "woman," the nearest representation in the vector space can be found to be "queen").</p> <p>In contrast to the topic distributions from LDA, word2vec representations are additive and can be more easily employed to measure semantic similarity. While LDA captures word polysemy in terms of word associations to multiple topics, it has a problem identifying the optimal number of corresponding topics. This, coupled with a bipartite matching of topics that is more rigid than transformations on word embeddings as found in word2vec, supports the use of word2vec over LDA in developing AoE 2.0. Word2vec models have also been shown to provide embeddings that are qualitatively superior to alternatives such as LSA (Mikolov, Chen, et al., [<reflink idref="bib57" id="ref112">57</reflink>]).</p> <p>There are numerous options for generating word embeddings besides using semantic models such as LSA, LDA, and word2vec. Ruas et al. ([<reflink idref="bib71" id="ref113">71</reflink>]), for example, proposed an algorithm for generating word embeddings that are able to capture multiple senses through context-aware disambiguation. Similarly, Li and Jurafsky ([<reflink idref="bib49" id="ref114">49</reflink>]) proposed that multi-sense embeddings can help improve natural language understanding and introduced an algorithm for training such embeddings. Furthermore, state-of-the-art natural language processing models that are based on transformers, such as BERT (Devlin et al., [<reflink idref="bib22" id="ref115">22</reflink>]) and XLNet (Yang et al., [<reflink idref="bib81" id="ref116">81</reflink>]), opt to learn embeddings at the level of sub-words which can greatly help with rare words by tokenizing them into statistically more common components. Our choice of using word2vec instead of other embedding methods, such as the multi-sense algorithms and the sub-word embedding schemes outlined previously, was motivated by the fact that word2vec is efficient for large datasets, which constitutes a major advantage in the context of training multiple word embeddings for a single corpus. Additionally, AoA scores are not typically separated by word senses, which potentially renders word2vec a better fit over multi-sense word embeddings. With this in mind, we elected to use word2vec, the most common neural embedding scheme, as a method for modelling holistic exposure trajectories for words.</p> <p>For each of the 10 incremental sets, a word2vec model was trained. Given the dataset split, each model encapsulates all of the words encountered by the previously trained models. The final word2vec model (i.e., trained on the "mature" dataset) holds the final word vectors, which are used as reference when analyzing all other previous models. We opted to rely on CBOW representations instead of Skip-gram because the overall performance of our model was better served by having higher quality word embeddings for common words. Each word2vec model is trained using the CBOW algorithm for five epochs, with a window size of 5 and a vector size of 300 (see Appendix B, <emph>Impact of Word Embedding Size</emph>, for a discussion on the impact the word embedding size has on the final performance). This means that a model uses 5-grams (i.e., five consecutive words) and is tasked to predict the central word, given the surrounding contextual words. For each word, a 300-dimensional embedding vector is learned over five passes over the texts in a subcorpus. A subcorpus consists of paragraphs assigned to a certain training stage, with the first training stage consisting of approximately 10% of the total corpus. Each consecutive training considers a larger corpus that includes all texts used in previous stages (i.e., the second stage will have around ~20% of the texts, including the 10% used in stage 1).</p> <hd id="AN0160647162-12">Generating word features</hd> <p>The cosine function between intermediate and mature representations of terms is used to model the evolution of a word across stages. The vocabulary and the number of word vectors in the word2vec vector space increase as the corpus increases and, as such, result in different vector space alignments at each intermediate step; in return, this leads to the need for a method of aligning the intermediate vector spaces to the mature vector space (i.e., the word2vec model trained on the entire corpus). This is achieved by using the Procrustes alignment (Gower, [<reflink idref="bib32" id="ref117">32</reflink>]; Krzanowski, [<reflink idref="bib43" id="ref118">43</reflink>]) performed using all words common between all intermediate models as pivots.</p> <p>The Procrustes alignment translates, scales, and rotates the word vectors of all words present in intermediate models to their corresponding mature values. The components of the Procrustes algorithm are an orthogonal rotation, a reflection, and a scaling transformation which are applied to the intermediate vector space, aligning it to the mature model. Aligning the two vector spaces is necessary primarily due to the stochastic nature of the word embeddings generated by the word2vec neural network. Because the intermediate models may not have seen certain words, 0-vectors are used to bring each intermediate vector space to the same dimensionality as the mature one. Another option might be to use the average vector of known word embeddings as a substitute for unseen words; however, this was not explored in this work. After the Procrustes transformation is applied to the intermediate vector space, the approximate directions and magnitudes of the vectors representing the same word in the two vector spaces should match. Inherent discrepancies between representations provide a measure of the dissimilarity between how a word is embedded in a certain intermediary stage, and how it is embedded when the entire corpus is considered.</p> <p>A demonstration of the process is displayed in Fig. 4 which illustrates the following steps: (a) the two vector spaces are translated so that both their origins are 0; (b) the spaces are scaled by their respective Frobenius norms; and then (c) the first vector space is transformed using an orthogonal matrix that minimizes the distance between the reference words using orthogonal Procrustes (Schönemann, [<reflink idref="bib72" id="ref119">72</reflink>]).</p> <p>Graph: Fig. 4 Procrustes rotation demonstration</p> <p>The formal definition of the Procrustes alignment is the following:</p> <p></p> <ulist> <item> Given two matrices A and B, our aim is to align B to A.</item> <p></p> <item> Normalize the matrices corresponding to the entire vector spaces and translate the data to the origin.</item> <item>ath display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mover&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;¯&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mfenced close="∥" open="∥"&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mover&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;¯&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;mfenced close="∥" open="∥"&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/math&gt; _ht_</item> </ulist> <p>Graph</p> <p></p> <ulist> <item> Find the orthogonal matrix R that most closely maps A to B using SVD.</item> <item>ath display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;U&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;Σ&lt;/mi&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mi&gt;V&lt;/mi&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/msup&gt;&lt;mspace width="4pt" /&gt;&lt;mfenced close=")" open="("&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mi&gt;u&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;V&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mi&gt;u&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/math&gt; _ht_</item> </ulist> <p>Graph</p> <p>5 <ephtml> &lt;math display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;U&lt;/mi&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mi&gt;V&lt;/mi&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Graph</p> <p></p> <ulist> <item> Apply R to B to align the two matrices.</item> <item>ath display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;/mrow&gt;&lt;mo&gt;′&lt;/mo&gt;&lt;/msup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;*&lt;/mo&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; _ht_</item> </ulist> <p>Graph</p> <p></p> <ulist> <item> Compute disparity.</item> <item>ath display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msubsup&gt;&lt;mfenced close="∥" open="∥"&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;′&lt;/mo&gt;&lt;/mfenced&gt;&lt;mi&gt;F&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;/math&gt; _ht_</item> </ulist> <p>Graph</p> <p>The disparity between each intermediate vector space and the mature vector space can be computed when performing the Procrustes rotation. Disparity measures the sum of the square post-alignment differences between the two vector spaces, which is also the metric minimized by the rotation itself. Our hypothesis is that earlier vector spaces have higher disparities than the later ones, which gradually tend to move closer to the mature model as more of the texts in the corpus are used for training. Figure 5 displays the disparities for the two scenarios introduced using the ordering functions (i.e., unsorted vs. sorted by readability).</p> <p>Graph: Fig. 5 Disparities of intermediate models for different corpora subsets</p> <p>Other methods for handling word embeddings to represent change over some dimension of the corpus exist. For example, temporal word embeddings with a compass (TWEC) is designed to aid in the analysis of diachronic shifts in word meaning, is based on word2vec, and uses atemporal reference embeddings, called "compasses," when training temporal-aware embeddings (Di Carlo et al., [<reflink idref="bib23" id="ref120">23</reflink>]). The hypothesis is that the majority of terms in a vocabulary do not change their meaning over time, while the ones that do present diachronic shifts will appear in the context windows of those stable ones. In our work, the trained word embeddings would not shift due to temporal changes, but because of their exposure to increasing quantities of texts. At each step, the word embeddings are trained on a section of the dataset that includes all documents used in previous steps. Because the embeddings at each level of exposure become increasingly more accurate (i.e., closer to the mature model as seen in Fig. 5), we elected to utilize the Procrustes method for aligning the vector spaces with the assumption that they represent human exposure trajectories, wherein an individual's mastery of a language increases cumulatively with exposure to that language.</p> <p>One way to interpret these disparities is that steeper slopes indicate a greater degree of change in terms of vector spaces, from the initial model until the mature one, which may potentially result in more useful word trajectory features. This would follow from the assumption that sorted models would ideally have a more natural progression, with less complex words appearing earlier than more complex ones. Unsorted models, on the other hand, may see most words in the initial training stages, which would result in the vector space being populated early, which may result in a reduction in the amount of change observed from one stage of training to another.</p> <p>After the intermediate vector spaces are rotated, the feature vector for a word is formed by measuring the cosine similarity between the vector representation of the word for each intermediate model in relation to the mature model. This can be defined as:</p> <p>8 <ephtml> &lt;math display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfenced close="|" open="{"&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mfenced close=")" open="("&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mi mathvariant="bold-italic"&gt;w&lt;/mi&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold-italic"&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mi mathvariant="bold-italic"&gt;w&lt;/mi&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="bold-italic"&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mfenced&gt;&lt;/mfenced&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mrow&gt;&lt;mo maxsize="1.623em" minsize="1.623em" stretchy="true"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Graph</p> <p>where <emph>T</emph> is the number of stages considered and <bold><emph>w</emph></bold><subs><bold><emph>t</emph></bold></subs> is the word2vec representation for a word at a given training step <emph>t</emph>.</p> <hd id="AN0160647162-13">Predicting age of acquisition</hd> <p>The last step consists of training a regressor model to predict the Kuperman AoA scores. If the cosine similarities between the intermediate word representations and the mature word representation are indicative of a word's evolution as language proficiency increases and learners are exposed to more texts, then they can be used as indicators of a word's AoA. Besides the cosine values for the nine intermediate models, Table 1 provides a list of six other informative features that were also reported in the initial AoE model by Dascalu et al. ([<reflink idref="bib20" id="ref121">20</reflink>]).</p> <p>Table 1 AoE 2.0 features extracted from word trajectories</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;p&gt;Feature name&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Formula&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Description&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Inverse average&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow xmlns=""&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mover&gt;&lt;mrow&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/mrow&gt;&lt;/math&gt;&lt;inline-graphic href="13428&amp;#95;2022&amp;#95;1797&amp;#95;Article&amp;#95;IEq1.gif" /&gt;&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The complement of the mean cosine similarity, showing the average dissimilarity of intermediate embeddings to the mature word embedding.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Highest cosine similarity&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow xmlns=""&gt;&lt;munder&gt;&lt;mo movablelimits="false"&gt;max&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/munder&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt;&lt;inline-graphic href="13428&amp;#95;2022&amp;#95;1797&amp;#95;Article&amp;#95;IEq2.gif" /&gt;&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The highest cosine similarity, denoting how well intermediary stages best match the mature model&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Inverse slope&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mfrac xmlns=""&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/mfrac&gt;&lt;/math&gt;&lt;inline-graphic href="13428&amp;#95;2022&amp;#95;1797&amp;#95;Article&amp;#95;IEq3.gif" /&gt;, for the linear regression fitting the cosine similarity values:&lt;/p&gt;&lt;p&gt;&lt;italic&gt;ax&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;t&lt;/italic&gt;&lt;/sub&gt; + &lt;italic&gt;b&lt;/italic&gt; = &lt;italic&gt;CosSim&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;t&lt;/italic&gt;&lt;/sub&gt;, &amp;#8704; &lt;italic&gt;t&lt;/italic&gt; &amp;#8712; {1.. &lt;italic&gt;T&lt;/italic&gt;}&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The inverse of the slope as measured while performing a linear interpolation between the cosine values. This feature approximates how quickly the intermediate word embedding models learn a representation of the word while matching the mature one&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index at threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;italic&gt;t&lt;/italic&gt;, &lt;italic&gt;where CosSim&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;t&lt;/italic&gt;&lt;/sub&gt; &amp;#8805; &lt;italic&gt;thresh&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The continuous index, or the index of the intermediate model that contains a representation of a given word with a cosine similarity above a certain threshold. Moreover, the successor model also has to exhibit the same property, for consistency. This is computed across multiple threshold values, ranging from 0.3 to 0.7, with increments of 0.05. This feature identifies the stage at which a word is well conceptualized&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Word-vocabulary integration&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;|{&lt;italic&gt;t&lt;/italic&gt;, &lt;italic&gt;CosSim&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;t&lt;/italic&gt;&lt;/sub&gt; &amp;#8805; 0.3}|, &amp;#8704; &lt;italic&gt;t&lt;/italic&gt; &amp;#8712; {1.. &lt;italic&gt;T&lt;/italic&gt;}&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The number of words with which the target word has a cosine similarity of at least 0.3 in the mature vector space, as well as the average of these cosine similarities&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Top 3 cosine similarities&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;italic&gt;top&lt;/italic&gt;&lt;sub&gt;3&lt;/sub&gt;{&lt;italic&gt;CosSim&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;t&lt;/italic&gt;&lt;/sub&gt;}, &amp;#8704; &lt;italic&gt;t&lt;/italic&gt; &amp;#8712; {1.. &lt;italic&gt;T&lt;/italic&gt;}}&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The cosine similarities of the three closest words in the mature vector space, together with their mean&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>In addition to the AoE features derived from word trajectories presented in Table 1, we also considered several word-level lexical features provided in Table 2. These features are frequently used to reflect word complexity and may impact the age at which a word is acquired (Nelson et al., [<reflink idref="bib65" id="ref122">65</reflink>]). Simple statistical measures, such as the number of syllables and the number of characters, are indicators of how difficult a word may be perceived, with longer words being acquired later in life. The number of synsets (i.e., word senses that are linked through synonym chains) from WordNet (Miller, [<reflink idref="bib60" id="ref123">60</reflink>]) can also impact an individual's ability to infer a word's meaning in a given context. Finally, the eccentricities of the word hyponym and hypernym trees derived from WordNet indicate the genericity of that concept. Words with high hypernym eccentricities are words that are very specific, whereas words with low hypernym eccentricities tend to be very generic terms that may be acquired earlier (e.g., the difference between learning the meaning of "bird" versus "owl"). Conversely, high hyponym eccentricities correspond to words having a wide semantic field, which means that their acquisition may occur later on (e.g., "lincoln" or "steed"). In addition, we also include the term frequency of words in the training corpus, normalized per 1 million terms, as term frequency was previously shown to be a strong determinant of acquisition (Brysbaert &amp; New, [<reflink idref="bib13" id="ref124">13</reflink>]).</p> <p>Table 2 Word-level features independent of the word2vec models</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;p&gt;Feature name&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Formula&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Description&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Average hyponym eccentricities&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mover xmlns=""&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mfenced close=")" open="("&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;word&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/math&gt;&lt;inline-graphic href="13428&amp;#95;2022&amp;#95;1797&amp;#95;Article&amp;#95;IEq4.gif" /&gt;&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The average eccentricity of the hyponym trees of the word. The eccentricity is the largest distance from the word to any of its hyponyms.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Average hypernym eccentricities&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mover xmlns=""&gt;&lt;mrow&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mfenced close=")" open="("&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;m&lt;/mi&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mrow&gt;&lt;mi mathvariant="italic"&gt;word&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/math&gt;&lt;inline-graphic href="13428&amp;#95;2022&amp;#95;1797&amp;#95;Article&amp;#95;IEq5.gif" /&gt;&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The average eccentricity of the hypernym trees of the word. The eccentricity is the largest distance from the word to any of its hypernyms.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;# synset&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;|synsets&amp;#95;word|&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The number of synonym sets of the word (i.e., word senses).&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;# syllables&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;|word&amp;#95;syllables|&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The number of syllables of the word.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;# chars&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;|word|&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;The length of the word expressed in number of characters.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term frequency &amp;#8211; Stage i&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;italic&gt;tf&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;i&lt;/italic&gt;&lt;/sub&gt;(&lt;italic&gt;word&lt;/italic&gt;)&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;Term frequency of the word in the training corpus, normalized per 1 million terms at a given intermediate stage.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term frequency - average&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mover xmlns=""&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;/math&gt;&lt;inline-graphic href="13428&amp;#95;2022&amp;#95;1797&amp;#95;Article&amp;#95;IEq6.gif" /&gt;&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;Average normalized term frequency for the intermediate models.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term frequency &amp;#8211; Standard Deviation&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;&amp;#963;[&lt;italic&gt;tf&lt;/italic&gt;&lt;sub&gt;&lt;italic&gt;i&lt;/italic&gt;&lt;/sub&gt;(&lt;italic&gt;word&lt;/italic&gt;)], &lt;italic&gt;i&lt;/italic&gt; = 1.. &lt;italic&gt;T&lt;/italic&gt;&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;Standard deviation of the normalized term frequency for the intermediate models.&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Some of these features, namely the word and syllable lengths, are also variables used in the Flesch Reading Ease equation. However, this does not introduce a circularity in our method due to a number of reasons. First, the Flesch Reading Ease uses the length information as aggregates together with sentence information, whereas the features our models use refer only to individual words. Second, even if there is a degree of overlap, these features are only used as word features for the regressors, while the Flesch Reading Ease will influence only the way the word2vec models are trained through the order in which they are shown the documents in the corpora.</p> <hd id="AN0160647162-14">Results</hd> <p></p> <hd id="AN0160647162-15">Internal reliability</hd> <p>The means, standard deviations, and internal reliability estimates are provided in Table 3 for the indices used within the analysis. A split-half correlation analysis was conducted to evaluate the internal reliability of the indices. We follow a linear growth scheme to generate two parallel datasets from our corpus. Each of these datasets contains half of the paragraphs described in the complete corpus, with paragraphs being assigned to only one of the two halves. This means that the two halves of the dataset are independent. All previous indices are recomputed for each half, thus resulting in two sets of indices per word. The split-half reliability for our feature building method uses the Spearman rank-order correlation between these two index observation vectors for all words in our vocabulary. All correlation coefficients from Table 3 are statistically significant (<emph>p &lt;</emph>.001) and denote high agreement.</p> <p>Table 3 Split-half reliability Spearman correlation coefficients</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;p&gt;Feature&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Mean&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Standard deviation&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Spearman correlation coefficient&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Average&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.50&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.29&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.94&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Inverse average&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.50&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.29&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.94&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Highest cosine word similarity&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.57&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.12&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.77&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Slope&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.09&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.037&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.62&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Inverse slope&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;12.78&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;30.54&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.59&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.30 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.43&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.43&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.86&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.35 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.53&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.43&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.85&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.40 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.67&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.45&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.85&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.45 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.86&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.49&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.85&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.50 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;3.11&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.56&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.86&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.55 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;3.42&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.65&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.86&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.60 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;3.81&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.76&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.87&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.65 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;4.27&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.85&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.88&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Continuous index above.70 threshold&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;4.81&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.89&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.89&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 1&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;37.87&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;451.06&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.80&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 2&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;203.18&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1713.13&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.93&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 3&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;393.38&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;3108.25&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.96&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 4&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;561.79&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;4271.31&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.97&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 5&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;722.88&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;5304.95&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.97&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 6&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;887.03&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;6244.37&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.98&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 7&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1065.98&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;7150.91&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.98&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 8&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1273.14&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;8141.60&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.98&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 9&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1488.82&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;9170.99&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.98&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Stage 10&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1638.18&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;9938.52&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.98&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Average&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;827.22&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;5487.43&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.98&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Term Frequency &amp;#8211; Standard Deviation&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;523.32&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;3118.65&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.98&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Word-vocabulary integration&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;709.60&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;744.61&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.79&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;2&lt;sup&gt;nd&lt;/sup&gt; highest cosine word similarity&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.53&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.11&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.77&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;3&lt;sup&gt;rd&lt;/sup&gt; highest cosine word similarity&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.51&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.11&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.77&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Top 3 cosine similarities&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.54&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.11&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.78&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 1&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.04&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.23&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.69&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 2&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.24&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.40&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.85&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 3&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.37&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.43&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.87&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 4&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.46&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.43&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.89&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 5&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.54&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.42&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.90&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 6&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.61&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.39&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.90&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 7&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.69&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.34&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.90&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 8&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.76&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.27&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.91&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Intermediate cosine similarity 9&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.81&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;0.20&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.92&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p> <emph>Note: all p</emph> &lt;.001</p> <hd id="AN0160647162-16">Multicollinearity analysis</hd> <p>We performed a thorough analysis of multicollinearities between the generated features to better understand their similarities and complementarities. This analysis of multicollinearity provides insight into the relationships between different features, highlighting linear dependencies between individual features used in follow-up models. First, Fig. 6 illustrates a heatmap of the linear relations between all pairs of features in terms of Pearson correlations. These are obtained from the features generated by the word2vec models trained on the corpus, sorted by readability.</p> <p>Graph: Fig. 6 Pearson correlation coefficient heatmap between word features</p> <p>We can observe that there are blocks of highly correlated features among the "continuous index above a threshold" type features; similarly, "intermediate cosine similarity" type features have lower, but still high, linear correlations between them and the term frequency features. Additionally, the average of the intermediate cosine similarities is correlated with the individual values, whereas the inverse of this mean is correlated with the "continuous index above a threshold" features. The highest, 2<sups>nd</sups> highest, 3<sups>rd</sups> highest, and the "average top3 cosine" features create a block of linearly correlated features. There are also correlations between the number of characters of a word and its number of syllables. These results are not surprising, given the definitions of these features. Interestingly, the first stage of training in all three correlation blocks (continuous index, intermediate cosine similarity, and term frequency) appears to have the lowest correlation to the rest of the features in the respective blocks. This may suggest that the first stage of training is the most distinct, a fact that may also be related to the significantly lower reliability for "Intermediate cosine similarity 1" and "Term Frequency – Stage 1" in Table 3, when compared to the counterpart features from later training stages.</p> <p>An automated method of measuring the degree of collinearity was required to remove features that suffer from multicollinearity. We employed the "variance inflation factor" (VIF; Craney &amp; Surles, [<reflink idref="bib16" id="ref125">16</reflink>]), which performs a linear regression on each feature using the other features as predictors. The associated VIF score for a variable is then given by the inverse of the complement of the coefficient of determination <emph>r</emph><sups>2</sups> measured for that feature. Craney and Surles ([<reflink idref="bib16" id="ref126">16</reflink>]) discuss the selection of a cutoff point for a feature's VIF value depending on the desired degree of tolerance. In our case, we opted to use a cutoff point of 5, corresponding to a variability of 80% in a feature being explained by the other features. Fifteen features remain using a VIF cutoff of 5. These relate to the Pearson correlation coefficient heatmap in Fig. 7 as the new correlation heatmap shows low linear relationships between the remaining features.</p> <p>Graph: Fig. 7 Pearson correlation coefficient heatmap after variance inflation factor analysis</p> <p>However, we found that performing a VIF analysis to select features for the nonlinear model leads to a loss in performance when it comes to the ability of the models to perform predictions of the Kuperman AoA scores. Because multicollinearity is a common issue that can impact linear models, we elected to utilize the traditional VIF cutoff of 5 for the linear regression model input features. For the other considered models, namely the support vector regressors and the random forest regressors, we did not perform a multicollinearity filtering on the entirety of the feature set because they are nonlinear models that should not be affected by the presence of multicollinearity in the input feature matrix. However, we did utilize the VIF analysis to minimize the redundancy found when it came to the "continuous index above a threshold" features. By using the same cutoff value of 5, we found that VIF retained only the first and last such indices, namely "continuous index above.3 threshold" and "continuous index above.7 threshold." Similarly, we reduced the term frequency features using VIF to four features: first and last stage term frequency normalized per 1 million words and the mean and average of the term frequencies per word. We elected to utilize this strategy of selectively applying VIF because of its simplicity and the robust results it provided, which slightly boosted the performance of the random forest and support vector regressors without a loss of performance. Random forests, in particular, have been shown to provide superb predictive accuracy, albeit at the cost of not being able to provide insight into the contributions of individual predictors and their interactions (Tomaschek et al., [<reflink idref="bib78" id="ref127">78</reflink>]). Given our primary goal of predicting AoA norms with high accuracy, we elected to favor performance over feature interpretability.</p> <hd id="AN0160647162-17">Predicting the Kuperman AoA scores</hd> <p>Our primary objective in this study is to predict the Kuperman AoA scores using the predictors described in Tables 1 and 2. The Kuperman AoA scores predicted by our AoE 2.0 model have a distribution (M = 10.98; SD = 3) that has a slight negative skew (−0.2) and a Pearson kurtosis of 2.62, which suggests a platykurtic distribution. The original Kuperman AoA scores form a relatively normal distribution (M = 11; SD = 3.04) with a similar negative skew (−0.2) and a Pearson kurtosis of 2.63.</p> <p>To predict these scores, a regressor was trained using the features derived from the cosine similarities. Results for the previously described scenarios are displayed in Table 4. The results are measured by averaging over a 10-fold cross-validation with random forest regression (Breiman, [<reflink idref="bib11" id="ref128">11</reflink>]), linear regression, and support vector regressor. We found that other models, such as multilayered perceptrons, produce weaker results and decision trees, and support vector regressors tend to outperform neural models for small-medium datasets. Moreover, decision trees and linear models are more interpretable than neural models. Additionally, we also measured the performance of the least-angle regression (LARS) models together with either the Bayes information or the Akaike information criterion for model selection in order to attempt to reduce the complexity of the resulting models (Zou et al., [<reflink idref="bib82" id="ref129">82</reflink>]). The mean absolutes errors (MAEs) and their normalized values from Table 4 are reported for the linear, support vector, LARS lasso, and random forest regressor models, as well as whether the texts were sorted by readability before they were split into subsets (i.e., whether simpler texts are seen first). For each experiment, we report the average result for a 10-fold cross validation when training a model using the features presented in Tables 1 and 2 to predict the AoA scores of words.</p> <p>Table 4 AoE v2 Kuperman AoA prediction results</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Texts sorted by readability&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Model&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;MAE&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Normalized MAE&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Normalized MAEstd. dev.&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;2&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Yes&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Linear regression&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.75&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.08&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00124&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.45&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Yes&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Lasso Lars (AIC)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.75&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.08&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00107&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.45&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Yes&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Lasso Lars (BIC)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.75&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.08&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00108&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.45&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;Yes&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;Random forest regressor&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;1.51&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.07&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00092&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.57&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Yes&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;SVR&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.58&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.08&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00127&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.54&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;No&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Linear regression&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.01&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.10&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00127&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.29&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;No&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Lasso Lars (AIC)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.01&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.10&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.00086&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.29&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;No&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Lasso Lars (BIC)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.01&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.10&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00105&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.29&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;No&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Random forest regressor&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.80&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.09&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00121&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.41&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;No&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;SVR&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.84&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.09&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.00117&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.39&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Bold marks the best results</p> <p>An ANOVA confirmed that there were significant differences between the models reported in Table 4, <emph>F</emph>(<reflink idref="bib2" id="ref130">2</reflink>, 48,<reflink idref="bib540" id="ref131">540</reflink>) = 1476, <emph>p</emph> &lt;.01, growth schemes, <emph>F</emph>(<reflink idref="bib1" id="ref132">1</reflink>, 26,<reflink idref="bib976" id="ref133">976</reflink>) = 34.5, <emph>p</emph> &lt;.01, and ordering schemes, <emph>F</emph>(<reflink idref="bib1" id="ref134">1</reflink>, 26,<reflink idref="bib976" id="ref135">976</reflink>) = 2594, <emph>p</emph> &lt;.01. Post hoc tests were conducted using a Bonferroni correction, and demonstrated that the random forest model had significantly less error than the linear regression model <emph>t</emph>(<reflink idref="bib107" id="ref136">107</reflink>,<reflink idref="bib907" id="ref137">907</reflink>) = 71.2, <emph>p</emph> &lt;.01, the SVR model <emph>t</emph>(<reflink idref="bib107" id="ref138">107</reflink>,<reflink idref="bib907" id="ref139">907</reflink>) = 22.5, <emph>p</emph> &lt;.01, the AIC model <emph>t</emph>(<reflink idref="bib107" id="ref140">107</reflink>, 907) = 62.1, <emph>p</emph> &lt;.01, and the BIC model <emph>t</emph>(<reflink idref="bib107" id="ref141">107</reflink>,<reflink idref="bib907" id="ref142">907</reflink>) = 62.4, <emph>p</emph> &lt;.01. Finally, the models that presorted text by readability had significantly less error than models that did not presort the texts, <emph>t</emph>(<reflink idref="bib269" id="ref143">269</reflink>,<reflink idref="bib770" id="ref144">770</reflink>) = 134.0, <emph>p</emph> &lt;.01. Thus, sorting the documents by their readability scores appears to improve performance.</p> <p>We also report the results obtained through using the features that were generated using the AoE 1.0 in Table 5. An ANOVA confirmed that there were significant differences between the AoE 1.0 models from Table 5, <emph>F</emph>(<reflink idref="bib2" id="ref145">2</reflink>, 33,<reflink idref="bib869" id="ref146">869</reflink>) = 779, <emph>p</emph> &lt;.01. Post hoc tests using a Bonferroni correction demonstrated that the SVR model, <emph>t</emph>(<reflink idref="bib21" id="ref147">21</reflink>, 997) = 38.7, <emph>p</emph> &lt;.01, and the random forest model, <emph>t</emph>(<reflink idref="bib21" id="ref148">21</reflink>, 197) = 23.1, <emph>p</emph> &lt;.01) resulted in significantly less error than the linear regression model.</p> <p>Table 5 Results for predicting Kuperman AoA scores using previous methods</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Model&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;MAE&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Normalized MAE&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;2&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;AoE 1.0 + Random forest regression&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.92&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.913&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.30&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;AoE 1.0 + Linear regression&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.02&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.096&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.23&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;AoE 1.0 + SVR&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;1.84&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.088&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.36&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Bold marks the best results</p> <p>Table 6 introduces a side-by-side comparison of the best AoE 1.0 and AoE 2.0 models, alongside the Word Maturity indices. Because the Word Maturity Index is a single numerical value, we report the squared correlation between the Word Maturity indices and the AoA word list in order to provide a direct comparison to the <emph>R</emph><sups>2</sups> of the other models (marked with "*" in the table). An ANOVA confirmed that there were significant differences between the models, <emph>F</emph>(<reflink idref="bib2" id="ref149">2</reflink>, 34,<reflink idref="bib599" id="ref150">599</reflink>) = 409, <emph>p</emph> &lt;.01. Post hoc tests using a Bonferroni correction demonstrated that the AoE v2 model had significantly less error than the Word Maturity model, <emph>t</emph>(<reflink idref="bib17" id="ref151">17</reflink>,<reflink idref="bib583" id="ref152">583</reflink>) = 12.2, <emph>p</emph> &lt;.01, and the AoE v1 model, <emph>t</emph>(<reflink idref="bib17" id="ref153">17</reflink>,<reflink idref="bib583" id="ref154">583</reflink>) = 28.6, <emph>p</emph> &lt;.01. Our method also marks a significant improvement over AoE 1.0, with <emph>R</emph><sups>2</sups> values improving by as much as.21, as well as having a higher correlation with the Kuperman AoA values than the Word Maturity Indices by up to.11.</p> <p>Table 6 Results for predicting Kuperman AoA using the best models available for each method</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Model&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;MAE&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Normalized MAE&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;2&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;AoE 2.0 + RF&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;1.51&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.070&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.57&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;AoE 1.0 + SVR&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.84&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.088&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.36&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Word Maturity indices&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;N/A&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;N/A&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.46&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Bold marks the best results</p> <p>In order to better understand how the regressor behaves for different AoA magnitudes, Fig. 8 introduces the scatterplots for a pair of training and test sets randomly sampled from the complete dataset. The blue line represents the ideal fit (i.e., the case where the predicted and the actual AoA scores are equal), while the orange line shows the observed linear fit between the predicted and the real AoA scores. Each data point has a hue that corresponds to the absolute error between the predicted score and the real score. These results suggest that the model produces reasonable approximations of the Kuperman AoA scores, with no perceivable bias towards over or under estimation. The <emph>R</emph><sups>2</sups> values indicate that our models are capable of explaining over half of the variance of the Kuperman AoA scores. Of note is that we predicted the mean AoA score for each word in the Kuperman word list, with each word in the study also having an associated standard deviation (SD = 3.04). This is introduced by the fact that the Kuperman AoA list, as well as other AoA lists, is generated by participants estimating the age at which they acquired a word and then aggregating these values. The variance in AoA scores introduces a level of noise in the distribution that limits the performance of models trained to predict such scores, which may explain to a certain degree the errors that the models produce.</p> <p>Graph: Fig. 8 Scatterplots showing the predicted AoA score and the Kuperman AoA score</p> <p>Figure 9 depicts feature importance within the random forest model in descending order, with higher importance reflecting features that are determined to be more relevant for splits in the constituent decision trees. The size of each horizontal bar is determined by the value of the feature importance. These are computed through the impurity decrease of each feature in the constituent trees of the random forest, for each decision node. For a given feature, the Pearson correlation between the feature and the Kuperman AoA is displayed, with the color of the bars corresponding to the absolute value of this correlation and their direction by the sign of the Pearson correlation. The features considered most important by the model are the continuous index values above various thresholds, which indicate that the model uses the inflection point at which an intermediate model generates a vector similar to the mature model. Other features with average importance values are the average and inverse average cosine similarities, various word-level statistical features such as the number of syllables, term frequency features, the number of hyponym eccentricities, the number of synsets to which it belongs, and the number of characters. Features that are deemed least useful by the model are the number of cosines above a certain threshold and specific cosine similarities at various intermediate stages (denoted as "intermediate cosine similarity X," where X is the selected stage), as well as the term frequency measured during the first training stage.</p> <p>Graph: Fig. 9 Importance of selected features within the best-performing random forest model</p> <p>The feature importance values are generated from the random forest regressor and represent the amount of error that each feature reduces across the decision trees that form the forest. Some of these values are intuitive, an example being that the number of synsets for a word is inversely correlated with its AoA score, meaning that words with more synonyms tend to have lower AoA scores and that words with higher term frequencies have lower AoA scores. However, the slope of the cosine similarities between the intermediate vector representations and the mature representation has a positive correlation with the AoA score, suggesting that words that have an initial lower cosine similarity tend to have higher acquisition ages. Continuous indices above certain thresholds are positively correlated with the AoA scores since they are a direct indicator of how many corpus subsets need to be used during the training of the word2vec model before the word representation becomes sufficiently aligned with the mature vector model. Negative correlations, which indicate lower AoA scores, include large hypernym/hyponym average eccentricities which indicate large semantic field tree hierarchies, large values of cosine similarities for the intermediate models, high inverse cosine similarity averages, inverse slopes, and normalized term frequencies.</p> <hd id="AN0160647162-18">Generalizing AoE 2.0 to predict other AoA word lists</hd> <p>Our method can easily be applied to alternative AoA word lists, for example: Bird (Bird et al., [<reflink idref="bib7" id="ref155">7</reflink>]), Bristol (Stadthagen-Gonzalez &amp; Davis, [<reflink idref="bib76" id="ref156">76</reflink>]), Cortese (Cortese &amp; Khanna, [<reflink idref="bib15" id="ref157">15</reflink>]) or Shock (Shock et al., [<reflink idref="bib74" id="ref158">74</reflink>]). In addition, we also attempt to model the objective AoA scores derived from children (Morrison et al., [<reflink idref="bib63" id="ref159">63</reflink>]). Training a random forest regressor and measuring the mean absolute error (MAE) across 10 folds, as before, we obtained the results from Table 7. The purpose of this experiment is to verify that the method is independent of the AoA word list used. Because the previous AoA metrics have differing scales, we also show the scatterplots to better illustrate the performance of the regressor. For each AoA word list, a different random forest model is trained using the same features that were used for predicting the Kuperman AoA scores.</p> <p>Table 7 Results for various AoA scores</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;p&gt;AoA metric&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Number of words in word list&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Normalized MAE of the random forest regressor&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;2&lt;/sup&gt; of the random forest regressor&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Bird&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1973&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.09&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.61&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Bristol&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;3274&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.08&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.66&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Cortese&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2816&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.06&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.74&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Shock&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2894&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.06&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.69&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Morrison - Objective AoA (75%) (months)&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;294&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.13&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.35&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>We can see that in the Bird case (see Fig. 10), the reduced number of terms leads to a scattered distribution of values in the test data. On average, the predictions do not appear to have a tendency of solely overestimating or underestimating the Bird scores. However, the model has a tendency to overestimate for low AoA scores, while underestimating high AoA scores.</p> <p>Graph: Fig. 10 Results for Bird score prediction</p> <p>In the case of the Bristol scores (see Fig. 11), the errors appear to be less scattered (corresponding to the higher <emph>R</emph><sups>2</sups> coefficient). Additionally, the errors do not appear to indicate the same degree of overestimation for low AoA scores and underestimation for high AoA scores observed in the case of the Bird word list.</p> <p>Graph: Fig. 11 Results for Bristol score prediction</p> <p>For the Cortese AoA and Shock datasets (see Figs. 12 and 13, respectively), the observed linear fit and the ideal fit have high overlap, confirming the high <emph>R</emph><sups>2</sups> coefficients of.74 and.69. We observe little bias towards overestimation or underestimation for these two datasets, albeit with some significant outliers. By contrast, the model resulted in substantially reduced performance for the objective AoA scores by Morrison et al., [<reflink idref="bib63" id="ref160">63</reflink>] (see Fig. 14; <emph>r</emph><sups>2</sups> =.35). The limited number of terms (i.e., almost an order of magnitude lower than the other word lists) in the latter dataset resulted in the model having a significantly reduced performance in comparison to the other experiments. Nonetheless, collectively, the results indicate that the AoE v2 model generalizes well to a wide variety of AoA scores. However, this method may only be applicable to AoA lists that comprise at least a few thousand words.</p> <p>Graph: Fig. 12 Results for Cortese score prediction</p> <p>Graph: Fig. 13 Results for Shock score prediction</p> <p>Graph: Fig. 14 Results for Morrison Objective AoA score prediction</p> <hd id="AN0160647162-19">Observing the evolution of AoE cosine similarities</hd> <p>Figure 15 illustrates the alignment discrepancy between the intermediate models at each stage and the mature model, as described by the cosine similarity between the intermediate vector representation of a word and its mature vector representation. The same words from Dascalu et al. ([<reflink idref="bib20" id="ref161">20</reflink>]) were selected. A cool-to-warm color gradient is used for each word based on its frequency (i.e., the rarest word is "singularity" and the most frequent word is "class," with "chocolate" having the average frequency between the seven selected words). Plots representing the words have an upward tendency, with the cosine similarity rapidly approaching 1, a value which signifies perfect overlap. These cosines are computed after the intermediate vector space was rotated to align it with the mature vector space. Of the seven words shown, there is a clear differentiation of more specialized words, such as "singularity," "clustering," and "virus" and the more common words "chocolate," "happy," "tech," and "class." Slopes tend to be smaller when the intermediate word representation is more similar to the one produced by the mature model. Overall, this figure illustrates that frequent words tend to have higher cosine similarities in early stages in comparison to rarer terms.</p> <p>Graph: Fig. 15 Word vector evolution over the 10 word2vec models</p> <hd id="AN0160647162-20">Correlation with text difficulty</hd> <p>In order to evaluate the potential utility of the predicted scores, we performed an experiment to measure the correlation between the generated AoE scores and text difficulty. To this end, we analyzed the StairStepper corpus (Balyan et al., [<reflink idref="bib3" id="ref162">3</reflink>]; Perret et al., [<reflink idref="bib67" id="ref163">67</reflink>]), which contains 162 expository texts rated in difficulty ranging from grade 1 to grade 12. The predicted AoE scores generated by the best-performing models from Table 5 for both AoE 1.0 and AoE 2.0, as well as the original Kuperman AoA scores and Word Maturity, were evaluated. In order to provide a single measurement of text difficulty, we aggregated the scores for the words of that text using a simple average, and then measured the Spearman rank correlation between the predicted scores of the 162 texts and their marked grades.</p> <p>The results reported in Table 8 indicate that the scores predicted by our method outperform the Word Maturity and AoE 1.0 indices, and even the Kuperman scores themselves. The results also show that the initial version of AoE had a substantially lower correlation with text difficulty. While it is perhaps unexpected for the predicted scores to achieve higher performance than the reference scores, the latter result may be due to the larger vocabulary being covered by AoE 2.0.</p> <p>Table 8 Correlation measurements with StairStepper text difficulty</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Word list&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Spearman rank correlations&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;AoE 1.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.21&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;AoE 2.0&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&lt;bold&gt;.84&lt;/bold&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Kuperman&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.76&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Word Maturity&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;.75&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Bold marks the best results</p> <hd id="AN0160647162-21">Impact of lexical, WordNet, and word trajectory features</hd> <p>In addition to the features derived from the exposure trajectories, we also added several features that are generated using external resources, such as WordNet (Miller, [<reflink idref="bib60" id="ref164">60</reflink>]). Because the number of syllables, the hyponym/hypernym trees, and the number of synsets depend on language-specific implementations, they may not be available in all cases. With this in mind, we performed an ablation study on the features used to predict the Kuperman AoA word list by incrementally adding features starting from the simple lexical features and then adding WordNet features and, finally, the word trajectory features. We performed this study on our best-performing experiment, namely the one in which the texts were sorted in increasing order of readability, and reported the results for the random forest regressor.</p> <p>The results presented in Table 9 show that the lexical features and the word trajectory features alone capture a significant amount of the Kuperman AoA scores variance, with the combination between lexical and word trajectory features resulting in a significant increase in performance. Additionally, the absence of the WordNet features has a negligible impact on the overall performance. As such, our method can be applied even in contexts wherein certain external word feature resources, such as those reported by WordNet, are not available—with minimal loss of accuracy.</p> <p>Table 9 Ablation study on AoE v.2 prediction features</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;p&gt;Feature set&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;MAE&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Normalized MAE&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;2&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Lexical features&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.612&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.077&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.513&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;WordNet features&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;2.225&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.106&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.126&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Word trajectory features&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.667&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.079&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.495&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Lexical + word trajectory features&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.558&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.074&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.545&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;All features&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.513&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.572&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0160647162-22">Discussion</hd> <p>The objective of this study was to generate AoA scores using a novel AoE 2.0 model in order to reduce the cost and effort to produce human AoA scores, increase the total number of words available in AoA databases, and reduce the error of adult-based corpora by simulating the word learning process. AoE 2.0 was able to model AoA word lists with <emph>R</emph><sups>2</sups> coefficients ranging from.35 for the objective, child derived, AoA norms (Morrison et al., [<reflink idref="bib63" id="ref165">63</reflink>]),.57 for the Kuperman AoA list (Kuperman et al., [<reflink idref="bib44" id="ref166">44</reflink>]), and up to.74 for the Cortese AoA norm (Cortese &amp; Khanna, [<reflink idref="bib15" id="ref167">15</reflink>]). Methodologically, we explored the impact of ordering the documents used to train the word vector models in AoE by using the Flesch Reading Ease score. Our findings suggest that sorting the texts in ascending order of their readability has a significant positive impact on the performance of the model. Despite the influence of ordering documents in the corpus, the results with a simple linear growth and an unsorted corpus are still relatively accurate, which means that good estimations of AoA can be obtained without the need of calculating readability scores.</p> <p>AoE 2.0 directly links the word trajectories described in Word Maturity and AoE 1.0 to various AoA scores by generating features from the word trajectories that are then used to train a regressor capable of predicting AoA values for words with low error. This leads to a more straightforward use case than previous work: our method produces scores that have direct interpretation, as given by AoA, and may be used to generalize to words for which studies were not previously performed. The addition of non-AoE lexical features was shown to increase the ability of our method to model human AoA ratings, with our analysis suggesting that our method can perform well even when omitting features that rely on external resources, such as WordNet. While the inclusion of these features can be seen as a step back from the original purpose of AoE models, we believe that the ability of our method to simulate human AoA ratings with low error can be of great use in extending existing AoA word lists with ratings for new words that align well with human ratings traditionally used in AoA.</p> <p>Moreover, we have investigated the influence that the ordering of documents has on the performance of AoE. We have found that simulating human word exposure by using increasingly more complex texts, as approximated using the Flesch Reading Ease, results in a significant increase in performance. One explanation relates to how humans are exposed to words. When adults speak to children, they typically use child-directed speech, which is simpler and more repetitive than adult-directed speech (Hills, [<reflink idref="bib34" id="ref168">34</reflink>]). Once children have been exposed to simpler words and sentences in their early childhood, they can construct a vocabulary (Hills et al., [<reflink idref="bib35" id="ref169">35</reflink>]) which affords them the ability to understand more complex words and sentences in adult-directed speech (Hills, [<reflink idref="bib34" id="ref170">34</reflink>]). By sorting the texts by readability, the model mimics a potential trajectory of exposure to simple words and sentences, followed by exposure to difficult words and sentences.</p> <p>AoE 2.0 may also be used for downstream tasks, such as measuring the complexity of a text. In our experiments, we have shown AoE 2.0 scores to be a more accurate predictor of textual complexity than Word Maturity and AoE 1.0. In fact, the scores that our method generates were found to have better correlations to the human textual complexity scores in the StairStepper corpus than to the Kuperman scores that they were initially trained to model. The increased accuracy, in comparison with previous methods, has implications for the creation and evaluation of educational materials. Students may benefit from materials that target their zone of proximal development (Shabani et al., [<reflink idref="bib73" id="ref171">73</reflink>]). In the context of reading and text complexity, more accurate AoA scores can potentially be used to match students to appropriate texts by providing a more accurate assessment of when students know, or are able to learn, the words in the text.</p> <p>There are several limitations to our method that should be noted. First, we did not check the accuracy of spaCy's POS taggers on the documents in our corpora, which may present issues especially for the CHILDES data. Second, the task of training predictors to infer AoA scores falls under the category of semantic norm extrapolation (Snefjella &amp; Blank, [<reflink idref="bib75" id="ref172">75</reflink>]) and could be interpreted as a missing data problem. Under this hypothesis, the usage of the inferred AoE scores could be subject to biases that would make them not interchangeable with the empirically obtained AoA scores. Nevertheless, our method shows clear improvements over the existing AoE method and the conducted text difficulty experiment indicates that the generated word scores are at least comparable in practice with their training AoA word list. Another limitation of our method is that it relies on an existing AoA word list to infer scores for words in the vocabulary that the list does not cover. When extending our method to other languages, the availability of AoA word lists, as well as their sizes and distributions, may cause issues. Finally, our method requires the training of word embedding models on increasingly larger subsets of a corpora. While word2vec is an efficient algorithm for such cases, the computational cost is still significant, and the availability of corpora that have sufficient diversity in text difficulty can also be an issue for other languages. However, word2vec can handle large datasets better than the LDA and LSA used in AoE 1.0 and WM because of its small memory footprint and iterative training procedure. Additionally, we assume that the use of alternative word embedding algorithms may improve the quality of the exposure trajectories and should be considered in future research.</p> <hd id="AN0160647162-23">Conclusions and future research directions</hd> <p>AoE 2.0 offers an automated way of generating AoA scores that significantly reduce the cost and effort in collecting human AoA scores. AoE 2.0 may also reduce error in AoA scores by simulating the word learning process through increasing word exposure instead of relying on adults' recollection of word learning. We show the AoA norms generated by the AoE 2.0 model can be used to predict a wide array of readability ratings. Our experiments strongly suggest that our method can be applied to different word lists with consistently good performance.</p> <p>Our AoE 2.0 takes advantage of many of the aspects outlined by Word Maturity and Age of Exposure 1.0. The word trajectories described by Landauer et al. ([<reflink idref="bib46" id="ref173">46</reflink>]) are modeled through word2vec instead of LSA which, in addition to being more computationally efficient, also has the advantage of generating word embeddings that better capture the relations between words. In contrast to the initial Age of Exposure model, the new version avoids the need for determining the optimal number of LDA topics, with word2vec being notably more robust to changes in vector dimensionalities. In addition, graph-based topic alignment via bipartite matching using Jensen-Shannon dissimilarity and max-flow algorithms is replaced with the Procrustes rotation used in Word Maturity, which is more straightforward and more robust because it does not rely on the distribution similarities measured using Jensen-Shannon dissimilarity.</p> <p>Because AoE 2.0 is flexible, accurate, and scalable, multilingual comparisons become possible and the AoE 2.0 approach could be used to generate scores and word trajectories to compare words across languages. The scores generated by our model and their corresponding features may be compared with equivalents in other languages in order to determine which terms are acquired later on in some languages. For example, an initial study of using AoE 2.0 for modeling multiple non-English languages (Botarleanu et al., [<reflink idref="bib9" id="ref174">9</reflink>]) shows that the word exposure trajectories are able to reflect some aspects of human word acquisition among various languages. Thus, specialized curricula suitable for particular language speakers may be constructed using AoE 2.0.</p> <p>Additionally, this method may be applied to other languages to investigate the way words evolve in a similar or dissimilar manner across languages. This may be achieved by training unsorted word embedding models on iterative corpora for different languages, and then analyzing the way equivalent words evolve across translations. Through this, the generation of AoA scores for new languages may be possible, thus offering a multilingual metric that does not require human crowdsourcing efforts. In addition, more specialized documents may be used to model the evolutions of domain-specific words by analyzing the word embedding trajectories generated by the iterative training process specific to the AoE and WM models, thus allowing for the evaluation of how different registers and genres as reflected in various corpora may impact a language user's ability to acquire various terms. Given the reliance of AoE 2.0 on data gathered from human estimations and child-based ratings, there is likely a series of optimal points at which the switch can be made from child-based ratings to adult estimates and then to automatically generated AoA scores when constructing a new AoA word list for a language. This constitutes a possible avenue for future research, in order to find the number of words needed that both minimize the data gathering costs, as well as provide data of sufficient quantity and quality for AoE 2.0 to take over and generate AoA scores for the entire vocabulary.</p> <p>A notable benefit of AoE 2.0 is that its models are available for public use (cf., the Word Maturity model was not publicly available at the time of our experiments). Our AoE 2.0 model is released as an open-source project available at: https://github.com/readerbench/Age-of-Exposure. The AoE lexical scores based on the random forest regressor model are also available within the repository.</p> <p>In conclusion, measuring AoA is important to research on language and discourse, including but not limited to lexical decisions, text comprehension, and writing quality. Our objective is to provide a means to estimate AoA scores with a publicly available and flexible model. We believe that this method of modeling word evolutions can have a variety of applications, such as helping to design better learning materials in both English and multilingual settings, improving the performance of downstream tasks such as measuring textual complexity and expanding existing AoA lists of limited size through model inference.</p> <hd id="AN0160647162-24">Acknowledgements</hd> <p>This research was supported by a grant of the Romanian National Authority for Scientific Research and Innovation, CNCS – UEFISCDI, project number TE 70 <emph>PN-III-P1-1.1-TE-2019-2209,</emph> ATES – "Automated Text Evaluation and Simplification," the Institute of Education Sciences (R305A180144 and R305A180261), and the Office of Naval Research (N00014-17-1-2300; N00014-20-1-2623). The opinions expressed are those of the authors and do not represent views of the IES or ONR. We would also like to thank Prof. Peter Foltz for providing the Word Maturity indices that were used as a baseline in this paper.</p> <hd id="AN0160647162-25">Appendix</hd> <p></p> <hd id="AN0160647162-26">Appendix 1. Impact of the Growth Scheme</hd> <p>In order to analyze the way in which the growth scheme of the corpus impacts the performance of the models we propose two growth functions: a linear and an exponential growth scheme (see Appendix Fig. 16). The former describes a constant rate of increase in the size of the corpus to which the model is exposed, while the latter describes an exponential increase.</p> <p>Graph: Fig. 16 Word complexity dataset growth schemes</p> <p>Both linear vocabulary growth and exponential vocabulary growth were considered because children learn words at variable rates. Previous research suggests a "vocabulary spurt," a period of exponential word growth beginning when children are 18–24 months old (Goldfield &amp; Reznick, [<reflink idref="bib31" id="ref175">31</reflink>]; Robinson &amp; Mervis, [<reflink idref="bib68" id="ref176">68</reflink>]). In contrast, other research has reported that children demonstrate variable word learning rates (Bates et al., [<reflink idref="bib5" id="ref177">5</reflink>]) and re-examination of the vocabulary spurt has demonstrated that not all children experience the same exponential growth (Ganger &amp; Brent, [<reflink idref="bib28" id="ref178">28</reflink>]). It is now widely accepted that the word learning rate is variable and depends on the individual differences of the learner (Fernald &amp; Marchman, [<reflink idref="bib25" id="ref179">25</reflink>]; Rowe et al., [<reflink idref="bib69" id="ref180">69</reflink>]).</p> <p>We split the dataset into 10 stages (9 intermediary models and the last, complete and mature model). The corpus sizes are cumulative and can be described mathematically with the formulas described below. The linear growth scheme can be described as:</p> <p>9 <ephtml> &lt;math display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfenced close="}" open="{"&gt;&lt;msub&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mfrac&gt;&lt;mfenced close="|" open="|"&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/mfenced&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8943;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mrow&gt;&lt;mfenced close=")" open="("&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mfenced&gt;&lt;mfrac&gt;&lt;mfenced close="|" open="|"&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/mfenced&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width="4pt" /&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;/mrow&gt;&lt;mfenced close="}" open="{"&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Graph</p> <p>where <emph>D</emph><subs><emph>t</emph></subs> is the subset at step <emph>t, C</emph> is the corpus and <emph>T</emph> represents the number of desired sets within the corpus. In our experiments, we use 10 stages: 9 intermediate models and 1 mature model, which is trained on the entire dataset.</p> <p>Similarly, for the exponential growth function, the subset at a timestep <emph>D</emph><subs><emph>t</emph></subs> can be described by the following equation:</p> <p>10 <ephtml> &lt;math display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfenced close="}" open="{"&gt;&lt;msub&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mfenced close="|" open="|"&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/mfenced&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mfrac&gt;&lt;mo&gt;*&lt;/mo&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mn&gt;10&lt;/mn&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo stretchy="false"&gt;/&lt;/mo&gt;&lt;mfenced close=")" open="("&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8943;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mfenced close="|" open="|"&gt;&lt;mi&gt;C&lt;/mi&gt;&lt;/mfenced&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mfrac&gt;&lt;mo&gt;*&lt;/mo&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mn&gt;10&lt;/mn&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mfenced close=")" open="("&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mfenced&gt;&lt;mo stretchy="false"&gt;/&lt;/mo&gt;&lt;mfenced close=")" open="("&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width="4pt" /&gt;&lt;mo stretchy="false"&gt;|&lt;/mo&gt;&lt;mspace width="4pt" /&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;/mrow&gt;&lt;mfenced close="}" open="{"&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mfenced&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Graph</p> <p>Training the best-performing models with the different growth schemes indicates that there is little difference in performance between the two methods (see Appendix Table 10). In addition, post hoc tests showed that exponential growth had significantly less error than linear growth,* <emph>t</emph>(<reflink idref="bib269" id="ref181">269</reflink>,<reflink idref="bib770" id="ref182">770</reflink>) = 8.96, <emph>p</emph> &lt; 0.01.</p> <p>Table 10 AoE v2 Kuperman AoA prediction results – growth schemes</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;p&gt;Growth scheme&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Model&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;MAE&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Normalized MAE&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;2&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Linear&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;Random forest regressor&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.513&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.571&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;Exponential&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;Random forest regressor&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.510&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.575&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0160647162-27">Appendix 2. Impact of Word Embedding Size</hd> <p>In order to better understand how the word2vec model configuration affects the performance of the AoA regressors, we performed an experiment on one of the more important hyperparameters of word2vec: the size of the word embeddings. Results from Appendix Table 11 indicate that the vector size has a minimal impact on the final performance for Kuperman AoA prediction. The default value, 300, appears to be an effective choice. All results are reported using the best-performing model: random forest regressor with sorted data. Given the increase in computational cost from a vector size of 300 to a vector size of 1000, the improvement in performance is marginal. As such, we elected to use the default vector size of 300.</p> <p>Table 11 AoE v2 Kuperman AoA prediction results – word2vec vector size</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;p&gt;Vector size&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;MAE&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;Normalized MAE&lt;/p&gt;&lt;/th&gt;&lt;th&gt;&lt;p&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;2&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;100&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.514&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.570&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;200&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.515&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.572&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;300 (default)&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.513&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.571&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;400&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.515&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.573&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;500&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.509&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.573&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;p&gt;1000&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;1.512&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.072&lt;/p&gt;&lt;/td&gt;&lt;td&gt;&lt;p&gt;.571&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>The data and materials for all experiments are available at https://github.com/readerbench/Age-of-Exposure and the experiment was not preregistered.</p> <hd id="AN0160647162-28">Publisher's note</hd> <p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p> <ref id="AN0160647162-29"> <title> References </title> <blist> <bibl id="bib1" idref="ref4" type="bt">1</bibl> <bibtext> Alonso MA, Fernandez A, Díez E. Subjective age-of-acquisition norms for 7,039 Spanish words. Behavior Research Methods. 2015; 47; 1: 268-274. 10.3758/s13428-014-0454-2</bibtext> </blist> <blist> <bibl id="bib2" idref="ref10" type="bt">2</bibl> <bibtext> Álvarez B, Cuetos F. Objective age of acquisition norms for a set of 328 words in Spanish. Behavior Research Methods. 2007; 39; 3: 377-383. 10.3758/BF03193006</bibtext> </blist> <blist> <bibl id="bib3" idref="ref162" type="bt">3</bibl> <bibtext> Balyan R, McCarthy KS, McNamara DS. Applying natural language processing and hierarchical machine learning approaches to text difficulty classification. International Journal of Artificial Intelligence in Education. 2020; 30; 3: 337-370. 10.1007/s40593-020-00201-7</bibtext> </blist> <blist> <bibl id="bib4" idref="ref88" type="bt">4</bibl> <bibtext> Baroni, M, Dinu, G, &amp; Kruszewski, G. (2014). Don't count, predict! A systematic comparison of context-counting vs. context-predicting semantic vectors. In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), (pp. 238-247).</bibtext> </blist> <blist> <bibl id="bib5" idref="ref177" type="bt">5</bibl> <bibtext> Bates E, Dale PS, Thal D. Individual differences and their implications for theories of language development. The Handbook of Child Language. 1995; 30: 96-151</bibtext> </blist> <blist> <bibl id="bib6" idref="ref57" type="bt">6</bibl> <bibtext> Biemiller A, Rosenstein M, Sparks R, Landauer TK, Foltz PW. Models of vocabulary acquisition: Direct tests and text-derived simulations of vocabulary growth. Scientific Studies of Reading. 2014; 18; 2: 130-154. 10.1080/10888438.2013.821992</bibtext> </blist> <blist> <bibl id="bib7" idref="ref22" type="bt">7</bibl> <bibtext> Bird H, Franklin S, Howard D. Age of acquisition and imageability ratings for a large set of words, including verbs and function words. Behavior Research Methods, Instruments, &amp; Computers. 2001; 33; 1: 73-79. 10.3758/BF03195349</bibtext> </blist> <blist> <bibl id="bib8" idref="ref83" type="bt">8</bibl> <bibtext> Blei DM, Ng AY, Jordan MI. Latent Dirichlet Allocation. Journal of Machine Learning Research. 2003; 3; 4-5: 993-1022</bibtext> </blist> <blist> <bibl id="bib9" idref="ref174" type="bt">9</bibl> <bibtext> Botarleanu, R.-M, Dascalu, M, Watanabe, M, McNamara, D. S, &amp; Crossley, S. A. (2021). Multilingual age of exposure. In 22nd International Conference on Artificial Intelligence in Education (AIED 2021). Utrecht, Netherlands (Online).</bibtext> </blist> <blist> <bibtext> Braginsky, M, Yurovsky, D, Marchman, V. A, &amp; Frank, M. (2016). From uh-oh to tomorrow: Predicting age of acquisition for early words across languages. In: Proceedings of the 38th Annual Conference of the Cognitive Science Society. Philadelphia.</bibtext> </blist> <blist> <bibtext> Breiman L. Random forests. Machine Learning. 2001; 45; 1: 5-32. 10.1023/A:1010933404324</bibtext> </blist> <blist> <bibtext> Brysbaert M, Biemiller A. Test-based age-of-acquisition norms for 44 thousand English word meanings. Behavior Research Methods. 2017; 49; 4: 1520-1523. 10.3758/s13428-016-0811-4</bibtext> </blist> <blist> <bibtext> Brysbaert, M, &amp; New, B. (2009). Moving beyond Kučera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English. Behavior Research Methods, 41(4), 977-990.</bibtext> </blist> <blist> <bibtext> Chalard M, Bonin P, Méot A, Boyer B, Fayol M. Objective age-of-acquisition (AoA) norms for a set of 230 object names in French: Relationships with psycholinguistic variables, the English data from Morrison et al. (1997), and naming latencies. European Journal of Cognitive Psychology. 2003; 15; 2: 209-245. 10.1080/09541440244000076</bibtext> </blist> <blist> <bibtext> Cortese MJ, Khanna MM. Age of acquisition ratings for 3,000 monosyllabic words. Behavior Research Methods. 2008; 40; 3: 791-794. 10.3758/BRM.40.3.791</bibtext> </blist> <blist> <bibtext> Craney TA, Surles JG. Model-dependent variance inflation factor cutoff values. Quality Engineering. 2002; 14; 3: 391-403. 10.1081/QEN-120001878</bibtext> </blist> <blist> <bibtext> Crossley SA, McNamara DS. Understanding expert ratings of essay quality: Coh-Metrix analyses of first and second language writing. International Journal of Continuing Engineering Education and Life Long Learning. 2011; 21; 2-3: 170-191. 10.1504/IJCEELL.2011.040197</bibtext> </blist> <blist> <bibtext> Crossley, S, Feng, S, Cai, Z, &amp; McNamara, D. S. (2013). Computer simulations of MRC Psycholinguistic Database word properties: Concreteness, familiarity, imageability. In S. Jarvis, &amp; M. Daller (Eds.), Vocabulary Knowledge: Human Ratings and Automated Measures. (pp. 135-156). John Benjamins Publishing Company</bibtext> </blist> <blist> <bibtext> Crossley SA, Skalicky S, Dascalu M, McNamara DS, Kyle K. Predicting text comprehension, processing, and familiarity in adult readers: New approaches to readability formulas. Discourse Processes. 2017; 54; 5-6: 340-359. 10.1080/0163853X.2017.1296264</bibtext> </blist> <blist> <bibtext> Dascalu M, McNamara DS, Crossley SA, Trausan-Matu S. Age of Exposure: A Model of Word Learning. 30th AAAI Conference on Artificial Intelligence. 2015; AAAI Press: 2928-2934</bibtext> </blist> <blist> <bibtext> Davies, M. (2008). The Corpus of Contemporary American English (COCA). Available online at https://<ulink href="http://www.english-corpora.org/coca/">www.english-corpora.org/coca/</ulink>. Accessed 10 Jan 2022.</bibtext> </blist> <blist> <bibtext> Devlin, J, Chang, M. W, Lee, K, &amp; Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171-4186). ACL.</bibtext> </blist> <blist> <bibtext> Di Carlo, V, Bianchi, F, &amp; Palmonari, M. (2019). Training temporal word embeddings with a compass. In: Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 33, No. 01, pp. 6326-6334). AAAI Press.</bibtext> </blist> <blist> <bibtext> Eckerth J, Tavakoli P. The effects of word exposure frequency and elaboration of word processing on incidental L2 vocabulary acquisition through reading. Language Teaching Research. 2012; 16; 2: 227-252. 10.1177/1362168811431377</bibtext> </blist> <blist> <bibtext> Fernald A, Marchman VA. Individual differences in lexical processing at 18 months predict vocabulary growth in typically developing and late-talking toddlers. Child Development. 2012; 83; 1: 203-222. 10.1111/j.1467-8624.2011.01692.x</bibtext> </blist> <blist> <bibtext> Flesch R. A new readability yardstick. Journal of Applied Psychology. 1948; 32; 3: 221-233. 10.1037/h0057532</bibtext> </blist> <blist> <bibtext> Frank MC, Braginsky M, Yurovsky D, Marchman VA. Wordbank: An open repository for developmental vocabulary data. Journal of Child Language. 2017; 44; 3: 677-694. 10.1017/S0305000916000209</bibtext> </blist> <blist> <bibtext> Ganger J, Brent MR. Reexamining the vocabulary spurt. Developmental Psychology. 2004; 40; 4: 621. 10.1037/0012-1649.40.4.621</bibtext> </blist> <blist> <bibtext> Ghyselinck M, Lewis MB, Brysbaert M. Age of acquisition and the cumulative-frequency hypothesis: A review of the literature and a new multi-task investigation. Acta Psychologica. 2004; 115; 1: 43-67. 10.1016/j.actpsy.2003.11.002</bibtext> </blist> <blist> <bibtext> Gilhooly, K. J, &amp; Logie, R. H. (1980). Age-of-acquisition, imagery, concreteness, familiarity, and ambiguity measures for 1,944 words. Behavior Research Methods &amp; Instrumentation, 12(4), 395-427.</bibtext> </blist> <blist> <bibtext> Goldfield BA, Reznick JS. Early lexical acquisition: Rate, content, and the vocabulary spurt. Journal of Child Language. 1990; 17; 1: 171-183. 10.1017/S0305000900013167</bibtext> </blist> <blist> <bibtext> Gower JC. Generalized procrustes analysis. Psychometrika. 1975; 40; 1: 33-51. 10.1007/BF02291478</bibtext> </blist> <blist> <bibtext> Grigoriev A, Oshhepkov I. Objective age of acquisition norms for a set of 286 words in Russian: Relationships with other psycholinguistic variables. Behavior Research Methods. 2013; 45; 4: 1208-1217. 10.3758/s13428-013-0319-0</bibtext> </blist> <blist> <bibtext> Hills T. The company that words keep: comparing the statistical structure of child-versus adult-directed language. Journal of Child Language. 2013; 40; 3: 586-604. 10.1017/S0305000912000165</bibtext> </blist> <blist> <bibtext> Hills TT, Maouene J, Riordan B, Smith LB. The associative structure of language: Contextual diversity in early word learning. Journal of Memory and Language. 2010; 63; 3: 259-273. 10.1016/j.jml.2010.06.002</bibtext> </blist> <blist> <bibtext> Hoff E. The specificity of environmental influence: Socioeconomic status affects early vocabulary development via maternal speech. Child Development. 2003; 74: 1368-1378. 10.1111/1467-8624.00612</bibtext> </blist> <blist> <bibtext> Hoff E, Naigles L. How children use input to acquire a lexicon. Child Development. 2002; 73; 2: 418-433. 10.1111/1467-8624.00415</bibtext> </blist> <blist> <bibtext> Ivens, S. H, &amp; Koslin, B. L. (1991). Demands for reading literacy require new accountability methods. Touchstone Applied Science Associates.</bibtext> </blist> <blist> <bibtext> Johnston RA, Barry C. Age of acquisition and lexical processing. Visual Cognition. 2006; 13; 7-8: 789-845. 10.1080/13506280544000066</bibtext> </blist> <blist> <bibtext> Justice LM, Petscher Y, Schatschneider C, Mashburn A. Peer effects in preschool classrooms: Is children's language growth associated with their classmates' skills?. Child Development. 2011; 82; 6: 1768-1777. 10.1111/j.1467-8624.2011.01665.x</bibtext> </blist> <blist> <bibtext> Kaufman, A.S, &amp; Kaufman, N.L. (1983). Kaufman Assessment Battery for Children. Circle Pines, MN: American Guidance Service.</bibtext> </blist> <blist> <bibtext> Kaufman AS, Kaufman NL. Kaufman Brief Intelligence Test. 1990; Pearson, Inc.</bibtext> </blist> <blist> <bibtext> Krzanowski, W. J. (2000). Principles of Multivariate Analysis, Revised Edition. Oxford University Press.</bibtext> </blist> <blist> <bibtext> Kuperman V, Stadthagen-Gonzalez H, Brysbaert M. Age-of-acquisition ratings for 30,000 English words. Behavior Research Methods. 2012; 44; 4: 978-990. 10.3758/s13428-012-0210-4</bibtext> </blist> <blist> <bibtext> Landauer TK, Dumais ST. A solution to Plato's problem: the Latent Semantic Analysis theory of acquisition, induction and representation of knowledge. Psychological Review. 1997; 104; 2: 211-240. 10.1037/0033-295X.104.2.211</bibtext> </blist> <blist> <bibtext> Landauer TK, Kireyev K, Panaccione C. Word maturity: A new metric for word knowledge. Scientific Studies of Reading. 2011; 15; 1: 92-108. 10.1080/10888438.2011.536130</bibtext> </blist> <blist> <bibtext> Lenci, A, Sahlgren, M, Jeuniaux, P, Gyllensten, A. C, &amp; Miliani, M. (2021). A comprehensive comparative evaluation and analysis of Distributional Semantic Models. arXiv. 2105.09825.</bibtext> </blist> <blist> <bibtext> Levy, O, Goldberg, Y, &amp; Dagan, I. (2015). Improving distributional similarity with lessons learned from word embeddings. Transactions of the Association for Computational Linguistics, 3, 211-225.</bibtext> </blist> <blist> <bibtext> Li, J, &amp; Jurafsky, D. (2015). Do Multi-Sense Embeddings Improve Natural Language Understanding? In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (pp. 1722-1732). ACL.</bibtext> </blist> <blist> <bibtext> Lund, K, &amp; Burgess, C. (1996) Producing high-dimensional semantic spaces from lexical co-occurrence. Behavior Research Methods, Instruments &amp; Computers 28(2) 203-208.</bibtext> </blist> <blist> <bibtext> Łuniewska M, Haman E, Armon-Lotem S. Ratings of age of acquisition of 299 words across 25 languages: Is there a cross-linguistic order of words?. Behavior Research Methods. 2016; 48; 3: 1154-1177. 10.3758/s13428-015-0636-6</bibtext> </blist> <blist> <bibtext> MacWhinney B. The CHILDES system. American Journal of Speech-Language Pathology. 1996; 5; 1: 5-14. 10.1044/1058-0360.0501.05</bibtext> </blist> <blist> <bibtext> Maddux CD. Peabody Picture Vocabulary Test (PPVT-III). Diagnostique. 1999; 24; 1-4: 221-228. 10.1177/153450849902401-419</bibtext> </blist> <blist> <bibtext> Mandera, P, Keuleers, E, &amp; Brysbaert, M. (2015). How useful are corpus-based methods for extrapolating psycholinguistic variables? Quarterly Journal of Experimental Psychology, 68(8), 1623-1642.</bibtext> </blist> <blist> <bibtext> McNamara, D. S, Graesser, A. C, McCarthy, P, &amp; Cai, Z. (2014). Automated evaluation of text and discourse with Coh-Metrix. Cambridge: Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Mikolov, T. (2013). Word2vec-toolkit [Online forum comment]. Retrieved from https://groups.google.com/forum/#!searchin/word2vec-toolkit/c-bow/word2vec-toolkit/NLvYXU99cAM/E5ld8LcDxlAJ. Accessed 10 Jan 2022.</bibtext> </blist> <blist> <bibtext> Mikolov, T, Chen, K, Corrado, G, &amp; Dean, J. (2013a). Efficient estimation of word representations in vector space. arXiv. 1301.3781.</bibtext> </blist> <blist> <bibtext> Mikolov, T, Sutskever, I, Chen, K, Corrado, G. S, &amp; Dean, J. (2013b). Distributed representations of words and phrases and their compositionality. In C.J.C. Burges, L. Bottou, M. Welling, Z. Ghahramani, &amp; K.Q. Weinberger (Eds.), Proceeding of the 26th International Conference on Neural Information Processing Systems (pp. 3111-3119).</bibtext> </blist> <blist> <bibtext> Mikolov, T, Le, Q. V, &amp; Sutskever, I. (2013c). Exploiting similarities among languages for machine translation. arXiv:1309.4168.</bibtext> </blist> <blist> <bibtext> Miller GA. WordNet: A lexical database for English. Communications of the ACM. 1995; 38; 11: 39-41. 10.1145/219717.219748</bibtext> </blist> <blist> <bibtext> Montefinese M, Vinson D, Vigliocco G, Ambrosini E. Italian age of acquisition norms for a large set of words (ItAoA). Frontiers in Psychology. 2019; 10: 278. 10.3389/fpsyg.2019.00278</bibtext> </blist> <blist> <bibtext> Moors, A, De Houwer, J, Hermans, D, Wanmaker, S, Van Schie, K, Van Harmelen, A. L,. &amp; Brysbaert, M. (2013). Norms of valence, arousal, dominance, and age of acquisition for 4,300 Dutch words. Behavior Research Methods, 45(1), 169-177.</bibtext> </blist> <blist> <bibtext> Morrison CM, Chappell TD, Ellis AW. Age of acquisition norms for a large set of object names and their relation to adult estimates and other variables. The Quarterly Journal of Experimental Psychology Section A. 1997; 50; 3: 528-559. 10.1080/027249897392017</bibtext> </blist> <blist> <bibtext> Nagy WE, Anderson RC, Herman PA. Learning word meanings from context during normal reading. American Educational Research Journal. 1987; 24; 2: 237-270. 10.3102/00028312024002237</bibtext> </blist> <blist> <bibtext> Nelson J, Perfetti C, Liben D, Liben M. Measures of text difficulty: Testing their predictive value for grade levels and student performance. 2012; Council of Chief State School Officers</bibtext> </blist> <blist> <bibtext> Pan BA, Rowe ML, Singer JD, Snow CE. Maternal correlates of growth in toddler vocabulary production in low-income families. Child Development. 2005; 76: 763-782</bibtext> </blist> <blist> <bibtext> Perret CA, Johnson AM, McCarthy KS, Guerrero TA, McNamara DSBoulay B, Baker R, Andre E. StairStepper: An adaptive remedial iSTART module. Proceedings of the 18th International Conference on Artificial Intelligence in Education (AIED). 2017; Springer: 557-560</bibtext> </blist> <blist> <bibtext> Robinson BF, Mervis CB. Disentangling early language development: Modeling lexical and grammatical acquisition using and extension of case-study methodology. Developmental Psychology. 1998; 34; 2: 363. 10.1037/0012-1649.34.2.363</bibtext> </blist> <blist> <bibtext> Rowe ML, Raudenbush SW, Goldin-Meadow S. The pace of vocabulary growth helps predict later vocabulary skill. Child Development. 2012; 83; 2: 508-525. 10.1111/j.1467-8624.2011.01710.x</bibtext> </blist> <blist> <bibtext> Roy BC, Frank MC, DeCamp P, Miller M, Roy D. Predicting the birth of a spoken word. Proceedings of the National Academy of Sciences. 2015; 112; 41: 12663-12668. 10.1073/pnas.1419773112</bibtext> </blist> <blist> <bibtext> Ruas T, Grosky W, Aizawa A. Multi-sense embeddings through a word sense disambiguation process. Expert Systems with Applications. 2019; 136: 288-303. 10.1016/j.eswa.2019.06.026</bibtext> </blist> <blist> <bibtext> Schönemann PH. A generalized solution of the orthogonal Procrustes problem. Psychometrika. 1966; 31; 1: 1-10. 10.1007/BF02289451</bibtext> </blist> <blist> <bibtext> Shabani, K, Khatib, M, &amp; Ebadi, S. (2010). Vygotsky's zone of proximal development: Instructional implications and teachers' professional development. English Language Teaching, 3(4), 237-248.</bibtext> </blist> <blist> <bibtext> Shock J, Cortese MJ, Khanna MM, Toppi S. Age of acquisition estimates for 3,000 disyllabic words. Behavior Research Methods. 2012; 44; 4: 971-977. 10.3758/s13428-012-0209-x</bibtext> </blist> <blist> <bibtext> Snefjella, B, &amp; Blank, I. (2020). Semantic Norm Extrapolation is a Missing Data Problem. ArXiv preprint.https://doi.org/10.31234/osf.io/y2gav</bibtext> </blist> <blist> <bibtext> Stadthagen-Gonzalez H, Davis CJ. The Bristol Norms for Age of Acquisition, Imageability and Familiarity. Behavior Research Methods. 2006; 38: 598-605. 10.3758/BF03193891</bibtext> </blist> <blist> <bibtext> Teng F. The effects of context and word exposure frequency on incidental vocabulary acquisition and retention through reading. The Language Learning Journal. 2019; 47; 2: 145-158. 10.1080/09571736.2016.1244217</bibtext> </blist> <blist> <bibtext> Tomaschek, F, Hendrix, P, &amp; Baayen, R. H. (2018). Strategies for addressing collinearity in multivariate linguistic data. Journal of Phonetics, 71, 249-267.</bibtext> </blist> <blist> <bibtext> Webb NM. Task related verbal interaction and mathematics learning in small groups. Journal for Research in Mathematics Education. 1991; 22: 366-389. 10.2307/749186</bibtext> </blist> <blist> <bibtext> Weisleder A, Fernald A. Talking to children matters: Early language experience strengthens processing and builds vocabulary. Psychological Science. 2013; 24; 11: 2143-2152. 10.1177/0956797613488145</bibtext> </blist> <blist> <bibtext> Yang, Z, Dai, Z, Yang, Y, Carbonell, J, Salakhutdinov, R, &amp; Le, Q. V. (2019). Xlnet: Generalized autoregressive pretraining for language understanding. In: The 33rd Conference on Neural Information Processing Systems (NeurIPS 2019). Vancouver, Canada.</bibtext> </blist> <blist> <bibtext> Zou, H, Hastie, T, &amp; Tibshirani, R. (2007). On the "degrees of freedom" of the lasso. The Annals of Statistics, 35(5), 2173-2192.</bibtext> </blist> </ref> <ref id="AN0160647162-30"> <title> Footnotes </title> <blist> <bibtext> In certain cases, the Flesch Reading Ease score may exceed the 0–100 range for ill-formed texts (i.e., missing sentence breaks or whitespaces).</bibtext> </blist> </ref> <aug> <p>By Robert-Mihai Botarleanu; Mihai Dascalu; Micah Watanabe; Scott Andrew Crossley and Danielle S. McNamara</p> <p>Reported by Author; Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib19" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib44" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib17" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib15" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib61" firstref="ref7"></nolink> <nolink nlid="nl6" bibid="bib62" firstref="ref8"></nolink> <nolink nlid="nl7" bibid="bib76" firstref="ref9"></nolink> <nolink nlid="nl8" bibid="bib12" firstref="ref11"></nolink> <nolink nlid="nl9" bibid="bib14" firstref="ref12"></nolink> <nolink nlid="nl10" bibid="bib27" firstref="ref13"></nolink> <nolink nlid="nl11" bibid="bib33" firstref="ref14"></nolink> <nolink nlid="nl12" bibid="bib63" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib46" firstref="ref19"></nolink> <nolink nlid="nl14" bibid="bib20" firstref="ref20"></nolink> <nolink nlid="nl15" bibid="bib74" firstref="ref25"></nolink> <nolink nlid="nl16" bibid="bib18" firstref="ref27"></nolink> <nolink nlid="nl17" bibid="bib30" firstref="ref28"></nolink> <nolink nlid="nl18" bibid="bib39" firstref="ref31"></nolink> <nolink nlid="nl19" bibid="bib29" firstref="ref39"></nolink> <nolink nlid="nl20" bibid="bib10" firstref="ref49"></nolink> <nolink nlid="nl21" bibid="bib54" firstref="ref53"></nolink> <nolink nlid="nl22" bibid="bib50" firstref="ref54"></nolink> <nolink nlid="nl23" bibid="bib59" firstref="ref55"></nolink> <nolink nlid="nl24" bibid="bib35" firstref="ref60"></nolink> <nolink nlid="nl25" bibid="bib37" firstref="ref61"></nolink> <nolink nlid="nl26" bibid="bib70" firstref="ref62"></nolink> <nolink nlid="nl27" bibid="bib80" firstref="ref63"></nolink> <nolink nlid="nl28" bibid="bib36" firstref="ref65"></nolink> <nolink nlid="nl29" bibid="bib66" firstref="ref66"></nolink> <nolink nlid="nl30" bibid="bib40" firstref="ref69"></nolink> <nolink nlid="nl31" bibid="bib79" firstref="ref70"></nolink> <nolink nlid="nl32" bibid="bib64" firstref="ref71"></nolink> <nolink nlid="nl33" bibid="bib77" firstref="ref72"></nolink> <nolink nlid="nl34" bibid="bib24" firstref="ref73"></nolink> <nolink nlid="nl35" bibid="bib45" firstref="ref76"></nolink> <nolink nlid="nl36" bibid="bib42" firstref="ref78"></nolink> <nolink nlid="nl37" bibid="bib53" firstref="ref79"></nolink> <nolink nlid="nl38" bibid="bib41" firstref="ref80"></nolink> <nolink nlid="nl39" bibid="bib26" firstref="ref86"></nolink> <nolink nlid="nl40" bibid="bib51" firstref="ref87"></nolink> <nolink nlid="nl41" bibid="bib47" firstref="ref89"></nolink> <nolink nlid="nl42" bibid="bib48" firstref="ref90"></nolink> <nolink nlid="nl43" bibid="bib57" firstref="ref91"></nolink> <nolink nlid="nl44" bibid="bib58" firstref="ref92"></nolink> <nolink nlid="nl45" bibid="bib11" firstref="ref94"></nolink> <nolink nlid="nl46" bibid="bib52" firstref="ref101"></nolink> <nolink nlid="nl47" bibid="bib38" firstref="ref102"></nolink> <nolink nlid="nl48" bibid="bib21" firstref="ref103"></nolink> <nolink nlid="nl49" bibid="bib55" firstref="ref106"></nolink> <nolink nlid="nl50" bibid="bib56" firstref="ref110"></nolink> <nolink nlid="nl51" bibid="bib71" firstref="ref113"></nolink> <nolink nlid="nl52" bibid="bib49" firstref="ref114"></nolink> <nolink nlid="nl53" bibid="bib22" firstref="ref115"></nolink> <nolink nlid="nl54" bibid="bib81" firstref="ref116"></nolink> <nolink nlid="nl55" bibid="bib32" firstref="ref117"></nolink> <nolink nlid="nl56" bibid="bib43" firstref="ref118"></nolink> <nolink nlid="nl57" bibid="bib72" firstref="ref119"></nolink> <nolink nlid="nl58" bibid="bib23" firstref="ref120"></nolink> <nolink nlid="nl59" bibid="bib65" firstref="ref122"></nolink> <nolink nlid="nl60" bibid="bib60" firstref="ref123"></nolink> <nolink nlid="nl61" bibid="bib13" firstref="ref124"></nolink> <nolink nlid="nl62" bibid="bib16" firstref="ref125"></nolink> <nolink nlid="nl63" bibid="bib78" firstref="ref127"></nolink> <nolink nlid="nl64" bibid="bib82" firstref="ref129"></nolink> <nolink nlid="nl65" bibid="bib540" firstref="ref131"></nolink> <nolink nlid="nl66" bibid="bib976" firstref="ref133"></nolink> <nolink nlid="nl67" bibid="bib107" firstref="ref136"></nolink> <nolink nlid="nl68" bibid="bib907" firstref="ref137"></nolink> <nolink nlid="nl69" bibid="bib269" firstref="ref143"></nolink> <nolink nlid="nl70" bibid="bib770" firstref="ref144"></nolink> <nolink nlid="nl71" bibid="bib869" firstref="ref146"></nolink> <nolink nlid="nl72" bibid="bib599" firstref="ref150"></nolink> <nolink nlid="nl73" bibid="bib583" firstref="ref152"></nolink> <nolink nlid="nl74" bibid="bib67" firstref="ref163"></nolink> <nolink nlid="nl75" bibid="bib34" firstref="ref168"></nolink> <nolink nlid="nl76" bibid="bib73" firstref="ref171"></nolink> <nolink nlid="nl77" bibid="bib75" firstref="ref172"></nolink> <nolink nlid="nl78" bibid="bib31" firstref="ref175"></nolink> <nolink nlid="nl79" bibid="bib68" firstref="ref176"></nolink> <nolink nlid="nl80" bibid="bib28" firstref="ref178"></nolink> <nolink nlid="nl81" bibid="bib25" firstref="ref179"></nolink> <nolink nlid="nl82" bibid="bib69" firstref="ref180"></nolink> CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=ED620060 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: ED620060 AccessLevel: 3 PubType: Report PubTypeId: report PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Age of Exposure 2.0: Estimating Word Complexity Using Iterative Models of Word Embeddings – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Botarleanu%2C+Robert-Mihai%22">Botarleanu, Robert-Mihai</searchLink><br /><searchLink fieldCode="AR" term="%22Dascalu%2C+Mihai%22">Dascalu, Mihai</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-4815-9227">0000-0002-4815-9227</externalLink>)<br /><searchLink fieldCode="AR" term="%22Watanabe%2C+Micah%22">Watanabe, Micah</searchLink><br /><searchLink fieldCode="AR" term="%22Crossley%2C+Scott+Andrew%22">Crossley, Scott Andrew</searchLink><br /><searchLink fieldCode="AR" term="%22McNamara%2C+Danielle+S%2E%22">McNamara, Danielle S.</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Grantee+Submission%22"><i>Grantee Submission</i></searchLink>. 2022. – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 28 – Name: DatePubCY Label: Publication Date Group: Date Data: 2022 – Name: SourceSuprt Label: Sponsoring Agency Group: SrcSuprt Data: Institute of Education Sciences (ED)<br />Office of Naval Research (ONR) (DOD) – Name: NumberContract Label: Contract Number Group: NumCntrct Data: R305A180144<br />R305A180261<br />N000141712300<br />N000142012623 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Age+Differences%22">Age Differences</searchLink><br /><searchLink fieldCode="DE" term="%22Vocabulary+Development%22">Vocabulary Development</searchLink><br /><searchLink fieldCode="DE" term="%22Correlation%22">Correlation</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+Comprehension%22">Reading Comprehension</searchLink><br /><searchLink fieldCode="DE" term="%22Word+Lists%22">Word Lists</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Adults%22">Adults</searchLink><br /><searchLink fieldCode="DE" term="%22Decision+Making%22">Decision Making</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Skills%22">Writing Skills</searchLink><br /><searchLink fieldCode="DE" term="%22Children%22">Children</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Error+Patterns%22">Error Patterns</searchLink><br /><searchLink fieldCode="DE" term="%22Computational+Linguistics%22">Computational Linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Reliability%22">Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Accuracy%22">Accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Language+Acquisition%22">Language Acquisition</searchLink><br /><searchLink fieldCode="DE" term="%22Linguistic+Input%22">Linguistic Input</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Software%22">Computer Software</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+Communication%22">Speech Communication</searchLink><br /><searchLink fieldCode="DE" term="%22Word+Frequency%22">Word Frequency</searchLink><br /><searchLink fieldCode="DE" term="%22Interpersonal+Relationship%22">Interpersonal Relationship</searchLink><br /><searchLink fieldCode="DE" term="%22Readability%22">Readability</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.3758/s13428-022-01797-5 – Name: Abstract Label: Abstract Group: Ab Data: Age of acquisition (AoA) is a measure of word complexity which refers to the age at which a word is typically learned. AoA measures have shown strong correlations with reading comprehension, lexical decision times, and writing quality. AoA scores based on both adult and child data have limitations that allow for error in measurement, and increase the cost and effort to produce. In this paper, we introduce Age of Exposure (AoE) version 2, a proxy for human exposure to new vocabulary terms that expands AoA word lists through training regressors to predict AoA scores. Word2vec word embeddings are trained on cumulatively increasing corpora of texts, word exposure trajectories are generated by aligning the word2vec vector spaces, and features of words are derived for modeling AoA scores. Our prediction models achieve low errors (from 13% with a corresponding R[superscript 2] of 0.35 up to 7% with an R[superscript 2] of 0.74), can be uniformly applied to different AoA word lists, and generalize to the entire vocabulary of a language. Our method benefits from using existing readability indices to define the order of texts in the corpora, while the performed analyses confirm that the generated AoA scores accurately predicted the difficulty of texts (R[superscript 2] of 0.84, surpassing related previous work). Further, we provide evidence of the internal reliability of our word trajectory features, demonstrate the effectiveness of the word trajectory features when contrasted with simple lexical features, and show that the exclusion of features that rely on external resources does not significantly impact performance. [This is the online first version of an article published in "Behavior Research Methods."] – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: CodeSource Label: IES Funded Group: SrcInfo Data: Yes – Name: DateEntry Label: Entry Date Group: Date Data: 2022 – Name: AN Label: Accession Number Group: ID Data: ED620060 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED620060 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.3758/s13428-022-01797-5 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 28 Subjects: – SubjectFull: Age Differences Type: general – SubjectFull: Vocabulary Development Type: general – SubjectFull: Correlation Type: general – SubjectFull: Reading Comprehension Type: general – SubjectFull: Word Lists Type: general – SubjectFull: Scores Type: general – SubjectFull: Adults Type: general – SubjectFull: Decision Making Type: general – SubjectFull: Writing Skills Type: general – SubjectFull: Children Type: general – SubjectFull: Prediction Type: general – SubjectFull: Error Patterns Type: general – SubjectFull: Computational Linguistics Type: general – SubjectFull: Reliability Type: general – SubjectFull: Accuracy Type: general – SubjectFull: Comparative Analysis Type: general – SubjectFull: Language Acquisition Type: general – SubjectFull: Linguistic Input Type: general – SubjectFull: Models Type: general – SubjectFull: Computer Software Type: general – SubjectFull: Speech Communication Type: general – SubjectFull: Word Frequency Type: general – SubjectFull: Interpersonal Relationship Type: general – SubjectFull: Readability Type: general Titles: – TitleFull: Age of Exposure 2.0: Estimating Word Complexity Using Iterative Models of Word Embeddings Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Botarleanu, Robert-Mihai – PersonEntity: Name: NameFull: Dascalu, Mihai – PersonEntity: Name: NameFull: Watanabe, Micah – PersonEntity: Name: NameFull: Crossley, Scott Andrew – PersonEntity: Name: NameFull: McNamara, Danielle S. IsPartOfRelationships: – BibEntity: Dates: – D: 15 M: 02 Type: published Y: 2022 Titles: – TitleFull: Grantee Submission Type: main |
| ResultId | 1 |