Predicting Contextual Informativeness for Vocabulary Learning
Saved in:
| Title: | Predicting Contextual Informativeness for Vocabulary Learning |
|---|---|
| Language: | English |
| Authors: | Kapelner, Adam, Soterwood, Jeanine, NessAiver, Shalev, Adlof, Suzanne |
| Source: | Grantee Submission. 2018. |
| Peer Reviewed: | Y |
| Page Count: | 17 |
| Publication Date: | 2018 |
| Sponsoring Agency: | Institute of Education Sciences (ED) |
| Contract Number: | R305A130467 |
| Document Type: | Reports - Research |
| Education Level: | High Schools Secondary Education |
| Descriptors: | Vocabulary Development, Databases, Training, Models, Statistical Analysis, Prediction, Performance, High School Students, Language Arts, Secondary School Teachers, Surveys, Context Effect |
| Geographic Terms: | South Carolina, Connecticut |
| DOI: | 10.1109/TLT.2018.2789900 |
| Abstract: | Vocabulary knowledge is essential to educational progress. High quality vocabulary instruction requires supportive contextual examples to teach word meaning and proper usage. Identifying such contexts by hand for a large number of words can be difficult. In this work, we take a statistical learning approach to engineer a system that predicts informativeness of a context for target words that span the range of difficulty from middle school to college level. Our database (released open source) includes 1,000 hand-selected words associated with approximately 70,000 contextual examples gathered from the Internet. Our training data included each context rated by 10 individuals on a four-point informativeness scale. We process the text of each context into a novel collection of approximately 600 numerical features that captures diverse linguistic information. We then fit a nonparametric regression model using Random Forests and compute out-of-sample prediction performance using cross-validation. Our system performs well enough that it can replace a human judge: for a target word not found in our dataset, we can provide curated contexts to a student learner such that most of the contexts (54 percent) feature rich contextual clues and confusing contexts are rare (<1 percent). The quality of our curated contexts was validated by an independent panel of high school language arts teachers. [This paper was published in "IEEE Transactions on Learning Technologies" v11 n1 p13-26 Jan-Mar 2018 (ISSN 1939-1382) (EJ1174702).] |
| Abstractor: | As Provided |
| IES Funded: | Yes |
| Entry Date: | 2018 |
| Accession Number: | ED589145 |
| Database: | ERIC |
| Abstract: | Vocabulary knowledge is essential to educational progress. High quality vocabulary instruction requires supportive contextual examples to teach word meaning and proper usage. Identifying such contexts by hand for a large number of words can be difficult. In this work, we take a statistical learning approach to engineer a system that predicts informativeness of a context for target words that span the range of difficulty from middle school to college level. Our database (released open source) includes 1,000 hand-selected words associated with approximately 70,000 contextual examples gathered from the Internet. Our training data included each context rated by 10 individuals on a four-point informativeness scale. We process the text of each context into a novel collection of approximately 600 numerical features that captures diverse linguistic information. We then fit a nonparametric regression model using Random Forests and compute out-of-sample prediction performance using cross-validation. Our system performs well enough that it can replace a human judge: for a target word not found in our dataset, we can provide curated contexts to a student learner such that most of the contexts (54 percent) feature rich contextual clues and confusing contexts are rare (<1 percent). The quality of our curated contexts was validated by an independent panel of high school language arts teachers. [This paper was published in "IEEE Transactions on Learning Technologies" v11 n1 p13-26 Jan-Mar 2018 (ISSN 1939-1382) (EJ1174702).] |
|---|---|
| DOI: | 10.1109/TLT.2018.2789900 |