Predicting Contextual Informativeness for Vocabulary Learning

Saved in:
Bibliographic Details
Title: Predicting Contextual Informativeness for Vocabulary Learning
Language: English
Authors: Kapelner, Adam, Soterwood, Jeanine, NessAiver, Shalev, Adlof, Suzanne
Source: Grantee Submission. 2018.
Peer Reviewed: Y
Page Count: 17
Publication Date: 2018
Sponsoring Agency: Institute of Education Sciences (ED)
Contract Number: R305A130467
Document Type: Reports - Research
Education Level: High Schools
Secondary Education
Descriptors: Vocabulary Development, Databases, Training, Models, Statistical Analysis, Prediction, Performance, High School Students, Language Arts, Secondary School Teachers, Surveys, Context Effect
Geographic Terms: South Carolina, Connecticut
DOI: 10.1109/TLT.2018.2789900
Abstract: Vocabulary knowledge is essential to educational progress. High quality vocabulary instruction requires supportive contextual examples to teach word meaning and proper usage. Identifying such contexts by hand for a large number of words can be difficult. In this work, we take a statistical learning approach to engineer a system that predicts informativeness of a context for target words that span the range of difficulty from middle school to college level. Our database (released open source) includes 1,000 hand-selected words associated with approximately 70,000 contextual examples gathered from the Internet. Our training data included each context rated by 10 individuals on a four-point informativeness scale. We process the text of each context into a novel collection of approximately 600 numerical features that captures diverse linguistic information. We then fit a nonparametric regression model using Random Forests and compute out-of-sample prediction performance using cross-validation. Our system performs well enough that it can replace a human judge: for a target word not found in our dataset, we can provide curated contexts to a student learner such that most of the contexts (54 percent) feature rich contextual clues and confusing contexts are rare (<1 percent). The quality of our curated contexts was validated by an independent panel of high school language arts teachers. [This paper was published in "IEEE Transactions on Learning Technologies" v11 n1 p13-26 Jan-Mar 2018 (ISSN 1939-1382) (EJ1174702).]
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2018
Accession Number: ED589145
Database: ERIC
Description
Abstract:Vocabulary knowledge is essential to educational progress. High quality vocabulary instruction requires supportive contextual examples to teach word meaning and proper usage. Identifying such contexts by hand for a large number of words can be difficult. In this work, we take a statistical learning approach to engineer a system that predicts informativeness of a context for target words that span the range of difficulty from middle school to college level. Our database (released open source) includes 1,000 hand-selected words associated with approximately 70,000 contextual examples gathered from the Internet. Our training data included each context rated by 10 individuals on a four-point informativeness scale. We process the text of each context into a novel collection of approximately 600 numerical features that captures diverse linguistic information. We then fit a nonparametric regression model using Random Forests and compute out-of-sample prediction performance using cross-validation. Our system performs well enough that it can replace a human judge: for a target word not found in our dataset, we can provide curated contexts to a student learner such that most of the contexts (54 percent) feature rich contextual clues and confusing contexts are rare (<1 percent). The quality of our curated contexts was validated by an independent panel of high school language arts teachers. [This paper was published in "IEEE Transactions on Learning Technologies" v11 n1 p13-26 Jan-Mar 2018 (ISSN 1939-1382) (EJ1174702).]
DOI:10.1109/TLT.2018.2789900