Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics
Saved in:
| Title: | Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics |
|---|---|
| Language: | English |
| Authors: | Yikai Lu, Teresa M. Ober (ORCID |
| Source: | Grantee Submission. 2022. |
| Peer Reviewed: | Y |
| Page Count: | 6 |
| Publication Date: | 2022 |
| Sponsoring Agency: | Institute of Education Sciences (ED) National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL) |
| Contract Number: | R305A180269 1350787 |
| Document Type: | Speeches/Meeting Papers Reports - Research |
| Education Level: | High Schools Secondary Education |
| Descriptors: | Prediction, Statistics Education, Data Analysis, Learning Analytics, Trend Analysis, Educational Trends, Learning Management Systems, High School Students, Advanced Placement, Scores, Mathematics Tests, Evaluation Criteria, Guidelines, Data Interpretation, Online Courses, Learning Problems |
| DOI: | 10.1109/ICALT55010.2022.00051 |
| Abstract: | Machine learning methods for predictive analytics have great potential for uncovering trends in educational data. However, simple linear models still appear to be most widely used, in part, because of their interpretability. This study aims to address the issues of interpretability of complex machine learning classifiers by conducting feature extraction by neighborhood components analysis (NCA). Our dataset comprises 287 features from both process data indicators (i.e., derived from log data of an online statistics learning platform) and self-report data from high school students enrolled in Advanced Placement (AP) Statistics (N=733). As a label for prediction, we use students' scores on the AP Statistics exam. We evaluated the performance of machine learning classifiers with a given feature extraction method by evaluation criteria including F1 scores, the area under the receiver operating characteristic curve (AUC), and Cohen's Kappas. We find that NCA effectively reduces the dimensionality of training datasets, stabilizes machine learning predictions, and produces interpretable scores. However, interpreting the NCA weights of features, while feasible, is not very straightforward compared to linear regression. Future research should consider developing guidelines to interpret NCA weights. |
| Abstractor: | As Provided |
| IES Funded: | Yes |
| Entry Date: | 2024 |
| Accession Number: | ED661382 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=ED661382 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: ED661382 AccessLevel: 3 PubType: Conference PubTypeId: conference PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Yikai+Lu%22">Yikai Lu</searchLink><br /><searchLink fieldCode="AR" term="%22Teresa+M%2E+Ober%22">Teresa M. Ober</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9698-9543">0000-0001-9698-9543</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cheng+Liu%22">Cheng Liu</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-2787-1653">0000-0002-2787-1653</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ying+Cheng%22">Ying Cheng</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-8654-1443">0000-0002-8654-1443</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Grantee+Submission%22"><i>Grantee Submission</i></searchLink>. 2022. – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 6 – Name: DatePubCY Label: Publication Date Group: Date Data: 2022 – Name: SourceSuprt Label: Sponsoring Agency Group: SrcSuprt Data: Institute of Education Sciences (ED)<br />National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL) – Name: NumberContract Label: Contract Number Group: NumCntrct Data: R305A180269<br />1350787 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Speeches/Meeting Papers<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22High+Schools%22">High Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Statistics+Education%22">Statistics Education</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Analytics%22">Learning Analytics</searchLink><br /><searchLink fieldCode="DE" term="%22Trend+Analysis%22">Trend Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Trends%22">Educational Trends</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Management+Systems%22">Learning Management Systems</searchLink><br /><searchLink fieldCode="DE" term="%22High+School+Students%22">High School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Advanced+Placement%22">Advanced Placement</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Tests%22">Mathematics Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Criteria%22">Evaluation Criteria</searchLink><br /><searchLink fieldCode="DE" term="%22Guidelines%22">Guidelines</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Interpretation%22">Data Interpretation</searchLink><br /><searchLink fieldCode="DE" term="%22Online+Courses%22">Online Courses</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Problems%22">Learning Problems</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1109/ICALT55010.2022.00051 – Name: Abstract Label: Abstract Group: Ab Data: Machine learning methods for predictive analytics have great potential for uncovering trends in educational data. However, simple linear models still appear to be most widely used, in part, because of their interpretability. This study aims to address the issues of interpretability of complex machine learning classifiers by conducting feature extraction by neighborhood components analysis (NCA). Our dataset comprises 287 features from both process data indicators (i.e., derived from log data of an online statistics learning platform) and self-report data from high school students enrolled in Advanced Placement (AP) Statistics (N=733). As a label for prediction, we use students' scores on the AP Statistics exam. We evaluated the performance of machine learning classifiers with a given feature extraction method by evaluation criteria including F1 scores, the area under the receiver operating characteristic curve (AUC), and Cohen's Kappas. We find that NCA effectively reduces the dimensionality of training datasets, stabilizes machine learning predictions, and produces interpretable scores. However, interpreting the NCA weights of features, while feasible, is not very straightforward compared to linear regression. Future research should consider developing guidelines to interpret NCA weights. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: CodeSource Label: IES Funded Group: SrcInfo Data: Yes – Name: DateEntry Label: Entry Date Group: Date Data: 2024 – Name: AN Label: Accession Number Group: ID Data: ED661382 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED661382 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1109/ICALT55010.2022.00051 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 6 Subjects: – SubjectFull: Prediction Type: general – SubjectFull: Statistics Education Type: general – SubjectFull: Data Analysis Type: general – SubjectFull: Learning Analytics Type: general – SubjectFull: Trend Analysis Type: general – SubjectFull: Educational Trends Type: general – SubjectFull: Learning Management Systems Type: general – SubjectFull: High School Students Type: general – SubjectFull: Advanced Placement Type: general – SubjectFull: Scores Type: general – SubjectFull: Mathematics Tests Type: general – SubjectFull: Evaluation Criteria Type: general – SubjectFull: Guidelines Type: general – SubjectFull: Data Interpretation Type: general – SubjectFull: Online Courses Type: general – SubjectFull: Learning Problems Type: general Titles: – TitleFull: Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Yikai Lu – PersonEntity: Name: NameFull: Teresa M. Ober – PersonEntity: Name: NameFull: Cheng Liu – PersonEntity: Name: NameFull: Ying Cheng IsPartOfRelationships: – BibEntity: Dates: – D: 17 M: 08 Type: published Y: 2022 Titles: – TitleFull: Grantee Submission Type: main |
| ResultId | 1 |