Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics

Saved in:
Bibliographic Details
Title: Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics
Language: English
Authors: Yikai Lu, Teresa M. Ober (ORCID 0000-0001-9698-9543), Cheng Liu (ORCID 0000-0002-2787-1653), Ying Cheng (ORCID 0000-0002-8654-1443)
Source: Grantee Submission. 2022.
Peer Reviewed: Y
Page Count: 6
Publication Date: 2022
Sponsoring Agency: Institute of Education Sciences (ED)
National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)
Contract Number: R305A180269
1350787
Document Type: Speeches/Meeting Papers
Reports - Research
Education Level: High Schools
Secondary Education
Descriptors: Prediction, Statistics Education, Data Analysis, Learning Analytics, Trend Analysis, Educational Trends, Learning Management Systems, High School Students, Advanced Placement, Scores, Mathematics Tests, Evaluation Criteria, Guidelines, Data Interpretation, Online Courses, Learning Problems
DOI: 10.1109/ICALT55010.2022.00051
Abstract: Machine learning methods for predictive analytics have great potential for uncovering trends in educational data. However, simple linear models still appear to be most widely used, in part, because of their interpretability. This study aims to address the issues of interpretability of complex machine learning classifiers by conducting feature extraction by neighborhood components analysis (NCA). Our dataset comprises 287 features from both process data indicators (i.e., derived from log data of an online statistics learning platform) and self-report data from high school students enrolled in Advanced Placement (AP) Statistics (N=733). As a label for prediction, we use students' scores on the AP Statistics exam. We evaluated the performance of machine learning classifiers with a given feature extraction method by evaluation criteria including F1 scores, the area under the receiver operating characteristic curve (AUC), and Cohen's Kappas. We find that NCA effectively reduces the dimensionality of training datasets, stabilizes machine learning predictions, and produces interpretable scores. However, interpreting the NCA weights of features, while feasible, is not very straightforward compared to linear regression. Future research should consider developing guidelines to interpret NCA weights.
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2024
Accession Number: ED661382
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=ED661382
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: ED661382
AccessLevel: 3
PubType: Conference
PubTypeId: conference
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yikai+Lu%22">Yikai Lu</searchLink><br /><searchLink fieldCode="AR" term="%22Teresa+M%2E+Ober%22">Teresa M. Ober</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9698-9543">0000-0001-9698-9543</externalLink>)<br /><searchLink fieldCode="AR" term="%22Cheng+Liu%22">Cheng Liu</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-2787-1653">0000-0002-2787-1653</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ying+Cheng%22">Ying Cheng</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-8654-1443">0000-0002-8654-1443</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Grantee+Submission%22"><i>Grantee Submission</i></searchLink>. 2022.
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 6
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2022
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Institute of Education Sciences (ED)<br />National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R305A180269<br />1350787
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Speeches/Meeting Papers<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22High+Schools%22">High Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Statistics+Education%22">Statistics Education</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Analytics%22">Learning Analytics</searchLink><br /><searchLink fieldCode="DE" term="%22Trend+Analysis%22">Trend Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Trends%22">Educational Trends</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Management+Systems%22">Learning Management Systems</searchLink><br /><searchLink fieldCode="DE" term="%22High+School+Students%22">High School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Advanced+Placement%22">Advanced Placement</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Tests%22">Mathematics Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Criteria%22">Evaluation Criteria</searchLink><br /><searchLink fieldCode="DE" term="%22Guidelines%22">Guidelines</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Interpretation%22">Data Interpretation</searchLink><br /><searchLink fieldCode="DE" term="%22Online+Courses%22">Online Courses</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Problems%22">Learning Problems</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1109/ICALT55010.2022.00051
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Machine learning methods for predictive analytics have great potential for uncovering trends in educational data. However, simple linear models still appear to be most widely used, in part, because of their interpretability. This study aims to address the issues of interpretability of complex machine learning classifiers by conducting feature extraction by neighborhood components analysis (NCA). Our dataset comprises 287 features from both process data indicators (i.e., derived from log data of an online statistics learning platform) and self-report data from high school students enrolled in Advanced Placement (AP) Statistics (N=733). As a label for prediction, we use students' scores on the AP Statistics exam. We evaluated the performance of machine learning classifiers with a given feature extraction method by evaluation criteria including F1 scores, the area under the receiver operating characteristic curve (AUC), and Cohen's Kappas. We find that NCA effectively reduces the dimensionality of training datasets, stabilizes machine learning predictions, and produces interpretable scores. However, interpreting the NCA weights of features, while feasible, is not very straightforward compared to linear regression. Future research should consider developing guidelines to interpret NCA weights.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: CodeSource
  Label: IES Funded
  Group: SrcInfo
  Data: Yes
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: ED661382
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED661382
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/ICALT55010.2022.00051
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 6
    Subjects:
      – SubjectFull: Prediction
        Type: general
      – SubjectFull: Statistics Education
        Type: general
      – SubjectFull: Data Analysis
        Type: general
      – SubjectFull: Learning Analytics
        Type: general
      – SubjectFull: Trend Analysis
        Type: general
      – SubjectFull: Educational Trends
        Type: general
      – SubjectFull: Learning Management Systems
        Type: general
      – SubjectFull: High School Students
        Type: general
      – SubjectFull: Advanced Placement
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Mathematics Tests
        Type: general
      – SubjectFull: Evaluation Criteria
        Type: general
      – SubjectFull: Guidelines
        Type: general
      – SubjectFull: Data Interpretation
        Type: general
      – SubjectFull: Online Courses
        Type: general
      – SubjectFull: Learning Problems
        Type: general
    Titles:
      – TitleFull: Application of Neighborhood Components Analysis to Process and Survey Data to Predict Student Learning of Statistics
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yikai Lu
      – PersonEntity:
          Name:
            NameFull: Teresa M. Ober
      – PersonEntity:
          Name:
            NameFull: Cheng Liu
      – PersonEntity:
          Name:
            NameFull: Ying Cheng
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 17
              M: 08
              Type: published
              Y: 2022
          Titles:
            – TitleFull: Grantee Submission
              Type: main
ResultId 1