Formative Feedback on Student-Authored Summaries in Intelligent Textbooks Using Large Language Models

Saved in:
Bibliographic Details
Title: Formative Feedback on Student-Authored Summaries in Intelligent Textbooks Using Large Language Models
Language: English
Authors: Wesley Morris (ORCID 0000-0001-6316-6479), Scott Crossley, Langdon Holmes, Chaohua Ou, Mihai Dascalu, Danielle McNamara
Source: International Journal of Artificial Intelligence in Education. 2025 35(3):1022-1043.
Availability: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/
Peer Reviewed: Y
Page Count: 22
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Higher Education
Postsecondary Education
High Schools
Secondary Education
Descriptors: Formative Evaluation, Feedback (Response), Textbooks, Artificial Intelligence, Natural Language Processing, Computer Uses in Education, Automation, Writing Evaluation, Reading Comprehension, Writing (Composition), College Students, High School Students
DOI: 10.1007/s40593-024-00395-0
ISSN: 1560-4292
1560-4306
Abstract: As intelligent textbooks become more ubiquitous in classrooms and educational settings, the need to make them more interactive arises. An alternative is to ask students to generate knowledge in response to textbook content and provide feedback about the produced knowledge. This study develops Natural Language Processing models to automatically provide feedback to students about the quality of summaries written at the end of intelligent textbook sections. The study builds on the work of Botarleanu et al. (2022), who used a Longformer Large Language Model (LLM) to develop a summary grading model. Their model explained around 55% of holistic summary score variance as assigned by human raters. This study uses a principal component analysis to distill summary scores from an analytic rubric into two principal components -- content and wording. This study uses two encoder-only classification large language models finetuned from Longformer on the summaries and the source texts using these principal components explained 82% and 70% of the score variance for content and wording, respectively. On a dataset of summaries collected on the crowd-sourcing site Prolific, the content model was shown to be robust although the accuracy of the wording model was reduced compared to the training set. The developed models are freely available on HuggingFace and will allow formative feedback to users of intelligent textbooks to assess reading comprehension through summarization in real time. The models can also be used for other summarization applications in learning systems.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1488267
Database: ERIC
Description
Abstract:As intelligent textbooks become more ubiquitous in classrooms and educational settings, the need to make them more interactive arises. An alternative is to ask students to generate knowledge in response to textbook content and provide feedback about the produced knowledge. This study develops Natural Language Processing models to automatically provide feedback to students about the quality of summaries written at the end of intelligent textbook sections. The study builds on the work of Botarleanu et al. (2022), who used a Longformer Large Language Model (LLM) to develop a summary grading model. Their model explained around 55% of holistic summary score variance as assigned by human raters. This study uses a principal component analysis to distill summary scores from an analytic rubric into two principal components -- content and wording. This study uses two encoder-only classification large language models finetuned from Longformer on the summaries and the source texts using these principal components explained 82% and 70% of the score variance for content and wording, respectively. On a dataset of summaries collected on the crowd-sourcing site Prolific, the content model was shown to be robust although the accuracy of the wording model was reduced compared to the training set. The developed models are freely available on HuggingFace and will allow formative feedback to users of intelligent textbooks to assess reading comprehension through summarization in real time. The models can also be used for other summarization applications in learning systems.
ISSN:1560-4292
1560-4306
DOI:10.1007/s40593-024-00395-0