GenAI Models as Keyword Rankers: A Learner-Centred Case Study for L2 Spanish

Saved in:
Bibliographic Details
Title: GenAI Models as Keyword Rankers: A Learner-Centred Case Study for L2 Spanish
Language: English
Authors: Jasper Degraeuwe (ORCID 0000-0003-0850-314X)
Source: The EUROCALL Review. 2025 32(2):63-75.
Availability: European Association for Computer-Assisted Language Learning (EUROCALL). EUROCALL Headquarters, School of Modern Languages, University of Ulster, Cromore Road, Coleraine BT52 1SA, Northern Ireland, UK. Tel: +34-67-943-1283; Web site: http://www.eurocall-languages.org/
Peer Reviewed: Y
Page Count: 13
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Descriptors: Second Language Learning, Spanish, Word Lists, Word Frequency, Vocabulary Development, Artificial Intelligence, Intellectual Disciplines, Indo European Languages, Computer Uses in Education, Educational Technology, Units of Study
ISSN: 1695-2618
Abstract: Frequency-based word lists form an important part of general-purpose vocabulary learning courses aimed at beginner and (lower-)intermediate learners of a foreign/second language (L2). For advanced learners and/or specific purposes, however, relying exclusively on these general word lists will be unlikely to lead to an adequate selection of vocabulary. As research in this latter area remains scarce (especially for languages other than English), the present study aims to fill (part of) the gap by investigating the use of Generative Artificial Intelligence (GenAI) models to automatically rank vocabulary items based on how typical they are of a given topic, focusing on Spanish as the target language. I compile a dataset containing four domain-specific subsets of 200 vocabulary items (for the topics economics, health, law, and migration) and analyse how well GenAI-based rankings of these vocabulary items (using zero-shot prompting) correlate with gold standard human rankings (provided by L2 learners). As the evaluation baseline, I use the rankings obtained by means of the Kullback-Leibler divergence (i.e., a statistical "keyness" measure based on word frequencies). With a top average Spearman's ? and Kendall's weighted [tau] of 0.73, this first-of-its-kind study demonstrates that the tested GenAI models (Gemma, Llama, and Mistral) outperform the baseline by a large margin, showing great potential for use in the real-life creation of domain-specific vocabulary lists for L2 learning purposes. [Note: The page range (63-74) shown on the PDF is incorrect. The correct page range is 63-75.]
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1494351
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1494351
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1494351
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: GenAI Models as Keyword Rankers: A Learner-Centred Case Study for L2 Spanish
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Jasper+Degraeuwe%22">Jasper Degraeuwe</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0850-314X">0000-0003-0850-314X</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22The+EUROCALL+Review%22"><i>The EUROCALL Review</i></searchLink>. 2025 32(2):63-75.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: European Association for Computer-Assisted Language Learning (EUROCALL). EUROCALL Headquarters, School of Modern Languages, University of Ulster, Cromore Road, Coleraine BT52 1SA, Northern Ireland, UK. Tel: +34-67-943-1283; Web site: http://www.eurocall-languages.org/
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 13
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Spanish%22">Spanish</searchLink><br /><searchLink fieldCode="DE" term="%22Word+Lists%22">Word Lists</searchLink><br /><searchLink fieldCode="DE" term="%22Word+Frequency%22">Word Frequency</searchLink><br /><searchLink fieldCode="DE" term="%22Vocabulary+Development%22">Vocabulary Development</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Intellectual+Disciplines%22">Intellectual Disciplines</searchLink><br /><searchLink fieldCode="DE" term="%22Indo+European+Languages%22">Indo European Languages</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Uses+in+Education%22">Computer Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Technology%22">Educational Technology</searchLink><br /><searchLink fieldCode="DE" term="%22Units+of+Study%22">Units of Study</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1695-2618
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Frequency-based word lists form an important part of general-purpose vocabulary learning courses aimed at beginner and (lower-)intermediate learners of a foreign/second language (L2). For advanced learners and/or specific purposes, however, relying exclusively on these general word lists will be unlikely to lead to an adequate selection of vocabulary. As research in this latter area remains scarce (especially for languages other than English), the present study aims to fill (part of) the gap by investigating the use of Generative Artificial Intelligence (GenAI) models to automatically rank vocabulary items based on how typical they are of a given topic, focusing on Spanish as the target language. I compile a dataset containing four domain-specific subsets of 200 vocabulary items (for the topics economics, health, law, and migration) and analyse how well GenAI-based rankings of these vocabulary items (using zero-shot prompting) correlate with gold standard human rankings (provided by L2 learners). As the evaluation baseline, I use the rankings obtained by means of the Kullback-Leibler divergence (i.e., a statistical "keyness" measure based on word frequencies). With a top average Spearman's ? and Kendall's weighted [tau] of 0.73, this first-of-its-kind study demonstrates that the tested GenAI models (Gemma, Llama, and Mistral) outperform the baseline by a large margin, showing great potential for use in the real-life creation of domain-specific vocabulary lists for L2 learning purposes. [Note: The page range (63-74) shown on the PDF is incorrect. The correct page range is 63-75.]
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1494351
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1494351
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 63
    Subjects:
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: Spanish
        Type: general
      – SubjectFull: Word Lists
        Type: general
      – SubjectFull: Word Frequency
        Type: general
      – SubjectFull: Vocabulary Development
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Intellectual Disciplines
        Type: general
      – SubjectFull: Indo European Languages
        Type: general
      – SubjectFull: Computer Uses in Education
        Type: general
      – SubjectFull: Educational Technology
        Type: general
      – SubjectFull: Units of Study
        Type: general
    Titles:
      – TitleFull: GenAI Models as Keyword Rankers: A Learner-Centred Case Study for L2 Spanish
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Jasper Degraeuwe
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-electronic
              Value: 1695-2618
          Numbering:
            – Type: volume
              Value: 32
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: The EUROCALL Review
              Type: main
ResultId 1