The Challenges of Large‐Scale, Web‐Based Language Datasets: Word Length and Predictability Revisited.

Saved in:
Bibliographic Details
Title: The Challenges of Large‐Scale, Web‐Based Language Datasets: Word Length and Predictability Revisited.
Authors: Meylan, Stephan C.1,2 (AUTHOR) smeylan@mit.edu, Griffiths, Thomas L.3 (AUTHOR)
Source: Cognitive Science. Jun2021, Vol. 45 Issue 6, p1-26. 26p.
Subject Terms: *Language research, *Word frequency, *Language & languages, *Vocabulary, Corpora
Abstract: Language research has come to rely heavily on large‐scale, web‐based datasets. These datasets can present significant methodological challenges, requiring researchers to make a number of decisions about how they are collected, represented, and analyzed. These decisions often concern long‐standing challenges in corpus‐based language research, including determining what counts as a word, deciding which words should be analyzed, and matching sets of words across languages. We illustrate these challenges by revisiting "Word lengths are optimized for efficient communication" (Piantadosi, Tily, & Gibson, 2011), which found that word lengths in 11 languages are more strongly correlated with their average predictability (or average information content) than their frequency. Using what we argue to be best practices for large‐scale corpus analyses, we find significantly attenuated support for this result and demonstrate that a stronger relationship obtains between word frequency and length for a majority of the languages in the sample. We consider the implications of the results for language research more broadly and provide several recommendations to researchers regarding best practices. [ABSTRACT FROM AUTHOR]
Copyright of Cognitive Science is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: ehh
DbLabel: Education Research Complete
An: 151133110
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: The Challenges of Large‐Scale, Web‐Based Language Datasets: Word Length and Predictability Revisited.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Meylan%2C+Stephan+C%2E%22">Meylan, Stephan C.</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> smeylan@mit.edu</i><br /><searchLink fieldCode="AR" term="%22Griffiths%2C+Thomas+L%2E%22">Griffiths, Thomas L.</searchLink><relatesTo>3</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Cognitive+Science%22">Cognitive Science</searchLink>. Jun2021, Vol. 45 Issue 6, p1-26. 26p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Language+research%22">Language research</searchLink><br />*<searchLink fieldCode="DE" term="%22Word+frequency%22">Word frequency</searchLink><br />*<searchLink fieldCode="DE" term="%22Language+%26+languages%22">Language & languages</searchLink><br />*<searchLink fieldCode="DE" term="%22Vocabulary%22">Vocabulary</searchLink><br /><searchLink fieldCode="DE" term="%22Corpora%22">Corpora</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Language research has come to rely heavily on large‐scale, web‐based datasets. These datasets can present significant methodological challenges, requiring researchers to make a number of decisions about how they are collected, represented, and analyzed. These decisions often concern long‐standing challenges in corpus‐based language research, including determining what counts as a word, deciding which words should be analyzed, and matching sets of words across languages. We illustrate these challenges by revisiting "Word lengths are optimized for efficient communication" (Piantadosi, Tily, & Gibson, 2011), which found that word lengths in 11 languages are more strongly correlated with their average predictability (or average information content) than their frequency. Using what we argue to be best practices for large‐scale corpus analyses, we find significantly attenuated support for this result and demonstrate that a stronger relationship obtains between word frequency and length for a majority of the languages in the sample. We consider the implications of the results for language research more broadly and provide several recommendations to researchers regarding best practices. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Cognitive Science is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=151133110
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/cogs.12983
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 26
        StartPage: 1
    Subjects:
      – SubjectFull: Language research
        Type: general
      – SubjectFull: Word frequency
        Type: general
      – SubjectFull: Language & languages
        Type: general
      – SubjectFull: Vocabulary
        Type: general
      – SubjectFull: Corpora
        Type: general
    Titles:
      – TitleFull: The Challenges of Large‐Scale, Web‐Based Language Datasets: Word Length and Predictability Revisited.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Meylan, Stephan C.
      – PersonEntity:
          Name:
            NameFull: Griffiths, Thomas L.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2021
              Type: published
              Y: 2021
          Identifiers:
            – Type: issn-print
              Value: 03640213
          Numbering:
            – Type: volume
              Value: 45
            – Type: issue
              Value: 6
          Titles:
            – TitleFull: Cognitive Science
              Type: main
ResultId 1