Word Sense Clustering and Clusterability.

Saved in:
Bibliographic Details
Title: Word Sense Clustering and Clusterability.
Authors: McCarthy, Diana1 diana@dianamccarthy.co.uk, Apidianaki, Marianna2 marianna.apidianaki@limsi.fr, Erk, Katrin3 katrin.erk@mail.utexas.edu
Source: Computational Linguistics. Jun2016, Vol. 42 Issue 2, p245-275. 31p.
Subjects: Terms & phrases, Sublanguage, Linguistics, Idiolect, Machine learning
Abstract: Word sense disambiguation and the related field of automated word sense induction traditionally assume that the occurrences of a lemma can be partitioned into senses. But this seems to be a much easier task for some lemmas than others. Our work builds on recent work that proposes describing word meaning in a graded fashion rather than through a strict partition into senses; in this article we argue that not all lemmas may need the more complex graded analysis, depending on their partitionability. Although there is plenty of evidence from previous studies and from the linguistics literature that there is a spectrum of partitionability of word meanings, this is the first attempt to measure the phenomenon and to couple the machine learning literature on clusterability with word usage data used in computational linguistics. We propose to operationalize partitionability as clusterability, a measure of how easy the occurrences of a lemma are to cluster. We test two ways of measuring clusterability: (1) existing measures from the machine learning literature that aim to measure the goodness of optimal k-means clusterings, and (2) the idea that if a lemma is more clusterable, two clusterings based on two different "views" of the same data points will be more congruent. The two views that we use are two different sets of manually constructed lexical substitutes for the target lemma, on the one hand monolingual paraphrases, and on the other hand translations. We apply automatic clustering to the manual annotations. We use manual annotations because we want the representations of the instances that we cluster to be as informative and "clean" as possible. We show that when we control for polysemy, our measures of clusterability tend to correlate with partitionability, in particular some of the type-(1) clusterability measures, and that these measures outperform a baseline that relies on the amount of overlap in a soft clustering. [ABSTRACT FROM AUTHOR]
Copyright of Computational Linguistics is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 116185546
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Word Sense Clustering and Clusterability.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22McCarthy%2C+Diana%22">McCarthy, Diana</searchLink><relatesTo>1</relatesTo><i> diana@dianamccarthy.co.uk</i><br /><searchLink fieldCode="AR" term="%22Apidianaki%2C+Marianna%22">Apidianaki, Marianna</searchLink><relatesTo>2</relatesTo><i> marianna.apidianaki@limsi.fr</i><br /><searchLink fieldCode="AR" term="%22Erk%2C+Katrin%22">Erk, Katrin</searchLink><relatesTo>3</relatesTo><i> katrin.erk@mail.utexas.edu</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Computational+Linguistics%22">Computational Linguistics</searchLink>. Jun2016, Vol. 42 Issue 2, p245-275. 31p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Terms+%26+phrases%22">Terms & phrases</searchLink><br /><searchLink fieldCode="DE" term="%22Sublanguage%22">Sublanguage</searchLink><br /><searchLink fieldCode="DE" term="%22Linguistics%22">Linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Idiolect%22">Idiolect</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Word sense disambiguation and the related field of automated word sense induction traditionally assume that the occurrences of a lemma can be partitioned into senses. But this seems to be a much easier task for some lemmas than others. Our work builds on recent work that proposes describing word meaning in a graded fashion rather than through a strict partition into senses; in this article we argue that not all lemmas may need the more complex graded analysis, depending on their partitionability. Although there is plenty of evidence from previous studies and from the linguistics literature that there is a spectrum of partitionability of word meanings, this is the first attempt to measure the phenomenon and to couple the machine learning literature on clusterability with word usage data used in computational linguistics. We propose to operationalize partitionability as clusterability, a measure of how easy the occurrences of a lemma are to cluster. We test two ways of measuring clusterability: (1) existing measures from the machine learning literature that aim to measure the goodness of optimal k-means clusterings, and (2) the idea that if a lemma is more clusterable, two clusterings based on two different "views" of the same data points will be more congruent. The two views that we use are two different sets of manually constructed lexical substitutes for the target lemma, on the one hand monolingual paraphrases, and on the other hand translations. We apply automatic clustering to the manual annotations. We use manual annotations because we want the representations of the instances that we cluster to be as informative and "clean" as possible. We show that when we control for polysemy, our measures of clusterability tend to correlate with partitionability, in particular some of the type-(1) clusterability measures, and that these measures outperform a baseline that relies on the amount of overlap in a soft clustering. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Computational Linguistics is the property of MIT Press and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=116185546
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1162/COLI_a_00247
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 31
        StartPage: 245
    Subjects:
      – SubjectFull: Terms & phrases
        Type: general
      – SubjectFull: Sublanguage
        Type: general
      – SubjectFull: Linguistics
        Type: general
      – SubjectFull: Idiolect
        Type: general
      – SubjectFull: Machine learning
        Type: general
    Titles:
      – TitleFull: Word Sense Clustering and Clusterability.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: McCarthy, Diana
      – PersonEntity:
          Name:
            NameFull: Apidianaki, Marianna
      – PersonEntity:
          Name:
            NameFull: Erk, Katrin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2016
              Type: published
              Y: 2016
          Identifiers:
            – Type: issn-print
              Value: 08912017
          Numbering:
            – Type: volume
              Value: 42
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Computational Linguistics
              Type: main
ResultId 1