SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption.

Saved in:
Bibliographic Details
Title: SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption.
Authors: Fernando, Thilak1 thilak.fernando@monash.edu, Webb, Geoffrey1
Source: Data Mining & Knowledge Discovery. Jan2017, Vol. 31 Issue 1, p264-286. 23p.
Subjects: Information retrieval software, Machine learning, Cluster analysis (Statistics)
Abstract: Similarity measures are central to many machine learning algorithms. There are many different similarity measures, each catering for different applications and data requirements. Most similarity measures used with numerical data assume that the attributes are interval scale. In the interval scale, it is assumed that a unit difference has the same meaning irrespective of the magnitudes of the values separated. When this assumption is violated, accuracy may be reduced. Our experiments show that removing the interval scale assumption by transforming data to ranks can improve the accuracy of distance-based similarity measures on some tasks. However the rank transform has high time and storage overheads. In this paper, we introduce an efficient similarity measure which does not consider the magnitudes of inter-instance distances. We compare the new similarity measure with popular similarity measures in two applications: DBScan clustering and content based multimedia information retrieval with real world datasets and different transform functions. The results show that the proposed similarity measure provides good performance on a range of tasks and is invariant to violations of the interval scale assumption. [ABSTRACT FROM AUTHOR]
Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 120690514
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Fernando%2C+Thilak%22">Fernando, Thilak</searchLink><relatesTo>1</relatesTo><i> thilak.fernando@monash.edu</i><br /><searchLink fieldCode="AR" term="%22Webb%2C+Geoffrey%22">Webb, Geoffrey</searchLink><relatesTo>1</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Data+Mining+%26+Knowledge+Discovery%22">Data Mining & Knowledge Discovery</searchLink>. Jan2017, Vol. 31 Issue 1, p264-286. 23p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Information+retrieval+software%22">Information retrieval software</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Cluster+analysis+%28Statistics%29%22">Cluster analysis (Statistics)</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Similarity measures are central to many machine learning algorithms. There are many different similarity measures, each catering for different applications and data requirements. Most similarity measures used with numerical data assume that the attributes are interval scale. In the interval scale, it is assumed that a unit difference has the same meaning irrespective of the magnitudes of the values separated. When this assumption is violated, accuracy may be reduced. Our experiments show that removing the interval scale assumption by transforming data to ranks can improve the accuracy of distance-based similarity measures on some tasks. However the rank transform has high time and storage overheads. In this paper, we introduce an efficient similarity measure which does not consider the magnitudes of inter-instance distances. We compare the new similarity measure with popular similarity measures in two applications: DBScan clustering and content based multimedia information retrieval with real world datasets and different transform functions. The results show that the proposed similarity measure provides good performance on a range of tasks and is invariant to violations of the interval scale assumption. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=120690514
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10618-016-0463-0
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 23
        StartPage: 264
    Subjects:
      – SubjectFull: Information retrieval software
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Cluster analysis (Statistics)
        Type: general
    Titles:
      – TitleFull: SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Fernando, Thilak
      – PersonEntity:
          Name:
            NameFull: Webb, Geoffrey
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: Jan2017
              Type: published
              Y: 2017
          Identifiers:
            – Type: issn-print
              Value: 13845810
          Numbering:
            – Type: volume
              Value: 31
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Data Mining & Knowledge Discovery
              Type: main
ResultId 1