SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption.
Saved in:
| Title: | SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption. |
|---|---|
| Authors: | Fernando, Thilak1 thilak.fernando@monash.edu, Webb, Geoffrey1 |
| Source: | Data Mining & Knowledge Discovery. Jan2017, Vol. 31 Issue 1, p264-286. 23p. |
| Subjects: | Information retrieval software, Machine learning, Cluster analysis (Statistics) |
| Abstract: | Similarity measures are central to many machine learning algorithms. There are many different similarity measures, each catering for different applications and data requirements. Most similarity measures used with numerical data assume that the attributes are interval scale. In the interval scale, it is assumed that a unit difference has the same meaning irrespective of the magnitudes of the values separated. When this assumption is violated, accuracy may be reduced. Our experiments show that removing the interval scale assumption by transforming data to ranks can improve the accuracy of distance-based similarity measures on some tasks. However the rank transform has high time and storage overheads. In this paper, we introduce an efficient similarity measure which does not consider the magnitudes of inter-instance distances. We compare the new similarity measure with popular similarity measures in two applications: DBScan clustering and content based multimedia information retrieval with real world datasets and different transform functions. The results show that the proposed similarity measure provides good performance on a range of tasks and is invariant to violations of the interval scale assumption. [ABSTRACT FROM AUTHOR] |
| Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 120690514 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Fernando%2C+Thilak%22">Fernando, Thilak</searchLink><relatesTo>1</relatesTo><i> thilak.fernando@monash.edu</i><br /><searchLink fieldCode="AR" term="%22Webb%2C+Geoffrey%22">Webb, Geoffrey</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Data+Mining+%26+Knowledge+Discovery%22">Data Mining & Knowledge Discovery</searchLink>. Jan2017, Vol. 31 Issue 1, p264-286. 23p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Information+retrieval+software%22">Information retrieval software</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Cluster+analysis+%28Statistics%29%22">Cluster analysis (Statistics)</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Similarity measures are central to many machine learning algorithms. There are many different similarity measures, each catering for different applications and data requirements. Most similarity measures used with numerical data assume that the attributes are interval scale. In the interval scale, it is assumed that a unit difference has the same meaning irrespective of the magnitudes of the values separated. When this assumption is violated, accuracy may be reduced. Our experiments show that removing the interval scale assumption by transforming data to ranks can improve the accuracy of distance-based similarity measures on some tasks. However the rank transform has high time and storage overheads. In this paper, we introduce an efficient similarity measure which does not consider the magnitudes of inter-instance distances. We compare the new similarity measure with popular similarity measures in two applications: DBScan clustering and content based multimedia information retrieval with real world datasets and different transform functions. The results show that the proposed similarity measure provides good performance on a range of tasks and is invariant to violations of the interval scale assumption. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Data Mining & Knowledge Discovery is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=120690514 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10618-016-0463-0 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 23 StartPage: 264 Subjects: – SubjectFull: Information retrieval software Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Cluster analysis (Statistics) Type: general Titles: – TitleFull: SimUSF: an efficient and effective similarity measure that is invariant to violations of the interval scale assumption. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Fernando, Thilak – PersonEntity: Name: NameFull: Webb, Geoffrey IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Text: Jan2017 Type: published Y: 2017 Identifiers: – Type: issn-print Value: 13845810 Numbering: – Type: volume Value: 31 – Type: issue Value: 1 Titles: – TitleFull: Data Mining & Knowledge Discovery Type: main |
| ResultId | 1 |