Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points.
Saved in:
| Title: | Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points. |
|---|---|
| Authors: | Yang, Yunlai1 (AUTHOR) yunlaiyang@hotmail.com, Wan, Zhenzhu2 (AUTHOR), Habib, Mohammad Rezwan (AUTHOR) mohabib@wiley.com |
| Source: | Journal of Applied Mathematics. 7/11/2026, Vol. 2026, p1-10. 10p. |
| Subjects: | Pearson correlation (Statistics) |
| Abstract: | In biomedical sciences, agricultural sciences, and geosciences, similarities are assessed among samples by using their compositional data for various applications. Pearson correlation coefficient and cosine similarity are commonly applied for these assessments. In this paper, we demonstrate that, for one type of samples with a data range over multiple orders of magnitude and clustered data points, it is not proper to use Pearson correlation coefficient or cosine similarity to measure the similarity. This is because the effect of individual data points in the cluster is suppressed, i.e., not equally treated when implementing Pearson correlation analysis; and the effect of low value data points is reduced in cosine similarity analysis. To properly assess the similarity for this special type of samples, based on the meaning of similarity, we propose a new similarity coefficient. Similarity is actually about the compositional proportion of two samples, the closer the compositional proportion among the data points, the higher the similarity between the two samples. Therefore, the similarity of the ratios (gradients), which are the measure of compositional proportion of the data points between two samples, can be used to measure their similarity. Because the gradients are independent of actual values of individual data points, the effect of each data point on the similarity coefficient is treated equally. This new similarity coefficient is thus scale‐independent and not affected by data clustering. Therefore, the new similarity coefficient can represent similarity more accurately than Pearson correlation coefficient or cosine similarity for this type of samples. The limitation of Pearson correlation coefficient and cosine similarity and advantage of the new similarity coefficient are demonstrated here by analyzing three sets of natural samples. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Applied Mathematics is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 195287950 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Yang%2C+Yunlai%22">Yang, Yunlai</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> yunlaiyang@hotmail.com</i><br /><searchLink fieldCode="AR" term="%22Wan%2C+Zhenzhu%22">Wan, Zhenzhu</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Habib%2C+Mohammad+Rezwan%22">Habib, Mohammad Rezwan</searchLink> (AUTHOR)<i> mohabib@wiley.com</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Applied+Mathematics%22">Journal of Applied Mathematics</searchLink>. 7/11/2026, Vol. 2026, p1-10. 10p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Pearson+correlation+%28Statistics%29%22">Pearson correlation (Statistics)</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: In biomedical sciences, agricultural sciences, and geosciences, similarities are assessed among samples by using their compositional data for various applications. Pearson correlation coefficient and cosine similarity are commonly applied for these assessments. In this paper, we demonstrate that, for one type of samples with a data range over multiple orders of magnitude and clustered data points, it is not proper to use Pearson correlation coefficient or cosine similarity to measure the similarity. This is because the effect of individual data points in the cluster is suppressed, i.e., not equally treated when implementing Pearson correlation analysis; and the effect of low value data points is reduced in cosine similarity analysis. To properly assess the similarity for this special type of samples, based on the meaning of similarity, we propose a new similarity coefficient. Similarity is actually about the compositional proportion of two samples, the closer the compositional proportion among the data points, the higher the similarity between the two samples. Therefore, the similarity of the ratios (gradients), which are the measure of compositional proportion of the data points between two samples, can be used to measure their similarity. Because the gradients are independent of actual values of individual data points, the effect of each data point on the similarity coefficient is treated equally. This new similarity coefficient is thus scale‐independent and not affected by data clustering. Therefore, the new similarity coefficient can represent similarity more accurately than Pearson correlation coefficient or cosine similarity for this type of samples. The limitation of Pearson correlation coefficient and cosine similarity and advantage of the new similarity coefficient are demonstrated here by analyzing three sets of natural samples. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Applied Mathematics is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=195287950 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1155/jama/6492494 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 10 StartPage: 1 Subjects: – SubjectFull: Pearson correlation (Statistics) Type: general Titles: – TitleFull: Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Yang, Yunlai – PersonEntity: Name: NameFull: Wan, Zhenzhu – PersonEntity: Name: NameFull: Habib, Mohammad Rezwan IsPartOfRelationships: – BibEntity: Dates: – D: 11 M: 07 Text: 7/11/2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 1110757X Numbering: – Type: volume Value: 2026 Titles: – TitleFull: Journal of Applied Mathematics Type: main |
| ResultId | 1 |