Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points.

Saved in:
Bibliographic Details
Title: Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points.
Authors: Yang, Yunlai1 (AUTHOR) yunlaiyang@hotmail.com, Wan, Zhenzhu2 (AUTHOR), Habib, Mohammad Rezwan (AUTHOR) mohabib@wiley.com
Source: Journal of Applied Mathematics. 7/11/2026, Vol. 2026, p1-10. 10p.
Subjects: Pearson correlation (Statistics)
Abstract: In biomedical sciences, agricultural sciences, and geosciences, similarities are assessed among samples by using their compositional data for various applications. Pearson correlation coefficient and cosine similarity are commonly applied for these assessments. In this paper, we demonstrate that, for one type of samples with a data range over multiple orders of magnitude and clustered data points, it is not proper to use Pearson correlation coefficient or cosine similarity to measure the similarity. This is because the effect of individual data points in the cluster is suppressed, i.e., not equally treated when implementing Pearson correlation analysis; and the effect of low value data points is reduced in cosine similarity analysis. To properly assess the similarity for this special type of samples, based on the meaning of similarity, we propose a new similarity coefficient. Similarity is actually about the compositional proportion of two samples, the closer the compositional proportion among the data points, the higher the similarity between the two samples. Therefore, the similarity of the ratios (gradients), which are the measure of compositional proportion of the data points between two samples, can be used to measure their similarity. Because the gradients are independent of actual values of individual data points, the effect of each data point on the similarity coefficient is treated equally. This new similarity coefficient is thus scale‐independent and not affected by data clustering. Therefore, the new similarity coefficient can represent similarity more accurately than Pearson correlation coefficient or cosine similarity for this type of samples. The limitation of Pearson correlation coefficient and cosine similarity and advantage of the new similarity coefficient are demonstrated here by analyzing three sets of natural samples. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Applied Mathematics is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:In biomedical sciences, agricultural sciences, and geosciences, similarities are assessed among samples by using their compositional data for various applications. Pearson correlation coefficient and cosine similarity are commonly applied for these assessments. In this paper, we demonstrate that, for one type of samples with a data range over multiple orders of magnitude and clustered data points, it is not proper to use Pearson correlation coefficient or cosine similarity to measure the similarity. This is because the effect of individual data points in the cluster is suppressed, i.e., not equally treated when implementing Pearson correlation analysis; and the effect of low value data points is reduced in cosine similarity analysis. To properly assess the similarity for this special type of samples, based on the meaning of similarity, we propose a new similarity coefficient. Similarity is actually about the compositional proportion of two samples, the closer the compositional proportion among the data points, the higher the similarity between the two samples. Therefore, the similarity of the ratios (gradients), which are the measure of compositional proportion of the data points between two samples, can be used to measure their similarity. Because the gradients are independent of actual values of individual data points, the effect of each data point on the similarity coefficient is treated equally. This new similarity coefficient is thus scale‐independent and not affected by data clustering. Therefore, the new similarity coefficient can represent similarity more accurately than Pearson correlation coefficient or cosine similarity for this type of samples. The limitation of Pearson correlation coefficient and cosine similarity and advantage of the new similarity coefficient are demonstrated here by analyzing three sets of natural samples. [ABSTRACT FROM AUTHOR]
ISSN:1110757X
DOI:10.1155/jama/6492494