Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points.

Saved in:
Bibliographic Details
Title: Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points.
Authors: Yang, Yunlai1 (AUTHOR) yunlaiyang@hotmail.com, Wan, Zhenzhu2 (AUTHOR), Habib, Mohammad Rezwan (AUTHOR) mohabib@wiley.com
Source: Journal of Applied Mathematics. 7/11/2026, Vol. 2026, p1-10. 10p.
Subjects: Pearson correlation (Statistics)
Abstract: In biomedical sciences, agricultural sciences, and geosciences, similarities are assessed among samples by using their compositional data for various applications. Pearson correlation coefficient and cosine similarity are commonly applied for these assessments. In this paper, we demonstrate that, for one type of samples with a data range over multiple orders of magnitude and clustered data points, it is not proper to use Pearson correlation coefficient or cosine similarity to measure the similarity. This is because the effect of individual data points in the cluster is suppressed, i.e., not equally treated when implementing Pearson correlation analysis; and the effect of low value data points is reduced in cosine similarity analysis. To properly assess the similarity for this special type of samples, based on the meaning of similarity, we propose a new similarity coefficient. Similarity is actually about the compositional proportion of two samples, the closer the compositional proportion among the data points, the higher the similarity between the two samples. Therefore, the similarity of the ratios (gradients), which are the measure of compositional proportion of the data points between two samples, can be used to measure their similarity. Because the gradients are independent of actual values of individual data points, the effect of each data point on the similarity coefficient is treated equally. This new similarity coefficient is thus scale‐independent and not affected by data clustering. Therefore, the new similarity coefficient can represent similarity more accurately than Pearson correlation coefficient or cosine similarity for this type of samples. The limitation of Pearson correlation coefficient and cosine similarity and advantage of the new similarity coefficient are demonstrated here by analyzing three sets of natural samples. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Applied Mathematics is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 195287950
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yang%2C+Yunlai%22">Yang, Yunlai</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> yunlaiyang@hotmail.com</i><br /><searchLink fieldCode="AR" term="%22Wan%2C+Zhenzhu%22">Wan, Zhenzhu</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Habib%2C+Mohammad+Rezwan%22">Habib, Mohammad Rezwan</searchLink> (AUTHOR)<i> mohabib@wiley.com</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Applied+Mathematics%22">Journal of Applied Mathematics</searchLink>. 7/11/2026, Vol. 2026, p1-10. 10p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Pearson+correlation+%28Statistics%29%22">Pearson correlation (Statistics)</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: In biomedical sciences, agricultural sciences, and geosciences, similarities are assessed among samples by using their compositional data for various applications. Pearson correlation coefficient and cosine similarity are commonly applied for these assessments. In this paper, we demonstrate that, for one type of samples with a data range over multiple orders of magnitude and clustered data points, it is not proper to use Pearson correlation coefficient or cosine similarity to measure the similarity. This is because the effect of individual data points in the cluster is suppressed, i.e., not equally treated when implementing Pearson correlation analysis; and the effect of low value data points is reduced in cosine similarity analysis. To properly assess the similarity for this special type of samples, based on the meaning of similarity, we propose a new similarity coefficient. Similarity is actually about the compositional proportion of two samples, the closer the compositional proportion among the data points, the higher the similarity between the two samples. Therefore, the similarity of the ratios (gradients), which are the measure of compositional proportion of the data points between two samples, can be used to measure their similarity. Because the gradients are independent of actual values of individual data points, the effect of each data point on the similarity coefficient is treated equally. This new similarity coefficient is thus scale‐independent and not affected by data clustering. Therefore, the new similarity coefficient can represent similarity more accurately than Pearson correlation coefficient or cosine similarity for this type of samples. The limitation of Pearson correlation coefficient and cosine similarity and advantage of the new similarity coefficient are demonstrated here by analyzing three sets of natural samples. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Applied Mathematics is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=195287950
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1155/jama/6492494
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 10
        StartPage: 1
    Subjects:
      – SubjectFull: Pearson correlation (Statistics)
        Type: general
    Titles:
      – TitleFull: Similarity Coefficient Based on Gradient Deviation for Samples With Data Ranges Over Multiple Orders of Magnitude and Clustered Data Points.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yang, Yunlai
      – PersonEntity:
          Name:
            NameFull: Wan, Zhenzhu
      – PersonEntity:
          Name:
            NameFull: Habib, Mohammad Rezwan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 11
              M: 07
              Text: 7/11/2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 1110757X
          Numbering:
            – Type: volume
              Value: 2026
          Titles:
            – TitleFull: Journal of Applied Mathematics
              Type: main
ResultId 1