When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values.

Saved in:
Bibliographic Details
Title: When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values.
Authors: Pierre Baldi1, Ramzi Nasr1
Source: Journal of Chemical Information & Modeling. Jul2010, Vol. 50 Issue 7, p1205-1222. 18p.
Subjects: Distribution (Probability theory), Gaussian distribution, Extreme value theory, Random variables, Intersection theory, Weibull distribution
Abstract: As repositories of chemical molecules continue to expand and become more open, it becomes increasingly important to develop tools to search them efficiently and assess the statistical significance of chemical similarity scores. Here, we develop a general framework for understanding, modeling, predicting, and approximating the distribution of chemical similarity scores and its extreme values in large databases. The framework can be applied to different chemical representations and similarity measures but is demonstrated here using the most common binary fingerprints with the Tanimoto similarity measure. After introducing several probabilistic models of fingerprints, including the Conditional Gaussian Uniform model, we show that the distribution of Tanimoto scores can be approximated by the distribution of the ratio of two correlated Normal random variables associated with the corresponding unions and intersections. This remains true also when the distribution of similarity scores is conditioned on the size of the query molecules to derive more fine-grained results and improve chemical retrieval. The corresponding extreme value distributions for the maximum scores are approximated by Weibull distributions. From these various distributions and their analytical forms, Z-scores, E-values, and p-values are derived to assess the significance of similarity scores. In addition, the framework also allows one to predict the value of standard chemical retrieval metrics, such as sensitivity and specificity at fixed thresholds, or receiver operating characteristic (ROC) curves at multiple thresholds, and to detect outliers in the form of atypical molecules. Numerous and diverse experiments that have been performed, in part with large sets of molecules from the ChemDB, show remarkable agreement between theory and empirical results. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Chemical Information & Modeling is the property of American Chemical Society and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 52862046
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Pierre+Baldi%22">Pierre Baldi</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Ramzi+Nasr%22">Ramzi Nasr</searchLink><relatesTo>1</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Chemical+Information+%26+Modeling%22">Journal of Chemical Information & Modeling</searchLink>. Jul2010, Vol. 50 Issue 7, p1205-1222. 18p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Distribution+%28Probability+theory%29%22">Distribution (Probability theory)</searchLink><br /><searchLink fieldCode="DE" term="%22Gaussian+distribution%22">Gaussian distribution</searchLink><br /><searchLink fieldCode="DE" term="%22Extreme+value+theory%22">Extreme value theory</searchLink><br /><searchLink fieldCode="DE" term="%22Random+variables%22">Random variables</searchLink><br /><searchLink fieldCode="DE" term="%22Intersection+theory%22">Intersection theory</searchLink><br /><searchLink fieldCode="DE" term="%22Weibull+distribution%22">Weibull distribution</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: As repositories of chemical molecules continue to expand and become more open, it becomes increasingly important to develop tools to search them efficiently and assess the statistical significance of chemical similarity scores. Here, we develop a general framework for understanding, modeling, predicting, and approximating the distribution of chemical similarity scores and its extreme values in large databases. The framework can be applied to different chemical representations and similarity measures but is demonstrated here using the most common binary fingerprints with the Tanimoto similarity measure. After introducing several probabilistic models of fingerprints, including the Conditional Gaussian Uniform model, we show that the distribution of Tanimoto scores can be approximated by the distribution of the ratio of two correlated Normal random variables associated with the corresponding unions and intersections. This remains true also when the distribution of similarity scores is conditioned on the size of the query molecules to derive more fine-grained results and improve chemical retrieval. The corresponding extreme value distributions for the maximum scores are approximated by Weibull distributions. From these various distributions and their analytical forms, Z-scores, E-values, and p-values are derived to assess the significance of similarity scores. In addition, the framework also allows one to predict the value of standard chemical retrieval metrics, such as sensitivity and specificity at fixed thresholds, or receiver operating characteristic (ROC) curves at multiple thresholds, and to detect outliers in the form of atypical molecules. Numerous and diverse experiments that have been performed, in part with large sets of molecules from the ChemDB, show remarkable agreement between theory and empirical results. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Chemical Information & Modeling is the property of American Chemical Society and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=52862046
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1021/ci100010v
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 18
        StartPage: 1205
    Subjects:
      – SubjectFull: Distribution (Probability theory)
        Type: general
      – SubjectFull: Gaussian distribution
        Type: general
      – SubjectFull: Extreme value theory
        Type: general
      – SubjectFull: Random variables
        Type: general
      – SubjectFull: Intersection theory
        Type: general
      – SubjectFull: Weibull distribution
        Type: general
    Titles:
      – TitleFull: When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Pierre Baldi
      – PersonEntity:
          Name:
            NameFull: Ramzi Nasr
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 26
              M: 07
              Text: Jul2010
              Type: published
              Y: 2010
          Identifiers:
            – Type: issn-print
              Value: 15499596
          Numbering:
            – Type: volume
              Value: 50
            – Type: issue
              Value: 7
          Titles:
            – TitleFull: Journal of Chemical Information & Modeling
              Type: main
ResultId 1