When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values.
Saved in:
| Title: | When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values. |
|---|---|
| Authors: | Pierre Baldi1, Ramzi Nasr1 |
| Source: | Journal of Chemical Information & Modeling. Jul2010, Vol. 50 Issue 7, p1205-1222. 18p. |
| Subjects: | Distribution (Probability theory), Gaussian distribution, Extreme value theory, Random variables, Intersection theory, Weibull distribution |
| Abstract: | As repositories of chemical molecules continue to expand and become more open, it becomes increasingly important to develop tools to search them efficiently and assess the statistical significance of chemical similarity scores. Here, we develop a general framework for understanding, modeling, predicting, and approximating the distribution of chemical similarity scores and its extreme values in large databases. The framework can be applied to different chemical representations and similarity measures but is demonstrated here using the most common binary fingerprints with the Tanimoto similarity measure. After introducing several probabilistic models of fingerprints, including the Conditional Gaussian Uniform model, we show that the distribution of Tanimoto scores can be approximated by the distribution of the ratio of two correlated Normal random variables associated with the corresponding unions and intersections. This remains true also when the distribution of similarity scores is conditioned on the size of the query molecules to derive more fine-grained results and improve chemical retrieval. The corresponding extreme value distributions for the maximum scores are approximated by Weibull distributions. From these various distributions and their analytical forms, Z-scores, E-values, and p-values are derived to assess the significance of similarity scores. In addition, the framework also allows one to predict the value of standard chemical retrieval metrics, such as sensitivity and specificity at fixed thresholds, or receiver operating characteristic (ROC) curves at multiple thresholds, and to detect outliers in the form of atypical molecules. Numerous and diverse experiments that have been performed, in part with large sets of molecules from the ChemDB, show remarkable agreement between theory and empirical results. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Chemical Information & Modeling is the property of American Chemical Society and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 52862046 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Pierre+Baldi%22">Pierre Baldi</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Ramzi+Nasr%22">Ramzi Nasr</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Chemical+Information+%26+Modeling%22">Journal of Chemical Information & Modeling</searchLink>. Jul2010, Vol. 50 Issue 7, p1205-1222. 18p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Distribution+%28Probability+theory%29%22">Distribution (Probability theory)</searchLink><br /><searchLink fieldCode="DE" term="%22Gaussian+distribution%22">Gaussian distribution</searchLink><br /><searchLink fieldCode="DE" term="%22Extreme+value+theory%22">Extreme value theory</searchLink><br /><searchLink fieldCode="DE" term="%22Random+variables%22">Random variables</searchLink><br /><searchLink fieldCode="DE" term="%22Intersection+theory%22">Intersection theory</searchLink><br /><searchLink fieldCode="DE" term="%22Weibull+distribution%22">Weibull distribution</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: As repositories of chemical molecules continue to expand and become more open, it becomes increasingly important to develop tools to search them efficiently and assess the statistical significance of chemical similarity scores. Here, we develop a general framework for understanding, modeling, predicting, and approximating the distribution of chemical similarity scores and its extreme values in large databases. The framework can be applied to different chemical representations and similarity measures but is demonstrated here using the most common binary fingerprints with the Tanimoto similarity measure. After introducing several probabilistic models of fingerprints, including the Conditional Gaussian Uniform model, we show that the distribution of Tanimoto scores can be approximated by the distribution of the ratio of two correlated Normal random variables associated with the corresponding unions and intersections. This remains true also when the distribution of similarity scores is conditioned on the size of the query molecules to derive more fine-grained results and improve chemical retrieval. The corresponding extreme value distributions for the maximum scores are approximated by Weibull distributions. From these various distributions and their analytical forms, Z-scores, E-values, and p-values are derived to assess the significance of similarity scores. In addition, the framework also allows one to predict the value of standard chemical retrieval metrics, such as sensitivity and specificity at fixed thresholds, or receiver operating characteristic (ROC) curves at multiple thresholds, and to detect outliers in the form of atypical molecules. Numerous and diverse experiments that have been performed, in part with large sets of molecules from the ChemDB, show remarkable agreement between theory and empirical results. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Chemical Information & Modeling is the property of American Chemical Society and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=52862046 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1021/ci100010v Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 18 StartPage: 1205 Subjects: – SubjectFull: Distribution (Probability theory) Type: general – SubjectFull: Gaussian distribution Type: general – SubjectFull: Extreme value theory Type: general – SubjectFull: Random variables Type: general – SubjectFull: Intersection theory Type: general – SubjectFull: Weibull distribution Type: general Titles: – TitleFull: When is Chemical Similarity Significant? The Statistical Distribution of Chemical Similarity Scores and Its Extreme Values. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Pierre Baldi – PersonEntity: Name: NameFull: Ramzi Nasr IsPartOfRelationships: – BibEntity: Dates: – D: 26 M: 07 Text: Jul2010 Type: published Y: 2010 Identifiers: – Type: issn-print Value: 15499596 Numbering: – Type: volume Value: 50 – Type: issue Value: 7 Titles: – TitleFull: Journal of Chemical Information & Modeling Type: main |
| ResultId | 1 |