One- to Four-Dimensional Kernels for Virtual Screening and the Prediction of Physical, Chemical, and Biological Properties.
Saved in:
| Title: | One- to Four-Dimensional Kernels for Virtual Screening and the Prediction of Physical, Chemical, and Biological Properties. |
|---|---|
| Authors: | Chloé-Agathe Azencott1, Alexandre Ksikes1, S. Joshua Swamidass1, Jonathan H. Chen1, Liva Ralaivola1, Pierre Baldi1 |
| Source: | Journal of Chemical Information & Modeling. May2007, Vol. 47 Issue 3, p965-974. 10p. |
| Subjects: | Logical prediction, High throughput screening (Drug development), Cheminformatics, Kernel functions, Machine learning, Molecular structure, Regression analysis, Properties of matter |
| Abstract: | Many chemoinformatics applications, including high-throughput virtual screening, benefit from being able to rapidly predict the physical, chemical, and biological properties of small molecules to screen large repositories and identify suitable candidates. When training sets are available, machine learning methods provide an effective alternative to ab initio methods for these predictions. Here, we leverage rich molecular representations including 1D SMILES strings, 2D graphs of bonds, and 3D coordinates to derive efficient machine learning kernels to address regression problems. We further expand the library of available spectral kernels for small molecules developed for classification problems to include 2.5D surface and 3D kernels using Delaunay tetrahedrization and other techniques from computational geometry, 3D pharmacophore kernels, and 3.5D or 4D kernels capable of taking into account multiple molecular configurations, such as conformers. The kernels are comprehensively tested using cross-validation and redundancy-reduction methods on regression problems using several available data sets to predict boiling points, melting points, aqueous solubility, octanol/water partition coefficients, and biological activity with state-of-the art results. When sufficient training data are available, 2D spectral kernels in general tend to yield the best and most robust results, better than state-of-the art. On data sets containing thousands of molecules, the kernels achieve a squared correlation coefficient of 0.91 for aqueous solubility prediction and 0.94 for octanol/water partition coefficient prediction. Averaging over conformations improves the performance of kernels based on the three-dimensional structure of molecules, especially on challenging data sets. Kernel predictors for aqueous solubility (kSOL), LogP (kLOGP), and melting point (kMELT) are available over the Web through: http://cdb.ics.uci.edu. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Chemical Information & Modeling is the property of American Chemical Society and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 35371851 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: One- to Four-Dimensional Kernels for Virtual Screening and the Prediction of Physical, Chemical, and Biological Properties. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Chloé-Agathe+Azencott%22">Chloé-Agathe Azencott</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Alexandre+Ksikes%22">Alexandre Ksikes</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22S%2E+Joshua+Swamidass%22">S. Joshua Swamidass</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Jonathan+H%2E+Chen%22">Jonathan H. Chen</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Liva+Ralaivola%22">Liva Ralaivola</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Pierre+Baldi%22">Pierre Baldi</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Chemical+Information+%26+Modeling%22">Journal of Chemical Information & Modeling</searchLink>. May2007, Vol. 47 Issue 3, p965-974. 10p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Logical+prediction%22">Logical prediction</searchLink><br /><searchLink fieldCode="DE" term="%22High+throughput+screening+%28Drug+development%29%22">High throughput screening (Drug development)</searchLink><br /><searchLink fieldCode="DE" term="%22Cheminformatics%22">Cheminformatics</searchLink><br /><searchLink fieldCode="DE" term="%22Kernel+functions%22">Kernel functions</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Molecular+structure%22">Molecular structure</searchLink><br /><searchLink fieldCode="DE" term="%22Regression+analysis%22">Regression analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Properties+of+matter%22">Properties of matter</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Many chemoinformatics applications, including high-throughput virtual screening, benefit from being able to rapidly predict the physical, chemical, and biological properties of small molecules to screen large repositories and identify suitable candidates. When training sets are available, machine learning methods provide an effective alternative to ab initio methods for these predictions. Here, we leverage rich molecular representations including 1D SMILES strings, 2D graphs of bonds, and 3D coordinates to derive efficient machine learning kernels to address regression problems. We further expand the library of available spectral kernels for small molecules developed for classification problems to include 2.5D surface and 3D kernels using Delaunay tetrahedrization and other techniques from computational geometry, 3D pharmacophore kernels, and 3.5D or 4D kernels capable of taking into account multiple molecular configurations, such as conformers. The kernels are comprehensively tested using cross-validation and redundancy-reduction methods on regression problems using several available data sets to predict boiling points, melting points, aqueous solubility, octanol/water partition coefficients, and biological activity with state-of-the art results. When sufficient training data are available, 2D spectral kernels in general tend to yield the best and most robust results, better than state-of-the art. On data sets containing thousands of molecules, the kernels achieve a squared correlation coefficient of 0.91 for aqueous solubility prediction and 0.94 for octanol/water partition coefficient prediction. Averaging over conformations improves the performance of kernels based on the three-dimensional structure of molecules, especially on challenging data sets. Kernel predictors for aqueous solubility (kSOL), LogP (kLOGP), and melting point (kMELT) are available over the Web through: http://cdb.ics.uci.edu. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Chemical Information & Modeling is the property of American Chemical Society and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=35371851 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1021/ci600397p Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 10 StartPage: 965 Subjects: – SubjectFull: Logical prediction Type: general – SubjectFull: High throughput screening (Drug development) Type: general – SubjectFull: Cheminformatics Type: general – SubjectFull: Kernel functions Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Molecular structure Type: general – SubjectFull: Regression analysis Type: general – SubjectFull: Properties of matter Type: general Titles: – TitleFull: One- to Four-Dimensional Kernels for Virtual Screening and the Prediction of Physical, Chemical, and Biological Properties. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Chloé-Agathe Azencott – PersonEntity: Name: NameFull: Alexandre Ksikes – PersonEntity: Name: NameFull: S. Joshua Swamidass – PersonEntity: Name: NameFull: Jonathan H. Chen – PersonEntity: Name: NameFull: Liva Ralaivola – PersonEntity: Name: NameFull: Pierre Baldi IsPartOfRelationships: – BibEntity: Dates: – D: 29 M: 05 Text: May2007 Type: published Y: 2007 Identifiers: – Type: issn-print Value: 15499596 Numbering: – Type: volume Value: 47 – Type: issue Value: 3 Titles: – TitleFull: Journal of Chemical Information & Modeling Type: main |
| ResultId | 1 |