PROTEIN FOLD CLASSIFICATION WITH GENETIC ALGORITHMS AND FEATURE SELECTION.

Saved in:
Bibliographic Details
Title: PROTEIN FOLD CLASSIFICATION WITH GENETIC ALGORITHMS AND FEATURE SELECTION.
Authors: CHEN, PENG1 pchen78@scs.howard.edu, LIU, CHUNMEI1 chunmei@scs.howard.edu, BURGE, LEGAND1, MAHMOOD, MOHAMMAD2, SOUTHERLAND, WILLIAM3, GLOSTER, CLAY4
Source: Journal of Bioinformatics & Computational Biology. Oct2009, Vol. 7 Issue 5, p773-788. 16p. 4 Diagrams, 2 Charts, 4 Graphs.
Subjects: Genetic algorithms, Support vector machines, Protein structure, Protein folding, Nuclear magnetic resonance
Abstract: Protein fold classification is a key step to predicting protein tertiary structures. This paper proposes a novel approach based on genetic algorithms and feature selection to classifying protein folds. Our dataset is divided into a training dataset and a test dataset. Each individual for the genetic algorithms represents a selection function of the feature vectors of the training dataset. A support vector machine is applied to each individual to evaluate the fitness value (fold classification rate) of each individual. The aim of the genetic algorithms is to search for the best individual that produces the highest fold classification rate. The best individual is then applied to the feature vectors of the test dataset and a support vector machine is built to classify protein folds based on selected features. Our experimental results on Ding and Dubchak's benchmark dataset of 27-class folds show that our approach achieves an accuracy of 71.28%, which outperforms current state-of-the-art protein fold predictors. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Bioinformatics & Computational Biology is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:Protein fold classification is a key step to predicting protein tertiary structures. This paper proposes a novel approach based on genetic algorithms and feature selection to classifying protein folds. Our dataset is divided into a training dataset and a test dataset. Each individual for the genetic algorithms represents a selection function of the feature vectors of the training dataset. A support vector machine is applied to each individual to evaluate the fitness value (fold classification rate) of each individual. The aim of the genetic algorithms is to search for the best individual that produces the highest fold classification rate. The best individual is then applied to the feature vectors of the test dataset and a support vector machine is built to classify protein folds based on selected features. Our experimental results on Ding and Dubchak's benchmark dataset of 27-class folds show that our approach achieves an accuracy of 71.28%, which outperforms current state-of-the-art protein fold predictors. [ABSTRACT FROM AUTHOR]
ISSN:02197200
DOI:10.1142/S0219720009004321