Generalized rough sets based feature selection.
Saved in:
| Title: | Generalized rough sets based feature selection. |
|---|---|
| Authors: | Mohamed Quafafou, Moussa Boussouf |
| Source: | Intelligent Data Analysis. 2000, Vol. 4 Issue 1, p3. 15p. |
| Subjects: | Selection theorems, Algorithms, Machine learning |
| Abstract: | The problem of feature subset selection can be defined as the selection of a relevant subset of features which allows a learning algorithm to induce small high-accuracy models. This problem is of primary important because irrelevant and redundant features may degrade the learner speed, especially in the context of high dimensionality, and reduce both the accuracy and comprehensibility of the induced model. Two main approaches have been developed, the first one is algorithm-independent (filter approach) which considers only the data, when the second approach which is algorithm-dependent takes into account both the data and a given learning algorithm (wrapper approach). Recent work was developed to study the interest of the rough set theory and more particularly its notions of reducts and core to deal with the problem of feature subset selection. Different methods were proposed to select features using both the core and the reduct concepts, whereas other researches show that useful feature subsets do not necessarily contain all features in cores. In this paper, we underline the fact that rough set theory is concerned with deterministic analysis of attribute dependencies which are at the basis of the two notions of reduct and core. We extend the notion of dependency which allows to find both deterministic and non-deterministic dependencies. A new notion of strong reducts is then introduced and leads to the definition of strong feature subsets (SFS). The interest of SFS is illustrated by the improvement of the accuracy of C4.5 on real-world datasets. Our study shows that generally the highest-accuracy-subset is not the best one as regards to the filter criteria. The highest accuracy subset is found by the new approach with minimum cost. The contribution of this work is four folds : (1) analysis of feature subset selection in the rough sets context, (2) introduction of new definitions based on a generalized rough set theory,.... [ABSTRACT FROM AUTHOR] |
| Copyright of Intelligent Data Analysis is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 4832210 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Generalized rough sets based feature selection. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Mohamed+Quafafou%22">Mohamed Quafafou</searchLink><br /><searchLink fieldCode="AR" term="%22Moussa+Boussouf%22">Moussa Boussouf</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Intelligent+Data+Analysis%22">Intelligent Data Analysis</searchLink>. 2000, Vol. 4 Issue 1, p3. 15p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Selection+theorems%22">Selection theorems</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: The problem of feature subset selection can be defined as the selection of a relevant subset of features which allows a learning algorithm to induce small high-accuracy models. This problem is of primary important because irrelevant and redundant features may degrade the learner speed, especially in the context of high dimensionality, and reduce both the accuracy and comprehensibility of the induced model. Two main approaches have been developed, the first one is algorithm-independent (filter approach) which considers only the data, when the second approach which is algorithm-dependent takes into account both the data and a given learning algorithm (wrapper approach). Recent work was developed to study the interest of the rough set theory and more particularly its notions of reducts and core to deal with the problem of feature subset selection. Different methods were proposed to select features using both the core and the reduct concepts, whereas other researches show that useful feature subsets do not necessarily contain all features in cores. In this paper, we underline the fact that rough set theory is concerned with deterministic analysis of attribute dependencies which are at the basis of the two notions of reduct and core. We extend the notion of dependency which allows to find both deterministic and non-deterministic dependencies. A new notion of strong reducts is then introduced and leads to the definition of strong feature subsets (SFS). The interest of SFS is illustrated by the improvement of the accuracy of C4.5 on real-world datasets. Our study shows that generally the highest-accuracy-subset is not the best one as regards to the filter criteria. The highest accuracy subset is found by the new approach with minimum cost. The contribution of this work is four folds : (1) analysis of feature subset selection in the rough sets context, (2) introduction of new definitions based on a generalized rough set theory,.... [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Intelligent Data Analysis is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=4832210 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.3233/IDA-2000-4102 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 15 StartPage: 3 Subjects: – SubjectFull: Selection theorems Type: general – SubjectFull: Algorithms Type: general – SubjectFull: Machine learning Type: general Titles: – TitleFull: Generalized rough sets based feature selection. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Mohamed Quafafou – PersonEntity: Name: NameFull: Moussa Boussouf IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 02 Text: 2000 Type: published Y: 2000 Identifiers: – Type: issn-print Value: 1088467X Numbering: – Type: volume Value: 4 – Type: issue Value: 1 Titles: – TitleFull: Intelligent Data Analysis Type: main |
| ResultId | 1 |