Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets.
Saved in:
| Title: | Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets. |
|---|---|
| Authors: | Barone, Matthew1, Ray, Jaideep2, Domino, Stefan3,4 |
| Source: | AIAA Journal. Mar2022, Vol. 60 Issue 3, p1332-1346. 15p. |
| Abstract: | This paper explores automated approaches for the analysis and categorization of turbulent flow data as a means of assessing the quality of a turbulence dataset used for constructing data-driven turbulence closures. Single-point statistics from several high-fidelity turbulent flow simulation datasets are differentiated into groups using a Gaussian mixture model clustering algorithm. Candidate features are proposed, and a feature selection algorithm is applied to the data in a sequential fashion, flow by flow, to identify a good feature set and an optimal number of clusters for each dataset. Clusters are first identified for plane channel flows, producing results that agree with existing theory and empirical observations. Further clusters are then identified in an incremental fashion for flow over a wavy-walled channel, flow over a bump in a channel, and flow past a square cylinder. Some clusters are closely identified with the anisotropy state of the turbulence, whereas others can be connected to physical phenomena, such as boundary-layer separation and free shear layers. Exemplar points from the clusters, or prototypes, are then identified using a prototype placement method. These exemplars effectively summarize the dataset using a greatly reduced collection of data points. The clusters and their prototypes are used to assess the quality of a training dataset constructed by simply pooling the four flows. We enumerate the dataset's shortcomings and state the limits of generalizability of any data-driven closure trained on it. [ABSTRACT FROM AUTHOR] |
| Copyright of AIAA Journal is the property of American Institute of Aeronautics & Astronautics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 175690020 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Barone%2C+Matthew%22">Barone, Matthew</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Ray%2C+Jaideep%22">Ray, Jaideep</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Domino%2C+Stefan%22">Domino, Stefan</searchLink><relatesTo>3,4</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22AIAA+Journal%22">AIAA Journal</searchLink>. Mar2022, Vol. 60 Issue 3, p1332-1346. 15p. – Name: Abstract Label: Abstract Group: Ab Data: This paper explores automated approaches for the analysis and categorization of turbulent flow data as a means of assessing the quality of a turbulence dataset used for constructing data-driven turbulence closures. Single-point statistics from several high-fidelity turbulent flow simulation datasets are differentiated into groups using a Gaussian mixture model clustering algorithm. Candidate features are proposed, and a feature selection algorithm is applied to the data in a sequential fashion, flow by flow, to identify a good feature set and an optimal number of clusters for each dataset. Clusters are first identified for plane channel flows, producing results that agree with existing theory and empirical observations. Further clusters are then identified in an incremental fashion for flow over a wavy-walled channel, flow over a bump in a channel, and flow past a square cylinder. Some clusters are closely identified with the anisotropy state of the turbulence, whereas others can be connected to physical phenomena, such as boundary-layer separation and free shear layers. Exemplar points from the clusters, or prototypes, are then identified using a prototype placement method. These exemplars effectively summarize the dataset using a greatly reduced collection of data points. The clusters and their prototypes are used to assess the quality of a training dataset constructed by simply pooling the four flows. We enumerate the dataset's shortcomings and state the limits of generalizability of any data-driven closure trained on it. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of AIAA Journal is the property of American Institute of Aeronautics & Astronautics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=175690020 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.2514/1.J060919 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 15 StartPage: 1332 Titles: – TitleFull: Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Barone, Matthew – PersonEntity: Name: NameFull: Ray, Jaideep – PersonEntity: Name: NameFull: Domino, Stefan IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 03 Text: Mar2022 Type: published Y: 2022 Identifiers: – Type: issn-print Value: 00011452 Numbering: – Type: volume Value: 60 – Type: issue Value: 3 Titles: – TitleFull: AIAA Journal Type: main |
| ResultId | 1 |