Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets.

Saved in:
Bibliographic Details
Title: Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets.
Authors: Barone, Matthew1, Ray, Jaideep2, Domino, Stefan3,4
Source: AIAA Journal. Mar2022, Vol. 60 Issue 3, p1332-1346. 15p.
Abstract: This paper explores automated approaches for the analysis and categorization of turbulent flow data as a means of assessing the quality of a turbulence dataset used for constructing data-driven turbulence closures. Single-point statistics from several high-fidelity turbulent flow simulation datasets are differentiated into groups using a Gaussian mixture model clustering algorithm. Candidate features are proposed, and a feature selection algorithm is applied to the data in a sequential fashion, flow by flow, to identify a good feature set and an optimal number of clusters for each dataset. Clusters are first identified for plane channel flows, producing results that agree with existing theory and empirical observations. Further clusters are then identified in an incremental fashion for flow over a wavy-walled channel, flow over a bump in a channel, and flow past a square cylinder. Some clusters are closely identified with the anisotropy state of the turbulence, whereas others can be connected to physical phenomena, such as boundary-layer separation and free shear layers. Exemplar points from the clusters, or prototypes, are then identified using a prototype placement method. These exemplars effectively summarize the dataset using a greatly reduced collection of data points. The clusters and their prototypes are used to assess the quality of a training dataset constructed by simply pooling the four flows. We enumerate the dataset's shortcomings and state the limits of generalizability of any data-driven closure trained on it. [ABSTRACT FROM AUTHOR]
Copyright of AIAA Journal is the property of American Institute of Aeronautics & Astronautics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 175690020
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Barone%2C+Matthew%22">Barone, Matthew</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Ray%2C+Jaideep%22">Ray, Jaideep</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Domino%2C+Stefan%22">Domino, Stefan</searchLink><relatesTo>3,4</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22AIAA+Journal%22">AIAA Journal</searchLink>. Mar2022, Vol. 60 Issue 3, p1332-1346. 15p.
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This paper explores automated approaches for the analysis and categorization of turbulent flow data as a means of assessing the quality of a turbulence dataset used for constructing data-driven turbulence closures. Single-point statistics from several high-fidelity turbulent flow simulation datasets are differentiated into groups using a Gaussian mixture model clustering algorithm. Candidate features are proposed, and a feature selection algorithm is applied to the data in a sequential fashion, flow by flow, to identify a good feature set and an optimal number of clusters for each dataset. Clusters are first identified for plane channel flows, producing results that agree with existing theory and empirical observations. Further clusters are then identified in an incremental fashion for flow over a wavy-walled channel, flow over a bump in a channel, and flow past a square cylinder. Some clusters are closely identified with the anisotropy state of the turbulence, whereas others can be connected to physical phenomena, such as boundary-layer separation and free shear layers. Exemplar points from the clusters, or prototypes, are then identified using a prototype placement method. These exemplars effectively summarize the dataset using a greatly reduced collection of data points. The clusters and their prototypes are used to assess the quality of a training dataset constructed by simply pooling the four flows. We enumerate the dataset's shortcomings and state the limits of generalizability of any data-driven closure trained on it. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of AIAA Journal is the property of American Institute of Aeronautics & Astronautics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=175690020
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.2514/1.J060919
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 15
        StartPage: 1332
    Titles:
      – TitleFull: Feature Selection, Clustering, and Prototype Placement for Turbulence Datasets.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Barone, Matthew
      – PersonEntity:
          Name:
            NameFull: Ray, Jaideep
      – PersonEntity:
          Name:
            NameFull: Domino, Stefan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: Mar2022
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 00011452
          Numbering:
            – Type: volume
              Value: 60
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: AIAA Journal
              Type: main
ResultId 1