Boosting neural network performance for high dimensional data through random projections.

Saved in:
Bibliographic Details
Title: Boosting neural network performance for high dimensional data through random projections.
Authors: Anagnostou, Panagiotis1 (AUTHOR) panagno@uth.gr, Tasoulis, Sotiris K.1 (AUTHOR) stasoulis@uth.gr, Vrahatis, Aristidis G.2 (AUTHOR) aris.vrahatis@ionio.gr, Georgakopoulos, Spiros V.3 (AUTHOR) spirosgeorg@uth.gr, Plagianakos, Vassilis P.1 (AUTHOR) vpp@uth.gr
Source: Pattern Recognition Letters. Jan2026, Vol. 199, p149-155. 7p.
Subjects: Random projection method, Data augmentation, Cell analysis, Dimensional reduction algorithms, Multivariate analysis, Artificial neural networks
Abstract: Advancements in molecular Biology have driven a paradigm shift in disease understanding, particularly with the rise of precision Medicine enabled by single-cell studies. These datasets are typically high-dimensional, posing several computational challenges. They are often sensitive to noise, and their sparse structure can lead to issues with overfitting. While Deep Neural Networks can address many of these limitations, they still struggle with high-dimensional tabular data. In response to these challenges, we propose a novel framework that leverages the concept of data augmentation for high-dimensional tabular data. By augmenting the samples and simultaneously reducing their dimensionality, we create a balanced environment that improves Deep Neural Network performance. This augmentation is achieved through the Random Projection method, combined with a stochastic filtering process for the randomly projected spaces. We validate our approach on several scRNA-seq datasets, showing that it not only enhances Deep Neural Network performance, but also outperforms state-of-the-art scRNA-seq classifiers. The proposed methodology offers new opportunities for addressing "small n, large p" problems across diverse domains. • Combine RP and PCA for dimensionality reduction that preserves relative structure. • RP for data augmentation to enhance NN training and Classification performance. • Majority voting applied to augmented test samples further enhances NN performance. [ABSTRACT FROM AUTHOR]
Copyright of Pattern Recognition Letters is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:Advancements in molecular Biology have driven a paradigm shift in disease understanding, particularly with the rise of precision Medicine enabled by single-cell studies. These datasets are typically high-dimensional, posing several computational challenges. They are often sensitive to noise, and their sparse structure can lead to issues with overfitting. While Deep Neural Networks can address many of these limitations, they still struggle with high-dimensional tabular data. In response to these challenges, we propose a novel framework that leverages the concept of data augmentation for high-dimensional tabular data. By augmenting the samples and simultaneously reducing their dimensionality, we create a balanced environment that improves Deep Neural Network performance. This augmentation is achieved through the Random Projection method, combined with a stochastic filtering process for the randomly projected spaces. We validate our approach on several scRNA-seq datasets, showing that it not only enhances Deep Neural Network performance, but also outperforms state-of-the-art scRNA-seq classifiers. The proposed methodology offers new opportunities for addressing "small n, large p" problems across diverse domains. • Combine RP and PCA for dimensionality reduction that preserves relative structure. • RP for data augmentation to enhance NN training and Classification performance. • Majority voting applied to augmented test samples further enhances NN performance. [ABSTRACT FROM AUTHOR]
ISSN:01678655
DOI:10.1016/j.patrec.2025.11.006