Enhanced U-Net architectures for accurate room impulse response generation via differential-phase learning.
Saved in:
| Title: | Enhanced U-Net architectures for accurate room impulse response generation via differential-phase learning. |
|---|---|
| Authors: | Martin-Salinas, Ignacio1 (AUTHOR) ignamart@ing.uc3m.es, Piñero, Gema2 (AUTHOR) gpinyero@iteam.upv.es, Belloch, Jose A.1 (AUTHOR) jbelloc@ing.uc3m.es, Amor-Martin, Adrian3 (AUTHOR) aamor@ing.uc3m.es |
| Source: | EURASIP Journal on Audio Speech & Music Processing. 11/17/2025, Vol. 2025 Issue 1, p1-15. 15p. |
| Subjects: | Phase estimation (Electronics), Deep learning, Loss functions (Statistics), Architectural acoustics, Latent variables |
| Abstract: | Generating accurate room impulse responses (RIRs) remains challenging, particularly regarding phase estimation. Building upon previous work utilizing encoder-decoder deep learning architectures, this paper investigates advanced techniques to improve phase prediction accuracy. We propose and evaluate several enhanced U-Net models, including variants with a variational autoencoder (VAE) bottleneck and differing input conditioning methods for spatial and room parameters (embedding layers vs. normalized dense layers). A key focus is the comparison between predicting direct phase and differential phase. Furthermore, we analyze the impact of using mean absolute error (MAE) versus mean squared error (MSE) for the magnitude component of the loss function. The study also explores the efficacy of applying the Griffin-Lim algorithm as a post-processing step to refine the phase estimated by the networks. Performance is evaluated on a real RIR dataset, comparing the different model architectures, information vector encoding strategies, phase targets (direct vs. differential), loss functions, and the contribution of phase recovery algorithms to overall RIR fidelity. Results provide insights into effective strategies for enhancing phase generation in data-driven RIR synthesis. [ABSTRACT FROM AUTHOR] |
| Copyright of EURASIP Journal on Audio Speech & Music Processing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
Be the first to leave a comment!