Event‐Based Spatiotemporal Disentangled Aggregation for Image Reconstruction.
Saved in:
| Title: | Event‐Based Spatiotemporal Disentangled Aggregation for Image Reconstruction. |
|---|---|
| Authors: | Wang, Yanwei1 (AUTHOR), Peng, Chubin1 (AUTHOR), Min, Zhenhui1 (AUTHOR) minzhenhui@usth.edu.cn, Tang, Qingju1 (AUTHOR), Cuccureddu, Floriano1 (AUTHOR) fcuccuredd@wiley.com |
| Source: | International Journal of Intelligent Systems. 5/4/2026, Vol. 2026, p1-22. 22p. |
| Subjects: | Image reconstruction, Signal denoising, Computer vision, Convolutional neural networks, Image sensors, Noise control |
| Abstract: | The dynamic characteristics of event cameras demonstrate unique advantages in extreme environments; however, their inherent spatiotemporal asynchrony and noise sensitivity instead become serious drawbacks during extreme image reconstruction of coal mine substations. To overcome the problems of nonconstant noise coupling and existing static convolutional networks in dealing with vibration‐induced motion artifacts, an event‐based spatiotemporal disentangled aggregated image reconstruction (ESDAR) method is proposed here. ESDAR is a method to enhance FireNet. The proposed method integrates three innovations. First, it introduces an asymptotic denoising approach inspired by the diffusion model, which effectively decouples high‐frequency noise and low‐frequency textures. Second, it employs a deformable convolutional mechanism to achieve spatiotemporal alignment between event streams and fuzzy grayscale images. Third, the recurrent unit is substituted with a time convolution network enhanced with deformable convolution (DF‐TCN), thereby achieving a significant reduction in the number of parameters, with a decrease of 48.32%. The effectiveness of this method was evaluated via 1000 synthetic sequences and 1670 real frames (captured at a coal mine substation). The findings of this study demonstrate that ESDAR is markedly superior to FireNet with respect to performance. The results revealed a 39.5% reduction in the mean square error (MSE), a 4.2 dB increase in the peak signal‐to‐noise ratio (PSNR), a 26% reduction in the local perceptual image similarity metric (LPIPS), and a 3.39% increase in the structural similarity index (SSIM). This method effectively achieves a balance between noise suppression and detail preservation under hardware constraints, providing a feasible solution to address image reconstruction in high‐noise, vibration‐intensive industrial environments. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal of Intelligent Systems is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | The dynamic characteristics of event cameras demonstrate unique advantages in extreme environments; however, their inherent spatiotemporal asynchrony and noise sensitivity instead become serious drawbacks during extreme image reconstruction of coal mine substations. To overcome the problems of nonconstant noise coupling and existing static convolutional networks in dealing with vibration‐induced motion artifacts, an event‐based spatiotemporal disentangled aggregated image reconstruction (ESDAR) method is proposed here. ESDAR is a method to enhance FireNet. The proposed method integrates three innovations. First, it introduces an asymptotic denoising approach inspired by the diffusion model, which effectively decouples high‐frequency noise and low‐frequency textures. Second, it employs a deformable convolutional mechanism to achieve spatiotemporal alignment between event streams and fuzzy grayscale images. Third, the recurrent unit is substituted with a time convolution network enhanced with deformable convolution (DF‐TCN), thereby achieving a significant reduction in the number of parameters, with a decrease of 48.32%. The effectiveness of this method was evaluated via 1000 synthetic sequences and 1670 real frames (captured at a coal mine substation). The findings of this study demonstrate that ESDAR is markedly superior to FireNet with respect to performance. The results revealed a 39.5% reduction in the mean square error (MSE), a 4.2 dB increase in the peak signal‐to‐noise ratio (PSNR), a 26% reduction in the local perceptual image similarity metric (LPIPS), and a 3.39% increase in the structural similarity index (SSIM). This method effectively achieves a balance between noise suppression and detail preservation under hardware constraints, providing a feasible solution to address image reconstruction in high‐noise, vibration‐intensive industrial environments. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 08848173 |
| DOI: | 10.1155/int/4139129 |