Bibliographic Details
| Title: |
TAIS-Net: Time adaptive implicit sampling diffusion model for arbitrary-scale UAV video super-resolution. |
| Authors: |
Li, Wenke1 (AUTHOR) ieulwk@163.com, Dai, Chenguang1 (AUTHOR), Fan, Huixin1 (AUTHOR), Zhao, Zhen1 (AUTHOR), Li, Kangrui1 (AUTHOR), Sun, Yifan1 (AUTHOR), Zhang, Yongsheng1 (AUTHOR), Wang, Longguang2 (AUTHOR), Wang, Hanyun1,3 (AUTHOR) wanghanyun@mail.sysu.edu.cn |
| Source: |
ISPRS Journal of Photogrammetry & Remote Sensing. Aug2026, Vol. 238, p176-190. 15p. |
| Subjects: |
Optical flow, Implicit functions, Probabilistic generative models, High resolution imaging |
| Abstract: |
Emerging applications urgently require unmanned aerial vehicle (UAV) videos with high clarity and rich structural details. However, limitations in imaging sensors and non-ideal conditions often result in blurred videos lacking sufficient details. Recent studies indicate diffusion models have strong potential for natural scene video super-resolution (SR). However, the application of diffusion models to UAV video SR remains underexplored, particularly in the context of arbitrary-scale reconstruction. To tackle the challenges in UAV video SR at arbitrary scales, this work introduces a time adaptive implicit sampling diffusion model TAIS-Net. First, to achieve robust temporal alignment under large camera motions and texture scarcity commonly encountered in UAV videos, we use a reference frame alignment module to compensate previously reconstructed frames by incorporating motion cues estimated from a pre-trained optical flow model. Second, to enable scale-flexible reconstruction while preserving fine geometric details for UAV videos under arbitrary upscaling factors, we introduce an implicit denoising U-Net to learn latent features by leveraging implicit neural representations and a temporal conditioning module. Third, to reduce the prediction error propagation during sampling, we introduce a historical sample gain module that dynamically corrects and refines latent features at each sampling step, thereby suppressing temporal artifacts such as flickering or drifting. Finally, an arbitrary-scale implicit decoder is used to reconstruct high-resolution videos directly from these enhanced latent features, avoiding scale-dependent blurring while preserving high-frequency details in UAV videos. To validate the effectiveness of TAIS-Net, we construct two visible and one infrared UAV video SR datasets. The experimental results on these datasets demonstrate that TAIS-Net attains competitive performance relative to existing methods in terms of reconstruction quality, perceptual quality and temporal consistency. Moreover, experiments on various motion intensities and additional Gaussian blur beyond conventional bicubic downsampling degradation without requiring model retraining also demonstrate the superiority of our TAIS-Net in maintaining perceptual and temporal consistency under real-world UAV scenarios. The code and datasets are available at https://github.com/cyber-lwk/TAIS-Net. [ABSTRACT FROM AUTHOR] |
|
Copyright of ISPRS Journal of Photogrammetry & Remote Sensing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) |
| Database: |
Engineering Source |