TRRHA: A two-stream re-parameterized refocusing hybrid attention network for synthesized view quality enhancement.

Saved in:
Bibliographic Details
Title: TRRHA: A two-stream re-parameterized refocusing hybrid attention network for synthesized view quality enhancement.
Authors: Cao, Ziyi1 (AUTHOR), Li, Tiansong1 (AUTHOR) tiansongli@cqnu.edu.cn, Wang, Guofen1 (AUTHOR), Yin, Haibing2 (AUTHOR), Wang, Hongkui2 (AUTHOR), Yu, Li3 (AUTHOR)
Source: Displays. Dec2024, Vol. 85, pN.PAG-N.PAG. 1p.
Subjects: Source code, Pyramids, Video coding, Videos
Abstract: In multi-view video systems, the decoded texture video and its corresponding depth video are utilized to synthesize virtual views from different perspectives using the depth-image-based rendering (DIBR) technology in 3D-high efficiency video coding (3D-HEVC). However, the distortion of the compressed multi-view video and the disocclusion problem in DIBR can easily cause obvious holes and cracks in the synthesized views, degrading the visual quality of the synthesized views. To address this problem, a novel two-stream re-parameterized refocusing hybrid attention (TRRHA) network is proposed to significantly improve the quality of synthesized views. Firstly, a global multi-scale residual information stream is applied to extract the global context information by using refocusing attention module (RAM), and the RAM can detect the contextual feature and adaptively learn channel and spatial attention feature to selectively focus on different areas. Secondly, a local feature pyramid attention information stream is used to fully capture complex local texture details by using re-parameterized refocusing attention module (RRAM). The RRAM can effectively capture multi-scale texture details with different receptive fields, and adaptively adjust channel and spatial weights to adapt to information transformation at different sizes and levels. Finally, an efficient feature fusion module is proposed to effectively fuse the extracted global and local information streams. Extensive experimental results show that the proposed TRRHA achieves significantly better performance than the state-of-the-art methods. The source code will be available at https://github.com/647-bei/TRRHA. • A two-stream re-parameterization refocusing hybrid attention network (TRRHA) for SVQE. • Design includes global multi-scale residual (GMR) and local feature pyramid attention (LFPA). • Proposed re-parameterized refocusing attention module (RRAM) for local multi-scale texture. • Captures multi-scale features with re-parameterized convolution (RC) branches. • Efficient feature fusion module (EFFM) significantly enhances SVQE performance. [ABSTRACT FROM AUTHOR]
Copyright of Displays is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:In multi-view video systems, the decoded texture video and its corresponding depth video are utilized to synthesize virtual views from different perspectives using the depth-image-based rendering (DIBR) technology in 3D-high efficiency video coding (3D-HEVC). However, the distortion of the compressed multi-view video and the disocclusion problem in DIBR can easily cause obvious holes and cracks in the synthesized views, degrading the visual quality of the synthesized views. To address this problem, a novel two-stream re-parameterized refocusing hybrid attention (TRRHA) network is proposed to significantly improve the quality of synthesized views. Firstly, a global multi-scale residual information stream is applied to extract the global context information by using refocusing attention module (RAM), and the RAM can detect the contextual feature and adaptively learn channel and spatial attention feature to selectively focus on different areas. Secondly, a local feature pyramid attention information stream is used to fully capture complex local texture details by using re-parameterized refocusing attention module (RRAM). The RRAM can effectively capture multi-scale texture details with different receptive fields, and adaptively adjust channel and spatial weights to adapt to information transformation at different sizes and levels. Finally, an efficient feature fusion module is proposed to effectively fuse the extracted global and local information streams. Extensive experimental results show that the proposed TRRHA achieves significantly better performance than the state-of-the-art methods. The source code will be available at https://github.com/647-bei/TRRHA. • A two-stream re-parameterization refocusing hybrid attention network (TRRHA) for SVQE. • Design includes global multi-scale residual (GMR) and local feature pyramid attention (LFPA). • Proposed re-parameterized refocusing attention module (RRAM) for local multi-scale texture. • Captures multi-scale features with re-parameterized convolution (RC) branches. • Efficient feature fusion module (EFFM) significantly enhances SVQE performance. [ABSTRACT FROM AUTHOR]
ISSN:01419382
DOI:10.1016/j.displa.2024.102843