Research on the Construction of Multimodal Large Models and Self-supervised Learning for Panoramic Perception of Power Equipment.

Saved in:
Bibliographic Details
Title: Research on the Construction of Multimodal Large Models and Self-supervised Learning for Panoramic Perception of Power Equipment.
Authors: ZHENG, Guoshun1 zhengguoshun_fj@163.com
Source: Technical Gazette / Tehnički Vjesnik. 2026, Vol. 33 Issue 3, p1175-1184. 10p.
Subjects: Digital twin, Data augmentation, Electric power equipment, Probabilistic generative models, Machine learning, Recommender systems, Visual perception
Abstract: The global information perception of power equipment is the key to supporting the efficient and stable operation of the new power system. This paper adopts the digital twin technology and constructs a new framework for panoramic perception of transformer vibration status. To address the difficulties in obtaining power defect samples, the dominance of normal samples, and the reliance on large-scale data of multi-modal large models, a power defect data enhancement method based on diffusion models is proposed. Under the premise of ensuring the rationality of the generated image structure, this method utilizes the trained multi-modal large model Qwen-VL-Max to extract the high-order semantic information of real power scene images and combines the prompt engineering technology to generate synthetic images with power defect features and high quality. Moreover, for the common problem of data sparsity in multi-modal recommendation systems, a multi-modal fusion recommendation algorithm based on collaborative self-supervised learning is proposed. This algorithm effectively enhances the representation ability of multi-modal data through the joint learning of the deep features of the data, thereby alleviating the problem of performance decline in recommendations caused by data sparsity. Experimental results show that, compared with the current mainstream multi-modal recommendation algorithms, this algorithm has significant improvements in multiple recommendation evaluation indicators. [ABSTRACT FROM AUTHOR]
Copyright of Technical Gazette / Tehnički Vjesnik is the property of Tehnicki Vjesnik and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:The global information perception of power equipment is the key to supporting the efficient and stable operation of the new power system. This paper adopts the digital twin technology and constructs a new framework for panoramic perception of transformer vibration status. To address the difficulties in obtaining power defect samples, the dominance of normal samples, and the reliance on large-scale data of multi-modal large models, a power defect data enhancement method based on diffusion models is proposed. Under the premise of ensuring the rationality of the generated image structure, this method utilizes the trained multi-modal large model Qwen-VL-Max to extract the high-order semantic information of real power scene images and combines the prompt engineering technology to generate synthetic images with power defect features and high quality. Moreover, for the common problem of data sparsity in multi-modal recommendation systems, a multi-modal fusion recommendation algorithm based on collaborative self-supervised learning is proposed. This algorithm effectively enhances the representation ability of multi-modal data through the joint learning of the deep features of the data, thereby alleviating the problem of performance decline in recommendations caused by data sparsity. Experimental results show that, compared with the current mainstream multi-modal recommendation algorithms, this algorithm has significant improvements in multiple recommendation evaluation indicators. [ABSTRACT FROM AUTHOR]
ISSN:13303651
DOI:10.17559/TV-20260203003363