Diffusion-Model-Based Data Augmentation for Target Detection in Side-Scan Sonar Images.
Saved in:
| Title: | Diffusion-Model-Based Data Augmentation for Target Detection in Side-Scan Sonar Images. |
|---|---|
| Authors: | Yang, Yuanxu1,2,3 (AUTHOR), Zhang, Tao1,2,3,4,5 (AUTHOR) zhangtao22@seu.edu.cn |
| Source: | Remote Sensing. Jul2026, Vol. 18 Issue 13, p2193. 26p. |
| Subjects: | Data augmentation, Sidescan sonar, Object recognition (Computer vision), Probabilistic generative models, Synthetic data, Deep learning, Automatic target recognition |
| Abstract: | Highlights: What are the main findings? FLUX+LoRA-generated side-scan sonar images improve YOLOv8n target detection under a fixed-detector and real-only validation/test protocol. The screened FLUX+LoRA subset achieves higher mAP@0.5 and mAP@0.5:0.95 than real-only training and traditional augmentation baselines. What are the implications of the main findings? Diffusion-based generation provides an effective data-augmentation route for sonar target detection when annotated real samples are limited. The released synthetic dataset and code repository support reproducible evaluation of generative augmentation for side-scan sonar imagery. Side-scan sonar images play an important role in underwater target detection, seabed mapping, and marine environment monitoring. However, the performance of deep learning-based detectors is often limited by the small scale of available sonar datasets, the high cost of data acquisition, and class imbalance among target categories. To address these issues, this paper proposes a diffusion-model-based data augmentation method for side-scan sonar target detection. A FLUX.1 diffusion model is adopted as the base generative framework and is fine-tuned using low-rank adaptation (LoRA) to adapt the pretrained model to the side-scan sonar image domain under limited training data conditions. The generated samples are further filtered and added only to the training set, while the validation and test sets are kept unchanged and contain only real sonar images. To ensure a fair evaluation of the augmentation strategy, all detection experiments are conducted using a fixed YOLOv8n (You Only Look Once version 8 nano) detector under the same training hyperparameters and three random seeds. Compared with training on the original dataset, the proposed FLUX+LoRA augmentation improves mean average precision (mAP)@0.5 from 0.7400 ± 0.0132 to 0.8582 ± 0.0328 and mAP@0.5:0.95 from 0.3994 ± 0.0187 to 0.5115 ± 0.0164. It also outperforms conventional augmentation methods under the same real-only validation/test protocol. In addition, Fréchet Inception Distance (FID)/Kernel Inception Distance (KID)-based image quality evaluation, generated-sample amount ablation, screening-strategy ablation, LoRA-rank sensitivity analysis, and a controlled 600-sample diffusion-backbone comparison are conducted. The results show that the 600-sample manually annotated FLUX+LoRA subset selected from generated samples achieves better image quality and detection performance than FLUX-base and SD1.5+LoRA under the same annotation budget. These findings demonstrate that FLUX+LoRA-generated sonar images can provide useful structural diversity for detector training and improve target detection performance under limited-data conditions. [ABSTRACT FROM AUTHOR] |
| Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Highlights: What are the main findings? FLUX+LoRA-generated side-scan sonar images improve YOLOv8n target detection under a fixed-detector and real-only validation/test protocol. The screened FLUX+LoRA subset achieves higher mAP@0.5 and mAP@0.5:0.95 than real-only training and traditional augmentation baselines. What are the implications of the main findings? Diffusion-based generation provides an effective data-augmentation route for sonar target detection when annotated real samples are limited. The released synthetic dataset and code repository support reproducible evaluation of generative augmentation for side-scan sonar imagery. Side-scan sonar images play an important role in underwater target detection, seabed mapping, and marine environment monitoring. However, the performance of deep learning-based detectors is often limited by the small scale of available sonar datasets, the high cost of data acquisition, and class imbalance among target categories. To address these issues, this paper proposes a diffusion-model-based data augmentation method for side-scan sonar target detection. A FLUX.1 diffusion model is adopted as the base generative framework and is fine-tuned using low-rank adaptation (LoRA) to adapt the pretrained model to the side-scan sonar image domain under limited training data conditions. The generated samples are further filtered and added only to the training set, while the validation and test sets are kept unchanged and contain only real sonar images. To ensure a fair evaluation of the augmentation strategy, all detection experiments are conducted using a fixed YOLOv8n (You Only Look Once version 8 nano) detector under the same training hyperparameters and three random seeds. Compared with training on the original dataset, the proposed FLUX+LoRA augmentation improves mean average precision (mAP)@0.5 from 0.7400 ± 0.0132 to 0.8582 ± 0.0328 and mAP@0.5:0.95 from 0.3994 ± 0.0187 to 0.5115 ± 0.0164. It also outperforms conventional augmentation methods under the same real-only validation/test protocol. In addition, Fréchet Inception Distance (FID)/Kernel Inception Distance (KID)-based image quality evaluation, generated-sample amount ablation, screening-strategy ablation, LoRA-rank sensitivity analysis, and a controlled 600-sample diffusion-backbone comparison are conducted. The results show that the 600-sample manually annotated FLUX+LoRA subset selected from generated samples achieves better image quality and detection performance than FLUX-base and SD1.5+LoRA under the same annotation budget. These findings demonstrate that FLUX+LoRA-generated sonar images can provide useful structural diversity for detector training and improve target detection performance under limited-data conditions. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 20724292 |
| DOI: | 10.3390/rs18132193 |