Infusing Multimodal Latent Embedding for Image Deblurring

Closed

Latansa Nury, Novanto Yudistira

2025 2025 IEEE International Conference on Automatic Control and Intelligent Systems, I2CACIS 2025 - Proceedings Conference paper Cited by 1 Quartile

Abstract

Blurred images significantly degrade the performance of vision-based systems. While some blur effects are intentional, such as bokeh, most applications prioritize structural restoration. We propose a novel multimodal framework for image deblurring that infuses high-level semantic features extracted from pretrained multimodal models of CLIP and Stable Diffusion into a two-stage restoration pipeline. Built upon the Adaptive Filter-Based Deblurring Module (AFDM), our approach combines low-level pixel refinement with semantic-aware feature fusion using a UNet2D-based diffusion model. Experiments on the GoPro dataset demonstrate significant improvements over the baseline Iterative Filter Adaptive Network (IFAN), with Peak Signal-to-Noise Ratio (PSNR) increasing from 8.28 to 34.51 and SSIM from 0.53 to 0.9702, confirming the efficacy of our method for context-aware image deblurring. © 2025 IEEE.

Affiliations

Universitas Brawijaya, Informatics Engineering Faculty of Computer Science, Malang, Indonesia; Universitas Brawijaya, Intelligent Systems Laboratory Faculty of Computer Science, Malang, Indonesia