Action-Aligned Video Pairing for Video Augmentation

Closed

Randy Cahya Wihandika, Israel Mendonca, Masayoshi Aritsugi

2025 IEEE Region 10 Annual International Conference, Proceedings/TENCON Conference paper Cited by 0 Quartile

Abstract

Video augmentation is an effective strategy for improving the performance of action recognition models. A recent video augmentation strategy addresses scene bias by mixing human regions from one video with the background from another. However, this often produces artifacts due to limitations in the video mixing process, which degrade training quality. This study proposes a video augmentation strategy that produces compatible action-scene video pairs rather than choosing them randomly, to improve the quality of mixed videos. To achieve this, two compatibility metrics are introduced to guide this selection to significantly reduce the occurrence of visual artifacts and generate higher-quality augmented videos. Our method improves alignment between actions which leads to more effective augmentation. The performances are further enhanced by applying a temporal morphological operation to improve object detection consistency. Experimental results on the UCF101, HMDB51, and Kinetics-100 datasets show that our approach improves classification performance. Code is available at https://github.com/rendicahya/video-action-alignment. © 2025 IEEE.

Affiliations

Graduate School of Science and Technology, Kumamoto University, Kumamoto, Japan; Brawijaya University, Faculty of Computer Science, Malang, Indonesia; Kumamoto University, Faculty of Advanced Science and Technology, Kumamoto, Japan