ACM MM Supplementary

Description of GIF

Video scene editing with audio-guided semantic change.

Fireworks

Rain

Waterfall

Abstract

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image editing, audio-driven video scene editing has not been extensively addressed. In this paper, we introduce AudioScenic, an audio-driven framework designed for video scene editing. AudioScenic integrates audio semantics into the visual scene through a temporal-aware audio semantic injection process. As our focus is on background editing, we further introduce a SceneMasker module, which maintains the integrity of the foreground content during the editing process. AudioScenic exploits the inherent properties of audio, namely, audio magnitude and frequency, to guide the editing process, aiming to control the temporal dynamics and enhance the temporal consistency. First, we present an audio Magnitude Modulator module that adjusts the temporal dynamics of the scene in response to changes in audio magnitude, enhancing the visual dynamics. Second, the audio Frequency Fuser module is designed to ensure temporal consistency by aligning the frequency of the audio with the dynamics of the video scenes, thus improving the overall temporal coherence of the edited videos. These integrated features enable AudioScenic to not only enhance visual diversity but also maintain temporal consistency throughout the video. We present a new metric named temporal score for more comprehensive validation of temporal consistency. We demonstrate substantial advancements of AudioScenic over competing methods on DAVIS and Audioset datasets.

Audio semantic diversity

Cracking fire clip 1

Cracking fire clip 2

Cracking fire clip 3

Sea wave clip 1

Sea wave clip 2

Sea wave clip 3

Rain clip 1

Rain clip 2

Rain clip 3

Comparison with text-driven method

Waterfall

Fireworks

Comparison with text-driven methods

Description of GIF

Comparing set (24frames, target semantic is "Sea wave").

Sea wave

Comparison with audio-driven method

Description of GIF

Comparing set, the target semantic is "Waterfall".

Waterfall

Robust to random variations

Description of GIF

Robust to random seeds.

Magnitude Control

Description of GIF

Magnitude-controlled temporal scene dynamics.

Emotional Scene Editing

Happy music clip1

Sad music clip1

Happy music clip1

Sad music clip2

Framework

Description

Framework of the proposed AudioScenic.