WorldSonus

WorldSonus:
Bringing Sound to Worlds

1 The Hong Kong University of Science and Technology
2 Noiz AI3 MetaX
4 Shanghai Jiao Tong University

† Corresponding authors

Abstract

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment.

Method Overview

WorldSonus architecture: causal streaming audiovisual generation, training-only temporal alignment, spatial stereo supervision, and time-varying prompt control.
Overview of WorldSonus. (a) Causal streaming pipeline: Video frames map to two-timescale visual tokens conditioning an AR transformer and a rectified-flow head. (b) Training-only ShiftNCE: A training-only frozen Synchformer teacher guides temporal alignment via contrastive window matching. (c) Spatial stereo supervision: In-the-wild stereo filtering and panoramic FOA view-decoding provide directional audio grounding. (d) Interactive prompt control: Text prompts update dynamically at chunk boundaries, preserving session state and acoustic continuity.

Results