ArXiv

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

Authors
Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng...
Categories
cs.CV
arXiv
https://arxiv.org/abs/2608.06231v1
PDF
https://arxiv.org/pdf/2608.06231v1

Brief

EmoWorld targets controllable emotional video generation by separating atmosphere, semantic affect cues, and temporal dynamics inside a frozen flow‑matching Video DiT. A one-time stage extracts layer-wise affect directions and a cue library from neutral and emotion-edited panoramas; at inference VAS, SAS, and TAS inject/scale those fields. Results on Wan2.2 show substantial gains across alignment, cue presence, and temporal consistency. Full paper text was not provided, summary based on the abstract.

Why it matters

EmoWorld decouples global atmosphere, affect-bearing semantic cues, and temporal progression in a frozen flow‑matching Video DiT via a one-time preparation that extracts layer-specific affect directions and a reusable cue library; it provides three controls: Visual Atmosphere Steering (VAS), Semantic Affective Steering (SAS), and Temporal Affective Steering (TAS).

Key details

  • On the Wan2.2 benchmark, VAS improves target-emotion alignment by 19% and reduces a temporal-fluctuation proxy by 48%; SAS boosts target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; TAS improves transition monotonicity by 15% versus the strongest baseline.
  • EmoWorld was evaluated across 27 emotion categories in text-to-video and image-to-video settings, is portable across multiple Video‑DiT backbones, and supports camera-conditioned composition without updating generator parameters.
Source evidence

Abstract

Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion transformer (Video DiT). A one-time preparation stage extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral and emotion-edited panoramas. At inference, Visual Atmosphere Steering (VAS) injects atmosphere directions into hidden states, Semantic Affective Steering (SAS) isolates a separately scalable prompt residual for semantic cues, and Temporal Affective Steering (TAS) interpolates endpoint residual fields across denoising and video time. On Wan2.2, VAS improves target-emotion alignment by 19% while reducing a temporal-fluctuation proxy by 48%; SAS improves target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; and TAS improves transition monotonicity by 15% over the strongest baseline. EmoWorld is evaluated across 27 emotion categories in text-to-video and image-to-video settings, demonstrates portability across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.