Twitter/X

FLUX 3 is a single multi-modal model handling Image, Video, Audio and…

Brief

FLUX 3 is a unified multi-modal model from @bfl_ai that covers image, video, audio and action-prediction. The announcement states FLUX 3 Video is available in early access and that the jointly trained architecture can be extended to robotics action prediction, with cited collaborations or demonstrations involving mimic and Audi.

Why it matters

FLUX 3 is a single multi-modal model handling Image, Video, Audio and Action-Prediction (named by @bfl_ai).

Key details

  • FLUX 3 Video entered early access as of the announcement, with the model jointly trained in one unified architecture.
  • The architecture can be extended to predict actions for robotics; the team cites work with mimic and Audi as examples.
Source evidence

Introducing FLUX 3.

One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.

FLUX 3 Video is now available in early access (link below).

Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.

Video