One architecture for image, video, audio, and action: FLUX 3 generates 20-second clips — Black Forest Labs
Black Forest Labs released FLUX 3 on July 23, its first unified multimodal model, jointly trained on image, video, audio, and action. A single prompt produces clips up to 20 seconds; video and action remain in gated early access.
Positive-sum angle: Joint training means each modality strengthens the others, so the gains accrue to everyone building on top. Adobe, Picsart, and agent harnesses already draw on it, and every studio or robotics team downstream inherits a shared visual substrate instead of stitching three separate models together.
What's the impact: SEA content studios and short-video shops should reprice production assuming one model covers stills, clips, and voice. Jakarta and Manila teams that rebuild pipelines around a single architecture this year will underbid the three-vendor stack still standard across the region.