Black Forest Labs launched FLUX 3 on July 23, and the pitch is bigger than another image model: one network generating image, video, audio, and even robot action prediction. Most of that isn't live yet. Here's what you can actually use today versus what's still a promise.
What FLUX 3 actually is
FLUX 3 is BFL's multimodal flow model, built to handle image, video, audio, and action-prediction from a single backbone instead of separate specialized models. The company is positioning it as infrastructure for "physical AI," not just a Midjourney competitor.
What's live right now, and what's still coming
| Capability | Status |
|---|---|
| FLUX 3 Video | Early access open now, application required |
| FLUX 3 Image | Early access coming in the following weeks |
| FLUX 3 Dev (open-weight backbone) | Confirmed for later release, no date given |
| Action-prediction / robotics | Live only through partners, starting with mimic robotics at Audi |
Video is the only piece in broad early access, and it's real: up to 20 seconds per clip at 720p, with audio (dialogue, sound effects, music) generated in the same pass rather than added afterward. It also supports text-to-video, image-to-video, video-to-video, keyframe control, and chaining multiple shots together.
Pricing hasn't been published for any of it. Access during early access is by request through BFL's site, and there's no public per-generation or subscription cost yet.
How good is it, actually
BFL's own internal comparisons claim human reviewers preferred FLUX 3's video output over Grok Imagine Video 69% of the time, over Kling v3 Pro 60%, over Runway Gen-4.5 77%, and over Luma Ray 3.2 93%. Those numbers sound decisive, but BFL hasn't published its evaluation methodology, sample size, or rater selection, and no independent benchmark site has run its own comparison yet. I'd treat that scorecard as a marketing claim until someone outside the company reproduces it.
The one piece of the launch that's verifiable outside BFL's own numbers is FLUX-mimic, the robotics model built on the FLUX 3 architecture that's reportedly running on production lines at Audi. A model actually deployed on a factory floor is a stronger signal than a self-reported preference study.
Should you apply for early access?
If you do creative video work and want to test a genuinely new generation approach, apply. The audio-in-the-same-generation-pass feature is worth seeing firsthand, especially next to tools like [Midjourney](/blog/midjourney-v8-2-is-live-what-changed-and-is-it-worth-switching) that are still image-first. If you need production-ready output today, wait. No pricing, no independent benchmarks, and an application-gated waitlist isn't a stack you can build a workflow on yet.


