Black Forest Labs released FLUX 3 on Thursday, July twenty third, a single model that generates images, video with synchronised audio, and robot actions from the same set of weights. That last item is the surprising one. FLUX started as an image generator, and the company has now folded video, sound, and physical action prediction into one architecture instead of shipping four specialised models. Video clips run up to twenty seconds with native dialogue, effects, and ambience. Only part of the family is reachable today, through early access, and the open weight release is scheduled for later this year. Here is what shipped, what did not, and which claims still need independent testing.
The short answer
On Thursday, July twenty third, Black Forest Labs released FLUX 3, its first natively multimodal model. One set of weights covers image generation and editing, video up to twenty seconds with synchronised dialogue, effects, and ambience, and action prediction for robots. FLUX 3 Video and FLUX 3 Action are in early access now, FLUX 3 Image follows in the coming weeks, and an open weight FLUX 3 Dev is planned for later in 2026. Audi is testing the robotics variant, built with mimic, for production manipulation tasks. The comparison numbers against other video models come from the company's own internal evaluations.
The interesting thing about FLUX 3 is not that it makes video. Plenty of models make video. It is that the company behind one of the most widely used open weight image generators decided its next release should also drive a robot arm, and shipped both behaviours out of the same weights. Black Forest Labs announced FLUX 3 on Thursday, July twenty third, and the architecture is the story.
One model instead of four
Earlier FLUX releases were image models, and the industry pattern has been to bolt modalities on as separate systems: one model for pictures, another for clips, a text to speech stack for the audio, something else entirely for robotics. FLUX 3 trains on images, video, audio, and action together in a single architecture, built on the company's Self Flow method for keeping generation and understanding aligned in the same system.
The practical argument for doing it that way is that these things are not actually separate problems. A model that has learned how objects move, how a sound relates to the event that caused it, and how a surface behaves when something touches it has learned physics it can reuse across all four outputs. Whether that generalisation holds up is exactly what independent testing will decide, but it explains why a company known for still images is suddenly talking about manipulation tasks.
What FLUX 3 Video ships with
Clips run up to twenty seconds, and the audio is generated alongside the picture rather than added afterwards, covering dialogue, sound effects, and ambient noise. Input can be text, an image, or an existing video, so text to video, image to video, and video to video all run through the same endpoint. Dialogue generation is multilingual.
The feature worth flagging for anyone doing real production work is keyframe controlled transitions. Being able to pin the first and last frame of a shot is the difference between a model you can cut into a timeline and a model you prompt repeatedly until something usable falls out. It is a small line in a launch announcement and a large line in a production budget.
The robotics part, and the claim inside it
FLUX 3 Action predicts what a machine should do rather than what a frame should look like. Black Forest Labs built a variant called FLUX-mimic with the robotics company mimic, and Audi is testing it on production manipulation tasks.
The number attached to it is the one to watch. The company says that depending on task difficulty, the model can be fine tuned for a specific manipulation task with as little as thirty minutes of robot data, against thirty hours or more for prior approaches. If that holds outside a partner deployment it changes the economics of robot teaching, because collecting demonstration data is usually the expensive part. For now it is a vendor claim from a small number of sites, published on launch day, and it deserves the same scepticism as any launch day number.
What you can actually use today
Less than the headline suggests, which is worth saying plainly. FLUX 3 Video and FLUX 3 Action are in early access. FLUX 3 Image was described as rolling out over the following weeks. Companies including Canva, Krea, Picsart, Burda, and Magnific are named as testing partners, which tells you the access list is short.
The date most readers here care about is the one that was not given. FLUX 3 Dev, the open weight variant, is planned for later in 2026 with no announced day. Black Forest Labs earned its position among developers partly by putting usable weights in public hands, and until that release lands FLUX 3 is a hosted preview. If your plan involves running this on your own GPUs, the honest status is that you are waiting.
About those comparison numbers
Black Forest Labs reports that FLUX 3 Video was preferred over Luma Ray 3.2 in 93 percent of head to head comparisons and over Runway Gen-4.5 in 77 percent. Full results were promised when the model reaches wider availability.
These are the company's own internal evaluations, released with its own launch, so read them as positioning rather than measurement. Preference rates on generative video swing hard on prompt selection and on who is doing the judging, and the field has no settled public benchmark to appeal to. The number that will matter to you is the one you get from your own prompts, on your own footage, once access opens up.
Sources and further reading
- Black Forest Labs press release: FLUX 3, a new multimodal frontier model for visual intelligence
- VentureBeat: Black Forest Labs launches FLUX 3, capable of images and 20 second video with audio, in limited release to start
- Crypto Briefing: Black Forest Labs launches FLUX 3 for video, audio, and robotics
- IBTimes: Which FLUX 3 features actually work today
Frequently asked questions
What is FLUX 3 and how is it different from FLUX.2?
FLUX 3 is Black Forest Labs' first natively multimodal architecture. Earlier FLUX families were image models. FLUX 3 trains on images, video, audio, and action prediction together, in one system, so a single set of weights can render a picture, produce a video clip with its own soundtrack, edit an existing image, and predict the movements a robot arm should make. The company describes the goal as learning spatial structure, movement, sound, and physical interaction jointly rather than treating each as a separate task, building on its Self Flow method for aligning generation and understanding in the same model.
What can FLUX 3 Video actually do?
FLUX 3 Video generates clips up to twenty seconds long with audio created alongside the visuals rather than dubbed on afterwards, covering dialogue, sound effects, and ambient noise. It accepts text, an image, or an existing video as input, so text to video, image to video, and video to video all work through the same model. It also supports keyframe controlled transitions, which matters if you need a shot to start and end on specific frames rather than hoping the model lands somewhere usable, and the company says dialogue generation is multilingual.
When do open weights arrive?
Black Forest Labs said an open weight variant, FLUX 3 Dev, is planned for later in 2026, along with faster versions of the model. As of the July twenty third announcement, nothing in the FLUX 3 family has open weights. FLUX 3 Video and FLUX 3 Action are in early access, and FLUX 3 Image was described as rolling out in the following weeks. For anyone planning to self host, that means the useful date is still unannounced, and the current release should be treated as a hosted preview rather than something you can run on your own hardware.
What is FLUX 3 Action and why is Audi involved?
FLUX 3 Action applies the same model to robotics, predicting the actions a machine should take rather than the pixels it should draw. Black Forest Labs co-developed a variant called FLUX-mimic with the robotics company mimic, and Audi is testing it for production manipulation tasks. The claim worth noting is the data efficiency. The company says that depending on task difficulty, the model can be fine tuned for a specific manipulation task with as little as thirty minutes of robot data, where prior approaches needed thirty hours or more. That is a vendor claim from a limited deployment, not an independently reproduced benchmark.
How much should I trust the benchmark numbers?
Treat them as vendor reported until someone else runs them. Black Forest Labs says internal testing found FLUX 3 Video preferred over Luma Ray 3.2 in 93 percent of head to head comparisons and over Runway Gen-4.5 in 77 percent. Those are the company's own evaluations, published alongside its own launch, and full results were promised when the model reaches broader availability. Preference rates on generative video are also sensitive to prompt selection and to who is judging. The numbers are a reasonable signal of where the company thinks it stands, and a poor substitute for testing your own prompts once access opens.