FLUX 3: what Black Forest Labs actually shipped

Rama Adi Nugraha
Written by

Rama Adi Nugraha

Katelin Teen
Reviewed by

Katelin Teen

Last edited August 4, 2026

Expert Verified
Illustration of one model producing image, video and audio panels, representing FLUX 3 from Black Forest Labs

What FLUX 3 actually is

Diagram of FLUX 3 as one shared backbone: images, video and audio feed in on the left, and image, video plus audio, robot action and open-weight outputs branch out on the right
Diagram of FLUX 3 as one shared backbone: images, video and audio feed in on the left, and image, video plus audio, robot action and open-weight outputs branch out on the right

FLUX 1 and FLUX 2 made images. FLUX 3 is a different kind of thing, and the naming undersells it.

It is trained jointly across images, video and audio inside one architecture, built on an approach Black Forest Labs calls Self-Flow. The company's framing is that each modality is a lossy projection of one reality, so learning all of them together forces constraints a single-modality model never sees. The line I keep coming back to from the announcement post: "the sound has to match the impact, the motion has to obey the mass."

That is a claim about representation, not about pixels. It is also the same claim under Gemini Robotics 2, just approached from the other end. DeepMind built a robot model and gave it world knowledge. Black Forest Labs built a video model and found the world knowledge was already sitting in there.

Concretely, from a single generation, FLUX 3 does:

  • Text-to-video and image-to-video, with audio, up to 20 seconds.
  • Video-to-video, carrying a character from a reference clip into a new scene.
  • Keyframe-to-video for controlled transitions.
  • Multilingual dialogue with lip-synced speech.
  • Agentic chaining of clips into multi-shot sequences.
  • Image synthesis and editing, with multilingual text rendering.
The FLUX 3 capability overview showing one model spanning image, video, audio and action, as published by Black Forest Labs
The FLUX 3 capability overview showing one model spanning image, video, audio and action, as published by Black Forest Labs

For anyone who has been juggling Midjourney for stills, a separate tool for clips and a third for voiceover, the consolidation is the pitch. One model, one bill, one prompt language. The diffusion lineage it came from was never going to stop at stills.

That squeeze is already showing up elsewhere. Every AI video script workflow that currently ends at a text file is a candidate for a model that can render the clip and the voiceover from the same prompt.

Video is the hard part, and the compute split proves it

Bar showing the FLUX 3 training compute split: video prediction takes over 95%, images a thin slice, audio under 0.5% of tokens
Bar showing the FLUX 3 training compute split: video prediction takes over 95%, images a thin slice, audio under 0.5% of tokens

The most useful engineering detail in the whole launch is a cost breakdown, and it is buried in the robotics companion post.

Video prediction accounts for over 95% of FLUX 3's total training compute. Audio is under 0.5% of the tokens in a 720p video with sound. Black Forest Labs calls audio "the easy modality", low-dimensional and tightly coupled to what the video already shows, so once the model has learned video it picks up lip sync and impact sounds almost as a side effect.

Actions turn out to have the same shape: a low-dimensional signal, tightly coupled to visual observation. That is the load-bearing argument of the entire release. If you accept it, then a video model is a world model that happens to output pixels, and robot control is just another decoder head.

They tested it the honest way, too. Adding action prediction to a live training run knocked human ratings on text-to-video and image-to-video down by up to 10%. Then, after 3,500 steps, video quality was back where it started, with action prediction now on top. A temporary dip, not a permanent tax. They made a falsifiable prediction and published the result for it, which is more than most launches bother with.

The numbers Black Forest Labs published

Chart of how often human raters preferred FLUX 3 over rival video models, from 52% against Gemini Omni Flash to 93% against Luma Ray 3.2, as taken from Black Forest Labs
Chart of how often human raters preferred FLUX 3 over rival video models, from 52% against Gemini Omni Flash to 93% against Luma Ray 3.2, as taken from Black Forest Labs

These are preference rates from 10-second, 720p text-to-video clips with audio. Read them with the vendor's own caveat attached: the model and the harness were still in development, and 50% means the vote split evenly.

Compared againstShare preferring FLUX 3Read
Luma Ray 3.293%Not close
Runway Gen-4.577%Clear
Grok Imagine Video69%Clear
Kling v3 Pro60%Real but modest
Happy Horse v159%Real but modest
Happy Horse 1.157%Marginal
Seedance 2.052%Coin flip
Gemini Omni Flash52%Coin flip

The spread is what makes this table worth reading rather than the headline. Against the two strongest current video models, FLUX 3 is a coin flip. Against Luma Ray 3.2 it is a rout. A press release could truthfully say "preferred in up to 93% of comparisons" and tell you almost nothing, which is the same trick a single blended deflection number plays in my own industry.

Black Forest Labs deserves credit for publishing the losing end of its own range. On the image side there are no comparative numbers yet, only samples, and the company says only that complex prompt handling and text rendering improved "significantly" over earlier FLUX versions.

Grid of FLUX 3 image samples across illustration, photographic and typographic styles, as published by Black Forest Labs
Grid of FLUX 3 image samples across illustration, photographic and typographic styles, as published by Black Forest Labs

If you want a like-for-like read on stills today, GPT Image 2 against Midjourney is the comparison that actually has both models available.

For editing rather than generation, Qwen Image Edit is the closest current equivalent. At the cheap-and-fast end, Nano Banana 2 Lite is what most teams reach for instead.

What you can actually get today

The FLUX 3 model page showing a Coming Soon label and a Request early access button, captured 4 August 2026 from Black Forest Labs
The FLUX 3 model page showing a Coming Soon label and a Request early access button, captured 4 August 2026 from Black Forest Labs

Here is the ladder, in the order Black Forest Labs published it.

Four-rung access ladder: FLUX 3 Video in early access, FLUX 3 Action partners only, FLUX 3 Image coming in weeks, FLUX 3 Dev open weights with no date, all sitting left of the line marking public self-serve access
Four-rung access ladder: FLUX 3 Video in early access, FLUX 3 Action partners only, FLUX 3 Image coming in weeks, FLUX 3 Dev open weights with no date, all sitting left of the line marking public self-serve access
CapabilityHow you get itStatus on 4 Aug 2026
FLUX 3 VideoAPI and private weight accessEarly access, request form
FLUX-mimic and FLUX 3 ActionSelected research and commercial partnersClosed, mimic robotics first
FLUX 3 ImageAPI and private weight access"In the following weeks"
FLUX 3 DevOpen-weight multimodal backboneAnnounced, no date

Now the pricing page, which is where the story lands.

The Black Forest Labs pricing page showing tabs for FLUX.2, FLUX Tools and FLUX.1 only, with no FLUX 3 row, as taken from Black Forest Labs
The Black Forest Labs pricing page showing tabs for FLUX.2, FLUX Tools and FLUX.1 only, with no FLUX 3 row, as taken from Black Forest Labs

FLUX 3 does not appear on it. Three tabs, none of them FLUX 3. The docs go further and say plainly that FLUX.2 "is our recommended model family for all use cases". So if you need to ship something this month, the real options are the ones with a rate card:

Model1st MPAdd'l MPRef imageFit
FLUX.2 [max]$0.07$0.03$0.03/MPHighest fidelity, up to 4 MP
FLUX.2 [pro]$0.03$0.015$0.015/MPQuality and speed balance
FLUX.2 [flex]$0.05$0.05$0.05/MPPure per-megapixel billing
FLUX.2 [klein] 9B$0.015$0.002$0.002/MPHigh throughput, low latency
FLUX.2 [klein] 4B$0.014$0.001$0.001/MPCheapest per image

Resolution rounds up to the next megapixel, separately for the output and for each reference image, and one megapixel is 1024×1024. Anything over 4 MP gets resized down. Which means the sticker price and your actual bill diverge fast once you start passing references in, so here is the arithmetic:

Two references on a 1080p [pro] job put a third of your bill in the inputs. That is the kind of detail that only shows up in month two, and it is the same lesson as every LLM rate card I have costed: the headline unit price is rarely the unit you actually buy.

The robot part is not a side quest

Chart of autonomous success rates on a soft-body kitting task, with FLUX-mimic at a 95% median against 55% for the pi-0.5 baseline, as taken from Black Forest Labs
Chart of autonomous success rates on a soft-body kitting task, with FLUX-mimic at a 95% median against 55% for the pi-0.5 baseline, as taken from Black Forest Labs

An image-generation company shipping a robotics model reads like a pivot. Black Forest Labs argues the opposite. If content creation and physical control run on the same world model, robotics is just a second decoder bolted onto work they had already paid for.

The evidence is a lightweight action decoder trained on intermediate features from the FLUX 3 video path, built with mimic robotics as FLUX-mimic. On a soft-body kitting task, over 20 autonomous trials each:

ModelMedian success
FLUX-mimic95%
Single-task flow matching70%
FLUX-mimic, frozen backbone65%
π0.555%
π0.5, frozen backbone0%

The frozen-backbone row is the one that matters. Freeze FLUX 3 entirely, train only the decoder, and you still get 65%. Do the same to the π0.5 baseline and you get zero. The world knowledge is not just present, it is legible enough to read actions straight out of the features. Two other numbers back that up: the backbone runs input-to-representation in under 80ms on a single RTX 5090, and mimic's full robot loop reacts in 101ms.

The deployment is real, not a lab demo. Audi has been running it on production and logistics tasks that conventional automation never touched, like seals and cables.

"We have seen these robots solve complex soft-body manipulation work that would have been simply impossible with conventional robotics."

Christoph Schneider, Audi Production Lab, quoted in Black Forest Labs' announcement

The sample-efficiency claim is the strategically interesting one. Black Forest Labs reports that better representations halved the training steps needed to hit a given success rate, and cites the mimic-video paper's up-to-10x sample efficiency for video-action models. If teaching a robot a new task stops being a data-collection project, the economics of factory automation change. That is a bigger deal than any 20-second clip, and it is the same thesis DeepMind is testing from the robotics side.

What Hacker News made of it

The launch thread hit 571 points and 133 comments, and the reaction split cleanly. The technical audience wants one thing, and it is not video length.

Hacker News

"Flux 2 Dev Klein has practically been the best you could use on most commercial hardware so I really hope Flux 3 has a comparable updated open-weights model to it. if not it'd be a great loss to most hobbyists."

Open weights came up in comment after comment, usually with the same worry attached: the open-weight line sits at the bottom of the launch plan with no date. Several people flagged VRAM as the deciding factor for whether FLUX 3 Dev is usable at home at all.

The skeptics were louder than the launch deserved, and one critique did land:

Hacker News

"Showed close to zero examples of people. Frivolous use of the term World Model. Claims 20 seconds of video, shows only jumpcuts."

That is a fair reading of a page whose strongest claims are about people and continuity. The counter-take got votes too:

Hacker News

"It's incredible how negative and dismissive the comments in here are while here I am thinking the model actually looks impressively capable."

Both things can be true at once. The result looks strong, the write-up oversells it, and the thing itself is behind a form. Personally I will take the compute breakdown and the frozen-backbone chart over another highlight reel any day, and Black Forest Labs published both of those.

What this means if you buy AI for a living

The eesel AI dashboard showing connected helpdesk activity and resolution reporting
The eesel AI dashboard showing connected helpdesk activity and resolution reporting

I do not generate video for a living. I build AI agents that answer support tickets, and FLUX 3 still changed how I will read the next three launches in my own category.

Read the access ladder before the demo reel. A launch has a capability list and an availability list, and they are rarely the same list. The FLUX 3 page is unusually honest about this, since it publishes the ladder right there. Most AI support vendors do not, which is why the useful question is always "which of these is on my plan, today?" rather than "can it do this?" It is the same gap I flag when comparing an AI agent to a chatbot: the demo shows the ceiling, the plan shows the floor.

Range beats headline. "Preferred in up to 93% of comparisons" and "a coin flip against the two best rivals" describe the same table. When a vendor gives you one number for AI quality, ask for the distribution behind it. That is the entire argument for per-intent metrics over one blended tier-1 deflection figure, and I make it about my own product too. It is also why confidence thresholds beat a single accuracy score.

One backbone, many jobs, is the actual trend. FLUX 3 does images and video and audio and actions from one model. On the language side, the same consolidation is why GPT-5.6 and DeepSeek V4 Flash now handle work that used to need three specialised models. Claude Opus 4.6 sits on the same curve.

Buying a tool with one narrow model inside it is a shorter-lived decision than it used to be. Whichever one you land on, check it against hallucination controls before you let it near a customer.

Open weights are a real fork in the road. FLUX 3 Dev will eventually let teams self-host a multimodal backbone. That is appealing right up until you own the fine-tuning, the GPUs and the on-call rota. A customer of ours who had every skill needed to build in-house put it plainly:

"We could try to write our own LLM application but we didn't want to invest our time into that. We wanted something that we would not have to maintain."

Karel, GENERAL BYTES

That is the trade, stated better than I could. The training and custom model work is the part nobody demos.

Want AI that actually answers your tickets, not a waitlist? eesel connects to your helpdesk in a few minutes, learns from your past tickets and help centre, and runs a simulation over your real history so you see the resolution rate before a customer does. No form, no early-access queue. Try eesel free.

eesel AI drafting and resolving tickets inside Zendesk

Where FLUX 3 sits right now

FLUX 3 is the most interesting model launch of the summer and the least buyable one. The technical case is strong: one backbone across four modalities, a published compute breakdown, an honest preference range including the losses, and a frozen-backbone robotics result that would be hard to fake. The Audi deployment gives it a floor of reality that most physical-AI announcements do not have.

What it does not have is a price, an endpoint or a date. FLUX.2 is what Black Forest Labs itself recommends for every current use case, and at $0.014 to $0.07 for the first megapixel it is the version with a rate card. If FLUX 3 is on your roadmap, request early access, and plan around FLUX.2 in the meantime.

If you are choosing generative tooling more broadly, my generative AI tools roundup covers what is buyable today. For the writing half of the same job, start with AI model for blogging.

On the content side, AI blog automation is where these image models actually get used day to day. An AI article writer is the piece that calls them, and generative AI SEO tools is where the output ends up.

For the support side, generative AI for support teams is the better starting point. If you are further back than that, generative AI for customer service covers the category, and which LLM fits support gets into the model choice itself.

Sources

Frequently Asked Questions

What is FLUX 3?
FLUX 3 is Black Forest Labs' multimodal foundation model, announced on 24 July 2026. Unlike FLUX.1 and FLUX.2, which generate images, FLUX 3 is trained jointly on images, video and audio in one architecture, and the same backbone also predicts robot actions. It generates video with native audio up to 20 seconds in a single pass. If the mechanism interests you, my explainer on diffusion models covers the family it grew out of.
How much does FLUX 3 cost?
There is no published FLUX 3 price. The Black Forest Labs pricing page lists FLUX.2, FLUX Tools and FLUX.1 only, and FLUX 3 has no row on it. The nearest real numbers are the FLUX.2 rates: $0.07 for the first megapixel on [max], $0.03 on [pro], and $0.014 on [klein] 4B. Anyone budgeting AI spend against outcomes rather than units will find cost per resolution a more useful frame.
Is FLUX 3 available yet?
Only partly. FLUX 3 Video is in early access behind a request form, and the FLUX 3 model page still reads "Coming Soon" as of 4 August 2026. FLUX 3 Image was described as arriving in "the following weeks", action prediction is limited to selected partners, and the open-weight FLUX 3 Dev release has no date. For self-serve work today the answer is still FLUX.2, the same way Nano Banana 2 and GPT Image 2 are usable now.
Is FLUX 3 better than Kling, Runway or Grok Imagine?
On Black Forest Labs' own preliminary numbers, human raters chose FLUX 3 over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69% and Kling v3 Pro in 60%. Against Seedance 2.0 and Gemini Omni Flash it was 52%, which is close to a coin flip. These are the vendor's own early evaluations of a model still in training, not third-party benchmarks, which is exactly the caution I apply to any AI performance metric.
Will FLUX 3 have open weights?
Black Forest Labs says yes, eventually: "FLUX 3 Dev" is listed as open-weight access to the multimodal backbone for both content creation and action prediction. No date, no parameter count and no licence terms have been attached to it. The existing pattern is a useful guide: FLUX.2 [klein] ships under Apache 2.0 with paid licence tiers starting at 10,000 images a month for open-weight deployment. Teams weighing self-hosting against an API should read my note on custom AI models first.
What does FLUX 3 have to do with robots?
The same backbone predicts actions. Black Forest Labs and mimic robotics built FLUX-mimic, a video-action model with a lightweight action decoder reading the FLUX 3 video path, and Audi has been testing it on soft-body manipulation work in production. Median autonomous success on a soft-body kitting task was 95% over 20 trials, versus 55% for the π0.5 baseline. It is the same architectural story as Gemini Robotics 2, told from the generation side.
Can I use FLUX 3 for blog images and support content?
Not yet, and FLUX.2 is the practical answer today. What FLUX 3 signals is that one model will soon cover the illustration, the short explainer clip and its voiceover, instead of three tools and three bills. That matters for anyone running an AI blog writer with images or building visual help-centre content, and it is worth reading alongside the best AI model for blogging.
What should support teams take from the FLUX 3 launch?
Read the access ladder, not the demo reel. FLUX 3 is a launch where the impressive part and the buyable part are different things, which is the exact shape of most AI support pitches too. Ask which capability is live today, on which plan, at which price, then test it on your own history before you commit. That is the whole point of agentic customer service software you can simulate first, and of running adversarial testing before launch rather than after.

Share this article

Rama Adi Nugraha

Article by

Rama Adi Nugraha

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
Illustration of a humanoid robot control model, representing Gemini Robotics 2
Trending

Gemini Robotics 2: what DeepMind's own numbers show

Gemini Robotics 2 ships three models and a per-task success table where most scores sit under 80%. That table is the most useful thing in the release.

Alicia Kirana UtomoAlicia Kirana UtomoAug 4, 2026
Illustration of image, video and document panels feeding a vision-language model, with the Qwen logo
Trending

Qwen 3.7 Flash: specs, pricing, and what it actually does

Qwen 3.7 Flash shipped with no blog post, no benchmarks and no weights. Here is the full spec sheet, the tiered pricing, and what Qwen never claimed.

Alicia Kirana UtomoAlicia Kirana UtomoJul 31, 2026
Skywork AI pricing breakdown illustration
Trending

Skywork AI pricing: what it really costs in 2026

A plain-English breakdown of Skywork AI pricing: the $1 trial, the credit system, the $19.99 Pro plan, and the billing gotchas to watch before you pay.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 20, 2026
Skywork AI review illustration showing a super-agent turning one prompt into slides, docs and websites
Trending

Skywork AI review (2026): capable agent, messy billing

An honest Skywork AI review: the super-agent makes real slides, docs and websites, but the trial-to-paid billing is where users get burned.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 20, 2026
Skywork AI super-agent workspace illustration
Trending

What is Skywork AI? The super-agent workspace, explained

Skywork AI is a general-purpose AI super-agent that builds slides, docs, sheets, sites and videos. Here's what it does, how it works, and what it costs.

Alicia Kirana UtomoAlicia Kirana UtomoJul 20, 2026
Editorial illustration representing a comparison of AI models as alternatives to Inkling
Trending

8 best Inkling alternatives in 2026

Inkling is open and interesting, but it's expensive for open weights and not the smartest model you can run. Here are the 8 alternatives I'd actually try instead, with real prices and where each one beats it.

Rama Adi NugrahaRama Adi NugrahaJul 20, 2026
Illustration of Inkling, Thinking Machines Lab's open-weights AI model under review
Trending

Inkling review: is Thinking Machines' open model worth it?

An honest Inkling review: what Thinking Machines Lab's first open-weights model is genuinely good at, where the price and benchmarks let it down, and who should actually run it.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 20, 2026
Illustration of Inkling, Thinking Machines Lab's open-weights AI model
Trending

Inkling explained: Thinking Machines' open-weights AI model

What Inkling actually is: Thinking Machines Lab's first open-weights model, its real benchmarks, what it costs to run, and whether it belongs anywhere near a support queue.

Alicia Kirana UtomoAlicia Kirana UtomoJul 20, 2026
Illustration of Google NotebookLM turning documents into a short vertical video overview
Trending

NotebookLM short Video Overviews: the 60-second format explained

NotebookLM's new Short Video Overviews turn your sources into a 60-second vertical clip. Here is what the format does, how it is built, and where it falls short.

Alicia Kirana UtomoAlicia Kirana UtomoJul 11, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free