FLUX 3 Action: how a video model learned to drive robots

Alicia Kirana Utomo
Written by

Alicia Kirana Utomo

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 28, 2026

Expert Verified
Illustration of a person watching a world-model diagram drive a robot arm, for FLUX 3 Action

What FLUX 3 Action actually is

Let me clear up the naming first, because it trips people up. FLUX 3 is a single multimodal foundation model from Black Forest Labs, the same lab behind the FLUX image models. It is trained jointly across image, video, and audio, and the company describes it as "one multimodal model for Image, Video, Audio and Action-Prediction."

Action-Prediction is that fourth capability. It takes visual observations plus a text instruction and predicts physical outcomes as robot control actions. So "FLUX 3 Action" is not a separate product you download. It is a capability of the FLUX 3 backbone, and the version built for real robots, in partnership with mimic, is called FLUX-mimic.

The framing BFL uses is the interesting part. In their words, a model that can both generate video and control robots "was never really only a content creation model. It is a model of how the world behaves, and content creation is one thing one can do with it." That is a big claim, so it is worth walking through why it might be true.

Why a video model can drive a robot

Here is the mechanism, because this is the whole idea. When you train a model to generate realistic video, it has no choice but to learn physics. Get the weight of an object wrong, or the way a hand closes around it, or what happens when two things collide, and the video looks obviously fake. So learning to render the world accurately is, under the hood, learning how the world behaves.

BFL puts real numbers on this. Video prediction accounts for over 95% of FLUX 3's total training compute. Audio, by comparison, is less than 0.5% of the tokens in a 720p clip, because once you understand video, matching sound to lip movement and impacts is comparatively easy. Actions follow the same shape: a robot's state is a low-dimensional signal tightly coupled to what the camera sees. Frames, audio, and actions are all partial views of one underlying physical reality, so once the model has learned that reality from video, action prediction is not a new departure.

How FLUX 3 Action turns a camera view and a text instruction into robot control actions through the shared FLUX 3 backbone and a lightweight action decoder
How FLUX 3 Action turns a camera view and a text instruction into robot control actions through the shared FLUX 3 backbone and a lightweight action decoder

The pipeline is deliberately lean. A camera view and an instruction go into the FLUX 3 backbone, which holds the learned world model. A small action decoder reads intermediate features straight out of the video-prediction path and turns them into control actions. The expensive part, understanding the world, is already done before the robot ever moves. This is a distinct approach from most vision-language-action models, which are trained end to end to spit out actions rather than reusing a generative backbone.

What it cost to add actions

If actions really are just another view of the same world, teaching FLUX 3 to predict them should be cheap and shouldn't damage what it already does. That is exactly what BFL reports. When they added action prediction to a large training run, human ratings on text-to-video and image-to-video initially fell by up to 10% as the model absorbed the new modality. After 3,500 steps, video quality was fully back to where it started, now with action prediction on top.

That recovery is the evidence for the "one backbone, two jobs" thesis. The model didn't permanently trade video quality for robot control. It just had to figure out how the new modality mapped onto the world it already knew, then the penalty vanished. The research this builds on is called Self-Flow, which unifies generation and representation learning so that improving one improves the other.

You can see the split in BFL's own architecture diagram. The text and video encoders feed the shared FLUX backbone, and the same layers produce future features two ways: through a video decoder to generate frames, and through a lightweight action decoder to output robot actions. One trunk, two heads.

FLUX-mimic architecture diagram: text and video encoders feed the shared FLUX backbone, whose future features branch to both a video decoder for frames and an action decoder for robot actions, as taken from Black Forest Labs
FLUX-mimic architecture diagram: text and video encoders feed the shared FLUX backbone, whose future features branch to both a video decoder for frames and an action decoder for robot actions, as taken from Black Forest Labs

What the benchmarks actually show

This is where FLUX 3 Action gets surprising, so I want to sit with the numbers rather than wave at them. On a soft-body kitting task (the kind of flexible-material handling conventional automation has always struggled with), BFL measured median success over 20 autonomous trials per model.

Benchmark chart of a soft-body kitting task showing median autonomous success rate over 20 trials: FLUX-mimic 95%, single-task flow matching 70%, FLUX-mimic frozen backbone 65%, pi-0.5 55%, pi-0.5 frozen backbone 0%, as taken from Black Forest Labs
Benchmark chart of a soft-body kitting task showing median autonomous success rate over 20 trials: FLUX-mimic 95%, single-task flow matching 70%, FLUX-mimic frozen backbone 65%, pi-0.5 55%, pi-0.5 frozen backbone 0%, as taken from Black Forest Labs

FLUX-mimic landed at 95%. A single-task flow-matching baseline hit 70%, and π0.5, a strong prior vision-language-action model, hit 55%. The number that matters most, though, is the frozen-backbone column. When you freeze FLUX-mimic's backbone (no fine-tuning, just the pretrained world model with a decoder on top), it still scores 65%. Do the same to π0.5 and it scores 0%. In other words, FLUX 3's pretrained representation is good enough to drive a robot with almost no task-specific training, which is not something earlier models could claim.

A few more findings worth pulling out, because they compound:

  • Sample efficiency. The mimic-video work reports up to 10x fewer examples than a vision-language-action model to reach the same capability, because the physics is already in the representation.
  • Self-recovery. A robot that misses a grasp corrects itself and tries again, even though no demonstration showed that recovery. It comes from the world model, not the training set.
  • Speed. The backbone runs from camera input to world representation in under 80ms on a single NVIDIA RTX 5090, and mimic's full robot loop reacts in 101ms, roughly human visual reaction time.

The speed point is subtle but important. Because better representations mean more capability per parameter, BFL can run a smaller (and therefore faster) backbone and still hit the accuracy bar. In robotics, latency is the constraint that kills deployments, so a smaller backbone that keeps up with the real world is a bigger deal than the raw success rate.

From the lab to the factory floor

None of this would matter if it only worked in simulation, so the Audi deployment is the proof point. Audi runs one of the most automated production networks in the car industry, which means it has a precise view of where conventional robots still stop: flexible parts, fine manipulation, and the variant diversity of premium production that makes hand-programmed robot cells too expensive to re-engineer per case.

"In partnership with mimic, Audi has been testing and deploying FLUX-mimic. We have seen these robots solve complex soft-body manipulation work that would have been simply impossible with conventional robotics."

The tasks BFL names are concrete: kitting parts into structured trays, inserting electronic control units into tight fixtures, assembling components, and handling soft, flexible materials like seals and cables. That last category is the one that has stayed stubbornly manual for decades, so it is a fair signal of where this approach earns its keep.

Can you actually use FLUX 3 Action?

Here is the honest status, because the excitement outruns the availability. Black Forest Labs is shipping FLUX 3 in phases, and action prediction is not one you can buy.

A four-step staircase showing FLUX 3's phased rollout: video and audio live now, action prediction for partners only, image generation soon, and an open-weight backbone with no date
A four-step staircase showing FLUX 3's phased rollout: video and audio live now, action prediction for partners only, image generation soon, and an open-weight backbone with no date

The rollout order BFL published is: video generation and editing first, action prediction with selected partners, then image synthesis, then open-weight access to the backbone. Today only video-with-audio is generally available. Action prediction is partner-gated (mimic has early access), with no public API, no published price, and no open weights. FLUX 3 image generation is still marked "soon," and the open-weight backbone has no date at all.

So if you want to touch FLUX 3 today, what you can actually pay for is video: text-to-video, image-to-video, and video-to-video up to 20 seconds, with native synchronized audio, on pay-as-you-go BFL API pricing with no subscription. For anything robot-related, you are watching a research preview and a partner program, not a product. That is not a criticism, it is just where the frontier is. Physical AI moves from selected partners to general availability slowly for good reasons.

Infrastructure versus the teammate you actually hire

I want to zoom out, because "FLUX 3 Action" and phrases like "AI agent" get thrown into the same bucket, and they belong on completely different shelves.

A two-column comparison: the model as infrastructure (raw capability, you build the app, you wire the tools) versus the teammate (arrives job-ready, knows your company, works the queue day one)
A two-column comparison: the model as infrastructure (raw capability, you build the app, you wire the tools) versus the teammate (arrives job-ready, knows your company, works the queue day one)

A foundation model like FLUX 3 is infrastructure. It is raw capability. To get value out of it you build the application, wire up the tools, handle deployment, and (in this case) attach it to physical robots. That is exactly the right shape for a robotics lab or an automotive production team with engineers. It is the wrong shape for a support lead who just needs tickets answered.

The other shelf is the AI teammate: something that arrives job-ready, already knows your company, and starts working on day one. You don't build an app around it, you brief it like a new hire. That is the layer I work on, and it is worth being clear that a model is not a teammate, any more than an engine is a car. FLUX 3 Action is a remarkable engine for physical AI. It is just not the thing you hire.

How this maps to eesel

At eesel, we take the "employee, not infrastructure" side of that line seriously. eesel is an AI teammate platform: instead of handing you a raw model to build on, you hire ready-to-work teammates for specific jobs. The current roster is an AI helpdesk teammate and an AI blog writer, each arriving with the skills, integrations, and company context its role needs.

The eesel reports dashboard showing task volume, trigger events by type, and approval usage for an AI teammate
The eesel reports dashboard showing task volume, trigger events by type, and approval usage for an AI teammate

Because the model underneath is a moving target (FLUX today, something else next quarter), the useful abstraction is the teammate, not the raw capability. The helpdesk teammate learns from your past tickets and help center, drafts replies for review, and you can simulate it against your own historical tickets before it ever goes live, because I have watched confident-sounding bots quietly give wrong answers and would rather catch that in a dry run. If you like driving things from a terminal, the eesel CLI and its MCP server let you (or a coding agent like Claude Code) set up integrations, upload knowledge, and approve actions as JSON, so the same teammate is programmable without the dashboard. It won't fold your laundry or man a factory line. For answering customers and writing content, though, it is the hire, not the infrastructure. It is free to try.

Frequently Asked Questions

What is FLUX 3 Action?
FLUX 3 Action is the action-prediction capability of FLUX 3, Black Forest Labs' multimodal foundation model. Instead of generating pixels, it decodes the model's learned world representation into robot control actions. The productized version, built with mimic robotics, is called FLUX-mimic and has been tested and deployed at Audi. It is closely related to other physical-AI models like Gemini Robotics 2.
Can I use FLUX 3 Action yet?
Not on your own. Action prediction is gated to selected partners (mimic first), with no public API, price, or open weights as of September 2026. What you can buy today from FLUX 3 is video-with-audio generation, on pay-as-you-go BFL API pricing. Image generation and an open-weight backbone are still on the roadmap.
How is FLUX 3 Action different from a normal robotics model?
Most vision-language-action models are trained specifically to output actions. FLUX 3 Action instead reuses a generative video model as the backbone and trains a small action decoder on top of its internal features. In benchmarks it beat prior approaches even with the backbone completely frozen, a setting where earlier models scored 0%.
What is FLUX-mimic?
FLUX-mimic is the video-action model that mimic robotics and Black Forest Labs built on the FLUX 3 backbone. It handles factory tasks like kitting soft parts, inserting control units into fixtures, and manipulating cables and seals, with a full robot reaction time of 101ms. Read the details on the FLUX 3 x mimic announcement.
Is FLUX 3 Action an AI agent I can hire for my business?
No. FLUX 3 Action is infrastructure for physical robots, not a business software agent. If you want a ready-to-work AI teammate for support or content, that is a different layer of the stack. eesel provides AI teammates that plug into your existing tools and start working from your company's history on day one.

Share this article

Alicia Kirana Utomo

Article by

Alicia Kirana Utomo

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Grok Bot and Claude Cowork compared, two AI agent teammates side by side
AI

Grok Bot vs Claude Cowork: which AI agent teammate actually fits your work?

A hands-on comparison of Grok Bot and Claude Cowork: how each general-purpose AI agent works, what they cost, the security trade-offs, and where neither one fits.

Alicia Kirana UtomoAlicia Kirana UtomoSep 22, 2026
Grok Bot and Lindy shown side by side as two AI teammates working at laptops
AI

Grok Bot vs Lindy: which AI teammate actually does the work?

A hands-on Grok Bot vs Lindy comparison: how each AI teammate works, what they really cost, the security trade-offs, and which one fits your team.

Alicia Kirana UtomoAlicia Kirana UtomoSep 22, 2026
Illustrated banner for a guide on automating customer support from the command line
AI

How to automate customer support from the command line in 2026

You can automate a lot of support from the terminal: routing, tagging, escalation, exports, scheduled sweeps. Here's the ladder of what's scriptable, and the one rung that isn't.

Alicia Kirana UtomoAlicia Kirana UtomoSep 7, 2026
Illustrated banner for a guide on running customer support from the command line
AI

A CLI for customer support: how to run support like code in 2026

A CLI for customer support isn't one magic binary. It's a way to make support programmable, testable, and versioned. Here's what actually works from the terminal.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 7, 2026
A developer at a terminal wiring code, an API connector, and a webhook into an AI support agent
AI

Programmatic AI agent access: which surface fits which job

Programmatic AI agent access is not one API. It is a spectrum of surfaces (REST, CLI, MCP, webhooks, Network Access), each fitting a different job.

Rama Adi NugrahaRama Adi NugrahaSep 8, 2026
Hand-drawn illustration of a team gathered around a laptop with an OpenClaw lobster agent connecting to several people
AI

OpenClaw 2.0 review: what actually changed, and is it worth it

An honest OpenClaw 2.0 review: the multiplayer shift, the 16,977-PR release, easier setup, and the catch nobody self-hosting can skip.

Rama Adi NugrahaRama Adi NugrahaSep 4, 2026
AI technology enhancing customer support operations
AI

The Future of AI in Customer Support

Exploring how AI is transforming customer support operations and what teams should know.

Stevia PutriStevia PutriAug 30, 2026
Illustration of a credit meter and three plan tiers, representing Gumloop's credit-based pricing
AI

Gumloop pricing 2026: what a credit really costs you

Gumloop pricing starts at $37/month for 20,000 credits. Here's what a credit actually is, the five meters on every agent chat, and where the bill jumps.

Rama Adi NugrahaRama Adi NugrahaAug 17, 2026
Illustration of the Buzz app: chat channels where people and AI agents collaborate, with a honeycomb motif
AI

What is Buzz? Jack Dorsey's AI agent workspace, explained

Buzz is Jack Dorsey's new open-source team chat app where humans and AI agents share the same channels. Here's what it is, who it's for, and the catch.

Alicia Kirana UtomoAlicia Kirana UtomoJul 23, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free