Black Forest Labs Unveils FLUX 3 AI: Ditches Stills for Video—And Robot Hands

FLUX 3 is the German AI lab’s first video model, and the same system is already teaching robots to work an Audi assembly line.

By Jose Antonio Lanz

3 min read

Black Forest Labs released FLUX 3 on Thursday, and for the first time, the company’s flagship model generates video instead of just still images. The German AI lab, known for the FLUX line of image generators, trained the new system on images, video, and audio at once, inside one shared system.

That’s what is known as multimodality: one model learning several types of information together instead of separate tools bolted side by side.

The video side is the headline feature. FLUX 3 produces clips up to 20 seconds long, with audio generated alongside the picture and synced to what’s happening on screen—dialogue, sound effects, ambient noise. In early evaluations, human reviewers preferred FLUX 3’s output over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93%. It seems to be slightly better than Gemini Omni and Seedance, beating those models in 52% of the evaluations.

Of course, that’s a preference test, not a fixed scoring rubric: evaluators simply watch two clips and pick the one that looks and sounds more convincing, and BFL counts how often FLUX 3 wins.

Other than that, the model seems to be very competent on still images too, following its legacy. BFL shared a few images, and FLUX 3 seems to be very versatile and capable of generating a broad variety of styles beyond photorealism.

BFL frames this as more than a content tool. “A model that only learns images can only generate images,” said co-founder and CEO Robin Rombach. The company’s bet is that learning to predict video also means learning the physics underneath it—weight, contact, timing—which is exactly what a machine needs to move through the physical world.

That bet has a name: FLUX-mimic. Built with Zurich-based mimic robotics, it takes FLUX 3’s video-prediction engine and adds a lightweight “decoder”—a small add-on component that translates the model’s internal sense of how things move into actual robot motions. Car maker Audi is already testing it on tasks like fitting flexible door seals, work that conventional automation has struggled to handle.

“Audi represents the kind of manufacturing partner we built FLUX-mimic for,” said mimic co-founder Stephan-Daniel Gravert. Audi’s Christoph Schneider said the robots now “solve complex soft-body manipulation work” that older machines couldn’t touch. BFL says the full system reacts in about 101 milliseconds, in the neighborhood of human visual reflexes.

FLUX’s rise didn’t happen in a vacuum. Founded in August 2024 by veteran researchers who’d helped build the original Stable Diffusion models at Stability AI, Black Forest Labs launched Flux models that beat MidJourney and outclassed Stability’s own underwhelming Stable Diffusion 3.

The open-source Flux Dev and Schnell models grabbed the “best open source image generator” title that AI artists had expected Stable Diffusion 3.5, Stability’s do-over, to eventually reclaim.

It never did. FLUX 1.1 Pro went on to top the Artificial Analysis image arena that October. That one wasn’t open source, though.

BFL released FLUX.2 in November 2025 but it wasn’t as popular. The open-source crown held by the original Flux lasted until Alibaba’s Z-Image Turbo dethroned it in late 2025, matching its quality on lower end consumer graphics cards. “This is what SD3 was supposed to be,” one CivitAI user wrote at the time.

FLUX 3 is BFL’s comeback, and it isn’t fully open yet. Video and Action are in early access now through APIs and select partners, mimic robotics among them, with image generation following “in the coming weeks,” per BFL. The open-weight Dev version, the only tier BFL plans to release for local use, isn’t due until later in 2026.

Get crypto news straight to your inbox--

sign up for the Decrypt Daily below. (It’s free).

Recommended News