Black Forest Labs launched FLUX 3 today, a single jointly trained multimodal model that generates images, video clips up to 20 seconds, and native audio from one prompt. This is BFL's first public video model and its most ambitious architectural move: not a pipeline of bolted-together specialists, but one backbone the Freiburg lab is positioning across creative generation, simulation, computer use, and robotics under the label 'visual intelligence.' Four product lines follow: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the forthcoming open-weight FLUX 3 Dev.

The benchmark numbers are preliminary, labeled by BFL itself as 'an early FLUX 3 candidate,' meaning they describe a pre-release checkpoint, not the model entering early access today. In internal preference tests on 10-second 720p text-to-video clips, FLUX 3 beat Luma Ray 3.2 in 93% of comparisons and Runway Gen-4.5 in 77%. It landed at 52% against Gemini Omni Flash, a statistical tie against the closest large-platform analogue to what FLUX 3 is attempting, a model already generally available at $1.00 per 10-second 720p clip. No pricing, SLA, benchmark methodology, sample sizes, or rater counts have been published for FLUX 3. Enterprise buyers cannot calculate cost of ownership or independently reproduce the video comparisons.

Two omissions define the release. First, no downloadable weights and no open-source license at launch. FLUX 3 Dev, which BFL describes as open-weight access covering video, audio, image, and action prediction, a far broader commitment than any prior FLUX Dev release, arrives later this year, last in the rollout sequence. Second, no API access yet for FLUX 3 Video or FLUX 3 Action, both gated behind an approval-required early access program. The full article is worth reading for the competitor pricing table, the legal backstory behind the Seedance 2.0 comparison, and the EU-specific restriction on Gemini Omni Flash that quietly gives a German lab like BFL a geographic opening.

[READ ORIGINAL →]