Model analysis · evidence checked

flux3

Early Access

flux3 is Black Forest Labs’ early-access multimodal foundation model for video with native audio, still-image generation and action prediction. Its central idea is one backbone trained across images, video and audio rather than a collection of separate generators.

Inputs

  • Text prompt / Image or keyframes
  • Video reference / Video + audio continuation
flux3 Unified multimodal backbone

Outputs and access

  • Video + native audio Early Access
  • Image Planned
  • Action prediction Partner access

Official launch film

The official Black Forest Labs launch film introducing the model’s image, video, audio and action-prediction direction.

46 seconds · Black Forest Labs
Official source media. Playback streams from Black Forest Labs’ media host. Open source page ↗

Our analysis of flux3

flux3 looks less like a routine video generator and more like a multimodal world-model platform. The announced controls are unusually broad, but today the evidence comes from Black Forest Labs and access remains limited. Treat it as a promising early-access model, not a finished open-source release.

Our analysis

This report analyzes the official announcement. FindMoreAI has not received model access or independently generated outputs.

Eight key capabilities

The translated social post supplied the outline; every point below is verified against the official BFL announcement.

  1. 01 T / I → V

    Text-to-video and image-to-video

    Early Access

    flux3 can generate video from text, animate a starting image, or use images as visual references.

  2. 02 V → V

    Reference-based video generation

    Early Access

    A reference clip can guide a new scene while carrying forward central elements, including the same character.

  3. 03 AV → AV+

    Video and audio continuation

    Early Access

    flux3 can extend an input video and its audio together rather than treating sound as a separate post-production step.

  4. 04 KF → V

    Keyframe-controlled generation

    Early Access

    Defined keyframes can steer transitions between selected moments, giving creators more control over motion and composition.

  5. 05 L → AV

    Multilingual dialogue and flexible formats

    Early Access

    Black Forest Labs reports multilingual dialogue, varied visual styles and a broad range of aspect ratios.

  6. 06 V₁…Vₙ

    Multi-shot sequence chaining

    Early Access

    Individual clips can be chained into longer multi-shot sequences. The source describes this as agentic chaining, not conventional timeline editing.

  7. 07 T → MOTION

    Typography and animated design

    Early Access

    The model is presented as capable of typography generation and animated design, an area many video systems struggle to render reliably.

  8. 08 T / I → I

    Image synthesis and editing

    Access planned

    flux3 is intended to synthesize and edit still images across styles, resolutions and aspect ratios; image access is not yet generally open.

How flux3 is intended to work

Confirmed in source

flux3 is Black Forest Labs’ first model presented as being trained around a unified representation of images, video and audio. The practical thesis is that each modality constrains the others: motion should follow physical structure, generated sound should match an event, and later frames should follow from earlier ones.

The announcement says the model builds on Self-Flow, Black Forest Labs’ method for aligning multimodal generation and understanding in the same architecture. That makes the model’s significance broader than its feature list: the intended foundation is shared across content creation and physical-world prediction.

Why action prediction matters

The same backbone is intended to predict actions as well as media. Black Forest Labs describes both native action prediction and specialized action models fine-tuned from the pretrained video backbone.

The company says mimic robotics received early access and used the backbone for a video-action system tested on production tasks at Audi. This is meaningful partner evidence, but it is not proof that general robotics access or an independently reproducible action model is available.

Preliminary benchmark claims

Vendor evaluation

For its preliminary comparison, Black Forest Labs says it generated 10-second, 720p text-to-video clips with audio. The announcement reports the following human preference rates for flux3:

Compared withReported preference for flux3
Grok Imagine VideoUp to 69%
Kling v3 Pro60%
Happy Horse v159%
Happy Horse 1.157%
Seedance 2.052%
Gemini Omni Flash52%
Runway Gen-4.577%
Luma Ray 3.293%
How to read these numbers

These are vendor-reported, preliminary results. The announcement does not disclose enough about evaluator count, prompt selection, blind-review procedure or statistical uncertainty to treat the figures as an independent ranking.

Availability and release plan

Announced does not always mean accessible. The rollout states are kept separate here.

Video and native audio

Early Access now

API delivery and private weight access are part of the announced rollout.

Image generation and editing

Early Access announced

Black Forest Labs says image access will begin in the following weeks.

Action prediction

Selected partners

The initial rollout is limited to research and commercial partners.

Dev backbone

Open weights planned

No public checkpoint or final license is included in the announcement.

Is flux3 open source?

flux3 is not publicly open source today. Black Forest Labs has announced future open-weight access to a Dev backbone, but the launch post does not provide downloadable weights, a public checkpoint or a license.

Open-weight and open-source are not interchangeable. The eventual license will determine whether commercial use, modification and redistribution are permitted. Until that license is published, describing flux3 as already open source overstates what is available.

What remains unknown

Not published

The launch announcement leaves several production-critical questions unanswered:

  • Public API pricing and rate limits
  • Parameter count and model sizes
  • Hardware and memory requirements
  • The Dev checkpoint release date
  • The Dev license and commercial-use terms
  • Independent benchmark results
  • Detailed training, safety and evaluation methodology

FindMoreAI conclusion

flux3 is notable because Black Forest Labs is combining media generation and action prediction around one multimodal backbone. The eight announced creation capabilities cover more control modes than a basic prompt-to-video system, while native audio and reference-driven generation point toward longer production workflows.

The limiting factor is evidence and access. The launch materials are promising, but they remain vendor evidence from an evolving Early Access model. For most people, the sensible next step is to follow the rollout rather than plan a production integration around specifications that have not yet been published.