Black Forest Labs
flux3
flux3 is Black Forest Labs’ early-access multimodal foundation model for video with native audio, still-image generation and action prediction. Its central idea is one backbone trained across images, video and audio rather than a collection of separate generators.
Inputs Text prompt / Image or keyframes · Video reference / Video + audio continuation
Outputs Video + native audio · Image · Action prediction