通量 3
Flux 3

原始链接: https://bfl.ai/blog/flux-3

FLUX 3 现已进入抢先体验阶段,标志着多模态人工智能的突破,超越了孤立的数据处理模式。该模型采用统一的“Self-Flow”架构,能够同时从图像、视频、音频和语言中进行学习。这种整体性的方法能够更深入地理解物理世界——包括物体如何移动、发声和互动——使模型能够将不同的模态视为同一现实的互补证据。 核心功能包括: * **视频与音频:** 高质量、多镜头生成(最长 20 秒)并支持原生音频,涵盖文本生成视频、图像生成视频以及多语言对话。 * **图像:** 增强了合成与编辑能力,在遵循复杂提示词方面表现更佳,且具备高精度的文字排版能力。 * **物理 AI:** 集成了动作预测功能,以与“FLUX-mimic”在机器人操控方面的合作为例。 早期评估显示,FLUX 3 在偏好测试中优于领先的竞争对手。该模型目前正通过 API 和私有权重访问权限进行推广,并计划在未来向开发者发布开放权重版本。FLUX 3 是该公司迈向构建统一模型目标的重要一步,旨在实现跨数字和物理环境的感知、预测与行动能力。

Black Forest Labs (BFL) 关于 **FLUX 3** 的公告在 Hacker News 上引发了激烈的讨论。BFL 计划发布一个用于内容创作和动作预测的开源权重“多模态主干网”,并同时提供闭源的 API 版本。 对该消息的反应呈两极分化: * **怀疑与批评:** 许多评论者持怀疑态度,指责 BFL 仅提供高调炒作,却缺乏实质性证据。批评者抱怨“世界模型”这一品牌名被滥用,且之前的开源权重模型往往逊色于专有模型(如 Ideogram 4 或 GPT-4),或受到限制性许可的束缚。 * **支持与乐观:** 支持者认为,BFL 之前的模型(如 Flux 2.dev)在消费级硬件上表现出色。他们看重在本地运行这些模型的能力,从而摆脱了集中式专有服务的束缚。 * **社会担忧:** 讨论进一步演变为关于人工智能对劳动力市场影响的广泛辩论,涉及“AI 垃圾内容”的伦理影响,以及技术社区是否变得过于愤世嫉俗或日益脱离自动化内容生成带来的社会后果。 总体而言,用户群体依然存在分歧:一部分人认为 FLUX 3 是一个充满希望的飞跃,而另一部分人则对行业营销周期感到疲惫。
相关文章

原文

FLUX 3 is now available in Early Access.

FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.

No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.

Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.

FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path.

FLUX 3: One model, multiple capabilities.

FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.

Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).

Capabilities & Early Evaluations

As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model’s key capabilities below.

Video

FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation.

Its core capabilities include the following (all outputs come with native audio generation):

  • Text-to-video generation.
  • Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references.
  • Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context.
  • Generative video-audio continuation from input video and audio.
  • Keyframe-to-video generation for controlled transitions between defined moments.
  • Multilingual dialogue.
  • A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output.
  • Agentic chaining of individual clips into longer, multi-shot sequences.
  • High style diversity -- FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.
  • Strong typography generation and animated designs.

For the preliminary analysis below, we generated 10-second text-to-video clips in 720p with audio.

Evaluations are early and we expect further improvements

As the model and the harness around it are still in development, these results are preliminary, and we expect further improvements during the early access phase. Across early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons.

While still in development, FLUX 3 Video is already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities. Furthermore, these capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that the characters remain consistent across all scenes.

FLUX 3 Video is now available in Early Access here

Image

FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations conducted during midtraining, FLUX 3 already shows a significant improvement over earlier versions of FLUX: its ability to handle complex prompts and text generation has improved significantly. The model produces a wide range of output styles (see the following samples), and is able to render high-accuracy text in multiple languages.

As with video evaluations, these are preliminary results, and we expect further improvements before release. We will open up an early access phase for FLUX 3 Image in the following weeks.

Action

FLUX 3's world understanding extends to action prediction. We have taken two routes to it: integrating native action prediction into FLUX 3 directly, scaling up our initial work in Self-Flow; and using the pretrained video backbone as a dynamics-aware foundation that specialized action models can be finetuned from with limited task-specific data.

For the second, mimic robotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation and production deployment. Read our thesis on why physical AI and content creation run on the same foundation, and how it's being tested on real production tasks at Audi.

Launch Plan

Over the next few weeks and months, we will make the following capabilities available, each after an early access phase for ensuring smooth rollout, collecting feedback and rigorous safety-testing. All capabilities are built from the same underlying multimodal flow matching model. These capabilities and models include:

  • Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
  • Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
  • Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
  • Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)

We will also release more technical details on the underlying approach.

Request early access here

What’s next?

We are only beginning to scratch the surface of versatile, capable, unified multimodal models, and what they will enable. From interactive image & video editing, simulation to computer use and physical AI, the frontier is wide open. While we gradually roll out these new capabilities, we are already working on the next generation models. Our goal is to unify perceptual, action and language prediction in the same unified model.

If you are interested in exploring and building with FLUX 3, get in touch here. If you are interested in contributing to our mission, join us! We are hiring in Germany and the US.

联系我们 contact @ memedata.com