你喜爱的模型——一站汇聚,无限生成。

限时享受最高 27% 优惠!›
OpenArt 更新

你下一个最爱的视频模型,也许还能撑起一条工厂生产线

O
Emily Watterson
2026年7月19日 · 阅读时长 7 分钟
Your Next Favorite Video Model Might Also Run a Factory Line

Black Forest Labs 开放了 FLUX 3 的抢先体验 on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.

这是一个实打实的架构选择,而非营销话术,并直接体现在 FLUX 3 的能力上。它能生成最长 20 秒的视频,音频在同一次生成中原生融入,而非事后叠加,同时还支持图像合成与编辑。更出人意料的是: 同一套底层架构还延伸到了动作预测领域。 Black Forest Labs 与 mimic robotics 达成合作 共同打造了 FLUX-mimic,这是一款视频动作模型,目前已在奥迪的生产线上运行。 创意生成与实体机器人控制出自同一个底层模型——这才是这里真正关键的细节,其分量胜过任何单一的基准测试成绩。

BFL 真正宣称的内容很容易被夸大,所以精确的说法很重要。 该公司称 FLUX 3 是完全基于这一联合训练理念构建的「首款模型」, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.

Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.

这一切并不意味着如今的品类标签会一夜之间消失。相当长一段时间里,仍会有大量工具以「视频那款」或「图像那款」的身份推向市场,因为落地的速度远慢于研究。但真正开始过时的,是那个底层前提——即视频、图像和音频是各自独立、需要各自独立模型的东西。而受益最大的创作者,不会是那些精通某一款特定视频模型的人。 他们会是那些不再围绕「先打开哪种媒介」来组织创作流程的人。

创作无极限

加入数百万创作者的行列,用 OpenArt 生成图像、视频、角色和故事——全都在一个平台上完成。

免费开始使用 →