期間限定オファー!年間プランが最大27%OFF、1年間の無限のクリエイティビティを解き放とう。

プランを見る ›
カルチャー対談

次にお気に入りになる動画モデルは、工場のラインも動かしているかもしれない

O
Emily Watterson
Jul 19, 2026 · 読了時間7分
Your Next Favorite Video Model Might Also Run a Factory Line

Black Forest LabsがFLUX 3の早期アクセスを開始 on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.

これはマーケティングの言葉ではなく本物のアーキテクチャ上の選択であり、FLUX 3ができることに直接表れています。最大20秒の動画を、後付けではなく同じ処理でネイティブオーディオを組み込んで生成でき、画像の合成や編集も可能です。意外なのは: 同じバックボーンがアクション予測へと拡張されています。 Black Forest Labsはmimic roboticsと提携しました FLUX-mimicを構築し、現在Audiの生産ラインで稼働している動画アクションモデルです。 クリエイティブ生成と物理ロボット制御が同じ基盤モデルから生まれること。これこそが、単一のベンチマーク以上に本当に重要なポイントです。

BFLが実際に主張していることは誇張されやすいので、正確な表現が重要です。 同社はFLUX 3を、このジョイントトレーニング原則に完全に基づいて構築した「初のモデル」と呼んでいます, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.

Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.

だからといって、今日のカテゴリーラベルが一夜にして消えるわけではありません。展開は研究より遅いため、しばらくは多くのツールが「動画用」「画像用」として出荷され続けるでしょう。しかし、動画・画像・音声は別々のものであり別々のモデルが必要だという根底にある前提こそが、実際に古びつつある部分です。そして最も恩恵を受けるクリエイターは、特定の動画モデルを使いこなした人ではありません。 彼らは、どのメディアを最初に開くかという問いを中心にプロセスを組み立てることをやめた人たちになるでしょう。

制限なく、思いのままに

OpenArtで画像・動画・キャラクター・ストーリーを生成する、数百万人のクリエイターに仲間入りしましょう。すべてが1つのプラットフォームに。

無料で始める →