عرض لفترة محدودة! أطلق عام كامل من الإبداع بلا حدود مع الخطط السنوية بخصم يصل إلى 27%.

عرض الخطة ›
محادثات ثقافية

قد يُدير نموذج الفيديو المفضّل التالي لديك خط إنتاج مصنع أيضًا

O
Emily Watterson
Jul 19, 2026 · قراءة لمدة 7 دقائق
Your Next Favorite Video Model Might Also Run a Factory Line

فتحت Black Forest Labs الوصول المبكر إلى FLUX 3 on July 23, and the company's own explanation of why it built the model is more interesting than any spec sheet. Images, video, and audio, in BFL's framing, are each an incomplete "projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train a model on only one of those projections and it learns that projection. Train it on all three at once, and the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past, and the model ends up learning something closer to how the world actually holds together.

هذا خيار معماري حقيقي، وليس لغة تسويقية، ويظهر مباشرةً فيما يمكن أن يفعله FLUX 3. فهو يولّد فيديو يصل طوله إلى 20 ثانية مع صوت أصلي مدمج من نفس التمريرة، لا مضاف لاحقًا، إلى جانب تركيب الصور وتحريرها. أما الأقل توقعًا: يمتد الأساس نفسه إلى التنبؤ بالحركة. تعاونت Black Forest Labs مع mimic robotics لبناء FLUX-mimic، وهو نموذج فيديو-حركة يعمل الآن على خطوط الإنتاج في Audi. إن التوليد الإبداعي والتحكم بالروبوت المادي، اللذين ينبعان من النموذج الأساسي نفسه، هما التفصيل الذي يهم فعلاً هنا، أكثر من أي معيار قياس منفرد.

من السهل المبالغة في وصف ما تدّعيه BFL فعلًا، لذا فإن الصياغة الدقيقة مهمة. تصف الشركة FLUX 3 بأنه "نموذجنا الأول" المبني بالكامل على مبدأ التدريب المشترك هذا, not the first multimodal model industry-wide, and Google's Veo 3 already generates audio and video jointly in a single pass. BFL's own early comparisons, preferring FLUX 3 over other video models in the majority of side-by-side tests, are preliminary and self-reported, an early signal rather than a settled result. The actual signal is that a serious lab spent its resources dissolving the boundary between "image model," "video model," and "audio model" into one shared representation, and that a growing list of others are converging on the same bet from different angles, not any single superlative claim.

Once that boundary stops being architecturally necessary, it stops being a useful way to think about the tools too. A creator working with something like this doesn't start by deciding "I'm making a video" or "I'm making an image." They start with a character, a world, a physical premise, and the medium that premise ends up taking, a still frame, a twenty-second clip, a sound design choice, becomes a downstream decision rather than the first one. That's a genuinely different working process than choosing a video generator and prompting it, and it's closer to how a director thinks about a scene than how someone thinks about operating a tool.

لا شيء من هذا يعني أن تسميات الفئات الحالية ستختفي بين ليلة وضحاها. كثير من الأدوات ستستمر في الظهور كـ"أداة الفيديو" أو "أداة الصور" لفترة، لأن الطرح أبطأ من البحث. لكن الفرضية الأساسية، أن الفيديو والصورة والصوت أشياء منفصلة تتطلب نماذج منفصلة، هي الجزء الذي بدأ يتقادم فعلاً، والمبدعون الأكثر استفادة لن يكونوا من أتقنوا نموذج فيديو معيّناً. سيكونون أولئك الذين توقفوا عن تنظيم سير عملهم حول سؤال أي وسيط يفتحونه أولًا.

أبدع بلا حدود

انضم إلى ملايين المبدعين الذين يستخدمون OpenArt لإنشاء الصور والفيديوهات والشخصيات والقصص - كلها في منصة واحدة.

ابدأ مجانًا →