Methodology

Creative professionals compare two outputs at a time without knowing which model produced either, voting on one named criterion at a time. The criteria, weights, judge counts and quality controls behind each ranking are documented below.

How a ranking is produced

  • BoardsBoards and criteria are defined

    Boards cover capability tracks and industry categories. Criteria are defined by people working in that field, then approved and frozen before scoring begins.

  • PromptsPrompts are procured, then run

    Prompts come from real creation use cases, in a mix of professional and amateur styles. OpenArt runs every model on the same brief; one output per model enters evaluation by a fixed rule, never hand-picked.

  • VotingJudges vote in blind pairs

    Model identities are hidden and left/right position randomized. Each comparison is a forced choice on one named criterion - no tie, no overall-preference question.

  • ResultsResults are compiled

    Votes are fit per criterion, combined into board scores with published weights, and reported with uncertainty. The scoring rules are set out below.

Boards in scope

From votes to published scores

  1. Fitting

    One fit per criterion

    Criterion scores use a Bradley-Terry estimator, so beating a strong model moves a score more than beating a weak one. One anchor model per board pins the scale, and scores are reported on an Elo-like scale relative to that anchor.

    General criteria are fit on the modality master from pooled industry comparisons. Industry and capability boards reuse those general fits, including scores, vote and judge counts, and bootstrap intervals, with ranks renumbered within their own model rosters.

  2. Composites

    Snapshot-weighted composites

    Industry and capability headline scores combine general capabilities with board-specific criteria using the weights recorded in each published snapshot. Overall boards use general criteria only. Snapshot weights apply identically to every eligible model; configuration changes do not rewrite earlier results.

    Votes are weighted by judge tier: the Creative Expert Council at 3× and Tastemakers at 1×. Weights are published before scoring begins and frozen after publication; any later change is a versioned ruleset update.

  3. Coverage

    Coverage before rank

    A model receives a board composite only after clearing every required criterion on that board. Criteria that have cleared the vote minimum are shown while the rest collect votes.

  4. Publishing

    Checked, versioned, reproducible

    Every score is published with a 95% confidence interval from judge-cluster bootstrap resamples, a rank, and the distinct-judge count behind it.

    Before release each board is re-run with the expert multiplier moved up and down, without its highest-volume judge, and on an alternate bootstrap seed. Results are published either way; an order change holds the board for review.

    Each scoring run creates an immutable snapshot of scores, weights, configuration and code version. Scores are comparable within a board only.

Reading the numbers

How to read an Arena score

  • It is a relative scale, not a grade

    Scores sit on an Elo-like scale pinned to one anchor model per board. A score says where a model sits relative to the others on that board, nothing more.

  • Overlapping intervals mean a tie

    Every score carries a 95% confidence interval and a rank. If two intervals overlap, the models are not distinguishable on that criterion, whatever the ranks say.

  • Never compare across boards

    Each board's composite reflects its own criteria and weights, even when general fits are shared. A 1,240 on Film and a 1,240 on Ads are not comparable headline scores.

General capability criteria

5 criteria · voted across video types and use cases

Prompt adherence

Which output follows the prompt more accurately?

Signature prompt

Referring to the scene in image 1, he remains still as the camera holds on the scene with a subtle natural handheld drift and zoom in slowly.

PreferredThe character stays still while the camera zooms in.
Not preferredThe character doesn't stay still, and there are other camera movements.

Subjective aesthetics

Which output is more aesthetically pleasing?

Signature prompt

Surreal fantasy landscape featuring floating bioluminescent islands in a twilight sky. Glowing magical flowers swaying in a light breeze, purple and emerald volumetric light beams slicing through soft clouds. Vertigo dolly zoom camera effect, ethereal atmosphere, fairytale aesthetics, photorealistic render.

PreferredThe good case has a much better atmosphere and hits that fantasy vibe really well.
Not preferredThe bad case has an AI-generated feel.

Physics & motion

Which output shows more convincing physics and motion?

Signature prompt

The girl in the reference picture looks at the shore in surprise [camera zooms out], then jumps into the pool without hesitation and begins to swim vigorously to the other side.

PreferredThe diving motion and the splash upon hitting the water look very natural.
Not preferredThe splash and ripples have a strong AI-generated look.

Consistency

Which output maintains stronger visual consistency?

Signature prompt

Using the reference image, slowly transform the woman in the first image into the woman in the second image.

PreferredThe characters in the video maintain consistency with the reference image.
Not preferredThe first frame shows changes to the reference character's face.

Audio quality

Which output has better audio quality?

Signature prompt

Visually explosive superpower action sequence, dynamic volumetric lighting, VFX masterclass. [0:00-0:03] A glowing energy-infused character floats an inch off the ground, hair and coat levitating in reverse gravity as bright blue plasma arcs around their fists. [0:03-0:06] The character thrusts both palms forward, releasing a massive kinetic plasma shockwave that blows back debris, car windows, and dust in a violent sphere explosion. Dynamic camera shake, blinding light flashes, smooth motion physics

PreferredThe audio matches the video content well, and the ambient sound effects feel natural and vivid.
Not preferredThe sound effects do not quite match the video content and lack realism and vividness.

Board-specific criteria

Overall VideoVideo
  • Prompt adherence20%
  • Subjective aesthetics20%
  • Physics & motion20%
  • Consistency20%
  • Audio quality20%
Video EditingVideo
  • General capabilities30%
  • Prompt adherence to the edit23.3%
  • Precise editing23.3%
  • Style adaptation23.3%
Lip SyncVideo
  • General capabilities30%
  • Talking lip sync17.5%
  • Singing lip sync17.5%
  • Multi-character lip sync17.5%
  • Multilingual lip sync17.5%
AdsVideo
  • General capabilities30%
  • Text and logo rendering accuracy17.5%
  • Product and brand consistency17.5%
  • Product realism17.5%
  • Scene-product logic17.5%
FilmVideo
  • General capabilities30%
  • Camera and lighting control14%
  • Character realism14%
  • Multi-shot continuity14%
  • Acting and performance14%
  • Style understanding14%
AnimationVideo
  • General capabilities30%
  • Motion quality17.5%
  • Character consistency17.5%
  • Shot continuity17.5%
  • Style understanding17.5%
Motion DesignVideo
  • General capabilities30%
  • Design fidelity23.3%
  • Motion control23.3%
  • Spatial composition23.3%
Overall ImageImage
  • Prompt adherence25%
  • Subjective aesthetics25%
  • Reference adherence25%
  • Creativity / variation25%
Image EditingImage
  • General capabilities30%
  • Prompt adherence to the edit23.3%
  • Precise editing23.3%
  • Style adaptation23.3%
Product / E-commerceImage
  • General capabilities30%
  • Text and logo rendering accuracy23.3%
  • Product and brand consistency23.3%
  • Product realism23.3%
FilmImage
  • General capabilities30%
  • Scene setup and lighting23.3%
  • Film texture23.3%
  • Character realism23.3%
Graphic DesignImage
  • General capabilities30%
  • Style adherence17.5%
  • Text accuracy17.5%
  • Layout17.5%
  • Typography17.5%

What the scores do not cover

  • Product specificationsSpecifications are not scored

    Cost, resolution, duration, speed and licence terms are published with their source and verification date. They are not folded into any quality score.

  • Quality controlsRepeat comparisons and vote caps

    For judges on a pre-assigned plan, 2–3% of comparisons repeat an earlier pair with sides reversed to measure self-consistency; these repeats never enter the model fit. No judge may exceed a set share of their tier's votes.

  • Cost efficiencyReported separately

    Quality relative to standardised cost is planned as a per-board view. It is not combined with the quality scores.

  • Provider participationIdentical treatment for all models

    Models are evaluated on the same briefs as every competitor, with outputs generated by the OpenArt team. Providers cannot submit, select or regenerate outputs.