Prompt adherence
“Which output follows the prompt more accurately?”
Referring to the scene in image 1, he remains still as the camera holds on the scene with a subtle natural handheld drift and zoom in slowly.


Creative professionals compare two outputs at a time without knowing which model produced either, voting on one named criterion at a time. The criteria, weights, judge counts and quality controls behind each ranking are documented below.
Boards cover capability tracks and industry categories. Criteria are defined by people working in that field, then approved and frozen before scoring begins.
Prompts come from real creation use cases, in a mix of professional and amateur styles. OpenArt runs every model on the same brief; one output per model enters evaluation by a fixed rule, never hand-picked.
Model identities are hidden and left/right position randomized. Each comparison is a forced choice on one named criterion - no tie, no overall-preference question.
Votes are fit per criterion, combined into board scores with published weights, and reported with uncertainty. The scoring rules are set out below.
Criterion scores use a Bradley-Terry estimator, so beating a strong model moves a score more than beating a weak one. One anchor model per board pins the scale, and scores are reported on an Elo-like scale relative to that anchor.
General criteria are fit on the modality master from pooled industry comparisons. Industry and capability boards reuse those general fits, including scores, vote and judge counts, and bootstrap intervals, with ranks renumbered within their own model rosters.
Industry and capability headline scores combine general capabilities with board-specific criteria using the weights recorded in each published snapshot. Overall boards use general criteria only. Snapshot weights apply identically to every eligible model; configuration changes do not rewrite earlier results.
Votes are weighted by judge tier: the Creative Expert Council at 3× and Tastemakers at 1×. Weights are published before scoring begins and frozen after publication; any later change is a versioned ruleset update.
A model receives a board composite only after clearing every required criterion on that board. Criteria that have cleared the vote minimum are shown while the rest collect votes.
Every score is published with a 95% confidence interval from judge-cluster bootstrap resamples, a rank, and the distinct-judge count behind it.
Before release each board is re-run with the expert multiplier moved up and down, without its highest-volume judge, and on an alternate bootstrap seed. Results are published either way; an order change holds the board for review.
Each scoring run creates an immutable snapshot of scores, weights, configuration and code version. Scores are comparable within a board only.
Scores sit on an Elo-like scale pinned to one anchor model per board. A score says where a model sits relative to the others on that board, nothing more.
Every score carries a 95% confidence interval and a rank. If two intervals overlap, the models are not distinguishable on that criterion, whatever the ranks say.
Each board's composite reflects its own criteria and weights, even when general fits are shared. A 1,240 on Film and a 1,240 on Ads are not comparable headline scores.
5 criteria · voted across video types and use cases
“Which output follows the prompt more accurately?”
Referring to the scene in image 1, he remains still as the camera holds on the scene with a subtle natural handheld drift and zoom in slowly.


“Which output is more aesthetically pleasing?”
Surreal fantasy landscape featuring floating bioluminescent islands in a twilight sky. Glowing magical flowers swaying in a light breeze, purple and emerald volumetric light beams slicing through soft clouds. Vertigo dolly zoom camera effect, ethereal atmosphere, fairytale aesthetics, photorealistic render.


“Which output shows more convincing physics and motion?”
The girl in the reference picture looks at the shore in surprise [camera zooms out], then jumps into the pool without hesitation and begins to swim vigorously to the other side.


“Which output maintains stronger visual consistency?”
Using the reference image, slowly transform the woman in the first image into the woman in the second image.


“Which output has better audio quality?”
Visually explosive superpower action sequence, dynamic volumetric lighting, VFX masterclass. [0:00-0:03] A glowing energy-infused character floats an inch off the ground, hair and coat levitating in reverse gravity as bright blue plasma arcs around their fists. [0:03-0:06] The character thrusts both palms forward, releasing a massive kinetic plasma shockwave that blows back debris, car windows, and dust in a violent sphere explosion. Dynamic camera shake, blinding light flashes, smooth motion physics


4 criteria · voted across image types and use cases
“Which output follows the prompt more accurately?”
A portrait of a stylish Asian teenage girl standing in a clean studio. She’s wearing a tied-up crop top, pigtails, an eye patch over her left eye, and asymmetrical jeans with design on her right leg, confident pose and facial expression, modern fashion photography, neutral background, soft lighting, highly detailed.


“Which output is more aesthetically pleasing?”
Photorealistic street-style photo of the same woman, wearing the exact same outfit as in the reference photo, taking a candid selfie-style shot of herself on a busy city street in broad daylight. Bright natural sunlight casting sharp shadows, urban backdrop with pedestrians and traffic blurred in motion, confident badass stance, phone held up as if capturing the shot herself, sharp direct gaze into the lens. Shot like a high-fashion street-style editorial: crisp daylight photography, natural contrast, ultra-realistic skin texture, candid energetic mood, 35mm film look, shallow depth of field. Manhattan skyline in the background


“Which output adheres to the reference more consistently?”
Keep the same woman from the reference image and render her in three consistent looks - casual daywear, elegant evening and a sporty outfit - each in a different setting but preserving her look - exact face, hair and identity throughout. Focus on identity consistency across all looks. Realistic portrait photography, clean lighting.


“Which output offers the more creative interpretation?”
A single frame of 2D hand-drawn theatrical anime, sakuga action style. A young woman warrior mid-battle in a mist-filled mountain gorge: crisp modern Shonen cel-shaded character — bold clean line art, sharp two-tone cel shadows — against an ink-wash watercolor background with visible paper grain and bleeding ink edges. She is a spear-wielder in dark indigo-and-black layered Chinese Wuxia robes with cyan accents, executing a sweeping spear strike; arcs of living black ink and water trail the spearhead like calligraphy brush strokes, flicked ink droplets suspended around her. Dynamic action pose, full body visible, low angle, strong rim light carving her silhouette out of the mist, hard shadow shapes. Muted ink-black and paper-white world, her cyan accent color burning against it. No 3D CGI or 3D particle simulation, nor photorealism. No text.


Cost, resolution, duration, speed and licence terms are published with their source and verification date. They are not folded into any quality score.
For judges on a pre-assigned plan, 2–3% of comparisons repeat an earlier pair with sides reversed to measure self-consistency; these repeats never enter the model fit. No judge may exceed a set share of their tier's votes.
Quality relative to standardised cost is planned as a per-board view. It is not combined with the quality scores.
Models are evaluated on the same briefs as every competitor, with outputs generated by the OpenArt team. Providers cannot submit, select or regenerate outputs.