One prompt, six video models, nine measured checks
Four models passed all eighteen checks. Every clip is below, with the rubric, the tools and the raw numbers.
- 12 min read
- Updated
The brief
The checks
Scores
All 12 clips
The misses
Beyond the checksWhat you make, in order. Tap one to jump to its step.
Most model comparisons show one clip per model and a verdict. We wanted a test you can check yourself: one prompt that names nine details a tool can confirm, two runs on each model, every clip shown, and the numbers behind every score.
What we sent
One prompt, sent unchanged to six text to video models on the same morning, from the same account. Every clip is 16:9 and 6 seconds long, on each model's 720p tier (768p on Hailuo 3, whose smallest tier that is). Sound was left at each model's default. Each model ran twice.
A quiet bakery storefront on a cobblestone street in the early morning. Above the door, a dark green wooden sign reads BAKERY in large white capital letters. A red bicycle leans against the wall to the left of the door. Exactly three terracotta potted plants stand in a row to the right of the door. A small round cafe table with an open yellow umbrella stands on the pavement at the right of the frame. A ginger cat walks steadily along the pavement in front of the shop, from the left edge of the frame to the right edge. The camera slowly pushes in toward the door for the whole shot. Soft low sunlight from the left. No people anywhere in the shot. Photorealistic.
The models: Seedance 2.5, Kling v3, Veo 3.1, Wan 3.0, Grok Imagine Video 1.5 and Hailuo 3. There are no people in the scene on purpose, so every check could be measured by a tool rather than judged.
How we scored it
Nine checks, pass or fail, so a run scores out of 9 and a model out of 18. Eight are measured by tools. One, the count of pots, is counted by eye, because the detector sometimes split one pot into two or counted a plant on a table.
| Check | The prompt said | Measured with | Passes when |
|---|---|---|---|
| Cat | a ginger cat walks along the pavement | object detector | a cat is found in at least 6 of 24 frames |
| Cat direction | from the left edge of the frame to the right edge | the cat's centre, frame by frame | it moves right by at least a quarter of the frame width |
| Red bicycle | a red bicycle leans against the wall | object detector and colour | a bicycle is found early in the clip and its most saturated pixels are red |
| Yellow umbrella | an open yellow umbrella | object detector and colour | an umbrella is found early in the clip and its canopy is yellow |
| Sign | a sign reads BAKERY | text recognition | BAKERY, spelled exactly, in 3 of the first 4 frames read |
| Three pots | exactly three terracotta potted plants in a row | counted by eye, first frame | the row beside the door holds exactly three |
| Layout | bicycle left of the door, pots right of it | detector plus the sign's position | the bicycle is left of the sign's centre and the pots are right of it |
| Camera | the camera slowly pushes in toward the door | feature tracking | the picture grows by at least 5% between 10% and 90% of the clip |
| No people | no people anywhere in the shot | object detector | a person is found in at most 1 of 24 frames |
The bicycle, umbrella, sign and layout are checked in the opening third of the clip. The prompt asks the camera to push in, which crops the edges of the scene later, so an umbrella that leaves the frame at second four has not failed.
Tip:Every probe was made to fail once. Played backwards, all twelve clips fail the direction check and the camera check. The text reader returns BAKFRY for a sign that says BAKFRY. We looked at every person the detector found: in both Kling v3 runs it is a real figure behind the door glass, and the single hits in the two Wan 3.0 runs were the cat's back and a table's iron foot, which is why one stray frame is allowed.
The scores
| Model | Run 1 | Run 2 | Out of 18 | What it missed |
|---|---|---|---|---|
| Veo 3.1 | 9 | 9 | 18 | nothing |
| Hailuo 3 | 9 | 9 | 18 | nothing |
| Wan 3.0 | 9 | 9 | 18 | nothing |
| Seedance 2.5 | 9 | 9 | 18 | nothing |
| Grok Imagine Video 1.5 | 8 | 9 | 17 | run 1: the camera slid sideways instead of pushing in |
| Kling v3 | 8 | 7 | 15 | both runs: a person inside the shop. Run 2: bicycle and pots on the wrong sides of the door |
The easy parts were easy for everyone. All twelve clips spelled BAKERY right while the sign was in frame, painted a red bicycle and a yellow umbrella, put a row of three pots on the ground, and walked a ginger cat from left to right. The misses were a person nobody asked for, a scene laid out the wrong way round, and a camera that moved the wrong way.
Every clip
Press play on any clip. Nothing loads until you do. First runs:






Second runs, same prompt, same settings:






What the misses looked like



What the checks do not catch
A rubric only sees what it measures. These came up while we checked the misses, and they matter more to some jobs than a pass or fail does.
- How slow is slow. The prompt asked for a slow push in. Measured from 10% to 90% of the clip, Veo 3.1's picture grew 1.8 and 2.6 times; the others grew between 1.10 and 1.45 times, apart from Grok's sideways first run. By our judgement, Veo's move is not slow, and in its second run the sign leaves the top of the frame.
- Lettering nobody asked for. Three of twelve clips added text that is not real words: window lettering in Seedance 2.5's first run, which the text reader read differently in every frame, and menu boards in Veo 3.1's and Wan 3.0's second runs. If your frame has to be clean of text, say so in the prompt and check.
- Open means open. Grok's first run shows a yellow umbrella that stays closed for the whole clip. Our detector finds umbrellas open or closed, so this passed the colour check. We only caught it by eye.
- Where the cat goes. In Veo 3.1's second run the cat turns toward the camera and fills the frame. It still ends up to the right, so it passed.

What you get back
Measured with ffprobe and ffmpeg on the files as delivered. Both runs of each model matched in size, frame rate and length.
| Model | Frame | Length | Video bitrate | Sound, mean loudness |
|---|---|---|---|---|
| Veo 3.1 | 1280 by 720, 24 fps | 6.0 s | 9.0 and 12.4 Mb/s | about minus 36 dB |
| Hailuo 3 | 1344 by 768, 24 fps | 6.6 s | 2.5 and 2.6 Mb/s | minus 18 and minus 30 dB |
| Wan 3.0 | 1280 by 720, 30 fps | 6.0 s | 16.6 and 15.7 Mb/s | minus 21 and minus 17 dB |
| Seedance 2.5 | 1280 by 720, 24 fps | 6.1 s | 10.4 and 8.3 Mb/s | about minus 36 dB |
| Grok Imagine Video 1.5 | 1280 by 720, 24 fps | 6.0 s | 8.1 and 11.4 Mb/s | about minus 61 dB, close to silent |
| Kling v3 | 1280 by 720, 24 fps | 6.0 s | 8.5 and 8.4 Mb/s | about minus 55 dB, close to silent |
Two things stand out. Hailuo 3 on its 768p tier returns a clip longer than asked for, at roughly a quarter of the bitrate most of the others use. Grok and Kling returned tracks so quiet they are close to silent; the prompt did not ask for any sound, so that is a fact to know rather than a fault.
How long each took
| Model | Run 1 | Run 2 |
|---|---|---|
| Grok Imagine Video 1.5 | 59 s | 49 s |
| Veo 3.1 | 84 s | 113 s |
| Hailuo 3 | 122 s | 125 s |
| Wan 3.0 | 164 s | 202 s |
| Seedance 2.5 | 179 s | 313 s |
| Kling v3 | 188 s | 183 s |
What it costs
Veo 3.1Pro and up
Per 5s clip
720p 112 credits · 1080p 112 credits
Hailuo 3Pro and up
Per 5s clip
768p 37 credits · 2K 65 credits
Wan 3.0Pro and up
Per 5s clip
480p 14 credits · 720p 28 credits · 1080p 56 credits
Seedance 2.5Pro and up
Per 5s clip
480p 26 credits · 720p 56 credits · 1080p 137 credits
1080p renders at 16:9 and 9:16
Grok Imagine Video 1.5Pro and up
Per 5s clip
480p 23 credits · 720p 41 credits · 1080p 71 credits
Kling v3Pro and up
Per 5s clip (audio on by default)
720p 35 credits · 1080p 47 credits · 4K 117 credits
Read live from the catalog when this page loads. Every price is shown before you generate. See every rate.
The rates above are read live from the price list, tier by tier, for a 5 second clip. Our clips ran 6 seconds on the 720p tier (768p on Hailuo 3), so each cost a fifth more than its line here. The price is always shown before you confirm a render.
What this test does not tell you
- It is one prompt and one kind of scene: a still street, one animal, one slow camera move. Hands, fast action, dialogue, several people or long clips can rank the models differently.
- It has no people in it, so it says nothing about how any model draws a person.
- Two runs per model is a small sample. A model that passed both could miss on a third.
- It ran each model at its 720p tier. Higher tiers can change detail, and on some models, the result.
- The detector, the colour tests and the text reader can be wrong. We looked at every miss ourselves and publish the raw numbers below.
- It does not score taste. Light, texture, how natural the cat walks and how the sound fits are for your own eyes and ears.
For movement, we ran a separate test: one standing backflip from the same frame on four models, in Make AI people move realistically. For a car that has to keep its shape, the drive-by test is in A car commercial with AI.
Method and raw numbers
Objects and people: Mask R-CNN trained on COCO, confidence 0.6 (0.7 for people), on 24 frames spread over the clip. Text: the macOS Vision text recognizer with language correction off, on 12 frames. Camera: SIFT features matched between 9 frames from 10% to 90% of the clip, with a RANSAC similarity fit, scales multiplied. Colour: for the bicycle, the most saturated fifth of the pixels inside its mask; for the umbrella, the median hue of the top 40% of its box inside its mask.
Two colour tests changed while we worked, and we say so because they changed results. The first bicycle test counted every saturated pixel in the mask, and the warm wall showing between the tubes outvoted the paint. The first umbrella test failed Seedance 2.5's mustard umbrella on the brick wall behind its pole. Both were fixed to measure the object, then every clip was scored again with the final code. The counts by eye matched the detector on 10 of 12 clips; on both Veo 3.1 clips it also counted a plant on a table or merged two pots.
| Run | Cat: frames, travel | Camera growth | Bicycle red share, umbrella hue | People frames |
|---|---|---|---|---|
| Veo 3.1, 1 | 23, 0.38 | 1.76 | 0.97, 23 | 0 |
| Veo 3.1, 2 | 16, 0.59 | 2.65 | 1.00, 20 | 0 |
| Hailuo 3, 1 | 13, 0.67 | 1.16 | 1.00, 20 | 0 |
| Hailuo 3, 2 | 11, 0.36 | 1.24 | 0.95, 20 | 0 |
| Wan 3.0, 1 | 20, 0.71 | 1.45 | 0.99, 17 | 1 |
| Wan 3.0, 2 | 24, 0.49 | 1.44 | 1.00, 15 | 1 |
| Seedance 2.5, 1 | 12, 0.59 | 1.22 | 0.89, 17 | 0 |
| Seedance 2.5, 2 | 20, 0.60 | 1.10 | 1.00, 18 | 0 |
| Grok Imagine Video 1.5, 1 | 24, 0.37 | 1.02 | 0.86, 18 | 0 |
| Grok Imagine Video 1.5, 2 | 24, 0.56 | 1.21 | 0.87, 18 | 0 |
| Kling v3, 1 | 13, 0.30 | 1.37 | 0.67, 24 | 24 |
| Kling v3, 2 | 18, 0.51 | 1.34 | 0.92, 21 | 24 |
Questions and answers
- Which model should I use?
- For a scene like this one, any of the four that passed all eighteen checks followed the brief. From there, pick on look, speed and price, which is why every clip and the live rates are on this page. If your brief says no people, check Kling v3's output closely.
- Why is there no person in the prompt?
- So that every check could be measured by a tool. A prompt with people in it is a different test, and we would rather publish that one when we can measure it properly.
- Can I run this test myself?
- Yes. The prompt above is exactly what we sent. Open any of the model pages, paste it, and pick 16:9, 6 seconds and the 720p tier.
- Why check the bicycle, umbrella and sign only early in the clip?
- The prompt asks the camera to push in toward the door, which crops the edges of the scene later on. Checking the opening third tests what the model drew, not where its camera ended up.
Run the same prompt on any model
Browse the video modelsRead next
Comparison
The best AI video model for each job, from our own tests
Which AI video model to use for a spoken line, movement, a long take, 4K, editing your own clip or a fast draft, picked from tests we ran and published.
· 9 min read
Comparison
Seedance 2.5 vs Kling 3.0: four tests on the same inputs
Seedance 2.5 and Kling v3 on the same prompts and frames: a scored street scene, a car drive-by, a backflip and a filmed look. What each one got right.
· 9 min read
Comparison
Seedance 2.5 at 1080p or Seedance 2.0 Pro at 4K: what the files contain
We rendered one prompt at 1080p and 4K and measured the files: Seedance 2.0 Pro and Kling v3 return real 4K detail. What else changes, and which to pick.
· 7 min read




