We spent $9 testing Kling, Seedance and MiniMax head to head. The cheap one won.

A real head-to-head under the hardest conditions in AI video: one recurring character, synced dialogue, continuity across a 14-minute episode. The cheapest model won, and here is the defect list nobody publishes.

Most model comparisons are one prompt, one clip, side by side, and a verdict. That tells you which model makes the prettiest five seconds. It tells you almost nothing about which one survives a project.

We ran ours differently: a bakeoff on identity stress, crowd occlusion and shot-to-shot continuity, then roughly forty-four clips of actual production. The test cost about $9. The production that followed cost considerably more, and taught us more.

A caveat that changes how you should read this

We were building the hardest thing in this format: one recurring photoreal character, holding her identity across hundreds of generations, with her mouth synced to scripted dialogue, cut into a continuous fourteen-minute episode.

Almost nobody needs that. If you are making b-roll, product shots, atmosphere, or anything where no single face has to be the same face in shot 40 as in shot 3, most of the failures below will never happen to you — and the cheap models get much more attractive.

Read the verdicts as worst-case behaviour under maximum stress, not as what a normal project looks like.

The verdicts

Model Cost per 5s Verdict
MiniMax H3
reference-to-video, 768p
~$0.38
25.84 credits
Won as primary engine. Cheapest of the three and the most real — least pretty, most human. Best scene stability. Faces wobbled in round one; walk gaits looked wrong until we prompted the gait explicitly.
Kling V3 / O3
image-to-video
~$0.40
27 credits
Specialist, not generalist. Best motion realism of anything tested, and the only model with a first-class reusable persona object. Close and medium framing only. Malformed hands, a background extra vanished mid-clip, identity dissolves when the character is small in frame.
Seedance 2.5
reference-to-video + audio
~$1.48
100.45 credits
Benched, then resurrected. The most alive performer and also the most melodramatic, at roughly 4Ă— the price. Too expensive for general shots. Then it solved lip-sync and became the talking-head engine at about $3 per fifteen-second shot.

Two models were rejected outright. A lip-sync post-processor produced a pasted-mouth look and padded short clips by playing footage backwards. Avatar generators of the HeyGen class sync perfectly but put a person in front of a backdrop — they cannot have the character touch the scene, which was the entire point.

Why the cheapest model won

This is the finding that surprised us, and it generalises.

Judged on a single clip, the expensive model wins. It is more expressive, more dynamic, more cinematic. Judged across forty clips that have to cut together, that same quality reads as melodrama — every take is a performance, and performances do not match each other.

The cheap model’s flat, slightly plain output is boring in isolation and consistent in aggregate. Consistency is what an edit needs. A run of unremarkable shots that match will cut into something watchable; a run of spectacular shots that do not match will not.

Motion realism is Kling’s advantage. Scene stability is MiniMax’s advantage. Neither model sweeps, and the one that looks best in a demo is rarely the one you should build on.

Nine ways they fail, and why you will not spot them

A malformed hand in a generated video clip, fingers fused and miscounted
Defect one, from our own footage. In a review strip this hand is roughly thirty
pixels across and looks fine. It is only wrong when you zoom, and only obviously wrong when it
moves.

Here is the part with real reuse value. Every one of these was caught by a human watching motion — none of them by reviewing frame strips. If your QC is looking at stills, you are checking the wrong thing.

  1. Malformed hands. Hands occupy about thirty pixels in a review strip. Invisible. Needs a dedicated hand-zoom pass.
  2. Background people vanishing mid-clip. A one-frame-per-second strip missed an extra disappearing entirely. Review every sixth frame or watch it properly.
  3. Unnatural walk gaits. Completely invisible in stills — it is a motion defect by definition. Fixed by prompting the gait explicitly, not by rerolling.
  4. Identity dissolution at distance. When the character gets small in frame, the face re-synthesises generic. Our rule became: close, medium, or true-wide. The mid-distance zone is banned.
  5. Silent dialogue drops. One take performed four words of a forty-word line and returned success. Every talking take now gets transcript-checked automatically; under 80% match triggers a retake.
  6. Crowd-density discontinuity between consecutive shots — a busy street becomes a quiet one across a cut.
  7. Dead interaction beats. Physical contact happens and nobody reacts to it.
  8. The camera paradox. Selfie-style shots render the character holding nothing, and phones appear in mirrors that should not be there.
  9. Freeze-feeling endings. Clip duration has to equal the measured audio plus about half a second, or the last frame hangs.

Three laws that held across every model

No model holds 3D room geometry between separate generations. Generate the same room twice and it is a different room. The fix is not a better prompt — it is shared boundary frames: the last frame of shot N becomes the first frame of shot N+1.

Identity references must be a neutral character sheet, never a scene image. One scene image used as a reference made every subsequent shot re-open on that scene’s bed. Generate everything from a single anchor image; never chain edits, because drift compounds.

Do not prompt for “cinematic”. These models are already biased toward a too-clean, too-composed look, and asking for cinematic makes it worse. Prompt the defects instead — focus hunting, exposure shifts, slight instability. A layered wobble added for free in post does more for realism than any prompt adjective.

What to actually use, by what you are making

No recurring character: almost none of this applies. Use the cheapest tier that looks acceptable and spend the savings on more takes.

A recurring character, no dialogue: a stable reference-to-video model like MiniMax, with a disciplined anchor image and reference pack. Ban mid-distance framing and you have removed the main failure mode.

A recurring character who talks: route by shot type rather than picking one model. Talking shots go to the expensive engine that does generation-native lip-sync; everything else goes to a cheap tier. Paying performance rates for b-roll is where budgets die.

Composed camera moves: Kling, in close or medium, with a reusable persona element. Not for wide shots, and check the hands.

Method

Tested June–August 2026 through an API aggregator, prices in that window. Bakeoff spend approximately $9 across three rounds; production spend across the wider project roughly $180. Credit-to-dollar conversion is an estimate rather than a published rate. Model behaviour changes with every version release — the defect taxonomy has outlasted the version numbers so far, but the verdicts have not.

Part of our AI video tools comparison — three categories, priced, and which one you are actually shopping in.