Which AI Video Model Should You Use? The 30-Second Picker

Two questions, one answer: which AI video engine to use and which platform to buy it through, priced from our own measured production spend.

There is no best AI video model. There is a best model for what you’re making — and the honest answer is usually two models, not one. Answer two questions; the routing below comes from our own measured production spend, not pricing pages.

1 · What are you making?




2 · How much of it talks?

 
 

Pick both and the routing appears here.

Why two questions and not a feature table

Every comparison of AI video models is a grid of resolutions, clip lengths and supported aspect ratios. We have read a lot of them, and none of them predicted a single problem we actually hit. The specs are not where these models differ. They differ in what they get wrong, and what they get wrong depends almost entirely on two things: whether a face has to talk, and whether that same face has to come back in the next shot.

Those two things also happen to be where the money goes. So the picker asks about them and ignores everything else.

Where the budget actually went

These figures are from our own production of a single episode between June and August 2026, not from vendor pricing pages. They are what we were charged, for the shots we kept.

Job What we used Measured cost
Talking shots, full frame Seedance 2.5 with audio references ~$0.30 per second
Talking shots, from a base take MiniMax base plus videoretalk ~$0.10 per second
Silent character shots MiniMax H3 ~$0.08 per second
B-roll with no recurring face Seedance mini via BytePlus ~$0.03 per second
Picture-in-picture reaction bubble OmniHuman-B not separately metered

Read the top and bottom rows together: a talking second costs about ten times what a b-roll second costs. That single ratio is the whole reason this page exists.

The number that matters is the ratio, not the rate. The talking seconds took roughly ninety percent of the money, and they were nowhere near ninety percent of the runtime. If you are budgeting a video with a presenter in it, budget the talking seconds separately and treat everything else as close to free by comparison.

The practical consequence is a structural one: it is far cheaper to write a script where the character talks in three places than one where the character talks throughout, and the second version does not look better.

The four routes, and why

Talking character, most shots talk. Seedance with audio references for the talking shots, MiniMax H3 for everything else. Two models, deliberately. Running the whole thing through the talking-capable tier is the single most expensive mistake available to you, and it buys nothing on the shots where nobody speaks.

Recurring character, silent. MiniMax H3 with a disciplined reference pack. The failure mode here is not what people expect — it is not the close-ups that break. It is the mid-distance zone, where the character is small in frame but not a true wide. That is where identity dissolves, and it dissolves quietly enough that you will not notice until you cut the shots together. Frame close, medium or genuinely wide, and stay out of the middle.

No recurring face at all. Take the cheapest tier that looks acceptable, because the expensive tiers are selling you face consistency you are not using. Seedance mini through BytePlus measured at roughly three cents a second for us, which is the cheapest acceptable output we found.

Composed camera moves. Kling, close or medium framing — and check the hands on every single take. Not most takes. Every one.

The defects worth knowing about before you spend anything

Hands. Still the first thing to fail and still the thing most likely to survive your own review, because you are watching the face. Check hands last, deliberately, on a second pass.

Identity drift at mid-distance. Described above. It is the defect that costs the most to fix, because you usually discover it at the edit rather than at generation, by which point the takes either side are already approved.

Face refusal. BytePlus refused to generate our host’s face at all. That cost nothing in money and a day in schedule, which is the more expensive currency. If your project depends on a specific recurring face, test that face on the platform before you plan around it — not a similar face, that face.

Version churn. Model behaviour changes with every release, and the verdicts above are pinned to June–August 2026. What has held up better than the version numbers is the defect taxonomy: hands, mid-distance identity, and refusals were the three failure modes throughout, across every model we tried.

What we have not tested

We have not run Veo, Runway or Pika through the same production, so they are absent here rather than rejected. We have not tested any of these at feature length, on live-action plates, or with more than one recurring character in frame. And every figure above comes from one episode’s worth of production — it is a real ledger, but it is one ledger.

Where a claim on this page is not something we paid for and watched, we have not made it.

Questions people actually ask

Which AI video generator is best?

Wrong question, and the reason every answer to it disagrees. Ask instead whether your video has a talking face in it, and whether that face recurs. Those two answers pick the model. If nothing talks and no face repeats, you are shopping on price alone and should take the cheapest output you can stand.

What is the cheapest AI video generator that is still usable?

The cheapest acceptable output we measured was Seedance mini through BytePlus at roughly three cents a second — against thirty cents a second for the talking tier. “Acceptable” is doing real work in that sentence: it is fine for b-roll and atmosphere, and it is not fine for anything where a specific face has to come back.

Do I need two models?

If your video has both talking shots and non-talking shots, yes, and it will cost you less rather than more. The talking-capable tier is priced for a job most of your shots are not doing.

Why does my character look different between shots?

Most likely the mid-distance problem. Look at the shots either side of the ones that bother you and check how large the character sits in frame. If they are small but not wide, reframe rather than regenerate — regenerating the same framing usually reproduces the same drift.

How current is this?

June–August 2026 production, one episode. Model versions have moved since; the failure modes have not.

This routing comes out of making something, not out of testing tools for the sake of it. The episode it came from is in production now — the footage behind these verdicts will be published with it.