Every AI video tool publishes a per-second price. None of them tell you what a finished video costs, because the number that matters is not the rate — it is how many seconds you generate to keep the ones you use, and which engine each of those seconds ran through.

This works it out from our own production ledger. The rates below are what we were charged making one episode between June and August 2026, not figures from a pricing page.

1 · How long is the finished video?


minutes

2 · How much of it is someone talking on camera?


35%

3 · How long is each generated clip?


8s (generate 2.00× over)

This is the slider that moves the bill. A fault anywhere means regenerating the whole clip, so longer clips throw away more each time.

4 · Does the same character appear across shots?

Enter a runtime to see the numbers.

Why an eight-second clip costs more than two four-second clips

It should not. The rate per second does not change. But it does, and the reason is retakes.

A fault anywhere in a clip means regenerating the whole clip. One bad hand in second seven and all eight seconds go back. Do the same work as two four-second clips and a fault only costs you four.

Put a number on it. If faults arrive at a roughly steady rate as you generate, the chance a clip comes back clean falls the longer it runs — and the seconds you pay for per usable second climb with it:

Clip length Seconds generated per usable second Cost against 4s
4s 1.41×
6s 1.68× +19%
8s 2.00× +42%
10s 2.38× +69%
12s 2.83× +101%

That curve is calibrated against our own production, which ran at about 2× on eight-second clips. It is a model rather than a measurement at every length — but the direction is not in doubt, and the practical rule falls straight out of it: generate short and cut, rather than generating long and hoping.

Our faults were almost always the same two: a hand that had gone wrong, or a face that had drifted. Neither is visible while you are generating. Both are obvious in the edit, which is the expensive place to find them.

The rates, and where they come from

Job Engine Measured
Talking shots, full frame Seedance 2.5 with audio references ~$0.30/s
Talking shots, from a base take MiniMax base + videoretalk ~$0.18/s combined
Silent recurring character MiniMax H3 ~$0.08/s
No recurring face Seedance mini via BytePlus ~$0.03/s

One episode’s production, June–August 2026. Model pricing moves, so treat these as the shape of the problem rather than a quote — the ratio between the tiers has held far better than the absolute numbers.

What the result usually shows

A talking second costs about ten times a b-roll second. That single ratio decides almost everything else, and it has two consequences worth acting on.

Running the whole video through the talking-capable tier is the most expensive mistake available, and it buys nothing on the shots where nobody speaks. Splitting by shot type is not an optimisation, it is the difference between a video you can afford and one you cannot.

And it is cheaper to write a script where the character speaks in three places than one where they speak throughout — and the second version does not look better anyway.

What this does not cover

Voice synthesis, music, editing time and your own hours are all outside it. So are the platforms that bundle generation into a subscription rather than billing per second, which change the arithmetic completely — if you are generating constantly, a flat monthly fee can beat any per-second rate.

The retake curve assumes faults arrive at a steady rate as a clip runs. In practice some faults are front-loaded and some models degrade towards the end of a generation, so treat the shape as right and the exact figures as ours rather than yours.

The picker answers the other half: which engine to use for what you are making.