Every AI video tool publishes a per-second price. None of them tell you what a finished video costs, because the number that matters is not the rate — it is how many seconds you generate to keep the ones you use, and which engine each of those seconds ran through.

This works it out from our own production ledger. The rates below are what we were charged making one episode between June and August 2026, not figures from a pricing page.

1 · How long is the finished video?


minutes

2 · How much of it is someone talking on camera?


35%

3 · How many takes per usable shot?


Ours ran about 2×. One take in three needed doing again, usually for a hand or a drifting face.

4 · Does the same character appear across shots?

Enter a runtime to see the numbers.

Why the retake slider is the important one

Every other calculator multiplies runtime by a rate. That is the one number guaranteed to be wrong, because you do not generate a finished video — you generate two or three times the footage and keep the parts that worked.

Our own ratio ran at about two. Roughly one take in three needed doing again, and it was almost always the same two faults: a hand that had gone wrong, or a face that had drifted between shots. Neither is visible while you are generating. Both are obvious in the edit, which is the expensive place to find them.

Set that slider to 1 and you get the number the vendor would quote you. Set it to 2 and you get the number you will actually pay.

The rates, and where they come from

Job Engine Measured
Talking shots, full frame Seedance 2.5 with audio references ~$0.30/s
Talking shots, from a base take MiniMax base + videoretalk ~$0.18/s combined
Silent recurring character MiniMax H3 ~$0.08/s
No recurring face Seedance mini via BytePlus ~$0.03/s

One episode’s production, June–August 2026. Model pricing moves, so treat these as the shape of the problem rather than a quote — the ratio between the tiers has held far better than the absolute numbers.

What the result usually shows

A talking second costs about ten times a b-roll second. That single ratio decides almost everything else, and it has two consequences worth acting on.

Running the whole video through the talking-capable tier is the most expensive mistake available, and it buys nothing on the shots where nobody speaks. Splitting by shot type is not an optimisation, it is the difference between a video you can afford and one you cannot.

And it is cheaper to write a script where the character speaks in three places than one where they speak throughout — and the second version does not look better anyway.

What this does not cover

Voice synthesis, music, editing time and your own hours are all outside it. So are the platforms that bundle generation into a subscription rather than billing per second, which change the arithmetic completely — if you are generating constantly, a flat monthly fee can beat any per-second rate.

It also assumes you keep what you generate at the ratio you set. If your standards are higher than ours, raise the slider.

The picker answers the other half: which engine to use for what you are making.