IronStratum

Whisper local vs API: the break-even math

Running Whisper yourself versus paying per audio hour is a volume question, and the answer is arithmetic. This guide gives the crossover formula in audio hours. It works the formula across two real GPU tiers at September 2026 street prices and the official United States electricity price. It reads the hosted rate ladder the way an invoice reads it. Two of the prices in that ladder turned out to be wrong in 2026, in ways that moved the answer by up to a hundredfold.

All amounts below are US dollar amounts written as bare numerals. Electricity is quoted in cents per kilowatt-hour. This guide types no competitor rate anywhere. Hosted prices enter as relations (how many times one tier sits above another) with their read dates. Rates drift, and the drift stories below are half the lesson. Rates for this platform's own audio lanes are never typed by hand. They render live from the catalog in the table further down, and the formula reads its hosted-rate variable straight off that table.

Three frames, and everyone selling you one

Every page that ranks for this question sells exactly one frame and hides the other two. The frames:

  1. Buy a card. The build-log frame: a used GPU, open weights, per-audio-hour spend at zero, done. What it hides: the card's price, the power draw, and the fact that you now run a serving stack.
  2. Rent a cloud GPU. The rental frame: a monthly instance as the "self-host floor". What it hides: that a rental is a subscription wearing ownership's clothes, and that at its own crossover volume you could have bought the card several times over.
  3. Pay per audio hour. The API frame: no hardware, a rate card, a bill that scales. What it hides: nothing about its own price, which is why its marketing leans on the other two frames' hidden costs.

The honest version puts all three in one formula and lets your volume pick. That is the rest of this guide.

What local Whisper actually costs

The model first, because the hardware bar is lower than most pages admit. The variant worth self-hosting for batch work is the turbo checkpoint of Whisper large v3. It is an approximately 809 million parameter model. Its weights fit in about 1.6 GB at fp16, released under MIT by the publisher and downloadable by anyone. Eight GB of VRAM runs it with headroom. That means nearly every modern card is turbo-capable. The interesting question is price per capability, not fit.

Hardware, at September 2026 street reads. A used RTX 3060 12 GB, the value floor, traded around 279 in a second-hand market guide this month. It launched at 329. A used RTX 3090 24 GB sold at listing averages near 1,000. Asking prices ran 1,287 to 1,411, and sold prices rose 11.3 percent over ninety days. It launched at 1,499 in September 2020. Use whatever your own listings show. The formula does not care which source you trust.

Power, from the manufacturer. NVIDIA's RTX 3060 specification page lists 170 W board power. The RTX 3090 page lists 350 W. Speech transcription is compute-bound in a way speech synthesis is not. So the worked tiers below use the full 350 W draw for the 3090-class box, not a partial-load discount.

Electricity, from the government. The EIA Short-Term Energy Outlook sets the 2026 United States residential average. The number is 18.2 cents per kilowatt-hour. A 350 W card at full duty draws 0.35 kWh per wall-clock hour, about 6.4 cents per hour at the national price. Your bill is the number that matters. 18.2 is the anchor.

Then the piece every guide guesses at: speed, as real-time factor. RTF is the ratio of processing time to audio length. An RTF of 0.10 means one hour of audio costs 0.10 hours of GPU time. Community testing of the turbo checkpoint on a 3090-class card under the common faster-whisper stack reports RTF in the range of roughly 0.05 to 0.15 for batch work. No in-house measurement exists yet. So this guide runs all three as explicitly assumed tiers, 0.05, 0.10, and 0.15, and labels them hypothetical until a measured pin replaces them. At those tiers the marginal electricity cost of local transcription is 0.003 to 0.010 per audio-hour. That is three-tenths of a cent to a full cent of power for every hour of speech, on a card you already own. Local marginal cost is close to nothing. The decision is almost entirely hardware amortization, exactly as with speech synthesis. The formula below is mostly a hardware formula.

The break-even formula

Rates for this platform's audio lanes render below from the live catalog. The formula reads its hosted-rate variable, P, straight off that table. Any provider's published per-audio-hour rate substitutes the same way.

ModelContext$/1M in$/1M out$/1M cached
Qwen
qwen3.8-27b262K$0.35$2.55$0.105
qwen3.6-35b131K$0.11$0.8$0.044
Minimax
minimax-m2.7197K$0.24$0.95$0.072
Muse
muse-glimmer-30b131K$0.28$1.2$0.084
Ornith
ornith-1.5-35b100K$0.35$2.55$0.105
ornith-1.5-9b100K$0.1$0.3$0.03
Glm
glm-5.3-flash1049K$0.11$0.35$0.033
Deepseek
deepseek-v4-flash-07311311K$0.15$0.42$0.045
Gemma
gemma-4-31b-it262K$0.22$0.49$0.066
Deepseek
deepseek-v4-pro1000K$1.13$2.21$0.339
deepseek-v4.1-flash1000K$0.19$0.71$0.06
Glm
glm-5.31000K$1.32$4.25$0.4
Flux
flux-1-schnell—$0.003/image
flux-pro-1.1—$0.04/image
Chatterbox
chatterbox-tts—$25/1M chars
Kokoro
kokoro-tts—$15/1M chars
Pocket
pocket-tts—$16/1M chars
Audio
audio8-tts—$8/1M chars
Hayamimi
hayamimi-stt—$0.6/audio-hr
Bge
bge-m3—$0.05/1M tokens
bge-reranker-v2-m3—$1.5/1k searches
Whisper
whisper—$0.12/audio-hr
Iab
iab-2x—$0.7/1k pages
iab-3x—$0.7/1k pages
Gliner
gliner-extract—$0.9/1k searches
Jev
jev-fast1K$0.03$0—
jev-causal1K$0.03$0—
jev-batch1K$0.03$0—
Wemm
wemm-embed—$0.09/1M tokens
hosted cost per day  =  D x P

local cost per day   =  G / (30 x A)  +  D x R x E x C / 100

crossover D*         =  G / (30 x A)  /  ( P  -  R x E x C / 100 )

The named variables:

  • D: audio hours you transcribe per day. The crossover D* is the volume where both sides cost the same.
  • P: hosted rate per audio hour, read off the fence above or off any provider's pricing page.
  • G: what the hardware cost you. 279 for the used 3060 tier, 1,000 for the used 3090 tier, zero if the box already exists.
  • A: amortization months; twenty-four is this guide's default.
  • R: real-time factor, the assumed speed tier: 0.05, 0.10, or 0.15, hypothetical until measured.
  • E: wall draw in kilowatts while transcribing: 0.35 for the 3090-class box.
  • C: electricity price in cents per kilowatt-hour: 18.2 is the 2026 national average.

Two worked shapes, stated in audio hours at the 0.10 tier. The used 3060 tier pays for itself in about 2,500 to 2,900 audio hours. That is against a hosted rate in this platform's recorded planning class for its removed Whisper lane. The class is about a tenth of a dollar per audio hour, a dated 2026-09-24 planning figure for the lane's return, not a live rate. Live rates render from the catalog. The used 3090 tier needs roughly 9,000 to 11,000 hours at the same hosted rate. Against a hosted rate three times higher, both horizons shrink by about three times. The card-rate floor of the hosted market is the harder case: entire hosted tiers there sell for about a hundredth of a dollar per audio hour. The card's own marginal electricity of 0.003 to 0.010 sits at or barely under the hosted price itself. The honest statement is that no owned card ever pays back on electricity alone at that floor. At prices that low, only volume you already own hardware for justifies the local road.

One sensitivity rule falls out of the shape. The denominator is the gap between the hosted rate and your marginal electricity. At premium hosted rates the electricity term is a rounding error. The crossover is amortization divided by price. At floor-class hosted rates the electricity term eats most of the gap. The crossover explodes by one to two orders of magnitude, and the naive build-log math quietly stops working. Cheap hosted rates are exactly where local math forgets electricity.

The rate ladder, read by the meter

The hosted batch-transcription market in September 2026 spans roughly an order of magnitude from its floor to the incumbent's list rate. The relations are stable even where the absolute numbers drift. Each is anchored on a dated read from the host's own page or a graded market record this month. A floor class of budget card-rate listings sits at the bottom. One speed-focused host sits at roughly three times the floor. A mid-market band around eight to ten times the floor is where most budget tiers cluster. It includes the early-bird pricing of new entrants. The incumbent's list rate sits at roughly thirty times the floor, about nine times the speed host. This platform's recorded planning class for its removed lane sits in the mid-market band, at roughly a third of the incumbent's list rate.

Now the drift stories. They are the reason this guide quotes relations and read dates instead of a price table.

First: an AI Overview panel caught in September 2026 attributing a router listing's per-second rate to the speed host. The speed host's own console page printed a rate nearly four times higher. The panel was not lying about a price. It was pasting the wrong party's number under the right party's name. This is the fourth independently recorded case of an AI summary printing a transcription rate that no host's own page carries. It is why this guide's law is that AI summaries are never rate sources. A rate that matters is read from the named host's own page and dated there.

Second, and material: a budget host whose own model page printed a card rate of about a hundredth of a dollar per audio hour. Its meter billed about one hundred times that. The discrepancy was caught not by reading the page. It was caught by running linear one-second, five-second, and sixty-second checks against a funded account and reconciling the wallet, in September 2026. The card and the meter disagreed by two orders of magnitude. Whatever the cause, the lesson prices the whole ladder. The rate that exists is the one the meter charges, and the only way to read a meter is to send measured audio at it and check the wallet. Any host's card is a claim. The invoice is the fact.

Two published break-evens, checked

Two of the ranking pages publish break-even math. Both were checked against street prices and arithmetic for this guide, both in September 2026 terms.

The first, an enterprise-stack guide published in October 2024, prices the self-host path with a line item of 10,000 to 12,000 for the RTX 3090. Its street price that year, and now, is about a tenth of that. The figure is data-center card money pasted onto a consumer card's name. Staffing is the rest of the story. The same stack's total, about 284,000 a year, is roughly 85 percent staffing. Those are real numbers for an enterprise building a transcription platform, presented as the cost of the self-host path generally. Its one durable figure is an electricity estimate for continuous 300 to 400 W operation. Its floor is plausible against the EIA math above, and its ceiling is padded. The correction, dated: the hardware premise is wrong by about ten times, and the staffing premise is wrong for anyone who is not hiring a team.

The second, a transcription vendor's guide read in September 2026, sets the self-host floor at a rented cloud GPU instance and states a monthly crossover in hours. The arithmetic is correct. The stated crossover hours times the incumbent's per-minute rate reproduces the rental price to the dollar. The frame is the problem twice over. The "self-host" floor is a rental, a monthly subscription that never ends and never becomes hardware you own. At the crossover volume it names, the rent paid by month four buys the card outright at street prices. The same vendor's per-file convenience tier sits in its own comparison table. It loses to the raw incumbent API at every file size. It loses by a factor of roughly twenty at the median job in that table. The correction, dated: the math is right, and the frame sells the middle option by pricing the extremes wrong.

This platform's lane: pulled, not repriced

This platform ran a hosted Whisper lane, batch shape over the standard transcription request, metered by the audio second. That lane is discontinued: it was listed, then removed from the catalog, and it returns only under named conditions. The reason is in the drift stories above. The provider behind it was the budget host whose meter billed at about a hundred times its own card rate. That is multiples of what batch transcription sells for anywhere. The lane was pulled from the catalog rather than repriced above the market, because selling above the market was a worse answer than no route. The return conditions are named. The lane comes back on self-hosted serving of the same open-weight checkpoint on platform GPU capacity. It keeps the same request shape and the same audio-second metering. That is the recorded plan, and it carries no date. When it returns, its pricing follows the platform's market-average discipline, the planning class quoted above. The rate renders from the catalog the day the row lands.

The metering law is the durable part, and it holds on every audio lane this platform runs. The hayamimi speech-to-text route answers calls today, metered per audio second from a prepaid wallet. It is the audio row that feeds the fence above. You are billed for the duration of the speech you actually sent, measured from the audio itself, never for upload wall-clock. A request turned away before transcription starts costs nothing. And spend draws down a wallet you charge in advance, so the ceiling applies to arriving work rather than to a month-end reconciliation. The speech-to-text category page carries the current state of every audio route. It updates as the catalog changes.

Work it with your own numbers

The formula above has seven inputs and you own five of them: your street price, your draw, your electricity price, your volume, your amortization window. Read the hosted rate, P, off any provider's own page, dated. If a rate decision matters, check the meter the way the drift stories did: measured audio in, wallet out. Per-unit rates for this platform's lanes sit on the pricing page. To put a small wallet behind a key and run the meter drill for yourself: sign up, create a key. Then, on whichever host's audio route you are pricing, send one measured minute of audio before you commit to a side of the crossover.

Last verified: 2026-09-25