ModelFare.ai

Gemini 3.1 Flash TTS: where to access it & real cost

live on the first-party API by Google t2a

Speech billed by the length of what comes out, not by the length of what goes in.

Quick answer prices verified
Best for our call
Speech billed by the length of what comes out, not by the length of what goes in.
Cheapest
$0.03 at Google Cloud · 60s / speech · derived
First-party
$0.03 at Google Cloud · 60s / speech · derived
Routes selling
1 of 1 tracked
Verified
2026-08-15

Gemini 3.1 Flash TTS, from Google, is purchasable on the one route we track. 60 seconds of speech costs $0.03 at Google Cloud, the cheapest published rate we can cite. This model is billed at 60s / speech, which nothing else on the index sells, so its price is not ranked against the other models.

The figure above buys 60 seconds of speech. The board below prices 2 units in all, each column named. Figures come from each seller's own published rates and carry the date we last read them.

Which route should you take

Each row is the table above read for one question, at the unit named beside the price. Rows marked our call are editorial judgement instead, and appear only where we have something specific to say.

Provider board

Each column is one unit, and every price in it buys the same thing. The figures quoted elsewhere on this page are 60 seconds of speech. That is a baseline for comparing providers, not a quote. What you are actually billed depends on the settings you send, and on some routes on how much you top up at once. Every figure is the provider's own published price; anything we derived rather than read off a price list is marked est.

1 provider tracked · sorted by status, then price
ProviderAccess60s1M inStatus
First-party API
official$0.03derived$1.00claimedlive · preview, billed per token

Swipe the table for price and status →

How these prices were read

How each figure above was read off the provider's own price list, and what that price list does not say.

Google Cloud
Google bills this in tokens on both sides and publishes the bridge from audio tokens to seconds, so the output figure here is a real duration rather than a conversion of ours. The script going in is priced separately, in text tokens, and stays on its own row.

What it can do

Max length
not verified
Tracks per generation
not verified
Named voices
not verified
Languages
not verified
Characters per request
not verified
Writes lyrics
not verified
Instrumental mode
not verified
Voice cloning
not verified

Where a row says not verified, we have not checked that figure yet; it does not mean the model lacks the feature. Every row is read off published documentation for the model, never from our own testing.

If this is not the right model