The mid-2026 text-to-speech decision usually comes down to three engines optimising three different things: Fish Audio (the value play), ElevenLabs (the quality-and-ecosystem leader), and Cartesia (the latency specialist). This leaf is the head-to-head with the discipline the vendor pages skip: every number dated and arena-named, the pricing gotcha nobody mentions (Fish bills per byte, not per character), the licence trap on "open" weights — and a South African accent test protocol you can run yourself instead of trusting anyone's demo reel.
S2.1 Pro: 83 languages (vendor-stated), ~90ms time-to-first-audio, and a paid API at $15 per million UTF-8 bytes. The API is currently free — a window Fish has extended to 31 August 2026 — with no SLA and inputs eligible for training. Backed by a $52M seed; the S2 family ships open weights (see §04 before you plan around that).
Flash (~75ms) at roughly $0.05/1K chars, Multilingual v2 / v3 around $0.10/1K — call it 3–7x Fish's paid rate, mediated by a credits system (the ElevenLabs leaf covers the pricing mechanics). Deepest voice library, most mature agent tooling, best docs. $11B Series D (Feb 2026), IPO trajectory — and a provenance flag (§04).
Sonic 3: ~40ms TTFA on Turbo (vendor-stated), the fastest commercial path; ~$39 per million on the Artificial Analysis listing, plans vary; 42 languages. SSM-based architecture (a Stanford spin-out). And a quality story that changed recently — see §03 before repeating the old "fast but lower quality" line.
The big labs (OpenAI, Google) ship credible TTS inside their platforms and are covered in the voice landscape leaf — they matter when you're already committed to that stack, but none of the three above requires one.
As of 1 August 2026. Latency figures are vendor-stated; quality positions are arena-cited in §03; pricing is from vendor pages and the Artificial Analysis listing. All of it moves — treat this as the shape, and re-verify the figures before a procurement decision.
| Dimension | Fish Audio S2.1 Pro | ElevenLabs (Flash / v2–v3) | Cartesia Sonic 3 |
|---|---|---|---|
| Paid price | $15 / M UTF-8 bytes | ~$50–$100 / M chars (credit-mediated) | ~$39 / M (AA listing; plans vary) |
| Latency (TTFA, vendor) | ~90ms | ~75ms (Flash) | ~40ms (Turbo) |
| Languages (vendor) | 83 | 70+ | 42 |
| Weights | Open — research licence (§04) | Closed | Closed |
| Free tier | API free to 31 Aug 2026; no SLA; trains on inputs | Limited free credits | Trial credits |
| Data policy note | ZDR on enterprise tier | BIPA class actions pending (§04) | Standard commercial terms |
| Distinct strength | Cost per unit of quality | Voice library, agent tooling, docs | Raw speed for turn-taking |
Fish's $15/M is per million UTF-8 bytes, not characters. For English text those are roughly equal — but non-Latin scripts encode at 2–3 bytes per character, so the effective per-character rate doubles or triples for isiZulu with diacritics, Arabic, or Chinese content. When you compare against ElevenLabs' per-character pricing, compare on your corpus, not on the headline.
TTS quality rankings are fragmented in a way most write-ups hide. Artificial Analysis alone runs more than one board: a Provider Voice Arena (each vendor's chosen voices) and a newer Controlled Voice Arena that clones the same eight voices across every model — the cleaner methodology, because it separates voice preference from model quality. The two boards crown different leaders, and positions move weekly.
On the Controlled Voice Arena (snapshot, late July 2026): Cartesia Sonic 3.5 led at Elo ~1,122, ahead of ElevenLabs Eleven v3 (~1,088) — while Fish S2 Pro (~1,128 in the same period's tracking) was the top open-weights entry, and the older Sonic 3 sat well down the table (~1,070). Three practical conclusions. First, the lazy line "Cartesia is fast but lower quality" is stale — its newest model currently tops the cleanest board. Second, Fish's quality-per-dollar claim is genuinely supported. Third, and most important: never quote a TTS rank without the arena name and a date — the boards disagree, the models version monthly, and any specific Elo in this paragraph may be wrong by the time you read it. Check the live board.
The S2 family's weights are published, with fine-tuning code and an inference engine — but under the Fish Audio Research License: free for research and non-commercial use, while commercial use requires a separate licence from Fish Audio. So the data-residency story — run the model in-country, POPIA-clean, no vendor dependency — is architecturally real but commercially priced: budget the licence conversation before promising a client "we'll just self-host Fish." It remains the only self-host path among the three; it just isn't a free one.
Fish's free API window (currently through 31 August 2026) carries no SLA, no latency guarantee, and permits training on your requests. That's fine for prototyping with non-sensitive text; it is not acceptable for client production or anything containing personal information — zero-data-retention sits on the enterprise tier. For POPIA-scoped work, the s.72 lens from the Data privacy & POPIA leaf applies to prompts sent to any of these hosted APIs.
In May 2026, journalists, podcasters, and voice actors filed class actions in Illinois federal court against ElevenLabs and eight other companies under the Biometric Information Privacy Act, alleging voices were scraped for training without the written consent BIPA requires ($1,000–$5,000 per violation). Allegations, not findings — but for clients whose brand sits on the output, training-data provenance is now a named, litigated risk, not a hypothetical. The commercial backdrop: an $11B Series D (Feb 2026) with a secondary sale reportedly near $22B and a stated IPO path — a vendor this size is stable, and also lawyered.
None of the three ships South African languages as first-class voices, and "sounds fine in the demo" tells you nothing about how an engine handles SA English, code-switching, or local proper nouns. This is a protocol a client can run in an afternoon, for free, on their own ears.
1 · Fix the test script. One paragraph, identical for every engine, containing the failure modes that matter locally: rand amounts said aloud ("R12,847.50"), a phone number, place names (Johannesburg, Umhlanga, Stellenbosch, Makhanda), personal names across languages (Thandiwe, Johan, Lerato, Pieter, Sipho), one code-switched sentence, and one long compound sentence for prosody.
2 · Generate blind. Run the identical script through each engine's default recommended voice (and one "South African English" voice where offered). Strip the filenames; label A/B/C.
3 · Score with local ears. At least three South African listeners score each clip 1–5 on three axes: intelligibility (every word recoverable?), naturalness (would this pass on a support call?), and proper-noun accuracy (names and places said correctly — the axis engines fail most). Average per engine, per axis.
4 · Re-run when it matters. Models version monthly and the rankings above will drift — re-run the protocol at procurement time and at renewal, not once. The point of the protocol is that it costs nothing and answers the only question the leaderboards can't: how it sounds to your users, on your names, in your accent.