What did the 9 October wire say?

PR Newswire’s 9 October note, source HeyGen, says the company “today launched HeyGen Voice, its new in-house voice model.” It says the model “debuts at #1 on Artificial Analysis’ Leaderboard, earning the top ranking as the world’s best model in independent evaluations of AI voice models.” A quoted line from CTO Rong Yan calls that ranking “independent proof that you don’t have to trade authenticity for quality.” Those are the company’s words.

Product highlights on the wire are the ranking, “authentic delivery,” an end-to-end stack with Avatar V, “Free Access” on the HeyGen platform and API, and a Professional Voice Clone “$99/month add-on” “trained, with explicit consent, on 30 minutes to three hours of the owner’s speech.” The same text cites a HeyGen survey of more than 1,000 small-business owners, 40 million users, 118 million videos, 85% of the Fortune 100, and “more than $200 million in ARR.” Those figures are HeyGen’s. We did not see the survey instrument or an audited revenue statement.

Claude Motion’s launch list already named HeyGen as a hand-off target. That is a different product. See Claude Motion and Dashboards. For how speech models are usually evaluated, see Whisper’s speech-model profile.

What does the Artificial Analysis board show tonight?

The page we opened is “Controlled Voice Arena Leaderboard.” The deck says it compares text-to-speech models “using the same 8 cloned voices (4 US, 4 UK).” That design is meant to isolate model quality from voice selection. It is not a test of HeyGen’s own cloned voices.

Row 1 is HeyGen Voice, range 1–2, Elo 1202±17, 95% interval 1185 to 1219, 1,482 samples, released Oct 2026, API pricing $30.0 / 1M chars. Row 2 is Qwen-Audio-3.1-TTS-Plus at 1186±15 (1171 to 1201), 1,856 samples, $19.3. Rows 3 and 4 are Eleven v4 Turbo at 1167±14 and Eleven v4 at 1160±14. The intervals for HeyGen and Qwen-Audio-3.1-TTS-Plus overlap. AA’s own range column prints 1–2 for HeyGen. That is the qualification the PR’s “#1” line leaves out.

A footnote says API pricing “reflects the cost to generate 1M characters on the model creator’s API with the model’s default settings.” We did not replay the arena or count the 1,482 samples.

Top of Artificial Analysis Controlled Voice Arena as opened on 9 October 2026, not a rerun
Rank / rangeModelElo (95% interval)SamplesAA API price
1 / 1–2HeyGen Voice1202±17 (1185–1219)1,482$30.0 / 1M chars
2 / 2Qwen-Audio-3.1-TTS-Plus1186±15 (1171–1201)1,856$19.3 / 1M chars
3 / 3–4Eleven v4 Turbo1167±14 (1153–1181)2,119$40.0 / 1M chars
4 / 3–4Eleven v41160±14 (1146–1174)2,079$80.0 / 1M chars

What do the developer docs actually ship?

The HeyGen Voice model page says the in-house model “learns one speaker’s voice, then speaks your text in that voice” in two modes. Instant takes one recording; “the first 3 minutes are used”; it is “usually ACTIVE within seconds”; it cannot be retrained. Professional takes one to ten recordings of the same speaker, “20+ minutes in total,” and is ready after training. Instant speech controls are expressiveness_boost. Professional keeps seed, speed, pitch_shift, pitch_variance and pause tags.

POST /v3/models/audio/voices creates either kind. POST /v3/models/audio/tts returns one 44.1 kHz WAV; /tts/stream streams parts. The default model id on speech is heygen-voice-1. Instant-clone docs say a recording may be an asset_id, a public URL or base64, up to 100 MiB (16 MB decoded for inline), and that a silent file fails. Supported base language codes include en, zh, ja, es, de, fr and a longer list on that page. A regional tag such as en-US “does not select a regional accent.”

Those endpoints are a clone-and-speak API, not a transcription pipeline. For uploaded audio used the other direction, see ChatGPT audio uploads. For the broader multimodal framing, see multimodal AI explained.

How do the price lists disagree?

The October 2026 API changelog, under “Instant voices on the HeyGen Voice model,” says instant speech “costs $15 per million submitted characters, rounded up per request to whole units of 1/60 API credit.” “Creating an instant voice is free during the preview.” Professional speech “stays at 0.6 API credits per generated minute.” The Voice model page says a professional voice occupies “one purchased voice clone slot,” with five pooled trainings per slot per month.

Artificial Analysis prints $30.0 per million characters. The PR says HeyGen Voice is “available free within the HeyGen platform and API” and names a $99-per-month Professional Voice Clone trained on 30 minutes to three hours. The docs’ professional floor is 20-plus minutes. Those are three documents. We did not buy a slot or generate a minute.

Treat platform “free” as the PR’s marketing line, $15 / $30 as two published per-character figures that we have not reconciled, and $99 as a wire add-on that is not the slot language on the API page.

What is still a vendor claim?

“World’s best model” is the PR. AA’s table is a preference Elo on a fixed eight-voice set, with an overlapping runner-up. Pronunciation-robustness and per-language ranks that appear in secondary write-ups were not on the leaderboard table we used as the independent source; we are not importing them from aggregators.

The September 2026 HeyGen blog, last updated 1 October, already described Professional Voice Clone on the API. The October changelog adds instant mode and makes mode required on create and retrain. The 9 October wire is the public ranking announcement. It is not evidence that the model first existed on that Friday.

Consent language on the wire — explicit consent, control of voice and likeness — is HeyGen’s policy sentence. The instant-clone docs do not reprint a consent flow. We did not create a voice.

What should a voice team verify before swapping Eleven or Qwen?

If you care about the arena, listen to HeyGen Voice and Qwen-Audio-3.1-TTS-Plus on the same eight AA voices. The Elo gap is smaller than the interval. If you care about product behaviour, clone one recording in instant mode and one 20-minute set in professional mode, then compare heygen-voice-1 against your current vendor on your own script, not on the PR’s “sounds like you” line.

Read the price that will actually hit the invoice: AA’s $30, the changelog’s $15 instant rate, credit-per-minute professional speech, or a $99 add-on. They are not one number. We did not submit characters or train a slot.

This is an evidence review of the 9 October PR Newswire note, the Controlled Voice leaderboard, the HeyGen Voice model page, the instant-clone guide, and the October 2026 changelog.

Common questions

Is HeyGen Voice a locked number-one on Artificial Analysis?

Not on the table we opened. It is row 1 with range 1–2. The 95% interval overlaps Qwen-Audio-3.1-TTS-Plus. AA’s Elo is a preference score on eight shared cloned voices.

Did HeyGen invent the model on 9 October?

The wire is dated 9 October. Professional Voice Clone was already in the September 2026 API notes. Instant mode is bucketed under “Added — October 2026” on the changelog, without a day stamp.

Is the API free, $15, $30 or $99 a month?

The PR says free plus a $99 professional add-on. The changelog says $15 per million characters for instant speech. AA prints $30.0. The docs bill professional speech in API credits. We did not receive an invoice.

THE TAKEAWAY

What to remember

Use Artificial Analysis for Elo 1202, the 1–2 range and the overlap with Qwen-Audio-3.1-TTS-Plus. Use HeyGen’s docs for instant versus professional modes. Keep “world’s best,” the $99 add-on and the ARR line in the PR’s column.

Sources & further reading

  1. HeyGen Launches HeyGen Voice, Debuting at #1 on Artificial Analysis’ Leaderboard ↗
  2. Controlled Voice Arena Leaderboard ↗
  3. HeyGen Voice ↗
  4. Instant Voice Clone ↗
  5. HeyGen API changelog ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories