Catalog voices only · No cloning

Vietnamese text to speech

Vietnamese is written in Latin script, which routinely tricks people into assuming it behaves like a European language. It does not. A set of six tones separates words otherwise identical to each other, and the marks carrying those tones stack on top of vowels that may already bear a separate mark modifying which vowel they are.

Try the live product

Type, choose a voice, then press play.

Live demo

Type a script, pick a voice, and hear it read back.

Free to try — no sign-up.

135/500
Voice

Tap a voice to hear it instantly, then press play to hear your script in it.

That stacking is where text handling usually breaks down. One vowel can hold a quality mark and a tone mark simultaneously, and losing either produces a different word rather than a slightly degraded version of the intended one. Text that has passed through careless processing arrives stripped and effectively unreadable.

Because the great majority of words are a single syllable, meaning rests heavily on those distinctions surviving intact from your document through to the audio. Northern and southern speech realise several tones noticeably differently, which is worth weighing when you choose a narrator for a specific audience.

No signup needed — the demo speaks up to 500 characters.

Regional delivery

  • Vietnamese

Every voice speaks every supported language, so one narrator can carry your whole catalogue across markets.

Vietnamese production details

  • Full tone set reproduced

    Tone is handled as lexical information, since it separates words sharing every other sound within the same syllable.

  • Stacked diacritics preserved

    A vowel bearing both a quality mark and a tone mark keeps both, because dropping either yields an entirely different word.

  • Single-syllable precision

    With most words one syllable long, small distinctions carry more weight than in languages spreading redundancy across syllables.

  • Regional realisation differs

    Northern and southern speech render several tones differently, which is worth considering when picking a narrator per market.

56 studio voices (30 premium, paid plans) · 20 output languages · SRT captions and bulk mode on every plan

Free to start

Use the free plan first, then move to a monthly plan when your production volume grows.

25,000 characters

Free plan monthly allowance

$7/month

About 3 hours on the Creator plan

Compare all plans →

Questions, answered

Is the complete tone set actually produced?

Yes. Tone separates words that share every other sound in Vietnamese, so reproducing the full set is a baseline requirement rather than a refinement to be layered on later.

What happens if my text has lost its diacritics?

Quality drops sharply and frequently becomes unusable. Stripped text is genuinely ambiguous, since a single bare syllable can correspond to several unrelated words distinguished only by those marks.

Why do some vowels carry two separate marks?

One identifies which vowel it is, the other carries tone. They perform different jobs, so both have to survive intact from your document into the narration to preserve the intended word.

Is the delivery northern or southern?

The voice you choose determines that. The two regions realise several tones differently, so audition narrators against the specific audience you actually publish to before committing.

Can I paste text copied out of a PDF?

Check the diacritics first. PDF extraction frequently mangles stacked marks, and text that looks acceptable on screen can be missing exactly the information narration depends on.

Your next project

Publish in Vietnamese this week

Start free