vocalove

Resemble AI · ultra-fast TTS

Chatterbox Turbo

Chatterbox Turbo is open-source text-to-speech built for low-latency, expressive speech — pick a narrator or clone from a short clip, then add inline tags for laughs, sighs, and other human reactions.

Chatterbox Turbo is open source (MIT) with official docs and public weights on Hugging Face — On Vocalove you never call the API yourself.

Create your story

Voice
Script
Voice model
0 / 1000

Free to preview in your browser — no credit card. New accounts start with 5 free credits that never expire, plus a free daily pool. Watermark-free HD exports use credits only when you're ready to download.

What is Chatterbox Turbo?

Chatterbox Turbo is Resemble AI’s speed-optimized TTS — made for voice agents, tributes, and any script that needs to feel immediate and human, not read from a teleprompter.

Type what should be heard, optionally embed tags like [chuckle], [sigh], or [gasp] in the text, and generate speech in the browser. Tags land in the same voice — no post-production needed for a natural pause or a gentle laugh.

The Chatterbox family includes Multilingual V3 (20+ languages). On Vocalove this page runs the English Chatterbox Turbo path — write English scripts for best results.

  • Ultra-low latency

    Built for real-time feel — quick enough for agents and interactive scripts.

  • Clone from a short clip

    About five to ten seconds of clear speech under Clone voice — optional when you need a specific timbre.

  • Paralinguistic tags

    [laugh], [sigh], [chuckle], [cough], and more — inline in your script, same voice.

  • Open source (MIT)

    Resemble AI’s Chatterbox family — preset narrators plus cloning in one turbo model.

Try Chatterbox Turbo on Vocalove in three steps

Pick a voice, paste your script, generate — no install. Start in the tool above.

  1. Step 1

    Pick Built-in or Clone voice

    Built-in for preset narrators. Clone records or uploads about five to ten seconds of clear speech.

  2. Step 2

    Type your script

    Write up to 1,000 characters per generation — the tool counter matches this limit. Add tags like [sigh] or [chuckle] where reactions belong. Use English for best results on this page.

  3. Step 3

    Generate and download

    Preview speech in the browser, download MP3, or continue to Talking Photo Videos when you want lip-sync.

Jump to the speech tool ↑

What Chatterbox Turbo is good for

When speech must sound reactive, not read from a teleprompter.

  • Voice agents & chatbots

    Low latency keeps turn-taking natural — agents can sigh, laugh, or pause like a person on a call.

  • Memorial & tribute lines

    Clone from a short sample and add [sigh] or gentle [chuckle] for lines that need emotional pacing.

  • Talking photo prep

    Generate expressive audio here, then upload the portrait on Talking Photo Videos for Kling Avatar lip-sync.

  • Memes & short-form content

    Preset voices plus tags — fast iteration for videos, games, and social clips without a studio session.

Why Chatterbox Turbo?

Most TTS reads cleanly but sounds flat. Chatterbox Turbo targets reactive speech — tags let you choreograph breath and reaction in the script instead of editing audio afterward.

It is also open source (MIT) with optional provenance watermarking on some hosted outputs — a different tradeoff from closed APIs when you care about traceability.

  • Optional clone from roughly five to ten seconds of reference speech.
  • Preset narrators when you do not have a sample yet.
  • Pair generated audio with Vocalove talking photo videos when the script is ready.

Tags, presets, and cloning — in plain terms

What makes this model different from a generic narrator read.

  1. Step 1

    Preset or cloned voice

    Pick a built-in narrator, or switch to Clone voice and upload a short sample so new lines follow that timbre.

  2. Step 2

    Script plus optional tags

    Paste up to 1,000 characters. Inline tags — [sigh] before bad news, [chuckle] after a warm line — perform in whichever voice you chose.

  3. Step 3

    Download MP3

    Generate in the browser and download audio when it sounds right. Longer projects: split into multiple ${SCRIPT_LIMIT}-character sections.

Chatterbox Turbo on Vocalove

What the browser tool accepts and returns — aligned with the counter in the tool card above.

On Vocalove
Browser tool — no repo clone or local GPU required
Text per generation
Up to 1,000 characters (matches the script counter)
Languages on Vocalove
English scripts — best results on this page
Chatterbox family
Multilingual V3 (20+ languages) exists upstream; not the default path here
Clone sample
About 5–10 seconds of clear speech under Clone voice
Expressive tags
[laugh], [chuckle], [sigh], [gasp], [cough], [clear throat], …
Output
MP3 audio for preview and download

Chatterbox Turbo vs other TTS approaches

How Chatterbox Turbo compares to one-step cloners and preset narration on Vocalove.

 Chatterbox TurboZero-shot clone (e.g. F5-TTS)Preset narrator
Primary strengthSpeed + paralinguistic tagsZero-shot clone (EN/ZH)Fast browser drafts
Clone sample~5–10 seconds (clone mode)3–30 secondsNone
Expressive tags[laugh], [sigh], etc. in scriptNatural writing onlyNatural writing only
Languages on VocaloveEnglish (Multilingual V3 elsewhere in family)English & ChineseEnglish & Chinese

Speech tools on Vocalove

Use the tool above to pick a built-in voice or clone from a sample, type your script, and download speech in the browser. Chatterbox Turbo fits tag-driven expressiveness and fast English TTS without leaving the page.

For developers (self-host & API hosts)(collapsed by default)

The notes below apply when you call Chatterbox Turbo on a hosted API or run weights yourself — not the Vocalove browser workflow above.

Preset voice names
Typical APIs expose narrators such as lucy, aaron, chloe, and brian
Clone reference field
audio_url — a 5–10 second reference clip; overrides the preset voice when set
Temperature
Optional inference parameter (~0.8 default on many hosts) for output variation
Longer scripts on APIs
Some hosts accept up to ~5,000 characters per request — Vocalove caps at 1,000 per generation
Output format
WAV or similar via URL on typical hosted inference
  • Preset voice selects a built-in narrator. Passing audio_url clones from your reference clip instead — same text and tags apply to either path.

What to do after speech audio

Download when the line sounds right. For lip-sync, open Talking Photo Videos.

Chatterbox Turbo FAQ

What is Chatterbox TTS?
Chatterbox TTS is Resemble AI’s open-source text-to-speech family — preset narrators, optional zero-shot cloning from a short clip, and inline paralinguistic tags like [laugh] and [sigh]. Chatterbox Turbo is the low-latency variant tuned for real-time, expressive speech.
Is Chatterbox Turbo the same as Chatterbox TTS?
Chatterbox is Resemble AI’s open-source TTS family; Turbo is the low-latency variant. On Vocalove you use it as Chatterbox Turbo — searches for "Chatterbox TTS" land on this same model.
Is Chatterbox TTS free?
You can preview speech free in your browser — no credit card. New accounts start with 5 credits that never expire, plus a free daily pool; watermark-free HD downloads use credits only when you download.
Is Chatterbox TTS good?
For reactive, human-sounding speech it is a strong choice: fast inference, optional short-sample cloning, and paralinguistic tags that beat post-production for agents, tributes, and short-form scripts. It is MIT-licensed and widely used where latency and expressiveness matter more than a generic narrator read.
How do I use Chatterbox TTS?
On this page: pick Built-in or Clone voice in the tool, type your script (tags like [sigh] optional), and click generate. Preview in the browser, then download MP3 — no install or repo setup required.
How long can my script be?
Up to 1,000 characters per generation on Vocalove — the counter in the tool matches this limit. Split longer narration into multiple sections and stitch the MP3 files in your editor.
What are paralinguistic tags?
Inline markers in your text — [chuckle], [sigh], [gasp], [cough], [clear throat], and others — that the model performs in the same voice. Useful for natural pacing in agents and tribute scripts.
How long should my clone sample be?
About five to ten seconds of clear single-speaker speech when using Clone voice on Vocalove.
Which languages does Chatterbox Turbo support?
The Chatterbox family includes Multilingual V3 (20+ languages). On Vocalove this page runs the English Chatterbox Turbo path — write English scripts for best results. Need Mandarin or other locales on Vocalove? Try F5-TTS or built-in Chinese voices on the homepage tool.
What is PerTh watermarking?
Resemble AI’s imperceptible audio watermark on some Chatterbox outputs — helps trace provenance without changing how speech sounds to listeners.
Does Chatterbox Turbo make talking photo videos?
No — it outputs speech audio only. Animate a portrait on Vocalove with Kling Avatar after you have an audio track you like.