Most text-to-speech engines give you one flat read and a language picker. Cartesia Sonic 3.5 gives you direction. Pick the emotion the line needs, set how fast or slow it lands, and hand it a voice to clone, and it delivers a take that actually fits the moment.
It is live now in the Text to Speech tool on Kolbo.AI.
What Cartesia Sonic 3.5 Does
Direct the emotional read. Choose from presets including neutral, calm, content, happy, excited, sad, angry, and scared. The same line reads completely differently depending on what you pick, tone, pacing, and delivery all shift together.
Control the pace. Slow a line down for weight and clarity, or speed it up for something punchy and urgent. The range runs from unhurried to brisk, so you can match the read to the edit instead of fixing the edit to match the read.
Speak in a wide range of languages. English, Spanish, Hebrew, and many more, all from the same tool, without switching models.
Clone a voice from one short sample. Provide a brief reference clip and Sonic 3.5 builds a voice you can generate from immediately. No training run to wait on. The clone is a separate, one-time step, then it is yours to reuse.
Clean MP3 output. Ready to drop straight into your edit.
Write long-form scripts in one pass. Each generation covers a generous stretch of text, so a full paragraph or script section does not need to be split into pieces.
What You Can Make With It
- Narration and voiceover for video projects, product demos, and explainers
- Character dialogue where the emotional read actually matches the line
- Multilingual versions of the same script without re-recording from scratch
- Branded intros and outros in a cloned or preset voice
- Accessibility audio for documents and interfaces
How to Try It
Cartesia Sonic 3.5 is already in your Kolbo workspace.
- Open Audio Tools from your dashboard
- Select Text to Speech
- In the model selector, choose Cartesia Sonic 3.5
- Pick an emotion, set your pace, write your line, and generate
To clone a voice, upload one short sample in the voice section. The clone is billed once, then it is ready to use right away.
Cartesia Sonic 3.5 runs at 5 credits per 100 characters generated, billed on the text you submit.
Frequently Asked Questions
Can I hear the same line read with a different emotion? Yes. Change the emotion preset and regenerate. The wording stays the same, the delivery changes.
Is the audio delivered as a live stream while it speaks? No. Cartesia Sonic 3.5 processes your text and hands back a finished MP3 file once the take is generated.
Can I write SSML or markup tags into my script? No. Direction comes from the emotion preset and the speed control, not markup tags in the text.
Does it output word-level timestamps for subtitles? No. If you need timed captions for the resulting audio, run it through Kolbo's transcription tool afterward.
Cartesia Sonic 3.5 is live in your Kolbo workspace now.
Try Cartesia Sonic 3.5Best, Zohar Founder, Kolbo.AI

