Audio production was once a specialty discipline. Composing a track, recording a voiceover, generating sound effects, transcribing a recording: each required specific software, specific training, or specific people.
In 2026, all of it is available in a single platform through AI. This guide covers every audio tool in Kolbo.AI and exactly when to reach for each one.
Music Generation
Suno v5.5
Suno is the leading AI music generation model available in Kolbo.AI. Give it a style description and optional lyrics, and it produces a complete song: melody, arrangement, vocals or instrumental, all generated from scratch.
What it does: generates full tracks from a text prompt. Handles a wide range of genres including pop, hip-hop, electronic, folk, country, classical, jazz, and fusion styles.
The two-field approach: Suno works best with two separate inputs:
- Style field: the musical style, mood, instruments, and tempo. Keep this factual and musical: "melancholic indie folk, fingerpicked acoustic guitar, soft male vocals, 80 BPM"
- Lyrics field: the actual song text in verse/chorus structure, with
[Verse],[Chorus], and[Bridge]markers
When to use Suno:
- Creating original background music for video content
- Producing demo versions of song ideas
- Generating music for a social media account that needs a consistent audio identity
- Clients who need original music without licensing costs
Commercial use note: commercial use of Suno-generated tracks requires a paid Suno subscription plan. Verify this before publishing monetized content.
Key limitation: Suno generates in a fixed structure. If you need something very specific (a 17-second jingle that matches an exact BPM for a video) the output may need trimming or looping. Suno does not currently support frame-accurate generation.
Google Lyria 3 Pro
Google Lyria 3 Pro is the second music generation model in Kolbo.AI and offers a different set of strengths.
Where Lyria 3 Pro excels:
- Instrumental music with complex orchestral or classical arrangements
- Extended compositions (longer tracks with more structural development)
- Cases where the musical technical quality matters more than the vocal performance
- Clean, consistent output for professional production use
How to choose between Suno and Lyria 3 Pro:
| Need | Recommended |
|---|---|
| Song with vocals and lyrics | Suno v5.5 |
| Instrumental background music | Lyria 3 Pro or Suno (both work well) |
| Classical or orchestral | Lyria 3 Pro |
| Pop, hip-hop, electronic with vocals | Suno v5.5 |
| Consistent technical quality for production | Lyria 3 Pro |
Minimax Music
Minimax Music is a third music generation option in Kolbo.AI, positioned as a fast, high-quality generator for shorter tracks and stingers.
Best for: intros, outros, short jingles, stingers, and background tracks for presentations or social content where you need something quickly and do not require the full song structure that Suno generates.
Text-to-Speech and Voice
ElevenLabs v3 TTS
ElevenLabs is the premier text-to-speech engine in Kolbo.AI and currently among the best in the world for naturalness and voice variety.
Capabilities:
- 3,000+ voices covering dozens of languages and accents
- Emotional range: voices that sound genuinely natural rather than robotic
- Custom voice cloning (see below)
- Multilingual: generate speech in English, Spanish, Arabic, Hebrew, French, German, Japanese, and dozens more from the same voice
The practical use case range is broad:
- Podcast episodes narrated by an AI voice
- YouTube video voiceovers
- Explainer video narration
- Audiobook production
- Language learning content
- Marketing and advertisement voiceovers
Choosing a voice: browse by gender, age, accent, and use case in Kolbo.AI. For a podcast, try voices tagged "conversational." For an advertisement, try voices tagged "promotional" or "commercial."
Tip: the combination of shorter sentences + natural punctuation produces more natural-sounding TTS output. Long run-on sentences reduce naturalness.
Voice Cloning
Custom Voice Cloning in Kolbo.AI (powered by ElevenLabs) lets you create a persistent AI version of any voice from audio samples.
How it works: upload 1 to 5 minutes of clean audio from the target speaker. Kolbo.AI creates a custom voice model you can reuse indefinitely for any script.
What you need for good results:
- Clean audio with no background noise or music
- The speaker talking continuously (not overlapping with others)
- Minimal editing artifacts or sudden cuts
- At least 1 minute of material; 3 to 5 minutes produces significantly better results
Important ethical requirement: you must have explicit consent from the person whose voice you are cloning. Cloning someone's voice without their consent is a violation of ElevenLabs' terms of service and potentially illegal in many jurisdictions.
Legitimate use cases:
- Content creators who want to produce narration in their own voice without re-recording for every project
- Companies that want a branded spokesperson voice across all audio content
- Localization: clone the original speaker's voice and generate dubbed versions in other languages while preserving the speaker's characteristic sound
Sound Effects (Kolbo Sound Studio)
Kolbo's Sound Effects tool generates custom sound effects from text descriptions. Instead of searching a stock library and licensing a track, you describe what you need and the AI creates it.
Examples:
- "footsteps on wet pavement, slow rhythmic walking"
- "notification chime, short, modern tech feel"
- "thunder crack followed by heavy rain"
- "coffee machine brewing, close microphone position, warm"
When to use generated SFX over a stock library:
- When you need something very specific that does not exist in standard libraries
- When the project requires a unique sound that differentiates it
- When you want to avoid any licensing complexity
When to use a stock library: for common sounds where a large selection at consistent quality saves time over custom generation.
Transcription
ElevenLabs Scribe v2
Scribe v2 is the transcription engine in Kolbo.AI. It converts audio and video recordings to text with high accuracy across 90+ languages.
Output formats:
- SRT: subtitle file for video captioning (YouTube, social, editing software)
- TXT: plain text transcript for notes, blog posts, show notes
- VTT: WebVTT for web-embedded video
- JSON: machine-readable with word-level timestamps and confidence scores
Speaker diarization: when enabled, Scribe v2 identifies different speakers in the recording and labels each segment. Essential for interviews, podcasts, meetings, and any multi-person recording.
Full guide: see the separate transcription guide for complete use-case workflows: AI Transcription and Subtitles Guide.
Audio Studio
The Audio Studio in Kolbo.AI is the post-production environment. Once you have audio (from any source: recorded, TTS-generated, music, SFX), the Audio Studio handles transformation and editing.
Voice-to-Voice
Feed one voice recording, receive it in a different voice. Use cases:
- Change the gender or age of a voice recording
- Apply a specific accent to a neutral recording
- Create variations of a voiceover from a single source recording
Voice-to-Voice preserves the prosody (rhythm and emphasis) of the original while transforming the voice characteristics. The timing of the speech stays the same; the sound of the voice changes.
Stem Separation
Stem Separation splits a mixed audio track into its components: vocals, drums, bass, and other instruments.
When this is genuinely useful:
- You have a recording with background music and need to extract just the speech
- You have a track with vocals and want to create an instrumental version
- You want to remix a track but only have the final mix
- You need to remove background music from a video recording before re-dubbing it
Practical example: a client provides a promotional video with music and voiceover combined on one track. They want the voiceover dubbed into Spanish but the original music kept. Stem Separation extracts the vocal track for re-dubbing and the music track for mixing back in.
Choosing the Right Audio Tool by Use Case
| What you need | Tool |
|---|---|
| Original music for a video | Suno v5.5 or Lyria 3 Pro |
| Short jingle or stinger | Minimax Music |
| Voiceover for any content | ElevenLabs v3 TTS |
| Narration in your own voice | Voice Cloning + ElevenLabs |
| Custom sound effect | Sound Effects generator |
| Subtitles for a video | Scribe v2 (SRT output) |
| Transcript of a recording | Scribe v2 (TXT output) |
| Remove background music | Audio Studio Stem Separation |
| Transform a voice recording | Audio Studio Voice-to-Voice |
| Multilingual audio from one script | ElevenLabs v3 TTS (multilingual voice) |
A Complete Audio Production Workflow
Scenario: a documentary series episode, 15 minutes, covering AI trends.
- Script: drafted in Kolbo Chat (Claude), edited by the producer
- Narration: generated with ElevenLabs v3 TTS, custom cloned voice of the show's regular narrator
- Background music: two tracks generated in Lyria 3 Pro: one for the intro (dramatic orchestral) and one for section transitions (ambient minimal)
- Sound effects: three custom effects generated in Kolbo's SFX tool: a digital data stream sound, a UI notification, and an ambient office environment
- Subtitles: audio fed to Scribe v2, SRT exported, uploaded to the video editor and to the publishing platform
- Final audio: narration, music, and SFX edited and mixed in the Audio Studio
Total audio production time: approximately 2 to 3 hours for a 15-minute episode. Traditional production equivalent: full day of recording plus post-production.
Get Started With Audio in Kolbo.AI
Every audio tool in this guide is available at app.kolbo.ai under a single subscription. No separate accounts for ElevenLabs, Suno, and your transcription service. One platform, one workflow, all the audio tools you need in 2026.
Try generating your first music track, voiceover, or transcription for free and see how much of your audio production you can move in-house.



