ElevenLabs Sound Effects v2 vs TTS
Choose Sound Effects v2 for non-speech events, ambience, impacts, and loops. Choose TTS Multilingual v2 when finished words must become a stock-voice MP3. Both use fal and the same private audio path, but their inputs and pricing units are deliberately separate.
ElevenLabs Sound Effects v2
Creates one private MP3 sound effect from a non-speech description through fal's v2 endpoint.
- Best fit
- Impacts, interfaces, transitions, foley, ambience, textures, and seamless background loops.
- Output
- One privately stored MP3 at 44.1 kHz and 128 kbps
- Cost
- About $0.002 per output second; 1-4 credits
ElevenLabs TTS Multilingual v2
Creates one private MP3 voice track from finished text with a stock voice and speed.
- Best fit
- Narration, product reads, accessibility audio, and spoken drafts.
- Output
- One privately stored MP3 voice track
- Cost
- About $0.10 per 1,000 characters; 1-17 credits
Compare the same decision criteria
The decision starts from whether the deliverable should contain spoken words. Duration and character count are different supplier meters and never share one quote.
| Criterion | ElevenLabs Sound Effects v2 | ElevenLabs TTS Multilingual v2 | Decision rule |
|---|---|---|---|
| Primary job | Generate non-speech sound | Generate spoken voice | Choose from the audible result, not the shared vendor mark. |
| Input | Sound description and output settings | Finished script, stock voice, and speed | Do not send dialogue text to the sound-effect route. |
| Duration | Explicit whole seconds from 1 through 22 | Determined by text, voice, and speed | Sound Effects prices the requested output duration. |
| Loop | One shot or seamless loop | No loop control | Use Sound Effects for repeatable ambience, then inspect the seam. |
| Credits | 1-4 from output seconds | 1-17 from input characters | Review the correct unit before comparing totals. |
| Endpoint | fal-ai/elevenlabs/sound-effects/v2 | fal-ai/elevenlabs/tts/multilingual-v2 | Deprecated Sound Effects v1 is not registered. |
On smaller screens, this matrix scrolls inside its own frame. The page itself does not scroll sideways.
Reserved audio proof without cross-model substitution
Both slots remain reserved until Kanvora approves source-tracked output from the exact endpoint. One model's audio is never presented as evidence for the other.
Sound Effects v2 direction
Heavy wooden door closing in a large stone hall, sharp impact, short low-frequency tail, no speech and no music.
Five seconds, balanced influence, and one-shot output are attached. The exact hold is 1 credit.
Open this directionTTS Multilingual v2 rehearsal
Welcome to Kanvora. This voice rehearsal checks pronunciation, pacing, names, numbers, and punctuation before the finished track is published.
Rachel at natural speed is attached. Character count sets the exact hold.
Open this directionPrepared directions do not prove event accuracy, loop quality, pronunciation, naturalness, loudness, or a universal winner.
Choose from what the audience must hear
These scenarios keep sound design and speech generation separate without bundling Noise Remover, Music, eleven-v3, dialogue, or cloning.
A game button needs a short confirmation tone.
Signal: No words should be audible.
Choose Sound Effects v2.Create the sound directionA concrete sound description, short duration, and one-shot ending match the endpoint.
A product script needs a calm spoken draft.
Signal: The exact words must be voiced.
Choose TTS Multilingual v2.Create the voice rehearsalTTS accepts finished text plus a stock voice and speed.
Rain ambience must repeat under a long scene.
Signal: The beginning and ending must meet cleanly.
Choose Sound Effects with Seamless loop.Open Sound Effects settingsThe v2 route exposes loop behavior, but the stored seam still needs review.
An existing noisy voice recording needs cleanup.
Signal: The speech already exists in an MP3 or WAV.
Choose the separate Noise Remover tool.Open Noise RemoverAudio isolation accepts the private recording and prices the request from server-verified input duration.
Switching audio jobs creates a separate paid request
Wait for active or uncertain jobs to settle. Changing between sound design and speech changes endpoint, input schema, supplier meter, and credit hold.
Switching and handoff costs
- Sound Effects bills requested output seconds at about $0.002 per second.
- TTS bills submitted input characters at about $0.10 per 1,000 characters.
- Kanvora rounds each supplier total up to a positive whole-credit hold under the same $0.012 coverage formula.
- Neither request automatically triggers Noise Remover, Music, eleven-v3, dialogue, cloning, or the deprecated v1 endpoint.
What the comparison cannot prove
- A matching prompt does not guarantee a specific recording, acoustic space, or loop seam.
- TTS does not guarantee every name, number, acronym, language, or emotional delivery.
- Reserved proof slots are not quality benchmarks.
- The shared ElevenLabs mark does not make the two endpoints interchangeable.
Design non-speech sound with Sound Effects. Voice finished words with TTS.
Both CTAs carry one legal live operation into the authenticated workbench. Safety checks, exact holds, and private storage stay at the real request boundary.