What it is
ElevenLabs is an AI audio platform built around its own foundational models. It has two product areas: ElevenCreative for generating speech, music, sound effects, images and video, and ElevenAgents for configuring, deploying and monitoring conversational agents. It also offers APIs for text to speech, speech to text and music. The homepage cites 90+ languages and a large voice library.
Who it's for
- Developers building voice, transcription or music features through the APIs
- Enterprises deploying voice or chat agents for customer experience
- Creators producing audiobooks, podcasts, voiceovers, ads, and social content
- Teams localizing content with dubbing and multilingual speech
Requirements
Requirements
- An ElevenLabs account (sign-up is offered on the site)
- An API key to use the API (the code samples use a YOUR_API_KEY placeholder)
- The @elevenlabs/elevenlabs-js client library, imported in the JavaScript samples
Setup
Create a client with your API key
The homepage's JavaScript sample imports ElevenLabsClient from @elevenlabs/elevenlabs-js and initializes it with an API key. The page gives no install command, so see the docs linked from the site (Explore docs).
typescriptimport { ElevenLabsClient } from "@elevenlabs/elevenlabs-js"; const client = new ElevenLabsClient({ apiKey: "YOUR_API_KEY" });
Examples
Convert text to speech
typescriptimport { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
const client = new ElevenLabsClient({ apiKey: "YOUR_API_KEY" });
await client.textToSpeech.convert("JBFqnCBsd6RMkjVDRZzb", {
outputFormat: "mp3_44100_128",
text: "The first move is what sets everything in motion.",
modelId: "eleven_multilingual_v2",
});What it does: Calls the Text to Speech API with a voice ID, an MP3 output format, the text to speak, and the Eleven Multilingual v2 model.
Create a music composition plan
typescriptimport { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
const { music } = new ElevenLabsClient();
const compositionPlan = await music.compositionPlan.create({
prompt: "Fast-paced electronic track for a video...",
musicLengthMs: 10000,
});What it does: Uses the Music API to create a composition plan from a natural-language prompt, with a length of 10000 ms.
Expressive speech text with inline cues
PromptIn the ancient land of Eldoria, where skies shimmered and forests, whispered secrets to the wind, lived a dragon named Zephyros. [sarcastically] Not the “burn it all down” kind... [giggles] but he was gentle, wise, with eyes like old stars. [whispers] Even the birds fell silent when he passed.Expected output: The homepage's speech demo text includes bracketed cues such as [sarcastically], [giggles] and [whispers], illustrating controllable, expressive speech.
Pros & cons
Pros
- Pro:Supports 90+ languages for speech, and the homepage cites 5,000+ voices
- Pro:Offers several TTS model options, including Eleven Flash at 75ms latency for conversational use
- Pro:ElevenAgents covers phone, chat, email and WhatsApp, with analytics, testing, guardrails and workflows
- Pro:One platform covers speech, music, sound effects, voice cloning, dubbing and transcription, with APIs and SDKs
Cons
- Con:Music commercial rights vary by subscription tier
- Con:The TTS API models are listed as supporting 29+ languages, fewer than the 90+ cited for the platform