Integrate eleven labs text-to-speech (#16573)

* Add elevenlaps tts support

* fix mypy error

* add simple usage in docs
This commit is contained in:
Sameer Kankute 2025-11-25 08:19:30 +05:30 • committed by GitHub
parent 35bfcac3bc
commit fc219c7db8
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
6 changed files with 723 additions and 5 deletions

View file

@ -7,10 +7,10 @@ ElevenLabs provides high-quality AI voice technology, including speech-to-text c
| Property | Details |
|----------|---------|
| Description | ElevenLabs offers advanced AI voice technology with speech-to-text transcription capabilities that support multiple languages and speaker diarization. |
| Description | ElevenLabs offers advanced AI voice technology with speech-to-text transcription and text-to-speech capabilities that support multiple languages and speaker diarization. |
| Provider Route on LiteLLM | `elevenlabs/` |
| Provider Doc | [ElevenLabs API ↗](https://elevenlabs.io/docs/api-reference) |
| Supported Endpoints | `/audio/transcriptions` |
| Supported Endpoints | `/audio/transcriptions`, `/audio/speech` |
## Quick Start
@ -228,4 +228,241 @@ ElevenLabs returns transcription responses in OpenAI-compatible format:
1. **Invalid API Key**: Ensure `ELEVENLABS_API_KEY` is set correctly
---
## Text-to-Speech (TTS)
ElevenLabs provides high-quality text-to-speech capabilities through their TTS API, supporting multiple voices, languages, and audio formats.
### Overview
| Property | Details |
|----------|---------|
| Description | Convert text to natural-sounding speech using ElevenLabs' advanced TTS models |
| Provider Route on LiteLLM | `elevenlabs/` |
| Supported Operations | `/audio/speech` |
| Link to Provider Doc | [ElevenLabs TTS API ↗](https://elevenlabs.io/docs/api-reference/text-to-speech) |
### Quick Start
#### LiteLLM Python SDK
```python showLineNumbers title="ElevenLabs Text-to-Speech with SDK"
import litellm
import os
os.environ["ELEVENLABS_API_KEY"] = "your-elevenlabs-api-key"
# Basic usage with voice mapping
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2",
input="Testing ElevenLabs speech from LiteLLM.",
voice="alloy", # Maps to ElevenLabs voice ID automatically
)
# Save audio to file
with open("test_output.mp3", "wb") as f:
f.write(audio.read())
```
#### Advanced Usage: Overriding Parameters and ElevenLabs-Specific Features
```python showLineNumbers title="Advanced TTS with custom parameters"
import litellm
import os
os.environ["ELEVENLABS_API_KEY"] = "your-elevenlabs-api-key"
# Example showing parameter overriding and ElevenLabs-specific parameters
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2",
input="Testing ElevenLabs speech from LiteLLM.",
voice="alloy", # Can use mapped voice name or raw ElevenLabs voice_id
response_format="pcm", # Maps to ElevenLabs output_format
speed=1.1, # Maps to voice_settings.speed
# ElevenLabs-specific parameters - passed directly to API
pronunciation_dictionary_locators=[
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
model_id="eleven_multilingual_v2", # Override model if needed
)
# Save audio to file
with open("test_output.mp3", "wb") as f:
f.write(audio.read())
```
### Voice Mapping
LiteLLM automatically maps common OpenAI voice names to ElevenLabs voice IDs:
| OpenAI Voice | ElevenLabs Voice ID | Description |
|--------------|---------------------|-------------|
| `alloy` | `21m00Tcm4TlvDq8ikWAM` | Rachel - Neutral and balanced |
| `amber` | `5Q0t7uMcjvnagumLfvZi` | Paul - Warm and friendly |
| `ash` | `AZnzlk1XvdvUeBnXmlld` | Domi - Energetic |
| `august` | `D38z5RcWu1voky8WS1ja` | Fin - Professional |
| `blue` | `2EiwWnXFnvU5JabPnv8n` | Clyde - Deep and authoritative |
| `coral` | `9BWtsMINqrJLrRacOk9x` | Aria - Expressive |
| `lily` | `EXAVITQu4vr4xnSDxMaL` | Sarah - Friendly |
| `onyx` | `29vD33N1CtxCmqQRPOHJ` | Drew - Strong |
| `sage` | `CwhRBWXzGAHq8TQ4Fs17` | Roger - Calm |
| `verse` | `CYw3kZ02Hs0563khs1Fj` | Dave - Conversational |
**Using Custom Voice IDs**: You can also pass any ElevenLabs voice ID directly. If the voice name is not in the mapping, LiteLLM will use it as-is:
```python showLineNumbers title="Using custom ElevenLabs voice ID"
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2",
input="Testing with a custom voice.",
voice="21m00Tcm4TlvDq8ikWAM", # Direct ElevenLabs voice ID
)
```
### Response Format Mapping
LiteLLM maps OpenAI response formats to ElevenLabs output formats:
| OpenAI Format | ElevenLabs Format |
|---------------|-------------------|
| `mp3` | `mp3_44100_128` |
| `pcm` | `pcm_44100` |
| `opus` | `opus_48000_128` |
You can also pass ElevenLabs-specific output formats directly using the `output_format` parameter.
### Supported Parameters
```python showLineNumbers title="All Supported Parameters"
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2", # Required
input="Text to convert to speech", # Required
voice="alloy", # Required: Voice selection (mapped or raw ID)
response_format="mp3", # Optional: Audio format (mp3, pcm, opus)
speed=1.0, # Optional: Speech speed (maps to voice_settings.speed)
# ElevenLabs-specific parameters (passed directly):
model_id="eleven_multilingual_v2", # Optional: Override model
voice_settings={ # Optional: Voice customization
"stability": 0.5,
"similarity_boost": 0.75,
"speed": 1.0
},
pronunciation_dictionary_locators=[ # Optional: Custom pronunciation
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
)
```
### LiteLLM Proxy
#### 1. Configure your proxy
```yaml showLineNumbers title="ElevenLabs TTS configuration in config.yaml"
model_list:
- model_name: elevenlabs-tts
litellm_params:
model: elevenlabs/eleven_multilingual_v2
api_key: os.environ/ELEVENLABS_API_KEY
general_settings:
master_key: your-master-key
```
#### 2. Make TTS requests
##### Simple Usage (OpenAI Parameters)
You can use standard OpenAI-compatible parameters without any provider-specific configuration:
```bash showLineNumbers title="Simple TTS request with curl"
curl http://localhost:4000/v1/audio/speech \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "elevenlabs-tts",
"input": "Testing ElevenLabs speech via the LiteLLM proxy.",
"voice": "alloy",
"response_format": "mp3"
}' \
--output speech.mp3
```
```python showLineNumbers title="Simple TTS with OpenAI SDK"
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:4000",
api_key="your-litellm-api-key"
)
response = client.audio.speech.create(
model="elevenlabs-tts",
input="Testing ElevenLabs speech via the LiteLLM proxy.",
voice="alloy",
response_format="mp3"
)
# Save audio
with open("speech.mp3", "wb") as f:
f.write(response.content)
```
##### Advanced Usage (ElevenLabs-Specific Parameters)
**Note**: When using the proxy, provider-specific parameters (like `pronunciation_dictionary_locators`, `voice_settings`, etc.) must be passed in the `extra_body` field.
```bash showLineNumbers title="Advanced TTS request with curl"
curl http://localhost:4000/v1/audio/speech \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "elevenlabs-tts",
"input": "Testing ElevenLabs speech via the LiteLLM proxy.",
"voice": "alloy",
"response_format": "pcm",
"extra_body": {
"pronunciation_dictionary_locators": [
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
"voice_settings": {
"speed": 1.1,
"stability": 0.5,
"similarity_boost": 0.75
}
}
}' \
--output speech.mp3
```
```python showLineNumbers title="Advanced TTS with OpenAI SDK"
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:4000",
api_key="your-litellm-api-key"
)
response = client.audio.speech.create(
model="elevenlabs-tts",
input="Testing ElevenLabs speech via the LiteLLM proxy.",
voice="alloy",
response_format="pcm",
extra_body={
"pronunciation_dictionary_locators": [
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
"voice_settings": {
"speed": 1.1,
"stability": 0.5,
"similarity_boost": 0.75
}
}
)
# Save audio
with open("speech.mp3", "wb") as f:
f.write(response.content)
```

View file

@ -103,6 +103,7 @@ litellm --config /path/to/config.yaml
| Azure AI Speech Service (AVA)| [Usage](../docs/providers/azure_ai_speech) |
| Vertex AI | [Usage](../docs/providers/vertex#text-to-speech-apis) |
| Gemini | [Usage](#gemini-text-to-speech) |
| ElevenLabs | [Usage](../docs/providers/elevenlabs#text-to-speech-tts) |
## `/audio/speech` to `/chat/completions` Bridge

View file

@ -0,0 +1,332 @@
"""
Elevenlabs Text-to-Speech transformation
Maps OpenAI TTS spec to Elevenlabs TTS API
"""
from typing import TYPE_CHECKING, Any, Dict, Optional, Tuple, Union
from urllib.parse import urlencode
import httpx
from httpx import Headers
import litellm
from litellm.types.utils import all_litellm_params
from litellm.llms.base_llm.chat.transformation import BaseLLMException
from litellm.llms.base_llm.text_to_speech.transformation import (
BaseTextToSpeechConfig,
TextToSpeechRequestData,
)
from litellm.secret_managers.main import get_secret_str
from ..common_utils import ElevenLabsException
if TYPE_CHECKING:
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
from litellm.types.llms.openai import HttpxBinaryResponseContent
else:
LiteLLMLoggingObj = Any
HttpxBinaryResponseContent = Any
class ElevenLabsTextToSpeechConfig(BaseTextToSpeechConfig):
"""
Configuration for ElevenLabs Text-to-Speech
Reference: https://elevenlabs.io/docs/api-reference/text-to-speech/convert
"""
TTS_BASE_URL = "https://api.elevenlabs.io"
TTS_ENDPOINT_PATH = "/v1/text-to-speech"
DEFAULT_OUTPUT_FORMAT = "pcm_44100"
VOICE_MAPPINGS = {
"alloy": "21m00Tcm4TlvDq8ikWAM", # Rachel
"amber": "5Q0t7uMcjvnagumLfvZi", # Paul
"ash": "AZnzlk1XvdvUeBnXmlld", # Domi
"august": "D38z5RcWu1voky8WS1ja", # Fin
"blue": "2EiwWnXFnvU5JabPnv8n", # Clyde
"coral": "9BWtsMINqrJLrRacOk9x", # Aria
"lily": "EXAVITQu4vr4xnSDxMaL", # Sarah
"onyx": "29vD33N1CtxCmqQRPOHJ", # Drew
"sage": "CwhRBWXzGAHq8TQ4Fs17", # Roger
"verse": "CYw3kZ02Hs0563khs1Fj", # Dave
}
# Response format mappings from OpenAI to ElevenLabs
FORMAT_MAPPINGS = {
"mp3": "mp3_44100_128",
"pcm": "pcm_44100",
"opus": "opus_48000_128",
# ElevenLabs does not support WAV, AAC, or FLAC formats.
}
ELEVENLABS_QUERY_PARAMS_KEY = "__elevenlabs_query_params__"
ELEVENLABS_VOICE_ID_KEY = "__elevenlabs_voice_id__"
def get_supported_openai_params(self, model: str) -> list:
"""
ElevenLabs TTS supports these OpenAI parameters
"""
return ["voice", "response_format", "speed"]
def _extract_voice_id(self, voice: str) -> str:
"""
Normalize the provided voice information into an ElevenLabs voice_id.
"""
normalized_voice = voice.strip()
mapped_voice = self.VOICE_MAPPINGS.get(normalized_voice.lower())
return mapped_voice or normalized_voice
def _resolve_voice_id(
self,
voice: Optional[Union[str, Dict[str, Any]]],
params: Dict[str, Any],
) -> str:
"""
Determine the ElevenLabs voice_id based on provided voice input or parameters.
"""
mapped_voice: Optional[str] = None
if isinstance(voice, str) and voice.strip():
mapped_voice = self._extract_voice_id(voice)
elif isinstance(voice, dict):
for key in ("voice_id", "id", "name"):
candidate = voice.get(key)
if isinstance(candidate, str) and candidate.strip():
mapped_voice = self._extract_voice_id(candidate)
break
elif voice is not None:
mapped_voice = self._extract_voice_id(str(voice))
if mapped_voice is None:
voice_override = params.pop("voice_id", None)
if isinstance(voice_override, str) and voice_override.strip():
mapped_voice = self._extract_voice_id(voice_override)
if mapped_voice is None:
raise ValueError(
"ElevenLabs voice_id is required. Pass `voice` when calling `litellm.speech()`."
)
return mapped_voice
def map_openai_params(
self,
model: str,
optional_params: Dict,
voice: Optional[Union[str, Dict]] = None,
drop_params: bool = False,
kwargs: Optional[Dict[str, Any]] = None,
) -> Tuple[Optional[str], Dict]:
"""
Map OpenAI parameters to ElevenLabs TTS parameters
"""
mapped_params: Dict[str, Any] = {}
query_params: Dict[str, Any] = {}
# Work on a copy so we don't mutate the caller's dictionary
params = dict(optional_params) if optional_params else {}
passthrough_kwargs: Dict[str, Any] = kwargs if kwargs is not None else {}
# Extract voice identifier
mapped_voice = self._resolve_voice_id(voice, params)
# Response/output format → query parameter
response_format = params.pop("response_format", None)
if isinstance(response_format, str):
mapped_format = self.FORMAT_MAPPINGS.get(response_format, response_format)
query_params["output_format"] = mapped_format
# ElevenLabs does not support OpenAI speed directly.
# Drop it to avoid sending unsupported keys unless caller already provided voice_settings.
speed = params.pop("speed", None)
if speed is not None:
speed_value: Optional[float]
try:
speed_value = float(speed)
except (TypeError, ValueError):
speed_value = None
if speed_value is not None:
if isinstance(params.get("voice_settings"), dict):
params["voice_settings"]["speed"] = speed_value # type: ignore[index]
else:
params["voice_settings"] = {"speed": speed_value}
# Instructions parameter is OpenAI-specific; omit to prevent API errors.
params.pop("instructions", None)
self._add_elevenlabs_specific_params(
mapped_voice=mapped_voice,
query_params=query_params,
mapped_params=mapped_params,
kwargs=passthrough_kwargs,
remaining_params=params,
)
return mapped_voice, mapped_params
def validate_environment(
self,
headers: dict,
model: str,
api_key: Optional[str] = None,
api_base: Optional[str] = None,
) -> dict:
"""
Validate Azure environment and set up authentication headers
"""
api_key = (
api_key
or litellm.api_key
or litellm.openai_key
or get_secret_str("ELEVENLABS_API_KEY")
)
if api_key is None:
raise ValueError(
"ElevenLabs API key is required. Set ELEVENLABS_API_KEY environment variable."
)
headers.update(
{
"xi-api-key": api_key,
"Content-Type": "application/json",
}
)
return headers
def get_error_class(
self, error_message: str, status_code: int, headers: Union[dict, Headers]
) -> BaseLLMException:
return ElevenLabsException(
message=error_message, status_code=status_code, headers=headers
)
def transform_text_to_speech_request(
self,
model: str,
input: str,
voice: Optional[str],
optional_params: Dict,
litellm_params: Dict,
headers: dict,
) -> TextToSpeechRequestData:
"""
Build the ElevenLabs TTS request payload.
"""
params = dict(optional_params) if optional_params else {}
extra_body = params.pop("extra_body", None)
request_body: Dict[str, Any] = {
"text": input,
"model_id": model,
}
for key, value in params.items():
if value is None:
continue
request_body[key] = value
if isinstance(extra_body, dict):
for key, value in extra_body.items():
if value is None:
continue
request_body[key] = value
return TextToSpeechRequestData(
dict_body=request_body,
headers={"Content-Type": "application/json"},
)
def _add_elevenlabs_specific_params(
self,
mapped_voice: str,
query_params: Dict[str, Any],
mapped_params: Dict[str, Any],
kwargs: Optional[Dict[str, Any]],
remaining_params: Dict[str, Any],
) -> None:
if kwargs is None:
kwargs = {}
for key, value in remaining_params.items():
if value is None:
continue
mapped_params[key] = value
reserved_kwarg_keys = set(all_litellm_params) | {
self.ELEVENLABS_QUERY_PARAMS_KEY,
self.ELEVENLABS_VOICE_ID_KEY,
"voice",
"model",
"response_format",
"output_format",
"extra_body",
"user",
}
extra_body_from_kwargs = kwargs.pop("extra_body", None)
if isinstance(extra_body_from_kwargs, dict):
for key, value in extra_body_from_kwargs.items():
if value is None:
continue
mapped_params[key] = value
for key in list(kwargs.keys()):
if key in reserved_kwarg_keys:
continue
value = kwargs[key]
if value is None:
continue
mapped_params[key] = value
kwargs.pop(key, None)
if query_params:
kwargs[self.ELEVENLABS_QUERY_PARAMS_KEY] = query_params
else:
kwargs.pop(self.ELEVENLABS_QUERY_PARAMS_KEY, None)
kwargs[self.ELEVENLABS_VOICE_ID_KEY] = mapped_voice
def transform_text_to_speech_response(
self,
model: str,
raw_response: httpx.Response,
logging_obj: LiteLLMLoggingObj,
) -> "HttpxBinaryResponseContent":
"""
Wrap ElevenLabs binary audio response.
"""
from litellm.types.llms.openai import HttpxBinaryResponseContent
return HttpxBinaryResponseContent(raw_response)
def get_complete_url(
self,
model: str,
api_base: Optional[str],
litellm_params: dict,
) -> str:
"""
Construct the ElevenLabs endpoint URL, including path voice_id and query params.
"""
base_url = (
api_base
or get_secret_str("ELEVENLABS_API_BASE")
or self.TTS_BASE_URL
)
base_url = base_url.rstrip("/")
voice_id = litellm_params.get(self.ELEVENLABS_VOICE_ID_KEY)
if not isinstance(voice_id, str) or not voice_id.strip():
raise ValueError(
"ElevenLabs voice_id is required. Pass `voice` when calling `litellm.speech()`."
)
url = f"{base_url}{self.TTS_ENDPOINT_PATH}/{voice_id}"
query_params = litellm_params.get(self.ELEVENLABS_QUERY_PARAMS_KEY, {})
if query_params:
url = f"{url}?{urlencode(query_params)}"
return url

View file

@ -5766,7 +5766,9 @@ def speech( # noqa: PLR0915
custom_llm_provider: Optional[str] = None,
aspeech: Optional[bool] = None,
**kwargs,
) -> HttpxBinaryResponseContent:
) -> Union[
HttpxBinaryResponseContent, Coroutine[Any, Any, HttpxBinaryResponseContent]
]:
user = kwargs.get("user", None)
litellm_call_id: Optional[str] = kwargs.get("litellm_call_id", None)
proxy_server_request = kwargs.get("proxy_server_request", None)
@ -5826,7 +5828,11 @@ def speech( # noqa: PLR0915
},
custom_llm_provider=custom_llm_provider,
)
response: Optional[HttpxBinaryResponseContent] = None
response: Union[
HttpxBinaryResponseContent,
Coroutine[Any, Any, HttpxBinaryResponseContent],
None,
] = None
if (
custom_llm_provider == "openai"
or custom_llm_provider in litellm.openai_compatible_providers
@ -5964,6 +5970,58 @@ def speech( # noqa: PLR0915
aspeech=aspeech,
litellm_params=litellm_params_dict,
)
elif custom_llm_provider == "elevenlabs":
from litellm.llms.elevenlabs.text_to_speech.transformation import (
ElevenLabsTextToSpeechConfig,
)
if text_to_speech_provider_config is None:
text_to_speech_provider_config = ElevenLabsTextToSpeechConfig()
elevenlabs_config = cast(
ElevenLabsTextToSpeechConfig, text_to_speech_provider_config
)
voice_id = voice if isinstance(voice, str) else None
if voice_id is None or not voice_id.strip():
raise litellm.BadRequestError(
message="'voice' must resolve to an ElevenLabs voice id for ElevenLabs TTS",
model=model,
llm_provider=custom_llm_provider,
)
voice_id = voice_id.strip()
query_params = kwargs.pop(
ElevenLabsTextToSpeechConfig.ELEVENLABS_QUERY_PARAMS_KEY, None
)
if isinstance(query_params, dict):
litellm_params_dict[
ElevenLabsTextToSpeechConfig.ELEVENLABS_QUERY_PARAMS_KEY
] = query_params
litellm_params_dict[
ElevenLabsTextToSpeechConfig.ELEVENLABS_VOICE_ID_KEY
] = voice_id
if api_base is not None:
litellm_params_dict["api_base"] = api_base
if api_key is not None:
litellm_params_dict["api_key"] = api_key
response = base_llm_http_handler.text_to_speech_handler(
model=model,
input=input,
voice=voice_id,
text_to_speech_provider_config=elevenlabs_config,
text_to_speech_optional_params=optional_params,
custom_llm_provider=custom_llm_provider,
litellm_params=litellm_params_dict,
logging_obj=logging_obj,
timeout=timeout,
extra_headers=extra_headers,
client=client,
_is_async=aspeech or False,
)
elif custom_llm_provider == "vertex_ai" or custom_llm_provider == "vertex_ai_beta":
generic_optional_params = GenericLiteLLMParams(**kwargs)

View file

@ -7865,6 +7865,12 @@ class ProviderConfigManager:
)
return AzureAVATextToSpeechConfig()
elif litellm.LlmProviders.ELEVENLABS == provider:
from litellm.llms.elevenlabs.text_to_speech.transformation import (
ElevenLabsTextToSpeechConfig,
)
return ElevenLabsTextToSpeechConfig()
elif litellm.LlmProviders.RUNWAYML == provider:
from litellm.llms.runwayml.text_to_speech.transformation import (
RunwayMLTextToSpeechConfig,

View file

@ -1,6 +1,8 @@
import os
import sys
from typing import Any, Dict
import pytest
from unittest.mock import patch, MagicMock
import httpx
@ -11,6 +13,8 @@ sys.path.insert(
import litellm
from base_audio_transcription_unit_tests import BaseLLMAudioTranscriptionTest
os.environ.setdefault("ELEVENLABS_API_KEY", "test-elevenlabs-key")
class TestElevenLabsAudioTranscription(BaseLLMAudioTranscriptionTest):
def get_base_audio_transcription_call_args(self) -> dict:
@ -108,4 +112,84 @@ class TestElevenLabsAudioTranscription(BaseLLMAudioTranscriptionTest):
except Exception as e:
print(f"❌ Test failed: {e}")
print(f"Captured request data: {captured_request_data}")
raise
raise
class TestElevenLabsTextToSpeechTransformation:
@pytest.fixture(scope="class")
def config(self):
from litellm.llms.elevenlabs.text_to_speech.transformation import (
ElevenLabsTextToSpeechConfig,
)
return ElevenLabsTextToSpeechConfig()
def test_map_openai_params_maps_voice_and_speed(self, config):
kwargs: Dict[str, Any] = {}
mapped_voice, mapped_params = config.map_openai_params(
model="eleven_multilingual_v2",
optional_params={
"response_format": "mp3",
"speed": 1.25,
"model_id": "eleven_multilingual_v2",
},
voice="alloy",
kwargs=kwargs,
)
assert mapped_voice == config.VOICE_MAPPINGS["alloy"]
assert mapped_params["voice_settings"]["speed"] == pytest.approx(1.25)
assert (
kwargs[config.ELEVENLABS_QUERY_PARAMS_KEY]["output_format"]
== "mp3_44100_128"
)
def test_transform_request_and_url(self, config):
kwargs: Dict[str, Any] = {}
voice_id, optional_params = config.map_openai_params(
model="eleven_multilingual_v2",
optional_params={
"response_format": "pcm",
"model_id": "eleven_multilingual_v2",
"pronunciation_dictionary_locators": [
{"pronunciation_dictionary_id": "dict_1"}
],
},
voice="alloy",
kwargs=kwargs,
)
litellm_params: Dict[str, Any] = {
config.ELEVENLABS_VOICE_ID_KEY: voice_id,
config.ELEVENLABS_QUERY_PARAMS_KEY: kwargs[
config.ELEVENLABS_QUERY_PARAMS_KEY
],
}
headers = config.validate_environment(
headers={}, model="eleven_multilingual_v2", api_key="test-key"
)
request_data = config.transform_text_to_speech_request(
model="eleven_multilingual_v2",
input="Hello world",
voice=voice_id,
optional_params=optional_params,
litellm_params=litellm_params,
headers=headers,
)
assert request_data["dict_body"]["text"] == "Hello world"
assert request_data["dict_body"]["model_id"] == "eleven_multilingual_v2"
assert request_data["dict_body"]["pronunciation_dictionary_locators"] == [
{"pronunciation_dictionary_id": "dict_1"}
]
url = config.get_complete_url(
model="eleven_multilingual_v2",
api_base=None,
litellm_params=litellm_params,
)
assert voice_id in url
assert "output_format=pcm_44100" in url