mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-23 00:41:40 +00:00
* feat(audio_transcription): add NVIDIA Riva STT provider Adds nvidia_riva as a new audio transcription provider, supporting both NVCF-hosted and self-hosted Riva ASR deployments via gRPC streaming. - Auto-resamples input audio to 16 kHz mono LINEAR_PCM (soundfile + numpy, audioread fallback) so callers can send any common format. - Maps OpenAI params: language (en -> en-US), response_format (text/json/ verbose_json), timestamp_granularities=["word"] -> enable_word_time_offsets, word offsets converted ms -> s for verbose_json. - Auth: NVCF when nvcf_function_id is set (SSL on by default), self-hosted otherwise (SSL off by default), with explicit use_ssl override. - gRPC errors wrapped via NvidiaRivaException -> litellm exception classes. - Optional deps gated behind [stt-nvidia-riva] extra (nvidia-riva-client, soundfile, audioread, numpy). Co-authored-by: Cursor <cursoragent@cursor.com> * fix(nvidia_riva): address PR review feedback - handler: forward call-level `timeout` to streaming_response_generator (kwarg-detected via inspect for older riva-client compat) so a stalled Riva server cannot block the caller indefinitely. - audio_utils: spill bytes to a tempfile before audioread.audio_open; most audioread backends (FFmpeg, GStreamer) require a real filesystem path and previously raised TypeError on BytesIO, breaking the mp3/m4a fallback path. - audio_utils: prefer soxr / scipy.signal.resample_poly for resampling (anti-aliased polyphase) when installed, falling back to linear only as a last resort. Avoids aliasing on 44.1/48 kHz -> 16 kHz downsamples. - transformation: bare `es` now maps to es-ES (Castilian) instead of es-US, matching BCP-47 conventions. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: trigger CI re-run [stabilize loop 1/3] * Update litellm/llms/nvidia_riva/audio_transcription/transformation.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * chore: trigger CI re-run [stabilize loop 1/3] * fix code qa * fix lint * fix mypy * fix mypy * Fix NVIDIA Riva ASR service lookup * Fix NVIDIA Riva transcription payload logging --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: oss-pr-review-agent-shin[bot] <281797381+oss-pr-review-agent-shin[bot]@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| audio_utils | ||
| llm_cost_calc | ||
| llm_response_utils | ||
| prompt_templates | ||
| specialty_caches | ||
| tokenizers | ||
| api_route_to_call_types.py | ||
| app_crypto.py | ||
| asyncify.py | ||
| cached_imports.py | ||
| cli_token_utils.py | ||
| cloud_storage_security.py | ||
| completion_timeout.py | ||
| core_helpers.py | ||
| coroutine_checker.py | ||
| credential_accessor.py | ||
| custom_logger_registry.py | ||
| dd_tracing.py | ||
| default_encoding.py | ||
| dot_notation_indexing.py | ||
| duration_parser.py | ||
| env_utils.py | ||
| exception_mapping_utils.py | ||
| fallback_utils.py | ||
| get_blog_posts.py | ||
| get_litellm_params.py | ||
| get_llm_provider_logic.py | ||
| get_model_cost_map.py | ||
| get_provider_specific_headers.py | ||
| get_supported_openai_params.py | ||
| health_check_helpers.py | ||
| health_check_utils.py | ||
| initialize_dynamic_callback_params.py | ||
| json_validation_rule.py | ||
| litellm_logging.py | ||
| llm_request_utils.py | ||
| logging_callback_manager.py | ||
| logging_utils.py | ||
| logging_worker.py | ||
| mock_functions.py | ||
| model_param_helper.py | ||
| model_response_utils.py | ||
| README.md | ||
| realtime_streaming.py | ||
| redact_messages.py | ||
| response_header_helpers.py | ||
| rules.py | ||
| safe_json_dumps.py | ||
| safe_json_loads.py | ||
| secret_redaction.py | ||
| sensitive_data_masker.py | ||
| streaming_chunk_builder_utils.py | ||
| streaming_handler.py | ||
| thread_pool_executor.py | ||
| token_counter.py | ||
| url_utils.py | ||
Folder Contents
This folder contains general-purpose utilities that are used in multiple places in the codebase.
Core files:
streaming_handler.py: The core streaming logic + streaming related helper utilscore_helpers.py: code used intypes/- e.g.map_finish_reason.exception_mapping_utils.py: utils for mapping exceptions to openai-compatible error types.default_encoding.py: code for loading the default encoding (tiktoken)get_llm_provider_logic.py: code for inferring the LLM provider from a given model name.duration_parser.py: code for parsing durations - e.g. "1d", "1mo", "10s"api_route_to_call_types.py: mapping of API routes to their corresponding CallTypes (e.g.,/chat/completions-> [acompletion, completion])