litellm/litellm/llms
Ishaan Jaff 29e3fd5d79
[Release Fix] (#22411)
* fix(lint): suppress PLR0915 for 3 complex methods that exceed 50-statement limit

- streaming_iterator.py: _process_event (84 statements)
- transformation.py: translate_messages_to_responses_input (51 statements)
- transformation.py: transform_realtime_response (54 statements)

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(mypy): resolve type errors in public_endpoints, user_api_key_auth, common_utils, transformation

- public_endpoints.py: fix _cached_endpoints type annotation
- user_api_key_auth.py: accept Optional[str] for end_user_id parameter
- common_utils.py: add NewProjectRequest/UpdateProjectRequest to Union type
- transformation.py: add ChatCompletionRedactedThinkingBlock and list[Any] to content type

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(proxy-extras): bump version to 0.4.50 and sync schema

- Bump litellm-proxy-extras from 0.4.49 to 0.4.50
- Sync schema.prisma with main proxy schema
- Includes new LiteLLM_ClaudeCodePluginTable model
- Includes new @@index([startTime, request_id]) on SpendLogs
- Update version references in requirements.txt and pyproject.toml

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(router): use string id in test_add_deployment and add defensive str() in register_model

- Change test to use string '100' instead of int 100 for model_info.id
- Add str() conversion in register_model to prevent AttributeError on non-string keys

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(security): update minimatch to 10.2.4 to fix CVE-2026-27903 and CVE-2026-27904

- Run npm audit fix in docs/my-website
- Updates minimatch from 10.2.1 to 10.2.4 (fixes HIGH severity ReDoS vulnerabilities)

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(test): update realtime guardrail test assertions to match actual guardrail behavior

- test_text_message_blocked_by_guardrail_no_ai_response: allow guardrail's own block
  message text in response.done (previously expected empty content)
- test_voice_transcript_blocked_by_guardrail: allow guardrail to send response.cancel
  + block message + response.create flow (previously expected no response.create)

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: revert proxy-extras version in requirements.txt and pyproject.toml

The litellm-proxy-extras 0.4.50 is not published to PyPI yet, so consumer
references must stay at 0.4.49. Only the source package pyproject.toml
should be bumped to 0.4.50 for the publish_proxy_extras CI job.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: make transcript delta check optional in voice guardrail test

The guardrail sends an error event (guardrail_violation) when blocking
voice transcripts; it does not always produce transcript deltas. Remove
the assertion requiring response.audio_transcript.delta since the error
event is the primary signal that blocked content was handled.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* Add missing env keys to documentation: LITELLM_MAX_STREAMING_DURATION_SECONDS and LITELLM_USE_CHAT_COMPLETIONS_URL_FOR_ANTHROPIC_MESSAGES

These two environment variables were used in code but not documented in the
environment variables reference section of config_settings.md, causing the
test_env_keys.py CI test to fail.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* Fix 13 mypy type errors across 6 files

- in_flight_requests_middleware.py: Fix type: ignore error codes from
  [union-attr] to [attr-defined], add [arg-type] for Gauge **kwargs
- transformation.py: Add [assignment] ignore for output_format reassignment,
  add fallback empty string for tool use id to fix arg-type
- responses/main.py: Remove redundant type annotation on second
  secret_fields assignment to fix no-redef
- streaming_iterator.py: Add [assignment] ignores for intermediate
  cache token assignments
- handler.py: Add [typeddict-item] ignore for AnthropicMessagesRequest
  construction from dict
- public_endpoints.py: Add [arg-type] ignore for _load_endpoints()
  return type mismatch with SupportedEndpoint model

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: add auth overrides to spend tracking tests, fix realtime guardrail assertion, update UI minimatch

- Add app.dependency_overrides for user_api_key_auth in 4 spend tracking tests
  that were returning 401 Unauthorized (error_code, error_message,
  error_code_and_key_alias, key_hash)
- Fix realtime guardrail test to check ANY error event for guardrail_violation
  instead of just the first (OpenAI may send its own errors first)
- Update ui/litellm-dashboard/package-lock.json to fix minimatch vulnerability

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* Fix failing MCP e2e and create_mcp_server UI tests

Test 1 (test_independent_clients_no_shared_session):
- Add allow_all_keys: true to MCP servers in test config. With master_key
  and no DB, get_allowed_mcp_servers returned empty, causing 0 tools and
  403 on tool calls. allow_all_keys bypasses per-key restrictions.
- Add asyncio.sleep(0.5) between client connections to allow MCP SDK
  TaskGroup cleanup and avoid ExceptionGroup on connection close (MCP #915).

Test 2 (create_mcp_server 'auth value is provided'):
- Use userEvent.setup({ delay: null }) for instant keystrokes to avoid
  timeout from default typing delay on CI.
- Increase per-test timeout to 15000ms for CI environments.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: stabilize proxy unit tests for parallel execution

- test_response_polling_handler: add xdist_group to prevent heavy import OOM
- test_db_schema_migration: use temp dir for worker isolation, sync schema.prisma index
- test_custom_tokenizer_bug: use lighter tokenizer to prevent OOM in parallel

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: add auth overrides to more spend tracking and model info tests

- Fix test_ui_view_spend_logs_pagination missing auth override (401)
- Fix test_view_spend_tags missing auth override (401)
- Fix test_view_spend_tags_no_database missing auth override (401)
- Fix test_empty_model_list.py to use app.dependency_overrides instead of patch()
  for FastAPI dependency injection auth

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(test): use patch.object for aiohttp transport test to work in parallel execution

The @patch decorator was not intercepting the static method call in parallel
xdist workers. Using patch.object on the directly-imported class is more reliable.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(security): update minimatch from 10.2.1 to 10.2.4 in Dockerfile

The Docker image was explicitly pinning minimatch@10.2.1 which has HIGH
severity ReDoS vulnerabilities (GHSA-7r86-cg39-jmmj, GHSA-23c5-xmqv-rm74).
Update to 10.2.4 which includes fixes for both CVEs.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(ui): prevent MCP and TeamInfo test timeouts on CI

- Add userEvent.setup({ delay: null }) to all tests using userEvent in both files
- Add timeout: 15000 to tests with significant user interaction (typing, multiple clicks)
- Fixes: create_mcp_server Bearer Token test, TeamInfo cancel button test

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix: stabilize parallel test execution and aiohttp transport test

- test_aiohttp_handler: rewrite transport test to not rely on static method mock
  (consistently fails in parallel xdist workers)
- test_proxy_cli: add xdist_group to prevent timeout during heavy imports
- test_swagger_chat_completions: add xdist_group to prevent timeout

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(security): add serialize-javascript override to fix GHSA-5c6j-r48x-rmvq

Add npm override for serialize-javascript>=7.0.3 in docs/my-website
to fix HIGH severity RCE vulnerability via RegExp.flags.
Also bump minimatch override to >=10.2.4.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* Fix flaky tests: remove broken Vertex model, add retries for Anthropic

- Remove vertex_ai/meta/llama-4-scout-17b-16e-instruct-maas from
  test_partner_models_httpx_streaming - consistently returns 400 BadRequest
- Add @pytest.mark.flaky(retries=6, delay=10) to test_function_call_parsing
  for transient Anthropic API overload errors
- Add @pytest.mark.flaky(retries=6, delay=10) to test_openai_stream_options_call
  for transient Anthropic InternalServerError

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(ci): add xdist_group(proxy_heavy) to prevent OOM in parallel proxy tests

- Add pytestmark = pytest.mark.xdist_group('proxy_heavy') to test_proxy_utils.py
- Change test_db_schema_migration.py from schema_migration to proxy_heavy group
- Add @pytest.mark.xdist_group('proxy_heavy') to test_proxy_server.py::test_health

Groups heavy proxy tests to run on same worker, avoiding worker OOM crashes.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* Fix vertex AI qwen global endpoint test to mock vertexai module import

The test_vertex_ai_qwen_global_endpoint_url test was failing because the
VertexAIPartnerModels.completion() method tries to 'import vertexai' before
any of the mocked code runs. In environments without google-cloud-aiplatform
installed, this import fails with a VertexAIError(status_code=400).

Fix by:
- Adding patch.dict('sys.modules', {'vertexai': MagicMock()}) to mock the
  vertexai module import
- Adding vertex_ai_location parameter to the acompletion call for completeness

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(ci): add xdist_group to health endpoint and watsonx tests for parallel stability

- test_health_liveliness_endpoint: add xdist_group('proxy_health') to prevent timeout
- test_watsonx_gpt_oss tests: add xdist_group('watsonx_heavy') to prevent mock interference

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(test): pre-populate WatsonX IAM token cache to prevent parallel test interference

The watsonx prompt transformation test was failing in parallel execution because
litellm.module_level_client.post mock was being interfered with by other tests.
Pre-populating the IAM token cache avoids the HTTP call entirely.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(test): add spend data polling with retries for e2e pass-through tests

- test_vertex_with_spend.test.js: Replace 15s fixed wait with polling loop
  (up to 6 attempts, 10s apart) for spend data to appear in DB
- Increase test timeout from 25s to 90s to accommodate polling
- base_anthropic_messages_tool_search_test.py: Add flaky(retries=3) for
  streaming test that depends on live Anthropic API

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(ci): reduce parallel workers from 8 to 4 for proxy tests to prevent OOM

- litellm_proxy_unit_testing_part2: -n 8 -> -n 4
- litellm_mapped_tests_proxy_part2: -n 8 -> -n 4, timeout 60 -> 120
- Worker crashes consistently caused by too many parallel proxy tests
  each loading the full FastAPI app and heavy dependency tree

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(db): add migration for SpendLogs composite index (startTime, request_id)

The @@index([startTime, request_id]) was added to schema.prisma but had no
corresponding migration. This caused test_aaaasschema_migration_check to fail
because prisma migrate diff detected the missing index.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(db): add migration for MCP available_on_public_internet default change to true

The schema.prisma changed the default for available_on_public_internet from
false to true, but no migration was created. This caused the schema migration
test to detect drift.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(test): increase server wait time and add retry to flaky external API tests

- test_basic_python_version.py: increase server startup wait from 60s to 90s
  for slower CI environments (fixes installing_litellm_on_python_3_13)
- test_a2a_agent.py: add flaky(retries=3, delay=5) for non-streaming test
  that depends on live A2A agent endpoint

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(test): add flaky retries to all intermittent external API tests for 0-fail CI

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* fix(test): add auth overrides to file endpoint tests that return 500

The test_target_storage tests were getting 500 because the FastAPI auth
dependency wasn't overridden. Added app.dependency_overrides for proper
auth bypass in test environment.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-02-28 09:46:35 -08:00
..
a2a Agent Guardrails - on streaming output (#21206) 2026-02-14 11:36:52 -08:00
ai21/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
aiml fix ai/ml api 2025-11-26 18:55:55 -08:00
aiohttp_openai/chat VertexAI non-jsonl file storage support (#9781) 2025-04-09 14:01:48 -07:00
amazon_nova amazon nova api fix 2025-12-06 16:09:27 -08:00
anthropic [Release Fix] (#22411) 2026-02-28 09:46:35 -08:00
aws_polly [Feat] New provider TTS - Add AWS polly API for TTS (#18326) 2025-12-22 18:19:34 +05:30
azure fix(realtime): guardrails with pre_call/post_call mode now work on realtime WebSocket (#22161) 2026-02-25 23:43:13 -08:00
azure_ai fix: support Azure AD token auth for non-Claude azure_ai models (#20981) 2026-02-11 10:48:44 -08:00
base_llm Enable local file support for OCR (#22133) 2026-02-27 10:50:02 -08:00
baseten lint 2025-08-19 13:36:41 -07:00
bedrock fix: use correct class name AmazonConverseConfig in helper method calls 2026-02-28 00:18:56 -03:00
brave/search add search provider for brave search api (#19433) 2026-01-20 19:23:29 -08:00
bytez ruff check ./litellm --fix 2025-07-12 11:04:02 -07:00
cerebras fix: add reasoning param support for GPT OSS cerebras 2026-02-02 17:20:04 +05:30
chatgpt fix(chatgpt): drop unsupported responses params for Codex 2026-02-15 00:18:02 +05:30
clarifai/chat UI - add arize on ui, LLMs - clarifai refactor to openai compatible route, added azure ai/grok-4 model family 2025-10-16 20:39:15 -07:00
cloudflare/chat VertexAI non-jsonl file storage support (#9781) 2025-04-09 14:01:48 -07:00
codestral/completion Codestral - return litellm latency overhead on /v1/completions + Add '__contains__' support for ChatCompletionDeltaToolCall (#10879) 2025-05-27 16:13:44 -07:00
cohere Feature/guardrail model argument (#19619) 2026-01-23 20:48:42 -08:00
cometapi feat(cometapi): Add CometAPI provider support (embeddings, image generation, docs) 2025-10-16 13:08:14 +08:00
compactifai Fix CompactifAI provider tests and implementation 2025-09-15 22:03:42 +02:00
custom_httpx fix(realtime): fix guardrails not firing for Gemini/Vertex AI and provider_config realtime WebSocket sessions (#22168) 2026-02-26 17:06:07 -08:00
dashscope fix: remove list-to-str transformation from dashscope 2026-02-19 07:30:13 +00:00
databricks Address Greptile review: fix SDK auth fallback and remove unused imports 2026-02-18 15:57:50 +09:00
dataforseo/search [Feat] UI - Search Tools, allow adding search tools on UI + testing search (#15871) 2025-10-23 17:59:29 -07:00
datarobot/chat Updated URL handling for DataRobot provider base 2025-08-21 19:42:46 -06:00
deepgram fix mypy lint 2025-11-03 18:02:19 -08:00
deepinfra fix mypy error 2026-01-08 15:38:28 +05:30
deepseek feat(deepseek): add native support for thinking and reasoning_effort params (#17712) 2025-12-11 15:28:43 -08:00
deprecated_providers fix(mypy): resolve type checking errors in 5 files (#20627) 2026-02-06 18:34:55 -08:00
docker_model_runner/chat [Feat] New LLM Provider - Docker Model Runner (#16948) 2025-11-21 16:09:32 -08:00
duckduckgo/search Add duckcukgo in model map 2026-02-18 16:13:20 +05:30
elevenlabs Integrate eleven labs text-to-speech (#16573) 2025-11-24 18:49:30 -08:00
empower/chat LiteLLM Common Base LLM Config (pt.3): Move all OAI compatible providers to base llm config (#7148) 2024-12-10 17:12:42 -08:00
exa_ai/search [Feat] UI - Search Tools, allow adding search tools on UI + testing search (#15871) 2025-10-23 17:59:29 -07:00
fal_ai fix QA check 2025-11-15 13:02:48 -08:00
featherless_ai/chat fix: fix model param mapping 2025-05-19 20:24:47 -07:00
firecrawl [Feat] /search API - add firecrawl search API support (#16257) 2025-11-04 17:52:12 -08:00
fireworks_ai Fix mypy type error in fireworks_ai transformation (#20391) 2026-02-04 15:44:45 -08:00
friendliai/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
galadriel/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
gemini fix(lint): suppress PLR0915 too-many-statements in complex transform methods 2026-02-27 23:18:06 -03:00
gigachat bugfix: Remove user messages merging 2026-02-03 12:58:28 +00:00
github/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
github_copilot Merge branch 'main' into ttl-prompt-caching-bedrock 2026-02-05 16:50:54 +05:30
google_pse/search [Feat] UI - Search Tools, allow adding search tools on UI + testing search (#15871) 2025-10-23 17:59:29 -07:00
gradient_ai/chat Add digitalocean provider (#12169) 2025-08-09 16:26:33 -07:00
groq Add grok reasoning content 2026-01-27 16:34:57 +05:30
heroku/chat fixes linter error 2025-07-28 09:44:02 -06:00
hosted_vllm Convert thinking_blocks to content blocks for hosted_vllm multi-turn 2026-02-19 12:25:18 +00:00
huggingface Use vertex creds passed via arguments (#16266) 2025-11-06 19:35:22 -08:00
hyperbolic Revert "Litellm dev 07 21 2025 p1 (#12848)" 2025-07-22 18:28:36 -07:00
infinity Use vertex creds passed via arguments (#16266) 2025-11-06 19:35:22 -08:00
jina_ai Use vertex creds passed via arguments (#16266) 2025-11-06 19:35:22 -08:00
lambda_ai Revert "Litellm dev 07 21 2025 p1 (#12848)" 2025-07-22 18:28:36 -07:00
langgraph [Fix] CI/CD – Clean Up Performance PR Changes & others (#17838) 2025-12-11 12:50:03 -08:00
lemonade Removing unecessary import 2025-09-30 12:12:24 -06:00
linkup [Feat] New Search API Provider - LinkUp Search (#18174) 2025-12-18 14:27:36 +05:30
litellm_proxy [Feat] Unified Skills API - works across Anthropic, Vertex, Azure, Bedrock (#18232) 2025-12-19 18:55:59 +05:30
llamafile/chat Add llamafile as a provider (#10203) (#10482) 2025-05-01 18:36:55 -07:00
lm_studio fix(lm_studio): resolve illegal Bearer header value issue 2025-09-12 22:41:30 +02:00
manus fix manus 2026-01-10 14:14:40 -08:00
meta_llama/chat [Feat] Enable Tool Calling for meta_llama (#11895) 2025-06-19 13:44:22 -07:00
milvus/vector_stores docs: document milvus endpoints 2025-11-01 12:17:02 -07:00
minimax Add Prompt caching and reasoning support for MiniMax, GLM, Xiaomi 2026-01-28 17:25:26 +05:30
mistral Added thinking streaming support for mistral (#16434) 2025-11-10 18:41:45 -08:00
moonshot/chat Fix MoonshotChatConfig to address limitations of kimi-thinking-preview model by excluding additional parameters (#12772) 2025-07-19 15:10:57 -07:00
morph fix morph api tests 2025-07-22 18:44:44 -07:00
nebius Integration with Nebius AI Studio added (#11143) 2025-05-27 11:05:22 -07:00
nlp_cloud VertexAI non-jsonl file storage support (#9781) 2025-04-09 14:01:48 -07:00
novita/chat Add new model provider Novita AI (#7582) (#9527) 2025-05-12 21:49:30 -07:00
nscale/chat Add nscale support for streaming (#10698) 2025-05-09 11:39:23 -07:00
nvidia_nim Fix nvdia and geminin tests 2025-12-10 22:05:11 +05:30
oci Fixes #20957 2026-02-11 11:20:18 +00:00
ollama fix(ollama): thread api_base to get_model_info + graceful fallback (#21970) 2026-02-23 21:00:37 -08:00
oobabooga VertexAI non-jsonl file storage support (#9781) 2025-04-09 14:01:48 -07:00
openai Merge pull request #22026 from ZeroClover/fix/img-extra-headers 2026-02-28 00:14:43 -03:00
openai_like docs: update AssemblyAI docs with Universal-3 Pro, Speech Understanding, and LLM Gateway (#21130) 2026-02-27 17:24:48 -08:00
openrouter Add Prompt caching and reasoning support for MiniMax, GLM, Xiaomi 2026-01-28 17:25:26 +05:30
ovhcloud [Refactor#2] litellm/init – Lazy-load utils to reduce memory + import time (#17171) 2025-12-03 11:40:16 -08:00
parallel_ai/search [Feat] UI - Search Tools, allow adding search tools on UI + testing search (#15871) 2025-10-23 17:59:29 -07:00
pass_through Feature/guardrail model argument (#19619) 2026-01-23 20:48:42 -08:00
perplexity lint 2026-02-23 18:05:13 +00:00
petals VertexAI non-jsonl file storage support (#9781) 2025-04-09 14:01:48 -07:00
pg_vector/vector_stores [Feat] UI Vector Stores - Allow adding Vertex RAG Engine, OpenAI, Azure (#12752) 2025-07-18 18:25:26 -07:00
predibase VertexAI non-jsonl file storage support (#9781) 2025-04-09 14:01:48 -07:00
ragflow Fix unused imports 2025-12-03 15:32:42 +05:30
recraft Fix: stability image optional para 2026-01-19 09:05:52 +05:30
replicate Fix Output None for replicate handler 2026-01-19 17:22:06 +05:30
runwayml fix(ollama): thread api_base to get_model_info + graceful fallback (#21970) 2026-02-23 21:00:37 -08:00
s3_vectors [Feat] RAG API - Add s3_vectors as provider on /vector_store/search API + UI for creating + PDF support for /rag/ingest (#19895) 2026-01-27 16:30:59 -08:00
sagemaker fix(sagemaker): Support TEI raw array response format for embeddings (#20487) 2026-02-12 19:39:05 +05:30
sambanova Fix : acompletion throws error with SambaNova models (#17217) 2025-11-27 21:59:24 -08:00
sap fix(sap): correct filtering of optional_params (#18716) 2026-01-07 21:48:29 +05:30
searxng [Feat] add serxng search API provider (#16259) 2025-11-04 17:56:07 -08:00
snowflake Snowflake provider support: added embeddings, PAT, account_id (#15727) 2025-11-17 20:27:46 -08:00
stability Fix: stability image optional para 2026-01-19 09:05:52 +05:30
tavily/search [Feat] UI - Search Tools, allow adding search tools on UI + testing search (#15871) 2025-10-23 17:59:29 -07:00
together_ai [Refactor#2] litellm/init – Lazy-load utils to reduce memory + import time (#17171) 2025-12-03 11:40:16 -08:00
topaz Add /vllm/* and /mistral/* passthrough endpoints (adds support for Mistral OCR via passthrough) 2025-04-14 22:06:33 -07:00
triton fix(triton/completion/transformation.py): remove bad_words / stop wor… (#10163) 2025-04-19 11:23:37 -07:00
v0 feat: add v0 provider support (#12751) 2025-07-18 18:26:44 -07:00
vercel_ai_gateway feat(vercel_ai_gateway): add embeddings support 2026-01-23 15:11:37 -03:00
vertex_ai Merge pull request #22223 from emerzon/feat/vertex-gemini-3-1-flash-image-preview-pricing 2026-02-27 18:10:04 +05:30
vllm fix vllm passthrough 2025-09-22 10:09:02 -03:00
volcengine feat (volcengine) : Support Volcengine responses api (#18508) 2026-01-19 19:02:29 -08:00
voyage [Fix] CI/CD – Clean Up Performance PR Changes & others (#17838) 2025-12-11 12:50:03 -08:00
wandb (feat): Add W&B Inference to LiteLLM 2025-09-11 00:07:30 +05:30
watsonx feat: Add IBM watsonx.ai rerank support (#21303) 2026-02-16 20:12:16 -08:00
xai Fix usage in xai 2026-02-19 18:48:30 +05:30
xinference/image_generation [Feat] Add XInference Image Generation API Provider (#12439) 2025-07-08 21:17:38 -07:00
zai Add Prompt caching and reasoning support for MiniMax, GLM, Xiaomi 2026-01-28 17:25:26 +05:30
__init__.py fix: move code from litellm/llms to the mcp_server dir 2026-01-05 12:05:16 +09:00
base.py Gemini - web search cost tracking + Update max output tokens for nova models 2025-06-05 23:25:18 -07:00
custom_llm.py Fix mypy issues 2026-01-16 14:47:56 +05:30
maritalk.py build(pyproject.toml): add new dev dependencies - for type checking (#9631) 2025-03-29 11:02:13 -07:00
README.md LiteLLM Minor Fixes and Improvements (09/13/2024) (#5689) 2024-09-14 10:02:55 -07:00

File Structure

August 27th, 2024

To make it easy to see how calls are transformed for each model/provider:

we are working on moving all supported litellm providers to a folder structure, where folder name is the supported litellm provider name.

Each folder will contain a *_transformation.py file, which has all the request/response transformation logic, making it easy to see how calls are modified.

E.g. cohere/, bedrock/.