Commit graph

34749 commits

Author SHA1 Message Date
Cesar Garcia
e4fddb9f24
Merge pull request #23093 from MaxwellCalkin/fix/thinking-blocks-interleave-23047
fix: preserve thinking block order with multiple web searches
2026-03-10 18:06:45 -03:00
Cesar Garcia
b905e1493b
Merge pull request #23201 from Chesars/claude/brave-ritchie
feat(images): support input_fidelity parameter for image edit API
2026-03-10 18:05:02 -03:00
Cesar Garcia
d34999900c
Merge pull request #23265 from Chesars/fix/vertex-gemini2-tool-schema-minimal-transform
fix(vertex): skip schema transforms for Gemini 2.0+ tool parameters
2026-03-10 18:04:46 -03:00
stevejaker
2341a38c08
fix(snowflake): transform tool_choice string to object format (#23268)
* fix(snowflake): transform tool_choice string to object format

Snowflake's Cortex API requires tool_choice to be an object, not a string.
For example, {"type": "auto"} instead of "auto".

Ref: https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/cortex-inference#post--api-v2-cortex-inference-complete-req-body-schema

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 01:41:24 +05:30
Chesars
3fe4829676 fix: include redacted_thinking in list-content thinking block detection
The _list_has_thinking guard only checked for type == "thinking" but
Anthropic can also return redacted_thinking blocks (safety-filtered).
These are also accumulated in thinking_blocks, so the same duplication
bug would occur with redacted thinking content.
2026-03-10 16:37:06 -03:00
Jason Roberts
70fca22f68
feat(panw-prisma-airs): PANW Prisma AIRS guardrail with apply_guardrail support (#22999)
* feat(panw-prisma-airs): PANW Prisma AIRS guardrail with apply_guardrail support

* fix(panw): honor masking and fallback behavior

* fix(panw): clean up apply_guardrail MCP metadata handling

* fix(panw): clean up apply_guardrail MCP metadata handling

* fix(panw): harden apply_guardrail edge cases

* fix(panw): apply MCP masked data on allow responses

* fix(panw): scan latest developer message in anthropic mode

* fix(panw): restore legacy user-only pre-call scanning

* fix(panw): record apply_guardrail in applied guardrails header

* fix(panw): scan developer role in legacy pre-call path

* fix(panw): harden SSE parsing and narrow MCP name fallback

* fix(panw): harden streaming attr lookup and document dual scans

* fix(panw): fail closed on permanent 4xx and cover streaming observability
2026-03-10 12:31:31 -07:00
xykong
810de556bd
fix(streaming): map unknown finish_reason values to finish_reason_unspecified to prevent ValidationError in stream_chunk_builder (#22673)
* fix(streaming): map unknown finish_reason values to finish_reason_unspecified

Some LLM providers return non-standard finish_reason values that are not
in the OpenAIChatCompletionFinishReason Literal (e.g. ZhipuAI/GLM returns
'network_error' when a streaming error occurs mid-response).

Previously map_finish_reason() fell through with return finish_reason,
passing the unknown value directly to Choices.__init__() which calls
Pydantic validation. This caused a ValidationError that was caught by
stream_chunk_builder() and re-raised as the misleading:
  litellm.APIError: Error building chunks for logging/streaming usage calculation

Fix: after all known provider-specific mappings, check if the value is in
the valid set (stop, length, tool_calls, content_filter, function_call,
guardrail_intervened, eos, finish_reason_unspecified, malformed_function_call).
Any value not in this set is mapped to 'finish_reason_unspecified' instead
of being returned as-is.

This is consistent with how other unknown stop reasons (e.g. Vertex AI's
FINISH_REASON_UNSPECIFIED) are already handled.

* refactor: use get_args(OpenAIChatCompletionFinishReason) for valid set

Per code review feedback: replace the hardcoded _valid_finish_reasons set
with a module-level frozenset derived dynamically from the source-of-truth
Literal type via typing.get_args(). This ensures the valid-reason check
stays in sync automatically when new finish reasons are added to the Literal,
and avoids recreating the set on every streaming chunk call.

* test(map_finish_reason): add unit tests and warning log for unknown finish reasons

- Add TestMapFinishReason class in test_core_helpers.py covering:
  - All known OpenAI-native values pass through unchanged (parametrized)
  - Provider-specific mappings: Anthropic, Cohere, Vertex AI
  - Unknown/provider-specific values map to 'finish_reason_unspecified'
  - Regression test for ZhipuAI/GLM-5 'network_error' case
- Add verbose_logger.warning() in map_finish_reason() when an unknown
  finish_reason is encountered, so operators can track which providers
  return non-standard values
2026-03-10 21:25:24 +05:30
Chesars
0680a97409 fix: handle list-content messages in thinking block interleaving
When assistant content is already a list containing thinking blocks
inline (not str/None), SEQUENTIAL MODE was still prepending all
thinking_blocks from provider_specific_fields, causing duplication
and breaking Anthropic's position-dependent signature verification.

Now detects if the content list already has thinking blocks and skips
the extend(thinking_blocks) to preserve the original interleaved order.

Addresses the correctness gap identified by Greptile review where
list-content messages bypass INTERLEAVED MODE.

Fixes: https://github.com/BerriAI/litellm/issues/23047
2026-03-10 12:48:30 -03:00
Carlo Alberto Ferraris
323b473835
fix: add missing indexes for top CPU-consuming queries (#23147)
* fix: add missing indexes for top CPU-consuming queries

Add indexes to eliminate full table scans on two of the top 5 queries
by CPU usage:

1. LiteLLM_VerificationToken(key_alias) — for ORDER BY key_alias ASC
   queries when listing verification tokens
2. LiteLLM_SpendLogs(user, startTime) — for WHERE user = $1 AND
   startTime BETWEEN $2 AND $3 GROUP BY queries on the spend logs

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: use CREATE INDEX CONCURRENTLY to avoid table locks

Both indexes are now created with CONCURRENTLY and IF NOT EXISTS
to avoid blocking writes on large production tables.
Uses -- SkipTransactionBlock for Prisma migrate compatibility.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 21:00:22 +05:30
Chesars
08d81f5d7c fix(vertex): shallow copy parameters before mutating in _build_vertex_schema_for_gemini_2
Avoids silently removing $defs from the caller's dict, which could
affect logging, caching, or retry logic referencing the same object.
2026-03-10 11:29:59 -03:00
Aarish Alam
2b093aa796
Merge pull request #23196 from CAFxX/docs/claude-md-db-performance-guidelines
docs: add DB query performance guidelines to CLAUDE.md
2026-03-10 19:51:51 +05:30
Chesars
a9c3095cc5 fix(vertex): skip harmful schema transforms for Gemini 2.0+ tool parameters
Gemini 2.0+ natively accepts JSON Schema in tool parameters, including
bare {} (TYPE_UNSPECIFIED), anyOf with null, and lowercase types. The
existing _build_vertex_schema pipeline was coercing {} to {"type": "object"},
breaking JsonValue/Any field semantics (issue #22391).

Add _build_vertex_schema_for_gemini_2() that only resolves $ref (which
Gemini doesn't support in tools) and filters unsupported fields. Use it
for Gemini 2.0+ models, keeping the full transform for Gemini 1.5.
2026-03-10 11:15:27 -03:00
Carlo Alberto Ferraris
9d2b0117d9
docs: add DB performance guidelines to CLAUDE.md
Extend the "Proxy database access" section with guidelines to prevent
common DB performance issues, tailored to actual Prisma usage patterns
in the litellm codebase: N+1 queries, client-side processing, batching
writes, bounding result sets, select on wide tables, index coverage,
and schema file sync.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 13:40:13 +09:00
yuneng-jiang
0e2aa7a5b2
Merge pull request #23213 from BerriAI/litellm_fix_flaky_audio_stream_test
[Fix] Flaky test_stream_chunk_builder_openai_audio_output_usage
2026-03-09 17:32:17 -07:00
yuneng-jiang
c1d042c2a3 Fix flaky test_stream_chunk_builder_openai_audio_output_usage
The test calls OpenAI's gpt-4o-audio-preview model which sometimes
doesn't return usage data in the streaming response. Fixed by:
- Adding @pytest.mark.flaky(retries=5, delay=2) for retry handling
- Fixing usage_obj loop to check chunk.usage is not None
- Skipping gracefully when OpenAI doesn't return usage data

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 17:18:00 -07:00
yuneng-jiang
6fe82d3886
Merge pull request #23211 from BerriAI/litellm_/sharp-keller
[Fix] Skills API test failing with duplicate skill name 500
2026-03-09 17:13:29 -07:00
yuneng-jiang
af8f91ef66 [Fix] Use unique skill names in Skills API test to avoid duplicate-name 500s
The test_create_skill test was consistently failing in CI with a 500 from
Anthropic because the SKILL.md frontmatter always used the same hardcoded
name (test-skill-litellm). Since test_delete_skill is permanently skipped,
skills accumulate in the CI account, and re-creating with a duplicate name
triggers an Internal Server Error on Anthropic's side.

Fix: pass a timestamp-based unique_suffix to create_skill_zip so each run
produces a distinct skill name in the zip's SKILL.md frontmatter.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 17:09:15 -07:00
yuneng-jiang
4c3f873bde
Merge pull request #23198 from BerriAI/litellm_fix_nova_pro_max_tokens
[Fix] Claude Agent SDK E2E Test Nova Pro max_tokens Limit
2026-03-09 15:54:00 -07:00
yuneng-jiang
4c5659ff30
Merge pull request #23199 from BerriAI/doc_per_model_input_tokenc_check
[doc improvement] input token check
2026-03-09 15:53:15 -07:00
yuneng-jiang
dda0146a66
Merge pull request #23197 from BerriAI/litellm_fix_flaky_watsonx_prompt_test
[Fix] Flaky test_watsonx_gpt_oss_prompt_transformation
2026-03-09 15:52:02 -07:00
Chesars
2ea2660ba5 feat(images): support input_fidelity parameter for image edit API
Add input_fidelity parameter ("high"/"low") to the image edit pipeline,
allowing users to control how much effort the model exerts to match
input image style and features. Fixes #22813.
2026-03-09 19:51:31 -03:00
yuneng-jiang
d719c8a53c
Merge branch 'main' into litellm_fix_nova_pro_max_tokens 2026-03-09 15:47:53 -07:00
yuneng-jiang
2a836c7103 Fix Claude Agent SDK E2E test for Nova Pro max_tokens limit
The Claude Agent SDK sends max_tokens=32000 for unrecognized model names
(like "bedrock-nova-pro"), which exceeds Nova Pro's 10,000 limit. Enable
modify_params in the test proxy config so LiteLLM clamps max_tokens to the
model's actual limit. Also swap nova-premier to nova-pro since premier
requires provisioned throughput unavailable in CI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:45:24 -07:00
shivam
5534f77314 doc improvement 2026-03-09 15:39:27 -07:00
yuneng-jiang
ffd1eb18e0 Merge remote main and resolve conflicts
Kept our sync test fix, accepted upstream's xdist_group marker on
the async handler test.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:34:50 -07:00
yuneng-jiang
74ed6a16ac Fix flaky test_watsonx_gpt_oss_prompt_transformation
The test was flaky under pytest-xdist parallel execution because it used
async acompletion (which runs completion() in a thread pool via
run_in_executor) and relied on shared global state (known_tokenizer_config,
iam_token_cache, module_level_client) that could be modified by other tests
running in parallel. Failures were silently swallowed by a broad try/except,
causing mock_post.call_count to remain 0.

Fix:
- Convert from async acompletion to sync completion, matching every other
  test in the file. The test's intent is verifying prompt transformation,
  not async behavior.
- Use monkeypatch.setitem for known_tokenizer_config to ensure proper
  teardown isolation.
- Remove unnecessary mock layers (async template fetchers, iam_token_cache
  pre-population, mock completion response) that were only needed for the
  async code path.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:32:30 -07:00
yuneng-jiang
a5ad414ae0
Merge pull request #23194 from BerriAI/litellm_/recursing-blackburn
[Fix] Batch retrieve missing model_id causing raw output_file_id
2026-03-09 15:28:59 -07:00
yuneng-jiang
4888a31e4f Fix batch retrieve not setting model_id, causing output_file_id to stay raw
When retrieving a batch via the unified batch ID path, only unified_batch_id
was set on _hidden_params but model_id was missing. The managed files hook
requires both to encode output_file_id into a managed ID.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:24:50 -07:00
yuneng-jiang
db77976b4e
Merge pull request #23192 from BerriAI/litellm_/pedantic-easley
[Test] Replace SearXNG integration tests with unit tests
2026-03-09 15:20:47 -07:00
yuneng-jiang
29ca052064 Merge remote main, resolve conflict keeping new unit tests
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:20:20 -07:00
yuneng-jiang
b7ac688b2b Replace SearXNG integration tests with unit tests for request/response transformation
The SearXNG search tests were failing in CI because they depend on a live
SearXNG instance that returns results. Since this provider is used by a
very small subset of customers, replace the flaky integration tests with
deterministic unit tests that validate request payloads, URL construction,
response parsing, and header configuration without requiring external infra.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 15:13:58 -07:00
yuneng-jiang
507c772370
Merge pull request #23190 from BerriAI/revert-22655-feat/prisma-metrics-collector
Revert "feat(proxy): add Prisma DB pool and engine health metrics to Prometheus"
2026-03-09 14:57:07 -07:00
yuneng-jiang
8ecac84789
Revert "feat(proxy): add Prisma DB pool and engine health metrics to Promethe…"
This reverts commit 0bb26c3f1b.
2026-03-09 14:55:11 -07:00
github-actions[bot]
c9434a8012
chore: regenerate poetry.lock to match pyproject.toml (#23189)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-03-09 21:47:52 +00:00
yuneng-jiang
a5e144c419
Merge pull request #23188 from BerriAI/mar9_bump_proxy_extras
[Infra] Bump proxy extras
2026-03-09 14:46:39 -07:00
yuneng-jiang
a9cc39b791 build artifacts 2026-03-09 14:46:03 -07:00
yuneng-jiang
bd914281e5 bump: version 0.4.52 → 0.4.53 2026-03-09 14:45:41 -07:00
yuneng-jiang
1a5e215f08
Merge pull request #23186 from BerriAI/litellm_doc_max_budget_per_session_ttl
[Docs] Add LITELLM_MAX_BUDGET_PER_SESSION_TTL to env vars reference
2026-03-09 14:41:52 -07:00
yuneng-jiang
b4e78ac7b4
Merge branch 'main' into litellm_doc_max_budget_per_session_ttl 2026-03-09 14:41:41 -07:00
yuneng-jiang
ea4e2bda8f Document LITELLM_MAX_BUDGET_PER_SESSION_TTL env var
Add missing env var to config_settings.md to fix test_env_keys CI check.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 14:40:05 -07:00
yuneng-jiang
be9d1798b2
Merge pull request #23182 from BerriAI/litellm_/exciting-swanson
[Fix] Model pricing schema test missing output_cost_per_image_token_batches
2026-03-09 14:26:21 -07:00
yuneng-jiang
379ce1aae5 [Fix] Add output_cost_per_image_token_batches to model pricing schema test
The gemini-3.1-flash-image-preview model introduced a new pricing field
that was missing from the test's validation schema and cost_fields list.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 14:17:52 -07:00
yuneng-jiang
729f32d6d5
Merge pull request #23179 from BerriAI/litellm/intelligent-wilbur
[Fix] Chocolatey v2.5.1 Interactive Prompt Blocking Windows CI
2026-03-09 14:09:58 -07:00
yuneng-jiang
4cc7e76fbe Fix Chocolatey v2.5.1 interactive prompt in Windows CI job
Chocolatey v2.5.1 introduced interactive prompts that block CI. Add
--no-progress, --force flags and CHOCOLATEY_CONFIRM_ALL env var to
fully suppress user input in non-interactive environments.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 14:06:46 -07:00
yuneng-jiang
5610945830
Merge pull request #23177 from BerriAI/litellm_fix_lint_error
[Fix] Remove duplicate jwt_key_mapping_router import
2026-03-09 13:59:41 -07:00
yuneng-jiang
169e76ccf9 Remove duplicate jwt_key_mapping_router import
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 13:58:04 -07:00
yuneng-jiang
1103a8c620
Merge pull request #23171 from BerriAI/litellm_survey_vitest_tests
[Test] UI - Survey: add Vitest unit tests for untested components
2026-03-09 12:13:24 -07:00
yuneng-jiang
994976ce6f [Test] UI - Survey: add Vitest tests for ClaudeCodeModal, ClaudeCodePrompt, SurveyPrompt, and SurveyModal
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-09 11:48:56 -07:00
michelligabriele
c47f77a348
fix(agentcore): handle JSON responses from agents using sync return (#23165)
* fix(agentcore): handle JSON responses from agents using sync return

BedrockAgentCoreApp agents that use synchronous `return` (instead of
async `yield`) respond with Content-Type: application/json instead of
text/event-stream. The streaming parser only handles SSE format, silently
discarding the JSON body and returning empty content to the client.

This adds Content-Type detection in both sync and async streaming
wrappers — when application/json is received, the response is parsed
and converted to a single-chunk stream. Also extends _parse_json_response
with a fallback chain supporting multiple agent response schemas (standard
AgentCore, Strands framework, plain string, raw JSON fallback).

* fix(agentcore): add dict-type guard to _parse_json_response

Prevent AttributeError when json.loads() returns a non-dict
(e.g. JSON array or primitive) by adding an isinstance check
at the top of _parse_json_response. Non-dict values fall back
to raw JSON string content.

* fix(agentcore): handle malformed JSON and split streaming chunks

- Wrap json.loads() in try/except in both sync and async streaming
  wrappers so malformed JSON bodies raise a structured BedrockError
  instead of a raw JSONDecodeError
- Split the JSON-fallback streaming path into two chunks (content
  chunk with finish_reason=None, then stop sentinel with empty delta)
  to match the SSE path convention

* fix(agentcore): catch IO errors in streaming JSON path + async error test

- Broaden except clause to catch both json.JSONDecodeError and IO-level
  exceptions (httpx.ReadError, etc.) from response.read()/aread(), so
  all failures surface as structured BedrockError
- Add async malformed-JSON test to mirror the sync test coverage
2026-03-09 10:22:36 -07:00
Aarish Alam
e21b06265a
fix fkey violation on deleting user (#23115) 2026-03-09 08:53:11 -07:00