Commit graph

37850 commits

Author SHA1 Message Date
Ryan Crabbe
7e50af9228 Migrate route_preview.tsx from Tremor to Ant Design
Replace Tremor Card/Title/Subtitle with antd Card/Typography equivalents.
2026-03-23 22:32:26 -07:00
ryan-crabbe
fb92ea21bc
Merge pull request #24475 from BerriAI/litellm_fix-sso-return-to-validation
fix(proxy): ignore return_to in SSO when control_plane_url is not con…
2026-03-23 22:15:45 -07:00
Ryan Crabbe
e40f68aec4 test(ui): add unit tests for 5 untested frontend components
- AntDLoadingSpinner: rendering, prop forwarding, icon styling
- MessageManager: static fallback, custom instance delegation
- claude_code_plugins/helpers: all pure utility functions (15 describe blocks, 55 tests)
- AgentSelector: fetch behavior, loading states, error handling, disabled state
- WorkerDropdown: conditional rendering, worker options, selection changes
2026-03-23 22:03:10 -07:00
Ryan Crabbe
0aadf51342 fix(proxy): ignore return_to in SSO when control_plane_url is not configured
Instead of returning a 400 error when return_to is passed without
control_plane_url configured, silently ignore it and proceed with
the normal same-origin SSO flow.
2026-03-23 21:54:29 -07:00
Sameer Kankute
80af635eb1 Fix docs 2026-03-24 09:44:04 +05:30
Sameer Kankute
4e6e566b4d docs(opencode): fix model prefix and clarify drop_params scope
- Use openai/gpt-5 prefix to match existing doc conventions
- Clarify that additional_drop_params must be added to every affected
  model entry, not just one

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 09:42:59 +05:30
Sameer Kankute
17e6e7a4cc docs(opencode): add guidance for dropping reasoningSummary param
OpenCode sends a `reasoningSummary` Responses API param with chat
completion requests. Document how to use `additional_drop_params` to
drop it and avoid 400 errors from the OpenAI API.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 09:33:28 +05:30
Krish Dholakia
3292d02aa4
Merge pull request #24460 from DmitriyAlergant/ci/skip-scheduled-workflows-on-forks
ci: skip scheduled workflows on forks
2026-03-23 19:54:50 -07:00
Krish Dholakia
14fffc2770
Merge pull request #24432 from BerriAI/krrishdholakia/project-id-tracking
feat(proxy): add project_alias tracking in callbacks
2026-03-23 19:24:44 -07:00
DmitriyAlergant
1310a275d2 ci: narrow codeql guard to schedule-only
Use event_name check so push/PR-triggered CodeQL scans still run on
forks — only the scheduled run is skipped.
2026-03-23 21:39:11 -04:00
DmitriyAlergant
91bc095e18 ci: skip scheduled workflows on forks
Add `if: github.repository == 'BerriAI/litellm'` guard to scheduled
jobs in stale.yml, codeql.yml, and create_daily_staging_branch.yml.

This matches the existing pattern in auto_update_price_and_context_window.yml
and prevents these workflows from running unnecessarily on fork repositories.
2026-03-23 21:29:00 -04:00
Krrish Dholakia
26d162ccf4 fix(test): add user_api_key_project_alias to spend logs expected keys
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 18:12:50 -07:00
Krrish Dholakia
742e176611 docs(reasoning_content.md): update guide 2026-03-23 17:23:14 -07:00
Krish Dholakia
8a3aa4d31c
Merge pull request #24434 from BerriAI/krrishdholakia/prometheus-spend-metadata
feat(prometheus): include spend_logs_metadata in custom labels
2026-03-23 16:52:59 -07:00
Josh
8a58281cbf Add org budget metrics initialization at startup 2026-03-23 19:33:57 -04:00
Josh
7fcf99ffaf Add org budget metrics to failure path 2026-03-23 19:07:01 -04:00
Lei Nie
1716956520 test(responses): update expected events and add mock test for content_part.added
- Update test_anthropic_via_responses_api expected_events to include
  CONTENT_PART_ADDED between OUTPUT_ITEM_ADDED and OUTPUT_TEXT_DELTA
- Add TestEnsureOutputItemContentPartAdded with 3 mock tests:
  message item emits content_part.added, reasoning item does not,
  and the event is only emitted once
2026-03-23 22:29:50 +00:00
Lei Nie
d560e4c009 fix(responses): emit content_part.added event for non-OpenAI models
LiteLLMCompletionStreamingIterator defined create_content_part_added_event()
but never called it, so non-OpenAI providers (Claude, Gemini, etc.) skipped
this spec-required event. Downstream parsers that process content_part.added
to initialize the text part structure would fail when output_text.delta
arrived before the text part existed.
2026-03-23 22:25:37 +00:00
Krrish Dholakia
dd0e7dcca8 test(prometheus): add tests for spend_logs_metadata in custom labels
Verify that spend_logs_metadata is correctly merged into combined_metadata
and flows through to Prometheus custom labels. Tests cover: basic extraction,
precedence when keys overlap, all three metadata sources combined, and None
handling.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 11:02:21 -07:00
Krrish Dholakia
7fa623df91 feat(prometheus): include spend_logs_metadata in custom labels
Add spend_logs_metadata to combined_metadata in Prometheus logger so
custom metadata from x-litellm-spend-logs-metadata header can be used
in Prometheus custom labels.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 10:59:18 -07:00
Krrish Dholakia
6809213957 feat(proxy): add project_alias tracking through callback metadata pipeline
Thread project_alias alongside project_id through the metadata pipeline so
callbacks receive the human-readable project name. DRY up duplicate metadata
dict construction in proxy_track_cost_callback and pass_through_endpoints by
reusing get_sanitized_user_information_from_key — future metadata fields only
need adding in one place.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 10:44:17 -07:00
Krish Dholakia
63425b4cb4
Merge pull request #23910 from michelligabriele/fix/guardrail-post-call-logging
fix(proxy): post-call guardrail response not captured for logging
2026-03-23 09:21:28 -07:00
michelligabriele
fa7ccf0893 fix(test): add request_data param to test mock + black formatting 2026-03-23 15:43:05 +01:00
michelligabriele
9a231bd758 fix(proxy): use real request_data in Responses API streaming fallback path 2026-03-23 15:39:23 +01:00
michelligabriele
4625ccbaa2 fix(proxy): anchor metadata dict in _process_response/_process_error so pop() mutates the real dict 2026-03-23 15:39:23 +01:00
michelligabriele
d8fd9a20ed fix(proxy): address Greptile review — streaming request_data, OCR backward compat, test coverage
- Pass request_data to end-of-stream process_output_streaming_response call
- Restore inputs.update() in OCR handler for third-party guardrail providers
- Add streaming end-to-end test for guardrail logging passthrough
2026-03-23 15:39:23 +01:00
michelligabriele
ae454fd700 fix(proxy): OpenAI Moderation post-call guardrail response not captured for logging
Two independent bugs prevented post-call OpenAI Moderation guardrail
results from reaching downstream logging callbacks (Langfuse, Datadog).

Bug 1: process_output_response() created a throwaway request_data dict,
so guardrail info written by @log_guardrail_information was discarded.
Fixed by threading the real request_data from the unified guardrail
dispatcher through all 13 BaseTranslation handlers, with litellm_metadata
injection preserved for third-party guardrails (Zscaler, Prompt Security).
Also extended to process_output_streaming_response for consistency.

Bug 2: The @log_guardrail_information decorator collapsed the full
moderation API response (categories, scores, flagged status) to "allow".
Fixed by overriding _process_response/_process_error on
OpenAIModerationGuardrail to stash and log the full response, following
the established Model Armor pattern.
2026-03-23 15:39:22 +01:00
Ben Langfeld
847c12e4f5
Update docs/my-website/docs/proxy/config_settings.md
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-23 11:06:09 -03:00
Ben Langfeld
5a846c2e64
Correct documentation of completion_model
See https://github.com/BerriAI/litellm/issues/21554
2026-03-23 10:59:18 -03:00
Cesar Garcia
d233d6694d
Merge pull request #24372 from Chesars/fix/gemini-web-search-cost
fix(gemini): read web search cost from model_info instead of hardcode
2026-03-22 18:48:08 -03:00
Chesars
21b9c68d42 fix: remove web_search_billing_unit from OpenRouter/Perplexity entries
These providers have their own web search systems and pricing,
independent of Google's grounding billing model.
2026-03-22 18:46:39 -03:00
Cesar Garcia
335f4e8c6e
Merge pull request #24370 from Chesars/fix/gemini-embedding-drop-unsupported-params
fix(gemini): filter unsupported params from embedding requests
2026-03-22 18:46:10 -03:00
Cesar Garcia
4e604d6723
Merge pull request #24368 from Chesars/fix/bedrock-converse-content-block-ordering
fix(bedrock): sort assistant content blocks so text precedes toolUse
2026-03-22 18:45:22 -03:00
Cesar Garcia
16c48b4a98
Merge pull request #24371 from Chesars/fix/responses-api-gpt5-temperature-drop-params
fix(responses-api): apply GPT-5 temperature validation
2026-03-22 18:44:12 -03:00
Cesar Garcia
de91bbb9ff
Merge pull request #24373 from Chesars/fix/zhipu-finish-reason-mapping
fix: map Zhipu GLM non-standard finish_reason values
2026-03-22 18:42:04 -03:00
Chesars
996d27b156 fix(vertex_ai): delegate web search cost to shared Gemini calculator
The vertex_ai cost calculator hardcoded $0.035 and charged for every
call with a PromptTokensDetailsWrapper (not just web search calls).

Delegate to the shared Gemini calculator which reads pricing and
billing unit from model_info, fixing both issues for vertex_ai models.
2026-03-22 18:39:12 -03:00
Chesars
bfee7f0b58 fix: improve error message for supports-none models with temperature 2026-03-22 18:37:34 -03:00
Chesars
6a466913fc fix: map Zhipu GLM non-standard finish_reason values
Zhipu GLM returns non-standard finish_reason values during streaming
when inference fails mid-request, causing Pydantic validation crash:
- "network_error" (inference interrupted) → map to "stop"
- "sensitive" (content policy violation) → map to "content_filter"

Fixes #23386
2026-03-22 18:36:13 -03:00
Chesars
b6079018cd docs: clarify web_search_billing_unit applies to Gemini models only 2026-03-22 18:32:03 -03:00
Chesars
e82d3f6d2e refactor(gemini): use web_search_billing_unit field instead of hardcoded model name check
Replace _is_gemini_3_model() substring check with a
web_search_billing_unit field in model_prices JSON:
- "per_query": each search query billed individually (Gemini 3.x)
- "per_prompt" (default): flat fee per grounded API call (Gemini 2.x)

Add web_search_billing_unit to 23 Gemini 3.x model entries.
Update docs and tests accordingly.
2026-03-22 18:29:38 -03:00
Chesars
a0d1d22bcf docs: add Web Search Cost Tracking section
Document how each provider bills for web search, the
search_context_cost_per_query field in model_prices JSON,
how to override pricing via proxy config, and how LiteLLM
extracts web_search_requests from each provider's response.
2026-03-22 18:25:01 -03:00
Chesars
f8a9bbd537 test: add supports_none branch coverage for Responses API GPT-5 temperature
Add tests for the gpt-5.1/5.2/5.4 reasoning.effort interaction:
- gpt-5.1 with no reasoning allows flexible temperature
- gpt-5.1 with effort='high' drops temperature
- gpt-5.4 with effort='none' allows flexible temperature
2026-03-22 18:22:11 -03:00
Chesars
4c99f3ddd8 fix(gemini): differentiate billing model and extract web search requests
- Gemini 2.x charges per grounded prompt (flat $0.035), clamped to 1
  regardless of internal query count
- Gemini 3.x charges per search query ($0.014 each)
- Extract web_search_requests from groundingMetadata in non-streaming
  responses (parity with streaming path)
- Add search_context_cost_per_query to vertex_ai and base Gemini entries
- Move tests to tests/test_litellm/ (CI directory)
2026-03-22 18:20:00 -03:00
Chesars
040c6fe920 fix(gemini): read web search cost from model_info instead of hardcode
The Gemini web search cost calculator hardcoded $0.035 per request,
which is only correct for Gemini 2.x models. Gemini 3.x models
charge $0.014 per request.

Read from search_context_cost_per_query in model_info (same field
used by Anthropic, OpenAI, and Perplexity) with fallback to the
legacy $0.035 for models not yet updated in the JSON.

Also add search_context_cost_per_query to all 25 Gemini models
that support web search in model_prices_and_context_window.json.

Fixes #24369
2026-03-22 18:00:13 -03:00
Chesars
fff83dd8a5 fix(responses-api): apply GPT-5 temperature validation in Responses API
The Responses API map_openai_params passed all params through without
applying model-specific validation. GPT-5 models (except gpt-5-chat)
only accept temperature=1 unless reasoning.effort="none" on models
that support it (5.1, 5.2, 5.4).

Reuse the existing OpenAIGPT5Config logic from chat completions to
validate temperature in the Responses API path. With drop_params=True,
unsupported temperature values are silently dropped; without it,
UnsupportedParamsError is raised.

Fixes #16090
2026-03-22 17:51:07 -03:00
Chesars
db0d85eefd fix(gemini): filter unsupported params from embedding requests
The Gemini batch embedding transformation was spreading all
optional_params into the request body via **gemini_params. Params
like max_tokens (injected by add_provider_specific_params_to_optional_params)
would reach the Gemini API and cause a 400 BadRequestError.

Extract _filter_embed_params() that maps dimensions/task_type and
keeps only the fields Gemini embeddings actually accept
(outputDimensionality, taskType, title). Applied to both
transform_openai_input_gemini_content and
transform_openai_input_gemini_embed_content.

This also fixes drop_params: true not preventing the error, since
the param was re-injected after the drop_params check.

Fixes #24293
2026-03-22 17:41:09 -03:00
Chesars
dd7269ee14 fix(bedrock): sort assistant content blocks so text precedes toolUse
When the Responses API converts function_call and message output items
into chat completion messages, they can become two consecutive assistant
messages. The Bedrock Converse transformer merges these into one, but
the merge preserves input order — so if function_call came first, the
toolUse block ends up before the text block.

Claude models (Sonnet 4, Haiku 3.5+) reject this ordering with:
"tool_use ids were found without tool_result blocks immediately after"

Add _sort_bedrock_assistant_content_blocks() that reorders content
blocks within assistant messages: reasoningContent → text → toolUse.
Applied in both sync and async Bedrock Converse transformation paths.

Fixes #24361
2026-03-22 17:18:40 -03:00
Cesar Garcia
2132db4f60
Merge pull request #24354 from Chesars/fix/azure-streaming-role-include-usage
fix: preserve role='assistant' in Azure streaming with include_usage
2026-03-22 11:19:53 -03:00
Cesar Garcia
5847c6166d
Merge pull request #24355 from Chesars/fix/anthropic-adapter-tool-args-dropped
fix: preserve tool_use input args in Anthropic adapter streaming
2026-03-22 11:19:12 -03:00
Chesars
e5de1ecd92 chore: remove unused _parse_sse_events helper in test 2026-03-22 11:18:50 -03:00