Commit graph

36177 commits

Author SHA1 Message Date
yuneng-jiang
1103a8c620
Merge pull request #23171 from BerriAI/litellm_survey_vitest_tests
[Test] UI - Survey: add Vitest unit tests for untested components
2026-03-09 12:13:24 -07:00
Ryan Crabbe
f334956fcf fixes for duplicate value + missing timezones 2026-03-09 12:11:03 -07:00
Ryan Crabbe
0da56beb79 feat: adding a timezone picker to the usage page, to be able to view by timezone, backend already supports this just ui change 2026-03-09 11:53:07 -07:00
mubashir1osmani
200910f7e6
Merge branch 'BerriAI:main' into main 2026-03-09 14:51:24 -04:00
yuneng-jiang
994976ce6f [Test] UI - Survey: add Vitest tests for ClaudeCodeModal, ClaudeCodePrompt, SurveyPrompt, and SurveyModal
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-09 11:48:56 -07:00
michelligabriele
c47f77a348
fix(agentcore): handle JSON responses from agents using sync return (#23165)
* fix(agentcore): handle JSON responses from agents using sync return

BedrockAgentCoreApp agents that use synchronous `return` (instead of
async `yield`) respond with Content-Type: application/json instead of
text/event-stream. The streaming parser only handles SSE format, silently
discarding the JSON body and returning empty content to the client.

This adds Content-Type detection in both sync and async streaming
wrappers — when application/json is received, the response is parsed
and converted to a single-chunk stream. Also extends _parse_json_response
with a fallback chain supporting multiple agent response schemas (standard
AgentCore, Strands framework, plain string, raw JSON fallback).

* fix(agentcore): add dict-type guard to _parse_json_response

Prevent AttributeError when json.loads() returns a non-dict
(e.g. JSON array or primitive) by adding an isinstance check
at the top of _parse_json_response. Non-dict values fall back
to raw JSON string content.

* fix(agentcore): handle malformed JSON and split streaming chunks

- Wrap json.loads() in try/except in both sync and async streaming
  wrappers so malformed JSON bodies raise a structured BedrockError
  instead of a raw JSONDecodeError
- Split the JSON-fallback streaming path into two chunks (content
  chunk with finish_reason=None, then stop sentinel with empty delta)
  to match the SSE path convention

* fix(agentcore): catch IO errors in streaming JSON path + async error test

- Broaden except clause to catch both json.JSONDecodeError and IO-level
  exceptions (httpx.ReadError, etc.) from response.read()/aread(), so
  all failures surface as structured BedrockError
- Add async malformed-JSON test to mirror the sync test coverage
2026-03-09 10:22:36 -07:00
yuneng-jiang
c1b3640310 feat: add audit log export to external callbacks (S3, Datadog, etc.)
Audit logs (CRUD events on keys, teams, users, models) were only stored in
the Prisma DB. This adds a pluggable callback system so audit logs can be
forwarded to external services like S3 for ingestion into security monitoring
tools.

New config key `audit_log_callbacks` under `litellm_settings` reuses the
existing callback infrastructure. Any CustomLogger subclass can opt in by
overriding `async_log_audit_log_event()`. S3Logger (s3_v2) is implemented
as the first handler, storing audit logs under `audit_logs/{date}/` prefix.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 09:48:10 -07:00
Varad Khonde
9b15f639a4
fix(responses): merge parallel function_call items into single assistant message (#23116) 2026-03-09 09:01:31 -07:00
Aarish Alam
e21b06265a
fix fkey violation on deleting user (#23115) 2026-03-09 08:53:11 -07:00
ohadgur
0bb26c3f1b
feat(proxy): add Prisma DB pool and engine health metrics to Prometheus (#22655)
* feat(proxy): add Prisma DB pool and engine health metrics to Prometheus

Add a PrismaMetricsCollector that periodically queries pg_stat_activity
and the Prisma engine process to expose connection pool and engine health
as Prometheus gauges/counters. Auto-enabled when prometheus_system is in
service_callback.

New metrics:
- litellm_db_pool_active_connections (Gauge)
- litellm_db_pool_idle_connections (Gauge)
- litellm_db_pool_total_connections (Gauge)
- litellm_db_pool_waiting_connections (Gauge)
- litellm_db_engine_up (Gauge)
- litellm_db_engine_restarts_total (Counter)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address Greptile review feedback

- Only increment engine_restarts counter on heavy reconnects (engine
  actually dead), not lightweight network-blip reconnects
- Fix potential KeyError in _get_or_create_gauge/counter fallback path
  when REGISTRY._names_to_collectors is absent
- Rename litellm_db_pool_waiting_connections to
  litellm_db_pool_lock_waiting_connections to clarify it measures lock
  contention, not pool slot queuing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: warn when prometheus_system enabled but watchdog disabled

Log a warning when users have prometheus_system in service_callback
but PRISMA_HEALTH_WATCHDOG_ENABLED=false, since DB pool and engine
metrics won't be collected in that configuration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* ci: retrigger CI checks

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: use labeled gauge for DB pool connection metrics

Replace 3 separate pool gauges (active, idle, total) with a single
`litellm_db_pool_connections` gauge using a `state` label. This is more
Prometheus-idiomatic and exposes all pg_stat_activity states (active,
idle, idle in transaction, etc.) without ambiguity about what "total"
includes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address Greptile review — stale labels and fallback re-registration

- Zero out known pg_stat_activity states that are absent from the current
  query result, preventing stale gauge values from persisting.
- Simplify _get_or_create_gauge/counter by removing the fallback loop
  that could re-register an already-registered metric (ValueError).
- Add test for stale label clearing across collection cycles.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: include "unknown" in _PG_STATES for stale label clearing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: collect immediately on start and consolidate into single query

- Move sleep to end of loop so metrics appear on /metrics immediately
  after startup instead of after a 30s delay.
- Combine pool state and lock waiting queries into a single SQL query
  using conditional aggregation, halving per-cycle DB overhead.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: prevent tight spin loop on collection error

Move asyncio.sleep outside the try/except so it always executes even
when _collect_engine_health() or _collect_pool_metrics() raises.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: add multiprocess_mode to _get_or_create_gauge initialization

- Include `multiprocess_mode` parameter to properly support multiprocessing in Gauge creation.
- Ensure consistent behavior for labeled and unlabeled Gauges.

* fix: handle invalid env var and document watchdog prerequisite

- Add try/except ValueError for PRISMA_METRICS_COLLECTION_INTERVAL_SECONDS
  to prevent proxy startup crash on non-numeric values (e.g. "30s")
- Document that DB metrics require both prometheus_system callback and
  PRISMA_HEALTH_WATCHDOG_ENABLED=true

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: use defensive null coalescing for query_raw row values

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: add invalid env var fallback test and fix mock signature

- Add test for non-numeric PRISMA_METRICS_COLLECTION_INTERVAL_SECONDS
- Add **kwargs to mock _patched_get_or_create_gauge for forward compat

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 08:49:46 -07:00
milan-berri
df2e1bca46
feat: allow JWT and OAuth2 auth to coexist on the same instance (#23153)
When both enable_jwt_auth and enable_oauth2_auth are True, the proxy now
routes tokens based on their format:
- JWT tokens (3 dot-separated parts) -> JWT auth handler
- Opaque tokens -> OAuth2 auth handler

This enables using JWT for human users and OAuth2 for M2M (machine) clients
on the same LiteLLM instance. Previously, enabling OAuth2 would intercept
all tokens on LLM API routes before JWT auth could run.

When only one auth method is enabled, behavior is unchanged (backward compatible).
2026-03-09 08:41:27 -07:00
Ihsan Soydemir
b1a6ba7711
feat(search): add Serper (serper.dev) as search provider (#23112)
* Add Serper (serper.dev) as a new search provider

* Add @greptileai fixes
2026-03-09 08:40:37 -07:00
Joe Reyna
36e04b6efe
fix(tests): restore litellm_params=None on mock agent in a2a invoke test (#23125) 2026-03-09 07:16:02 -07:00
Joe Reyna
0bc1bd6871
fix(tests): use AsyncMock for prisma find_unique in agent get-by-id test (#23122) 2026-03-09 07:13:50 -07:00
Sameer Kankute
ee3ecb5994 fix(openai): preserve reasoning_effort summary + fix xhigh/none guards for dict inputs
- Add _get_effort_level() to extract effective effort from string or dict
- Use effective_effort for xhigh validation, tool-drop, sampling, temperature guards
- Preserve dict format when it has summary/generate_summary for Responses API
- Add tests: xhigh-dict validation, none-dict for tools/sampling/temperature
- Update tests: dict-with-summary now preserved (not normalized)

Made-with: Cursor
2026-03-09 18:44:05 +05:30
Sameer Kankute
8cf80a14d9 fix(openai): preserve reasoning_effort summary field for Responses API
When reasoning_effort is passed as a dict with additional fields like 'summary' or 'generate_summary', preserve the full dict format instead of normalizing it to a string. This ensures that when requests are routed to the OpenAI Responses API, all reasoning parameters are correctly included.

The normalization to string format now only happens for simple dicts with just the 'effort' key, which is appropriate for the Chat Completions API.

Fixes issue where summary field was being dropped when routing gpt-5.4+ requests with tools + reasoning to Responses API.

Made-with: Cursor
2026-03-09 18:28:11 +05:30
Sameer Kankute
ca4d4a0188
Merge pull request #23143 from giulio-leone/fix/gpt-5-4-pro-support
fix(models): set gpt-5.4-pro mode to responses — fixes #23014
2026-03-09 17:36:06 +05:30
Sameer Kankute
20d911599c
Merge pull request #23145 from BerriAI/litellm_publish_enterprise_pr_workflow
Fix enterpise bump yml
2026-03-09 17:28:07 +05:30
Sameer Kankute
b28d6eca67
Merge pull request #23144 from BerriAI/bump/enterprise-0.1.34
bump: litellm-enterprise 0.1.33 → 0.1.34
2026-03-09 16:44:12 +05:30
Sameer Kankute
0ee4d90d7e Fix enterpise bump yml 2026-03-09 16:43:40 +05:30
github-actions[bot]
6ff693149d bump: litellm-enterprise 0.1.33 → 0.1.34 2026-03-09 11:12:05 +00:00
Giulio Leone
556c64875e fix(models): set gpt-5.4-pro mode to responses instead of chat
gpt-5.4-pro and gpt-5.4-pro-2026-03-05 do not support the
/v1/chat/completions endpoint — OpenAI returns a 404 with
"This is not a chat model". These models are responses-only,
like o3-pro and o1-pro.

Changes:
- Set mode from "chat" to "responses" for both model entries
- Update supported_endpoints to ["/v1/responses", "/v1/batch"]
- Add regression test for responses API bridge routing

Fixes BerriAI/litellm#23014
2026-03-09 12:10:08 +01:00
Sameer Kankute
7c668a8021
Merge pull request #23142 from BerriAI/litellm_publish_enterprise_pr_workflow
fix(enterprise): create PR for version bump instead of pushing to protected main
2026-03-09 16:39:58 +05:30
Sameer Kankute
4d92c720c7 Fix enterpise bump yml 2026-03-09 16:39:38 +05:30
Yong woo Song
0d8880ab9f chore: fix 2026-03-09 11:03:39 +00:00
Sameer Kankute
a52a4fd28a fix(enterprise): create PR for version bump instead of pushing to protected main
Made-with: Cursor
2026-03-09 16:31:27 +05:30
Sameer Kankute
ba25d652e3
Merge pull request #23133 from BerriAI/litellm_fix_cicd_090326
Litellm fix cicd 090326
2026-03-09 16:14:16 +05:30
Yong woo Song
2683fa714c fix: typo on max_output_tokens and max_tokens from qwen3.5 series 2026-03-09 09:11:31 +00:00
Yong woo Song
9c07325396 feat: add qwen3.5 series for openrouter 2026-03-09 08:58:10 +00:00
Sameer Kankute
a8301d5614 Fix: varaitions endpoint geting 401 2026-03-09 12:51:21 +05:30
Sameer Kankute
4b1929ce93 Fix mistral ocr failing test 2026-03-09 11:29:33 +05:30
Sameer Kankute
b20c0afb64 Fix test_anthropic_messages_openai_model_streaming_cost_injection & openrouter image gen 2026-03-09 11:29:04 +05:30
Sameer Kankute
4dc277e427 fix(vertex_ai): strip LiteLLM-internal keys from extra_body before merging to Gemini request
PR #20950 added extra_body forwarding to Vertex AI Gemini. LiteLLM-internal
keys (cache, tags) were being merged into the request body, causing Vertex AI
to reject with 400: 'Unknown name "cache": Cannot find field.'

- Add _LITELLM_INTERNAL_EXTRA_BODY_KEYS frozenset (cache, tags)
- Skip these keys in _pop_and_merge_extra_body before merging
- Add regression tests for cache and tags stripping

Fixes regression from 1.79.3 → 1.81.12 when using proxy cache with
extra_body={"cache": {"use-cache": True, "ttl": 86400}}

Made-with: Cursor
2026-03-09 10:21:29 +05:30
Sameer Kankute
28b312f87a
Merge pull request #22382 from davidvpe/add-gemini-3.1-flash-image-preview
[Feature] Add Gemini 3.1 Flash Image Preview pricing details
2026-03-09 09:02:31 +05:30
netbrah
ffc6d84f27 fix: shallow copy input_schema to avoid caller mutation + add mutation guard test
Addresses Greptile review:
- dict(_input_schema) before mutation prevents cross-provider state leakage
- Test asserts original tool parameters dict is unchanged after call
2026-03-08 10:04:29 -04:00
netbrah
5d1106f018 fix(anthropic): deduplicate tool_result messages by tool_call_id
Anthropic requires exactly one tool_result per tool_use. When
conversation history (e.g. from session resume/checkpoint restore)
contains duplicate tool result messages with the same tool_call_id,
the API rejects with: 'each tool_use must have a single result.
Found multiple tool_result blocks with id: <id>'.

This is already handled for Bedrock via _deduplicate_bedrock_tool_content()
but was missing from the Anthropic direct and Vertex AI partner paths,
which share sanitize_messages_for_tool_calling().

Fix: Add Case D to sanitize_messages_for_tool_calling() — after the
existing orphan detection passes, scan for duplicate tool_call_ids
and keep only the last occurrence (most complete result).

Added 3 unit tests: dedup with duplicates, no-op with unique IDs,
and behavior when modify_params=False.

Related issues: #11804, #11029, #6836, #1782, #151
2026-03-08 10:02:45 -04:00
netbrah
78159212d9 fix(anthropic): enforce type:'object' on tool input schemas
Anthropic's API requires all tool input_schema to have type:'object'
at the root level. When OpenAI-format tools have parameters with a
missing or non-'object' type field (common with MCP tool servers),
the schema was passed through unchanged, causing Anthropic to reject
with: 'tools.N.custom.input_schema.type: Input should be object'.

The existing default handles the case where parameters is entirely
missing, but does not normalize schemas that ARE provided with a
wrong or absent type field.

Fix: After extracting _input_schema in _map_tool_helper(), ensure
type is set to 'object' and properties exists. This matches the
normalization already done implicitly by the Bedrock handler.

Added 4 unit tests covering: missing type, wrong type, valid schema
(no-op), and entirely missing parameters.

Related issues: #12020, #64, #1671
2026-03-08 07:52:07 -04:00
yuneng-jiang
160e2d9642
Merge pull request #23098 from BerriAI/litellm_org_admin_team_fix
[Fix] UI - Agents: fix stale test assertions after AgentsPanel table refactor
2026-03-07 23:51:04 -08:00
yuneng-jiang
81b223138a fix(test): update agents.test.tsx for AgentsPanel table refactor
AgentsPanel no longer renders AgentCardGrid — it uses a Table directly.
The two tests that looked for agent-card-grid are updated to instead verify
the Actions column header appears for admins and is absent for non-admins.
2026-03-07 23:50:06 -08:00
yuneng-jiang
5acb25fd9e
Merge pull request #23095 from BerriAI/litellm_org_admin_add_user_e2e
[Feature] Org Admin Access to Team Management Endpoints
2026-03-07 23:29:38 -08:00
yuneng-jiang
3a1ac964f7 fix: pass organization_ids=None in get_users test calls
When calling get_users() directly (not via FastAPI), Query() defaults
are not resolved. Pass organization_ids=None explicitly to avoid
'Query' object has no attribute 'split' error.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 23:15:16 -08:00
yuneng-jiang
70066426d0 fix: missing closing paren in agent_endpoints get_agents Query()
The `health_check` Query() call was missing its closing parenthesis,
causing a SyntaxError that blocked all proxy_server imports.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 23:13:42 -08:00
yuneng-jiang
ac5128493e fix: repair test regressions from org admin auth changes
- test_get_users_*: pass proxy admin user_api_key_dict since get_users
  now calls _authorize_user_list_request which checks user_role
- test_validate_team_member_add_permissions_non_admin: set
  organization_id on mock team since _is_user_org_admin_for_team
  accesses it

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 23:10:53 -08:00
yuneng-jiang
063198cf8e fix: push org filter into DB query for /team/list, fix agents.tsx build
- _authorize_and_filter_teams now uses Prisma WHERE clause
  (organization_id IN ...) instead of fetching the full team table
- Add missing Switch import in agents.tsx (pre-existing build error)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 23:04:08 -08:00
yuneng-jiang
8f33983389 Merge remote-tracking branch 'origin' into litellm_org_admin_add_user_e2e 2026-03-07 22:58:57 -08:00
yuneng-jiang
f77b5a24c3 fix: resolve ruff PLR0915 and F401 lint errors
Extract _authorize_user_list_request and _authorize_and_filter_teams
helpers to reduce statement count in get_users and list_team. Remove
unused _is_user_team_admin import.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 22:49:26 -08:00
yuneng-jiang
ce317148b9 feat: org admin access to team management — backend auth, UI visibility, tests
- Add _is_user_org_admin_for_team() reusable helper to common_utils.py
- Grant org admins access to /team/list, /team/info, /team/member_add,
  /team/member_delete, /team/member_update, /team/model/add,
  /team/model/delete, /team/permissions_list, /team/permissions_update
- Make validate_membership async with org admin fallback
- Add /user/list to self_managed_routes (endpoint handles own auth)
- UI: org admins see Members, Member Permissions, Settings tabs in team view
- UI: CreateUserButton uses useOrganizations() for org dropdown
- UI: org admin delete-member respects disable_team_admin_delete_team_user
- Add 16 unit tests for _is_user_org_admin_for_team, validate_membership,
  _user_is_org_admin route check, and privilege escalation prevention

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 22:43:09 -08:00
Maxwell Calkin
95ecf9f46d test: add tests for thinking block interleaving with server tool calls 2026-03-08 01:35:03 -05:00
Maxwell Calkin
e3717c0067 fix: preserve thinking block order when interleaved with server tool calls
When using extended thinking with web search, Anthropic interleaves thinking
blocks between server_tool_use/tool_result blocks. The previous code prepended
all thinking blocks first, then appended tool calls last, breaking Anthropic's
thinking block signature verification on round-trip.

This change detects when both thinking_blocks and server tool calls (srvtoolu_*)
are present, and interleaves them in the original order: each thinking block
precedes its corresponding server tool use group. This preserves the signature
positions that Anthropic validates.

Fixes #23047
2026-03-08 01:33:33 -05:00
yuneng-jiang
8bf3c0c67f fix: org admin invite user — multi-org selector, organizations list in POST body, auth check
- Thread org objects {organization_id, organization_alias} instead of bare IDs from
  users/page.tsx → view_users.tsx → CreateUserButton so the selector can show aliases
- Replace single-select org dropdown with multi-select; always shown when organizationIds
  is non-null; disabled/pre-selected for single-org admins; displays "Alias (id)"
- handleCreate: maps organization_ids → organizations before POST, removes redundant
  organizationMemberAddCall (backend _add_user_to_organizations handles it)
- _user_is_org_admin: also checks organizations list field in addition to singular
  organization_id so /user/new succeeds for org admins
- Add 5 backend unit tests for _user_is_org_admin and 2 frontend tests for new form behavior

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-07 20:34:12 -08:00