Commit graph

37751 commits

Author SHA1 Message Date
harish-berri
7021eb41c7 Merge branch 'litellm_gcp-iam-redis-token-caching' of https://github.com/BerriAI/litellm into litellm_gcp-iam-redis-token-caching 2026-04-24 17:26:27 +00:00
harish-berri
fbf1846b4d refactor(redis): improve documentation for GCPIAMCredentialProvider class
Updated the docstring for the GCPIAMCredentialProvider class to clarify its purpose and the caching mechanism for GCP IAM tokens. The changes enhance readability and maintainability by providing a more concise explanation of the token caching strategy and its benefits for Redis authentication.
2026-04-24 17:20:46 +00:00
harish-berri
5ae3d1df69 refactor(redis): improve documentation for GCPIAMCredentialProvider class
Updated the docstring for the GCPIAMCredentialProvider class to clarify its purpose and the caching mechanism for GCP IAM tokens. The changes enhance readability and maintainability by providing a more concise explanation of the token caching strategy and its benefits for Redis authentication.
2026-04-24 17:20:33 +00:00
harish-berri
3c0ba0e835 refactor(redis): remove unused Optional import from _redis_credential_provider.py 2026-04-24 17:12:34 +00:00
harish-berri
037b4f476b fix(redis): cache GCP IAM token to prevent async event loop blocking
## Problem

GCPIAMCredentialProvider.get_credentials() calls _generate_gcp_iam_access_token
on every Redis connection establishment. This function performs synchronous HTTP
and gRPC calls (google-auth + google-cloud-iam) which block Python's asyncio
event loop while running.

Under concurrent load (e.g. connection pool warm-up, parallel health checks),
multiple connections are established simultaneously, each triggering an
independent blocking IAM token refresh. These refreshes serialise behind each
other inside the single-threaded event loop, causing individual Redis spans to
take 20-25 seconds instead of milliseconds.

Observed in production via Datadog APM: a single INCRBYFLOAT Redis span took
25.6 seconds (90% of a 28.4s trace), with GCP metadata + GenerateAccessToken
gRPC calls visible inside the span. This cascaded into aiohttp SocketTimeoutError
on upstream LLM API calls — not because the upstream was slow, but because the
event loop was frozen and the 30-second sock_read timer fired on a connection
that was never given CPU time.

## Fix

Add a module-level token cache (dict keyed by service account, value is
(token, expiry_monotonic)). _get_cached_gcp_iam_token() returns the cached
token on cache hit (no I/O), and refreshes only when expired using
double-checked locking so only one thread performs the network round-trip.

GCP IAM tokens are valid for 1 hour; the cache TTL is set to 55 minutes
(_GCP_IAM_TOKEN_TTL_SECONDS = 3300) to refresh safely before expiry.

The cache is shared across all GCPIAMCredentialProvider instances for the same
service account, so N concurrent Redis connections on the same pod share a
single token and avoid N concurrent blocking refreshes.

get_credentials_async() already used asyncio.to_thread (non-blocking), and is
updated to call _get_cached_gcp_iam_token so it also benefits from caching.

## Tests

- Updated existing test that expected a fresh token on every call to reflect
  the new caching behaviour.
- Added tests for: cache hit (no redundant I/O), cache expiry and refresh,
  and cache sharing across multiple provider instances.
- Added autouse fixture to clear the module-level cache between tests.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-04-24 17:01:42 +00:00
yuneng-jiang
8dda834cf9
Merge pull request #25842 from BerriAI/litellm_docs-gemini3-thinking-defaults
docs(gemini): Gemini 3 thinking_level defaults and release note
2026-04-24 09:45:24 -07:00
yuneng-jiang
023dad5bde
Merge pull request #25932 from BerriAI/litellm_docs-code-block-padding-parity
feat(docs): align fenced code block padding on blog and doc pages
2026-04-24 09:45:08 -07:00
yuneng-jiang
61ad127a75
Merge pull request #25935 from BerriAI/litellm_anthropic-stream-strip-gemini-thought-tool-id
fix(anthropic): strip Gemini thought suffix from streaming tool_use id
2026-04-24 09:43:27 -07:00
yuneng-jiang
4e3feda952
Merge pull request #26221 from BerriAI/litellm_responses_strip_custom_tool_call_namespace
feat(responses): strip custom_tool_call namespace for all providers
2026-04-24 09:42:55 -07:00
yuneng-jiang
d73b790cae
Merge pull request #26248 from BerriAI/litellm_anthropic_messages_call_type_fix
fix(proxy): preserve anthropic_messages call type for /v1/messages logging
2026-04-24 09:42:36 -07:00
Yuneng Jiang
4d5c3476a4
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_docs-gemini3-thinking-defaults 2026-04-24 09:40:04 -07:00
Yuneng Jiang
b2afc70080
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_docs-code-block-padding-parity 2026-04-24 09:39:06 -07:00
yuneng-jiang
78d171b2b9
Merge pull request #26115 from BerriAI/litellm_gpt54_mini_nano_versioned_models
feat(models): add versioned GPT-5.4 mini/nano snapshots
2026-04-24 09:34:55 -07:00
Sameer Kankute
4dbea4e957
fix(responses): enforce spec object on completion bridge (#26327)
Ensure Chat Completions -> Responses bridge always emits object="response" so non-native providers return the same top-level schema as native OpenAI Responses.

Made-with: Cursor
2026-04-24 09:29:06 -07:00
Yuneng Jiang
55ea431c05
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gpt54_mini_nano_versioned_models 2026-04-24 09:28:54 -07:00
Sameer Kankute
e1466be825
feat(pricing): gemini-embedding-2 GA cost map, blog, and test (#26391)
* feat(pricing): gemini-embedding-2 GA cost map, blog, and test

- Add model_prices entries for gemini-embedding-2 (Gemini + Vertex paths)
- Add docs blog gemini_embedding_2_ga with LiteLLM proxy curl examples
- Add test_gemini_embedding_2_ga_in_cost_map in test_utils

Made-with: Cursor

* Fix greptile reviews
2026-04-24 09:28:18 -07:00
Shivam Rawat
9dcb2bd528
fix(proxy): respect object-level permissions for managed vector store endpoints (#26351)
* fix(proxy): honor object_permission for managed vector store access

* perf(proxy): preload team object_permission on UserAPIKeyAuth

Populate team_object_permission during virtual-key and JWT auth when the
team is loaded, so can_user_access_vector_store uses it in memory first
and only falls back to get_object_permission by id when missing.

Made-with: Cursor
2026-04-24 09:21:13 -07:00
ishaan-berri
863f922be8
fix(team_endpoints): auto-add SSO team members to org on move (proxy admin only) (#26377)
* fix(team_endpoints): auto-add SSO team members to org for proxy admins

* test: proxy_admin vs team_admin security boundary for team→org move

* screenshots: before/after for team-org SSO fix

* fix(team_endpoints): restore staging security features dropped in SSO commit

Co-Authored-By: Ishaan Jaff <ishaan@berri.ai>

* style: black formatting for team_endpoints
2026-04-24 08:36:25 -07:00
Sameer Kankute
1720903bda
Merge pull request #25346 from BerriAI/litellm_Sameerlite/responses-bridge-optin
feat(responses): add use_chat_completions_api flag for openai/ models with custom api_base
2026-04-24 20:55:22 +05:30
ryan-crabbe-berri
0992bf2271
Merge pull request #26367 from BerriAI/litellm_/split-mcp-routes-management-vs-inference
Split MCP routes into inference vs management (unblock Admin UI on DISABLE_LLM_API_ENDPOINTS nodes)
2026-04-23 22:05:48 -07:00
Sameer Kankute
3c1b27e155
Merge pull request #26381 from BerriAI/litellm_internal_staging
merge main
2026-04-24 09:22:28 +05:30
Sameer Kankute
2378ef7f8c
FIx black formatinig 2026-04-24 09:15:13 +05:30
shin-berri
8e652d129d
Merge pull request #26356 from BerriAI/litellm_cci_gha_dedup_and_shard
Some checks are pending
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / schema-migration (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Security / security (push) Waiting to run
[Infra] Remove CCI/GHA test duplication and semantically shard proxy DB tests
2026-04-23 18:17:56 -07:00
yuneng-jiang
654b688c8f
Merge pull request #25746 from BerriAI/litellm_vector-store-team-byok-model-none
fix(router): restore BYOK key injection for vector store endpoints with team-scoped deployments
2026-04-23 18:16:47 -07:00
shivam
e982fe85e9
Merge branch 'litellm_internal_staging' into litellm_vector-store-team-byok-model-none 2026-04-23 17:41:11 -07:00
Ryan Crabbe
35eef7d92c
chore: apply black formatting to _types.py management_routes block 2026-04-23 17:35:35 -07:00
shivam
812044a805
rerun tests 2026-04-23 17:34:19 -07:00
shivam
b217ad44d3
rerun tests 2026-04-23 17:31:37 -07:00
yuneng-jiang
1f6ce45702
Merge pull request #26370 from BerriAI/litellm_version_bump
[Infra] Bump version 1.83.12 → 1.83.13
2026-04-23 17:30:45 -07:00
yuneng-jiang
08cc1e66cf
Merge pull request #26207 from BerriAI/litellm_team_member_total_spend_frontend
Surface per-member budget cycle in Teams > Members tab
2026-04-23 17:24:03 -07:00
Ryan Crabbe
6b6b8c7418
restore budget_reset_at on TeamMembership type
Members tab column reads this field; dropping it from the type in the
previous revert broke the type check without affecting the reverted
render logic.
2026-04-23 17:07:21 -07:00
Ryan Crabbe
fbaedc36dc
revert TeamInfo budget reset display changes
Out of scope for the members-tab feature and regressed legacy teams
whose budget_reset_at is null (duration was previously shown as a
fallback).
2026-04-23 17:01:32 -07:00
Yuneng Jiang
ffaeff54cd
add uv 2026-04-23 17:00:20 -07:00
Yuneng Jiang
29e30d9ddb
bump: version 1.83.12 → 1.83.13 2026-04-23 16:58:17 -07:00
Ryan Crabbe
4d2acafa43
Split MCP routes into inference vs management categories
MCP server CRUD endpoints (/v1/mcp/server*) were bundled with MCP
tool-call / passthrough endpoints under llm_api_routes, so setting
DISABLE_LLM_API_ENDPOINTS=true on admin-only nodes also blocked the
Admin UI from listing, adding, or attaching MCP servers.

Separate mcp_inference_routes (data-plane, gated by
DISABLE_LLM_API_ENDPOINTS) from mcp_management_routes (control-plane,
gated by DISABLE_ADMIN_ENDPOINTS). Keep mcp_routes as a union for
backward compat with allowed_routes=["mcp_routes"] virtual key configs.

Upgrade is_management_route to pattern-aware matching so
/v1/mcp/server/{path:path} resolves for concrete IDs.
2026-04-23 16:52:45 -07:00
Ryan Crabbe
ea626d9fb8
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_team_member_total_spend_frontend 2026-04-23 16:47:41 -07:00
yuneng-jiang
87e120d958
Merge pull request #26346 from BerriAI/litellm_reset_budget_is_not_null
[Fix] Reset budget windows failing due to Prisma Json? null filter
2026-04-23 16:37:09 -07:00
Yuneng Jiang
66bf890226
[Infra] Stop attaching push-only postgres workflows to a GHA environment
The `_test-unit-services-base.yml` reusable workflow attached every job
to the `integration-postgres` GHA environment to read three "secrets":
DATABASE_URL, POSTGRES_USER, POSTGRES_PASSWORD. These are not secrets —
the postgres service container is spawned per-job on localhost and
destroyed with the job, so the user/password are bootstrap values for a
throwaway container and the URL is always `postgresql://…@localhost:…`.

Each environment attachment produces a "temporarily deployed to
integration-postgres" deployment record, which the PR timeline renders
as a message per matrix shard per push. With 14 proxy-db shards that's
~14 notifications per push, drowning the PR conversation.

Changes:
* Hardcode POSTGRES_USER/POSTGRES_PASSWORD/POSTGRES_DB and the derived
  DATABASE_URL in `_test-unit-services-base.yml`.
* Delete the `environment: integration-postgres` attachment.
* Delete the `secrets:` declarations on the reusable workflow and on
  the two callers (test-unit-proxy-db.yml, test-unit-security.yml).
* The `services:` container still starts a fresh postgres per job;
  the connection string now matches what the container boots up with.

Security review: no regression. The environment wasn't gating anything
real — no protection rules configured, no approval gates, and the
branch restriction is already enforced by `on: push: branches: [...]`
on both caller workflows. Zizmor pedantic-mode findings are identical
before and after (same 6 pre-existing findings, zero new ones).

The `integration-postgres` environment and its three "secrets" in repo
settings are now unreferenced and can be deleted from repo admin.
2026-04-23 16:32:18 -07:00
Yuneng Jiang
21e08b0bb5
[Infra] Run schema-migration shard serially (workers: 0)
test_db_schema_migration.py has exactly one test, and that test is mostly
waiting on prisma subprocesses (~170s: prisma migrate deploy + prisma
migrate diff). No CPU-bound Python work inside the test body, and only
one test in the file means xdist's parallelism is unused regardless.

Previous run on commit 5df9f397e6: 10.0m wall-clock for the shard, of
which 4:56 was silence between step start and pytest banner — the cost
of 4 xdist workers each cold-starting (pytest plugin load + litellm
import + pytest-cov instrumentation) so that exactly one of them could
pick up the single test.

Switching to workers: 0 takes the serial pytest branch in the base
workflow, which already handles this case correctly (no -n, no --dist).
Single-process startup instead of 4. Expected wall-clock: ~6m.
2026-04-23 16:24:40 -07:00
milan-berri
2001d91b27
fix(mcp): share temporary MCP OAuth sessions across instances via Redis (#26162) (#26318)
Temporary MCP OAuth sessions were kept in process-local memory, so on
multi-instance/LB proxy deployments a session created on instance A could
not be found when the follow-up /server/oauth/{server_id}/... request
landed on instance B.

Persist temporary session records to Redis (encrypted with the existing
proxy encryption helpers) as a best-effort L2 cache alongside the current
in-memory L1. Convert get_cached_temporary_mcp_server to async and await
it from the authorize/token/register OAuth endpoints.

Made-with: Cursor
2026-04-23 16:21:27 -07:00
shin-berri
7c69262279
Merge pull request #26349 from BerriAI/litellm_deflakeSpendTests
[Fix] Deflake spend tracking tests
2026-04-23 16:12:19 -07:00
shin-berri
bb94144111
Merge pull request #26359 from BerriAI/litellm_fixCreateReleasePerms
[Fix] Infra: grant contents:write to create-release-branch caller job
2026-04-23 15:58:42 -07:00
milan-berri
b6d0f6b649
fix(vertex_ai): use aiplatform.{geo}.rep.googleapis.com for multi-region locations (#26281)
Vertex multi-region endpoints (e.g. us, eu) use the rep host pattern, not
{geo}-aiplatform.googleapis.com. Regional IDs still contain a hyphen.

common_utils.get_vertex_base_url centralizes the rule for SDK/API URL building.
Proxy pass-through duplicates the same branching in a local get_vertex_base_url
(with trailing slashes) to avoid importing from common_utils there; live
WebSocket passthrough uses the same multi-region host logic for wss://.

Tests cover us/eu for the common_utils helper.

Made-with: Cursor
2026-04-23 15:58:02 -07:00
Ryan Crabbe
1f6e01802d
Show absolute date in Budget Reset column
Relative labels ("today", "in 2 days", "on May 12, 2026") mixed three
shapes in one column, breaking scannability. Always render MMM D, YYYY
for consistency and easier at-a-glance comparison across members.
2026-04-23 15:57:22 -07:00
Yuneng Jiang
5df9f397e6
[Infra] Match xdist workers to runner cores; revert test_proxy_utils -k split
Two changes:

1. workers: 8 -> 4 on every non-serial proxy-db shard. ubuntu-latest is a
   4-core runner; -n 8 oversubscribes 2x and workers block each other
   during their cold-start imports (pytest-cov instruments every litellm
   module per worker). Measured ~441% CPU locally with -n 8 on 8 cores
   (i.e. ~55% effective). Matching -n to physical cores should give
   ~2x faster worker startup, which is where most of the ~9m wall-clock
   per shard goes (7+ minutes is plugin load + xdist imports before any
   test runs).

2. Revert the -k split on test_proxy_utils.py. It was split into
   proxy-utils-a-h / proxy-utils-i-z as a semantic-adjacent hack; merge
   back to a single proxy-utils shard. Still uses --dist=worksteal so
   xdist can balance the 188 parametrized cases across workers.

Also drops the now-unused `keyword` input from _test-unit-services-base.yml
and its matching matrix field across all proxy-db entries.

Shard count: 14 -> 13 (+ the assert-shard-coverage guard).
2026-04-23 15:56:27 -07:00
yuneng-jiang
9bdd447891
Merge pull request #26355 from BerriAI/litellm_fixFlakyTpmRoutingTest
[Fix] Tests - drain logging worker in test_router_caching_ttl to fix flakiness
2026-04-23 15:45:19 -07:00
Yuneng Jiang
584a7cd40f
[Infra] Clean up proxy-db matrix job display names
Default GHA matrix job names join every matrix field, producing unreadable
check labels like:
  'proxy-db (logging-misc, tests/proxy_unit_tests/test_proxy_reject_logging.py
   tests/proxy_unit_tests/test_audit_logs_proxy.py ..., 8, loadscope, "", 15)'

Set the job's display name to '${{ matrix.test-group }}' so each check
shows just 'logging-misc', 'proxy-utils-a-h', etc.
2026-04-23 15:29:42 -07:00
Yuneng Jiang
e0201ece1e
[Infra] Split slow proxy-db shards to hit 7m wall-clock target
Previous run (13.8m total) was bottlenecked by shards with 9-12m wall-clock.
Setup + xdist spawn + coverage teardown is ~3m per shard, so each shard's
pytest runtime must stay under ~4m to fit inside 7m total.

Observed per-shard pytest times (before split):
  db-and-spend            9:08   (170s outlier: test_aaaasschema_migration_check)
  proxy-server            7:15
  logging-and-callbacks   6:45
  guardrails-budget-hooks 6:37
  proxy-utils             6:23
  auth-and-jwt            6:54

Split 6 shards into 12, keeping key-generation and endpoints-and-responses
(already <7m). Adds a `keyword` input to _test-unit-services-base.yml so
test_proxy_utils.py can be split by -k expression (same file, two runners).
New matrix entries:

  auth-and-jwt           -> auth-checks + jwt-and-keys
  proxy-server           -> proxy-server-core + proxy-runtime
  logging-and-callbacks  -> custom-logging + logging-misc
  db-and-spend           -> schema-migration (isolated 170s test) + db-and-spend
  guardrails-budget-hooks-> guardrails-hooks + budgets
  proxy-utils            -> proxy-utils-a-h + proxy-utils-i-z (-k split)

The -k expression split is verified to cover every one of the 64 test
functions in test_proxy_utils.py exactly once. The assert-shard-coverage
guard still catches any file not in any shard.
2026-04-23 15:25:37 -07:00
Yuneng Jiang
4a2deae92c
[Fix] Infra: grant contents:write to create-release-branch caller job
The create-branch job in create-release.yml calls the reusable
create-release-branch.yml workflow, which requires contents: write.
The top-level permissions: {} blocks the inherited default, and only
the release job overrode it, so the nested call failed with:

  The nested job 'create-branch' is requesting 'contents: write',
  but is only allowed 'contents: none'.

Add the permission at the calling job level so the reusable
workflow is granted what it needs.
2026-04-23 15:11:12 -07:00
Yuneng Jiang
c14a73fa59
fix: make LoggingWorker.flush() wait for in-flight callbacks
The previous `while not self._queue.empty(): await self._queue.join()`
pattern skipped the join entirely when the worker had already dequeued a
task but not yet called task_done(). asyncio.Queue.join() tracks
_unfinished_tasks (incremented by put, decremented by task_done), not
queue depth, so it already handles that case on its own.
2026-04-23 15:06:33 -07:00