Commit graph

38079 commits

Author SHA1 Message Date
Krrish Dholakia
ca0eae194c address greptile review feedback (greploop iteration 6 followup)
- Gate all report_task_result error-path writes on status=pending to
  prevent late error reports from mutating terminal-state tasks

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-01 13:26:45 -07:00
Krrish Dholakia
98aa9dc84f address greptile review feedback (greploop iteration 6)
- Exclude expired tasks from claim_due return value to prevent agents
  from receiving expired task payloads

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-01 13:26:14 -07:00
Krrish Dholakia
de849bff75 address greptile review feedback (greploop iteration 5)
- Sweep expired tasks before counting active to prevent false 429s

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-01 12:51:15 -07:00
Krrish Dholakia
cbc0cb5196 address greptile review feedback (greploop iteration 4)
- Override _require_feature_enabled in test harness to prevent 404 in CI

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-01 12:45:24 -07:00
Krrish Dholakia
0cc6787286 address greptile review feedback (greploop iteration 3)
- Guard naive expires_at in CREATE path to prevent 500 TypeError
- Auto-recompute next_run_at when schedule fields change via PATCH
- Reject schedule_kind='once' with fire_once=False to prevent infinite loop

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-01 12:39:02 -07:00
Krrish Dholakia
621d74a478 address greptile review feedback (greploop iteration 2)
- Fix TOCTOU race in report_task_result using update_many with status guard
- Fix test fake agent_id filter to match production SQL semantics

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-01 12:33:16 -07:00
Krrish Dholakia
933f38362c address greptile review feedback (greploop iteration 1)
- Add feature flag guard (LITELLM_SCHEDULED_TASKS_ENABLED) to all endpoints
- Validate expires_at in PATCH path to reject past timestamps
- Fix TOCTOU race in cancel/update by using update_many with status guard
- Default fire_once to False for cron/interval, True for once
- Fix agent_id=NULL bypass in claim_due SQL to prevent intra-owner leakage

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-01 12:25:45 -07:00
Krrish Dholakia
c6f21d0d84 fix kind='once' silently parked next_run_at at year 9999
Bug: POST /v1/tasks with schedule_kind='once' accepted any spec, ignored
it, and stored next_run_at = 9999-01-01T00:00:00Z. Row was visible in
GET /v1/tasks but never returned by GET /v1/tasks/due — task never
fired. No validation error at create time, no client-visible signal
that the spec was discarded.

Three fixes:

1. compute_next_run('once', spec, ...) now parses spec as an ISO-8601
   timestamp and returns it verbatim (in UTC). Handles trailing 'Z',
   explicit offsets, microseconds, and naive timestamps (treated as
   UTC).

2. validate_schedule('once', spec, ...) now exercises the same parser
   so unparseable specs fail fast at create/update time with HTTP 400
   instead of being silently stored.

3. Doc comment on schedule.py corrected — kind='once' uses an absolute
   ISO-8601 timestamp, not a parked-far-future sentinel.

Tests: 12 new across test_schedule.py and test_endpoints.py covering
ISO with Z / offset / microseconds / naive, rejection of relative
durations / garbage / empty, and end-to-end POST behavior (200 + correct
next_run_at on valid spec, 400 on bad spec).

Reported with reproduction by user against PR head 4a78324.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 17:58:28 -07:00
Krrish Dholakia
58105c1c9c Merge remote-tracking branch 'origin/litellm_internal_staging' into claude/funny-lamarr-6de68c
# Conflicts:
#	litellm-proxy-extras/litellm_proxy_extras/schema.prisma
#	litellm/proxy/schema.prisma
#	schema.prisma
2026-04-29 17:24:50 -07:00
ishaan-berri
4a7af1ff68
feat(proxy): durable agent workflow run tracking via /v1/workflows/runs (#26793)
* feat(schema): add workflow run tracking tables (LiteLLM_WorkflowRun, LiteLLM_WorkflowEvent, LiteLLM_WorkflowMessage)

* feat(proxy): add /v1/workflows/runs endpoints for durable agent workflow tracking

* feat(proxy): register workflow management router in proxy_server

* docs(workflows): add README for workflow run tracking API

* test(workflows): add unit tests for /v1/workflows/runs endpoints

* fix(workflows): atomic event+status update via tx(), run_id 404 guard, sequence retry on collision

* test(workflows): add tx mock, 404 on unknown run_id, retry-on-collision tests

* fix(workflows): constrain status to Literal enum, rename total→count in list responses

* add tenant isolation and bounded limits to workflow endpoints

* add created_by column and index to LiteLLM_WorkflowRun

* add ownership and bounded-limit tests for workflow endpoints

* Fix workflow run ownership for null owners

* guard prisma import in workflow_management_endpoints

* sync schema.prisma copies with workflow run models

* black: format workflow_management_endpoints.py

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-04-29 17:12:18 -07:00
Krrish Dholakia
4a783242f2 fix /v1/tasks/due cross-tenant claim bypass
Greptile/Veria flagged a 7/10 in claim_due:

  WHERE status = 'pending'
    AND ($1::text IS NULL OR agent_id = $1)
    AND ($2::text[] IS NULL OR action = ANY($2))

Most LiteLLM keys have no agent_id set, so $1 was NULL and the IS NULL
branch made the predicate always true. Any authenticated key with NULL
agent_id could call GET /v1/tasks/due and receive every pending task in
the table — including check_prompt and action_args content — and have
those rows mutated (next_run_at advanced or status flipped to 'fired').

Fix: claim_due now requires owner_token (the calling key's hashed token)
and the SQL filters on it unconditionally. agent_id remains an optional
metadata filter applied on top, never the authorization scope.

The /v1/tasks/due endpoint passes user_api_key_dict.token through to
claim_due. Caller never supplies owner_token from the request body — same
contract as create/list/get/update/delete/report.

Regression tests in TestDueTenantIsolation cover both branches:
  - caller with NULL agent_id does NOT see foreign rows, and foreign
    rows are not mutated by the bypass attempt;
  - caller with NULL agent_id still claims its own rows.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 16:40:22 -07:00
Krrish Dholakia
8587f3b770 add error tracking + lazy expiry to scheduled tasks
Three additions to match the local task-scheduler semantics that the
proxy was missing:

1. consecutive_errors + last_error columns on LiteLLM_ScheduledTaskTable.
   New POST /v1/tasks/{task_id}/report endpoint:
       {"result": "success" | "error", "reason": "..."}
   success → resets counter, clears last_error.
   error   → increments counter. On the Nth consecutive error
             (store.MAX_CONSECUTIVE_ERRORS = 3), status flips to 'failed'
             and /due stops emitting the task. Caller's next list() shows
             it as failed with last_error populated; agent renders its
             own user-facing notification — proxy stays out of the
             notification channel.

2. Lazy expiry sweep on every list/get. Previously a task that expired
   before its next fire window sat in 'pending' indefinitely (claim_due
   only flips status when the row is also otherwise due). Added
   sweep_expired_for_owner() called from list_tasks_for_owner and
   get_task_for_owner — flips pending rows past expires_at to 'expired'
   before any read returns them. Cheap update_many gated on the
   (owner_token, status, expires_at) index.

3. Status CHECK constraint widened to include 'failed'.

Tests:
- 6 new tests in TestReport / TestLazyExpiry
- Existing 39 still pass
- /Users/krrishdholakia/Documents/temp_py_folder/test_tasks.py smoke
  script extended with sections 4b (report success/error/failure flip),
  4c (10-task cap), and 6 (lazy expiry)

The per-key 10-task cap was already enforced server-side via
store.MAX_ACTIVE_TASKS_PER_KEY → 429; smoke test now exercises it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 16:35:03 -07:00
Krrish Dholakia
cda2a9cc80 add metadata Json? column to LiteLLM_ScheduledTaskTable
Lets agents stash arbitrary state on a task — session ids, tags,
correlation refs, anything they want surfaced back through /due.
Same shape as action_args (Json?), same encoder path, same omit-when-
None semantics on create + update.

Surfaced in:
- POST /v1/tasks                accepts optional metadata
- PATCH /v1/tasks/{task_id}     can update metadata
- GET /v1/tasks{,/{task_id}}    returns metadata
- GET /v1/tasks/due             returns metadata in claim payload

Schema applied to all three schema.prisma files plus the migration.
Tests cover metadata round-trip and omit-when-not-supplied.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 16:20:09 -07:00
Krrish Dholakia
254352f28e omit action_args from create/update payload when None
Cleaner than encoding null. prisma-client-python rejects None on Json?
columns, and json.dumps(None) round-trips through "null" only to come
back as None anyway. Just don't send the key.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 16:11:09 -07:00
yuneng-jiang
fc0cc9c581
Merge pull request #26225 from BerriAI/litellm_dbReconnectNonBlocking
[Fix] Proxy: reconnect Prisma DB without blocking the event loop
2026-04-29 16:09:22 -07:00
Krrish Dholakia
cb51e0438e always json.dumps action_args for Prisma Json column
Prisma rejects bare None on Json? columns. Switch _serialize_json_for_prisma
to json.dumps unconditionally — None becomes the string "null", which
Postgres jsonb accepts as JSON null and prisma-client-python deserialises
back to None on read.

Mirrors the pattern in memory_endpoints._serialize_metadata_for_prisma.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 16:02:52 -07:00
Krrish Dholakia
e871514c1b drop FK on owner_token to support master-key callers
The proxy's master key is not stored in LiteLLM_VerificationToken, so a
hard FK from LiteLLM_ScheduledTaskTable.owner_token blocked every
master-key task creation with RecordNotFoundError on the nested connect.

Decision: keep owner_token as a plain string column. Cleanup of orphaned
rows when a key is deleted is now an operator concern (acceptable —
rows are small, deletions of keys are infrequent, and the alternative
locks out the most common local-dev auth path).

This change:
- removes the @relation declaration from all three schema.prisma files
- removes the AddForeignKey statement from the migration
- switches create_task back to the scalar owner_token write (no nested
  connect required)
- documents the rationale inline on the model

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 15:56:03 -07:00
Krrish Dholakia
93f8bec80c fix scheduled-tasks create/update against Prisma input shape
Two bugs surfaced on first real proxy invocation:

1. owner_token: prisma-client-python rejected the bare scalar with
   "owner_key: A value is required but not set". Prisma exposes the FK
   via the relation field, not the scalar column. Switched create_task
   to {"owner_key": {"connect": {"token": owner_token}}}.

2. action_args: prisma-client-python rejected the bare Python dict on
   the Json? column ("Invalid argument type. action_args should be of
   any of the following types: NullableJsonNullValueInput, Json"). Added
   _serialize_json_for_prisma helper (same pattern as
   memory_endpoints._serialize_metadata_for_prisma) and applied to both
   create and update paths. Read path round-trips back to native Python.

Updated the in-memory fake in tests to mirror real Prisma write
semantics: owner_key.connect.token unfolds to the owner_token column,
and Json strings are deserialised back on read.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 15:50:36 -07:00
Krrish Dholakia
0f73a091dd remove LITELLM_SCHEDULED_TASKS_ENABLED flag
Always-on. Migration ships the table either way (prisma migrate deploy
runs at startup), endpoints are pure-additive, no existing behavior
changes. Flag was leftover caution from earlier plan iterations that
had a startup ticker loop — that loop was removed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 15:39:08 -07:00
Mateo Wang
9bc317b4d0
Merge pull request #26584 from BerriAI/litellm_mcp-oauth-azure-entra-discovery2
[Feat]Add support for azure entra discovery endpoint
2026-04-29 14:28:41 -07:00
Michael-RZ-Berri
08d35f6b42
Merge pull request #26662 from BerriAI/litellm_spendLogsErrorRedaction
[Fix] Redact spend logs error message
2026-04-29 14:26:48 -07:00
Yuneng Jiang
4f6192a49e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_dbReconnectNonBlocking_local
# Conflicts:
#	tests/test_litellm/proxy/db/test_prisma_self_heal.py
2026-04-29 13:57:35 -07:00
yuneng-jiang
602a6cff81
Merge pull request #26756 from BerriAI/litellm_prisma_reconnect_hardening
fix(proxy): self-heal Prisma read paths + harden reconnect state machine
2026-04-29 13:49:40 -07:00
Mateo Wang
97a3bd5ff4
Merge pull request #26733 from BerriAI/litellm_mcp-short-prefix-id-0e42
feat(mcp): opt-in short-ID tool prefix to keep MCP tool names under the 60-char limit
2026-04-29 13:48:01 -07:00
Krrish Dholakia
e5dec5e809 add /v1/tasks scheduled-task scheduler endpoints
Add Postgres-backed scheduler primitive to the proxy so external agents
can offload "fire this work at time X" without running their own
APScheduler + tasks table. Six endpoints under /v1/tasks behind
LITELLM_SCHEDULED_TASKS_ENABLED:

- POST   /v1/tasks            create
- GET    /v1/tasks            list
- GET    /v1/tasks/{id}       fetch
- PATCH  /v1/tasks/{id}       update (pending only)
- DELETE /v1/tasks/{id}       cancel / stop early
- GET    /v1/tasks/due        atomic claim + schedule advance

Identity (user_id, team_id, agent_id) and ownership (owner_token, FK to
LiteLLM_VerificationToken) are stamped from auth — never accepted from
the request body. /due reads agent_id from the calling key.

claim_due() uses one approved raw-SQL exception per CLAUDE.md
(SELECT FOR UPDATE SKIP LOCKED is required for multi-pod safety and is
not expressible via Prisma model methods). All other writes go through
Prisma model methods, batched inside the same transaction.

Migration adds:
- Partial index on (next_run_at) WHERE status='pending' so the ticker
  scans only candidate rows.
- CHECK constraints for schedule_kind, status, and the
  "action='check' implies check_prompt IS NOT NULL" invariant.

Schedule kinds: interval ('5m'/'2h'/'1d'), cron (5-field crontab with
optional IANA tz), once. compute_next_run is ported from the original
test_local_agent_2/task_runner.py.

Behind LITELLM_SCHEDULED_TASKS_ENABLED=false default — router only
mounts when the flag is on.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-29 12:59:34 -07:00
Mateo Wang
295a36aa69
Merge pull request #26685 from BerriAI/litellm_bedrock_retrievalconfig_passthrough2
feat(vector-stores): support Bedrock retrievalConfiguration passthrough
2026-04-29 12:36:14 -07:00
Sameer Kankute
4cecfec9f9
feat(proxy): LiteLLM headers on Google native generateContent routes (#25500)
* feat(proxy): return LiteLLM headers on Google native generateContent routes

Wire build_litellm_proxy_success_headers_from_llm_response for :generateContent
and :streamGenerateContent so x-litellm-*, rate limit, and provider headers
match the OpenAI-style proxy path. Add unit test.

Annotate httpx.HTTPStatusError branch so pyright accepts .response after optional
exception transform. Remove unused variable in streaming tracer test (Ruff F841).

Made-with: Cursor

* fix(proxy): prefill Google GenAI stream _hidden_params for proxy headers

- Pass model_id, api_base, and process_response_headers output into streaming
  iterators so streamGenerateContent gets the same x-litellm-* headers as
  non-streaming paths.
- Drop request_data deployment mutation from build_litellm_proxy_success_headers_from_llm_response.
- Avoid logging raw request key names in oversized debug payload (code scanning).
- Extend tests for streaming iterator shape, metadata fallback, and helper.

Made-with: Cursor

* Update litellm/proxy/common_request_processing.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* remove unused key count

* Fix greptile review

* Update litellm/proxy/common_request_processing.py

Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com>

---------

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com>
2026-04-29 12:34:14 -07:00
Yassin Kortam
9b3cd5ca25
Merge pull request #26730 from yassinkortam/fix/http-handler-keepalive
fix: add optional TCP SO_KEEPALIVE support to aiohttp's TCPConnector
2026-04-29 10:10:59 -07:00
ishaan-berri
ea275659ac
remove /ui/chat page (#26739)
* remove /ui/chat static page from dashboard build

* add screenshot showing /ui/chat 404

* update screenshots: swagger working, /ui/chat broken

* remove screenshots from repo

* restore screenshots from previous PR
2026-04-29 09:28:57 -07:00
Yassin Kortam
848b79acb5 fix: added keepalive args for aiohttp tcpconnector 2026-04-29 09:14:57 -07:00
Yuneng Jiang
06d9a69444 docs(proxy): clarify _kill_engine_process is on the routine reconnect path
Greptile review on #26225 (P2): the docstring said "Called when disconnect()
fails", and the SIGTERM warning log read "after failed disconnect", but
both were stale — `_kill_engine_process` is now invoked on every routine
reconnect (via the unified `recreate_prisma_client` path), not as a
disconnect-failure recovery branch. The misleading wording would have
produced confusing log lines on every reconnect cycle in production.

Update the docstring to explain the actual reason (avoiding the blocking
`disconnect()` event-loop freeze) and reword the SIGTERM warning to "during
reconnect" so it matches reality.

No behavior change; logs only.
2026-04-28 23:57:28 -07:00
Yuneng Jiang
aa2ef41200 fix(proxy): preserve original transport error if reconnect itself raises
Greptile review on #26756 (P2): if `attempt_db_reconnect` itself raises
(e.g. lock cancellation, timer error, unexpected internal failure), the
original `httpx.ReadError` / transport error was lost — `failure_handler`
and `db_exceptions` alerts then logged the reconnect exception instead of
the actual DB transport problem, masking the root cause.

Wrap the reconnect call in a try/except. On reconnect failure, re-raise
the *original* `first_exc` and chain the reconnect error as `__cause__`
so it remains visible for debuggability without becoming the primary
exception observers see.

Adds `test_call_with_db_reconnect_retry_preserves_original_error_when_reconnect_raises`
asserting (a) the propagated exception is the original transport error
and (b) the reconnect exception is attached as `__cause__`.
2026-04-28 23:55:46 -07:00
Yuneng Jiang
1c9c219a74 fix(proxy): self-heal Prisma read paths + harden reconnect state machine
Two related fixes layered on top of the existing reconnect plumbing:

1. Restore reconnect-and-retry on `PrismaClient.get_generic_data` (issue
   #25143). 1.83.x lost the transport-reconnect-and-retry-once branch that
   1.82.6 had on this method, so transient `httpx.ReadError` flaps now
   surface immediately as `db_exceptions` alerts. `_update_config_from_db`
   fans out four concurrent `get_generic_data` reads, so a single transport
   blip used to mark four alerts and a stale config window.

   Adds `call_with_db_reconnect_retry` to `litellm/proxy/db/exception_handler.py`
   — a single canonical "try DB read, on transport error reconnect once and
   retry once" wrapper. Mirrors the inline pattern in
   `auth_checks._fetch_key_object_from_db_with_reconnect` so we have one
   implementation rather than three drifting copies, and gives future read
   paths a clean opt-in.

2. Fix the `_engine_confirmed_dead` flag-reset bug in
   `_run_reconnect_cycle`. The flag was cleared before `_do_heavy_reconnect()`
   ran, so any failure inside the heavy reconnect (timeout, missing
   DATABASE_URL, recreate failure) left the flag False — and the next
   attempt could silently demote to the lightweight path even though the
   engine was genuinely dead. Move the reset into the success branch so the
   flag stays True across heavy-reconnect failures and the next attempt
   re-enters the heavy branch.

Tests:

- `tests/test_litellm/proxy/db/test_exception_handler_reconnect_retry.py`
  (new) — 9 tests covering the helper's contract: happy path, retry on
  transport error, no retry on data-layer errors, propagation when reconnect
  fails, propagation after second transport error, `hasattr` guard for
  partial mocks, fresh-coroutine-per-call invariant, explicit timeout
  override, default timeouts read off the prisma_client.
- `tests/test_litellm/proxy/db/test_prisma_self_heal.py` — adds:
  - `test_get_generic_data_retries_on_transport_error_for_config_table`
  - `test_get_generic_data_propagates_when_reconnect_fails`
  - `test_engine_confirmed_dead_persists_across_failed_heavy_reconnect`
    (regression test for the flag-reset bug).

All 16 self-heal tests + 9 helper tests + 535 auth/exception-handler tests
pass locally.
2026-04-28 23:44:34 -07:00
Yuneng Jiang
8c91c8b2c4 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_dbReconnectNonBlocking 2026-04-28 23:29:36 -07:00
Mateo Wang
6e6b2ca2d8
Merge pull request #26741 from BerriAI/litellm_fix-model-alias-flake-c5db 2026-04-28 21:28:13 -07:00
Cursor Agent
3fb5056305
fix(mcp): address greptile review on short tool prefix
- server.py: drop the redundant server_id append in
  _get_filtered_mcp_servers_from_mcp_server_names. iter_known_server_prefixes
  already yields server_id unconditionally, so the manual append (and its
  misleading comment) was a no-op duplicate.
- utils.py: rewrite the SHORT_MCP_TOOL_PREFIX docstring to accurately
  describe the collision behaviour. The previous wording said collisions
  were 'cosmetic only', but a natural-hash collision IS a routing-correctness
  issue, which is precisely why we already added _assign_unique_short_prefix
  to rehash deterministically. The new comment cross-references that path.
- utils.py: restrict the first character of the short prefix to [A-Za-z]
  via a 52-char alphabet for position 0 only. The remaining two positions
  still use the full base62 alphabet. This keeps prefixes valid identifiers
  on every backend and gives 52*62*62 = 199_888 distinct prefixes (still
  comfortably more than any realistic deployment).
- tests: add coverage proving the first character of the prefix is always
  alphabetic across many server_ids and rehash attempts.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-04-29 03:59:40 +00:00
Cursor Agent
6b3f07ba25
fix(mcp): register OpenAPI tools after short prefix collision resolution in reload 2026-04-29 03:53:03 +00:00
Sameer Kankute
af5b7be51d
Merge pull request #26742 from BerriAI/litellm_internal_staging
merge main
2026-04-29 09:20:12 +05:30
Cursor Agent
3215874e40
fix(test): scope ERROR log assertion to LiteLLM logger in test_model_alias_map
The test was flaking on unrelated asyncio ERROR records (e.g. "Unclosed
client session" from background tasks in other tests). Restrict the
assertion to records emitted by LiteLLM loggers so the test only fails
on errors actually produced by the code under test.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-04-29 03:48:41 +00:00
Cursor Agent
df3dbd18d6
feat(mcp): rehash short tool prefix on collision and cache per server
Two MCP servers can natural-hash to the same three-character base62
prefix. With 62**3 = 238_328 slots the birthday bound is ~488 servers
for 50% collision probability, so a single proxy hosting more than
~100 MCP servers has a non-trivial chance of seeing a collision in
practice — and a collision means tool names from two different servers
share a routing key, causing silent mis-routing.

Mitigation:

- compute_short_server_prefix(server_id, attempt=N) folds an attempt
  counter into the SHA-256 seed, so rehashes are deterministic and
  produce a fresh three-char prefix space per attempt.
- New MCPServer.short_prefix field caches the resolved (post-dedup)
  prefix on the model so it stays stable across the process lifetime.
- MCPServerManager._assign_unique_short_prefix walks attempts 0..N
  until it finds a prefix not already used by another server in the
  combined registry. Logs an INFO line when a rehash happens so
  operators have a breadcrumb if it ever does.
- Wired into every registration path: load_servers_from_config,
  add_server, update_server, reload_servers_from_database. The
  database reload path also carries the previously-resolved prefix
  forward so reloads don't churn it.
- get_server_prefix prefers the cached short_prefix when set, so the
  resolved value (not the raw natural hash) is used everywhere.
- iter_known_server_prefixes yields the cached short_prefix too, so
  reverse-lookup tolerance covers the rehashed form.

No-op when LITELLM_USE_SHORT_MCP_TOOL_PREFIX is disabled — the field
stays None and behaviour is unchanged.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-04-29 03:43:34 +00:00
Sameer Kankute
e0cd536eaa
Fix lint 2026-04-29 09:06:34 +05:30
Sameer Kankute
b516120036
Merge pull request #26737 from BerriAI/litellm_internal_staging
merge internal staging
2026-04-29 08:50:12 +05:30
xinrui
44ab016743
feat(provider): add AIHubMix as an OpenAI-compatible provider (#24294)
* feat: add AIHubMix provider to providers.json

* fix: add aihubmix to provider_endpoints_support.json for CI check

---------

Co-authored-by: yuneng-jiang <yuneng@berri.ai>
2026-04-28 20:18:30 -07:00
ishaan-berri
4ae2996f08
Add gpt-image-2 support (#26644) (#26705)
* Add gpt-image-2 support

* Address gpt-image-2 PR feedback

Co-authored-by: Emerson Gomes <emerson.gomes@thalesgroup.com>
2026-04-28 20:10:42 -07:00
Sameer Kankute
cf74f55b79
Fix extra body error 2026-04-29 08:34:31 +05:30
mateo-berri
4e827446d2 fix: type error 2026-04-28 19:58:56 -07:00
yuneng-jiang
804e7c0c7b
Merge pull request #26734 from BerriAI/yj/create-release-pep440-tags
Some checks are pending
Unit Tests: Caching (Redis) / caching-redis (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / schema-migration (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Security / security (push) Waiting to run
ci(release): accept PEP 440 tag forms in create-release workflow
2026-04-28 19:44:58 -07:00
Yuneng Jiang
3a5980804c ci(release): mark rc / dev / nightly tags as GitHub pre-releases
`prerelease: false` was hardcoded, so dispatching create-release with
`1.84.0rc1`, `1.84.0.dev42`, or legacy `v1.83.13-nightly` would publish
them as stable releases on the GitHub Releases page. Derive the flag
from the tag instead.

The detector matches `rc`, `.dev`, `nightly`, `alpha`, `beta`. PEP 440
post-releases (`1.84.0.post1`) and legacy `-stable[.patch.N]` are
stable maintenance releases per PEP 440, so they intentionally do not
match.
2026-04-28 19:38:13 -07:00
Yuneng Jiang
1da1eb661b ci(release): accept PEP 440 tag forms in create-release workflow
The tag validator required a leading `v`, so dispatching create-release
with `1.84.0` (or `1.84.0rc1`, `1.84.0.dev42`, `1.84.0.post1`) failed
even though those are the new naming convention. Make the leading `v`
optional in both create-release.yml and create-release-branch.yml so
both legacy (`v1.83.10-stable`, `v1.83.14.rc.1`, `v1.82.3.dev.9`,
`v1.82.3-stable.patch.4`, `v1.83.13-nightly`) and new PEP 440 forms are
accepted during the transition. Refresh the input descriptions to show
the new examples.
2026-04-28 19:33:18 -07:00
Cursor Agent
fc49c181bc
feat(mcp): opt-in short-ID tool prefix to stay under 60-char tool name limit
Adds LITELLM_USE_SHORT_MCP_TOOL_PREFIX. When enabled, tool / prompt /
resource / resource-template names emitted from MCP servers are prefixed
with a deterministic three-character base62 ID derived from the server's
server_id (SHA-256 → base62) instead of the (potentially long)
alias / server_name. This keeps namespaced tool names well under the
60-character upper bound enforced by some model APIs while still letting
us distinguish MCP-routed tools from local tools.

Behavioural notes:

- Default off — when the env var is unset, the long-prefix behaviour
  is unchanged. The plan is to flip the default in a future release
  and remove the gate after a deprecation window.
- Prefix derivation is deterministic, so it is stable across processes,
  workers and restarts without any persistence layer.
- Reverse-lookup is tolerant: _create_prefixed_tools registers every
  known prefix form (alias / server_name / server_id / short ID) in
  the routing map and _get_mcp_server_from_tool_name resolves any of
  them. Old clients holding cached long-prefixed names continue to
  route correctly even after the flag is enabled.
- _get_allowed_mcp_servers_from_mcp_server_names accepts the short
  prefix in /mcp/{server_name}-style URLs.
- The OpenAPI tool-listing path now filters by the active server
  prefix instead of server.name so spec-backed servers benefit too.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-04-29 01:41:24 +00:00