- Guard naive expires_at in CREATE path to prevent 500 TypeError
- Auto-recompute next_run_at when schedule fields change via PATCH
- Reject schedule_kind='once' with fire_once=False to prevent infinite loop
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Fix TOCTOU race in report_task_result using update_many with status guard
- Fix test fake agent_id filter to match production SQL semantics
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add feature flag guard (LITELLM_SCHEDULED_TASKS_ENABLED) to all endpoints
- Validate expires_at in PATCH path to reject past timestamps
- Fix TOCTOU race in cancel/update by using update_many with status guard
- Default fire_once to False for cron/interval, True for once
- Fix agent_id=NULL bypass in claim_due SQL to prevent intra-owner leakage
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Bug: POST /v1/tasks with schedule_kind='once' accepted any spec, ignored
it, and stored next_run_at = 9999-01-01T00:00:00Z. Row was visible in
GET /v1/tasks but never returned by GET /v1/tasks/due — task never
fired. No validation error at create time, no client-visible signal
that the spec was discarded.
Three fixes:
1. compute_next_run('once', spec, ...) now parses spec as an ISO-8601
timestamp and returns it verbatim (in UTC). Handles trailing 'Z',
explicit offsets, microseconds, and naive timestamps (treated as
UTC).
2. validate_schedule('once', spec, ...) now exercises the same parser
so unparseable specs fail fast at create/update time with HTTP 400
instead of being silently stored.
3. Doc comment on schedule.py corrected — kind='once' uses an absolute
ISO-8601 timestamp, not a parked-far-future sentinel.
Tests: 12 new across test_schedule.py and test_endpoints.py covering
ISO with Z / offset / microseconds / naive, rejection of relative
durations / garbage / empty, and end-to-end POST behavior (200 + correct
next_run_at on valid spec, 400 on bad spec).
Reported with reproduction by user against PR head 4a78324.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(schema): add workflow run tracking tables (LiteLLM_WorkflowRun, LiteLLM_WorkflowEvent, LiteLLM_WorkflowMessage)
* feat(proxy): add /v1/workflows/runs endpoints for durable agent workflow tracking
* feat(proxy): register workflow management router in proxy_server
* docs(workflows): add README for workflow run tracking API
* test(workflows): add unit tests for /v1/workflows/runs endpoints
* fix(workflows): atomic event+status update via tx(), run_id 404 guard, sequence retry on collision
* test(workflows): add tx mock, 404 on unknown run_id, retry-on-collision tests
* fix(workflows): constrain status to Literal enum, rename total→count in list responses
* add tenant isolation and bounded limits to workflow endpoints
* add created_by column and index to LiteLLM_WorkflowRun
* add ownership and bounded-limit tests for workflow endpoints
* Fix workflow run ownership for null owners
* guard prisma import in workflow_management_endpoints
* sync schema.prisma copies with workflow run models
* black: format workflow_management_endpoints.py
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Greptile/Veria flagged a 7/10 in claim_due:
WHERE status = 'pending'
AND ($1::text IS NULL OR agent_id = $1)
AND ($2::text[] IS NULL OR action = ANY($2))
Most LiteLLM keys have no agent_id set, so $1 was NULL and the IS NULL
branch made the predicate always true. Any authenticated key with NULL
agent_id could call GET /v1/tasks/due and receive every pending task in
the table — including check_prompt and action_args content — and have
those rows mutated (next_run_at advanced or status flipped to 'fired').
Fix: claim_due now requires owner_token (the calling key's hashed token)
and the SQL filters on it unconditionally. agent_id remains an optional
metadata filter applied on top, never the authorization scope.
The /v1/tasks/due endpoint passes user_api_key_dict.token through to
claim_due. Caller never supplies owner_token from the request body — same
contract as create/list/get/update/delete/report.
Regression tests in TestDueTenantIsolation cover both branches:
- caller with NULL agent_id does NOT see foreign rows, and foreign
rows are not mutated by the bypass attempt;
- caller with NULL agent_id still claims its own rows.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three additions to match the local task-scheduler semantics that the
proxy was missing:
1. consecutive_errors + last_error columns on LiteLLM_ScheduledTaskTable.
New POST /v1/tasks/{task_id}/report endpoint:
{"result": "success" | "error", "reason": "..."}
success → resets counter, clears last_error.
error → increments counter. On the Nth consecutive error
(store.MAX_CONSECUTIVE_ERRORS = 3), status flips to 'failed'
and /due stops emitting the task. Caller's next list() shows
it as failed with last_error populated; agent renders its
own user-facing notification — proxy stays out of the
notification channel.
2. Lazy expiry sweep on every list/get. Previously a task that expired
before its next fire window sat in 'pending' indefinitely (claim_due
only flips status when the row is also otherwise due). Added
sweep_expired_for_owner() called from list_tasks_for_owner and
get_task_for_owner — flips pending rows past expires_at to 'expired'
before any read returns them. Cheap update_many gated on the
(owner_token, status, expires_at) index.
3. Status CHECK constraint widened to include 'failed'.
Tests:
- 6 new tests in TestReport / TestLazyExpiry
- Existing 39 still pass
- /Users/krrishdholakia/Documents/temp_py_folder/test_tasks.py smoke
script extended with sections 4b (report success/error/failure flip),
4c (10-task cap), and 6 (lazy expiry)
The per-key 10-task cap was already enforced server-side via
store.MAX_ACTIVE_TASKS_PER_KEY → 429; smoke test now exercises it.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Lets agents stash arbitrary state on a task — session ids, tags,
correlation refs, anything they want surfaced back through /due.
Same shape as action_args (Json?), same encoder path, same omit-when-
None semantics on create + update.
Surfaced in:
- POST /v1/tasks accepts optional metadata
- PATCH /v1/tasks/{task_id} can update metadata
- GET /v1/tasks{,/{task_id}} returns metadata
- GET /v1/tasks/due returns metadata in claim payload
Schema applied to all three schema.prisma files plus the migration.
Tests cover metadata round-trip and omit-when-not-supplied.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Cleaner than encoding null. prisma-client-python rejects None on Json?
columns, and json.dumps(None) round-trips through "null" only to come
back as None anyway. Just don't send the key.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Prisma rejects bare None on Json? columns. Switch _serialize_json_for_prisma
to json.dumps unconditionally — None becomes the string "null", which
Postgres jsonb accepts as JSON null and prisma-client-python deserialises
back to None on read.
Mirrors the pattern in memory_endpoints._serialize_metadata_for_prisma.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The proxy's master key is not stored in LiteLLM_VerificationToken, so a
hard FK from LiteLLM_ScheduledTaskTable.owner_token blocked every
master-key task creation with RecordNotFoundError on the nested connect.
Decision: keep owner_token as a plain string column. Cleanup of orphaned
rows when a key is deleted is now an operator concern (acceptable —
rows are small, deletions of keys are infrequent, and the alternative
locks out the most common local-dev auth path).
This change:
- removes the @relation declaration from all three schema.prisma files
- removes the AddForeignKey statement from the migration
- switches create_task back to the scalar owner_token write (no nested
connect required)
- documents the rationale inline on the model
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two bugs surfaced on first real proxy invocation:
1. owner_token: prisma-client-python rejected the bare scalar with
"owner_key: A value is required but not set". Prisma exposes the FK
via the relation field, not the scalar column. Switched create_task
to {"owner_key": {"connect": {"token": owner_token}}}.
2. action_args: prisma-client-python rejected the bare Python dict on
the Json? column ("Invalid argument type. action_args should be of
any of the following types: NullableJsonNullValueInput, Json"). Added
_serialize_json_for_prisma helper (same pattern as
memory_endpoints._serialize_metadata_for_prisma) and applied to both
create and update paths. Read path round-trips back to native Python.
Updated the in-memory fake in tests to mirror real Prisma write
semantics: owner_key.connect.token unfolds to the owner_token column,
and Json strings are deserialised back on read.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Always-on. Migration ships the table either way (prisma migrate deploy
runs at startup), endpoints are pure-additive, no existing behavior
changes. Flag was leftover caution from earlier plan iterations that
had a startup ticker loop — that loop was removed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add Postgres-backed scheduler primitive to the proxy so external agents
can offload "fire this work at time X" without running their own
APScheduler + tasks table. Six endpoints under /v1/tasks behind
LITELLM_SCHEDULED_TASKS_ENABLED:
- POST /v1/tasks create
- GET /v1/tasks list
- GET /v1/tasks/{id} fetch
- PATCH /v1/tasks/{id} update (pending only)
- DELETE /v1/tasks/{id} cancel / stop early
- GET /v1/tasks/due atomic claim + schedule advance
Identity (user_id, team_id, agent_id) and ownership (owner_token, FK to
LiteLLM_VerificationToken) are stamped from auth — never accepted from
the request body. /due reads agent_id from the calling key.
claim_due() uses one approved raw-SQL exception per CLAUDE.md
(SELECT FOR UPDATE SKIP LOCKED is required for multi-pod safety and is
not expressible via Prisma model methods). All other writes go through
Prisma model methods, batched inside the same transaction.
Migration adds:
- Partial index on (next_run_at) WHERE status='pending' so the ticker
scans only candidate rows.
- CHECK constraints for schedule_kind, status, and the
"action='check' implies check_prompt IS NOT NULL" invariant.
Schedule kinds: interval ('5m'/'2h'/'1d'), cron (5-field crontab with
optional IANA tz), once. compute_next_run is ported from the original
test_local_agent_2/task_runner.py.
Behind LITELLM_SCHEDULED_TASKS_ENABLED=false default — router only
mounts when the flag is on.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Greptile review on #26225 (P2): the docstring said "Called when disconnect()
fails", and the SIGTERM warning log read "after failed disconnect", but
both were stale — `_kill_engine_process` is now invoked on every routine
reconnect (via the unified `recreate_prisma_client` path), not as a
disconnect-failure recovery branch. The misleading wording would have
produced confusing log lines on every reconnect cycle in production.
Update the docstring to explain the actual reason (avoiding the blocking
`disconnect()` event-loop freeze) and reword the SIGTERM warning to "during
reconnect" so it matches reality.
No behavior change; logs only.
Greptile review on #26756 (P2): if `attempt_db_reconnect` itself raises
(e.g. lock cancellation, timer error, unexpected internal failure), the
original `httpx.ReadError` / transport error was lost — `failure_handler`
and `db_exceptions` alerts then logged the reconnect exception instead of
the actual DB transport problem, masking the root cause.
Wrap the reconnect call in a try/except. On reconnect failure, re-raise
the *original* `first_exc` and chain the reconnect error as `__cause__`
so it remains visible for debuggability without becoming the primary
exception observers see.
Adds `test_call_with_db_reconnect_retry_preserves_original_error_when_reconnect_raises`
asserting (a) the propagated exception is the original transport error
and (b) the reconnect exception is attached as `__cause__`.
Two related fixes layered on top of the existing reconnect plumbing:
1. Restore reconnect-and-retry on `PrismaClient.get_generic_data` (issue
#25143). 1.83.x lost the transport-reconnect-and-retry-once branch that
1.82.6 had on this method, so transient `httpx.ReadError` flaps now
surface immediately as `db_exceptions` alerts. `_update_config_from_db`
fans out four concurrent `get_generic_data` reads, so a single transport
blip used to mark four alerts and a stale config window.
Adds `call_with_db_reconnect_retry` to `litellm/proxy/db/exception_handler.py`
— a single canonical "try DB read, on transport error reconnect once and
retry once" wrapper. Mirrors the inline pattern in
`auth_checks._fetch_key_object_from_db_with_reconnect` so we have one
implementation rather than three drifting copies, and gives future read
paths a clean opt-in.
2. Fix the `_engine_confirmed_dead` flag-reset bug in
`_run_reconnect_cycle`. The flag was cleared before `_do_heavy_reconnect()`
ran, so any failure inside the heavy reconnect (timeout, missing
DATABASE_URL, recreate failure) left the flag False — and the next
attempt could silently demote to the lightweight path even though the
engine was genuinely dead. Move the reset into the success branch so the
flag stays True across heavy-reconnect failures and the next attempt
re-enters the heavy branch.
Tests:
- `tests/test_litellm/proxy/db/test_exception_handler_reconnect_retry.py`
(new) — 9 tests covering the helper's contract: happy path, retry on
transport error, no retry on data-layer errors, propagation when reconnect
fails, propagation after second transport error, `hasattr` guard for
partial mocks, fresh-coroutine-per-call invariant, explicit timeout
override, default timeouts read off the prisma_client.
- `tests/test_litellm/proxy/db/test_prisma_self_heal.py` — adds:
- `test_get_generic_data_retries_on_transport_error_for_config_table`
- `test_get_generic_data_propagates_when_reconnect_fails`
- `test_engine_confirmed_dead_persists_across_failed_heavy_reconnect`
(regression test for the flag-reset bug).
All 16 self-heal tests + 9 helper tests + 535 auth/exception-handler tests
pass locally.
- server.py: drop the redundant server_id append in
_get_filtered_mcp_servers_from_mcp_server_names. iter_known_server_prefixes
already yields server_id unconditionally, so the manual append (and its
misleading comment) was a no-op duplicate.
- utils.py: rewrite the SHORT_MCP_TOOL_PREFIX docstring to accurately
describe the collision behaviour. The previous wording said collisions
were 'cosmetic only', but a natural-hash collision IS a routing-correctness
issue, which is precisely why we already added _assign_unique_short_prefix
to rehash deterministically. The new comment cross-references that path.
- utils.py: restrict the first character of the short prefix to [A-Za-z]
via a 52-char alphabet for position 0 only. The remaining two positions
still use the full base62 alphabet. This keeps prefixes valid identifiers
on every backend and gives 52*62*62 = 199_888 distinct prefixes (still
comfortably more than any realistic deployment).
- tests: add coverage proving the first character of the prefix is always
alphabetic across many server_ids and rehash attempts.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
The test was flaking on unrelated asyncio ERROR records (e.g. "Unclosed
client session" from background tasks in other tests). Restrict the
assertion to records emitted by LiteLLM loggers so the test only fails
on errors actually produced by the code under test.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Two MCP servers can natural-hash to the same three-character base62
prefix. With 62**3 = 238_328 slots the birthday bound is ~488 servers
for 50% collision probability, so a single proxy hosting more than
~100 MCP servers has a non-trivial chance of seeing a collision in
practice — and a collision means tool names from two different servers
share a routing key, causing silent mis-routing.
Mitigation:
- compute_short_server_prefix(server_id, attempt=N) folds an attempt
counter into the SHA-256 seed, so rehashes are deterministic and
produce a fresh three-char prefix space per attempt.
- New MCPServer.short_prefix field caches the resolved (post-dedup)
prefix on the model so it stays stable across the process lifetime.
- MCPServerManager._assign_unique_short_prefix walks attempts 0..N
until it finds a prefix not already used by another server in the
combined registry. Logs an INFO line when a rehash happens so
operators have a breadcrumb if it ever does.
- Wired into every registration path: load_servers_from_config,
add_server, update_server, reload_servers_from_database. The
database reload path also carries the previously-resolved prefix
forward so reloads don't churn it.
- get_server_prefix prefers the cached short_prefix when set, so the
resolved value (not the raw natural hash) is used everywhere.
- iter_known_server_prefixes yields the cached short_prefix too, so
reverse-lookup tolerance covers the rehashed form.
No-op when LITELLM_USE_SHORT_MCP_TOOL_PREFIX is disabled — the field
stays None and behaviour is unchanged.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* feat: add AIHubMix provider to providers.json
* fix: add aihubmix to provider_endpoints_support.json for CI check
---------
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
`prerelease: false` was hardcoded, so dispatching create-release with
`1.84.0rc1`, `1.84.0.dev42`, or legacy `v1.83.13-nightly` would publish
them as stable releases on the GitHub Releases page. Derive the flag
from the tag instead.
The detector matches `rc`, `.dev`, `nightly`, `alpha`, `beta`. PEP 440
post-releases (`1.84.0.post1`) and legacy `-stable[.patch.N]` are
stable maintenance releases per PEP 440, so they intentionally do not
match.
The tag validator required a leading `v`, so dispatching create-release
with `1.84.0` (or `1.84.0rc1`, `1.84.0.dev42`, `1.84.0.post1`) failed
even though those are the new naming convention. Make the leading `v`
optional in both create-release.yml and create-release-branch.yml so
both legacy (`v1.83.10-stable`, `v1.83.14.rc.1`, `v1.82.3.dev.9`,
`v1.82.3-stable.patch.4`, `v1.83.13-nightly`) and new PEP 440 forms are
accepted during the transition. Refresh the input descriptions to show
the new examples.
Adds LITELLM_USE_SHORT_MCP_TOOL_PREFIX. When enabled, tool / prompt /
resource / resource-template names emitted from MCP servers are prefixed
with a deterministic three-character base62 ID derived from the server's
server_id (SHA-256 → base62) instead of the (potentially long)
alias / server_name. This keeps namespaced tool names well under the
60-character upper bound enforced by some model APIs while still letting
us distinguish MCP-routed tools from local tools.
Behavioural notes:
- Default off — when the env var is unset, the long-prefix behaviour
is unchanged. The plan is to flip the default in a future release
and remove the gate after a deprecation window.
- Prefix derivation is deterministic, so it is stable across processes,
workers and restarts without any persistence layer.
- Reverse-lookup is tolerant: _create_prefixed_tools registers every
known prefix form (alias / server_name / server_id / short ID) in
the routing map and _get_mcp_server_from_tool_name resolves any of
them. Old clients holding cached long-prefixed names continue to
route correctly even after the flag is enabled.
- _get_allowed_mcp_servers_from_mcp_server_names accepts the short
prefix in /mcp/{server_name}-style URLs.
- The OpenAPI tool-listing path now filters by the active server
prefix instead of server.name so spec-backed servers benefit too.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>