* fix(azure_sentinel): split batches under the 1MB ingestion cap and keep undelivered records queued
Azure Monitor rejects any Logs Ingestion body over 1MB with a 413. The Sentinel logger
posted the whole queue as one body and cleared it in a finally block, so an oversize
batch, a transient 5xx, or a failed token call dropped every queued record, and records
logged while a send was in flight were cleared with it. Both the standard and the audit
queue share the sender.
Move Datadog's proactive size split and 413 halving into a shared helper,
litellm/integrations/batch_utils.send_batch_with_413_split, and route Sentinel through it
with a 1MB size check. A lone record that still 413s is dropped, everything a transient
failure leaves undelivered goes back to the front of its queue, and the retry queue is
capped at max_queue_size so an unreachable workspace cannot grow memory without bound
* fix(azure_sentinel): retry undelivered records on the flush timer only
Requeued records made every later event cross the batch_size threshold, so a
down ingestion endpoint got one full-queue resend per request. Threshold sends
now go through flush_queue, so they take the flush lock instead of racing the
timer, and they stand down while records are awaiting retry.
A record that cannot be serialized raised out of the size probe and killed the
periodic flush task. The probe now runs inside the failure handling, so the
batch is split and only the record that cannot be serialized is dropped.
* fix(azure_sentinel): decide threshold sends under the flush lock
Concurrent callbacks all read logs_awaiting_retry before the first send
finished, so each one resent the whole queue once that send failed. The
flag and the batch_size threshold are now rechecked while holding the
flush lock, and each queue sends only itself instead of going through
flush_queue, which was retrying the other queue too.
* test(azure_sentinel): cover successful threshold waiters
* fix(azure_sentinel): preserve cancelled batches for retry
* fix(azure_sentinel): requeue only the undelivered part of a cancelled split
A batch over the ingestion cap goes out in pieces, so a cancellation partway
through requeued pieces the destination had already accepted and sent them a
second time on the next flush
The split helper now raises a cancellation carrying the records it never
delivered, and Azure Sentinel requeues those instead of the whole batch
* fix(azure_sentinel): drop batches a permanent rejection will never accept
A non-413 4xx from the ingestion endpoint or from the OAuth token call means the request
will fail the same way on every retry, so requeueing it held the batch, and every record
logged behind it, until the queue cap dropped them. Retryable statuses (5xx, 408, 429)
still keep the whole batch, and a shared classifier gives Datadog the same rule
The serialization probe now catches any exception, not just TypeError and ValueError,
because safe_dumps hands pydantic models to model_dump and can raise anything. It also
splits on record count, so a recovery flush sends batch_size records per request instead
of serializing the whole requeued queue to measure it
Both integrations re-raise a cancelled send as exactly asyncio.CancelledError. Python
3.12's asyncio.wait_for only translates the exact class into TimeoutError, so the
BatchSendCancelled subclass escaped the logging worker as an unhandled error
The awaiting-retry flag now follows the queue that survived the max_queue_size trim, so
a deployment with the cap at zero is not left waiting for a timer flush with nothing
queued to retry
* chore(logging): document mutable queue ownership
Annotate the queue detach and requeue constructions required by the logger's appendable queue contract so the type-discipline budget stays clean
* fix(datadog): preserve non-413 retry behavior
Keep Datadog's existing contract of requeuing every non-413 HTTP failure while Azure Sentinel applies its permanent-client-error policy through the shared splitter
* fix(batch_utils): requeue by default and let Sentinel opt into dropping
The shared splitter's default non-success handler is now requeue_after_http_error, the behavior Datadog had before the extraction, so a caller that omits the argument keeps its records. Azure Sentinel passes undelivered_after_http_error explicitly to drop permanent 4xx rejections
Also drops an explicit return None the strict ruff gate flags in the test helper
A gpt-6-astra deployment on a Foundry project reached through the
azure_ai route had no cost map entry of its own, so it resolved to the
OpenAI gpt-6-astra card: missing from the azure_ai/* wildcard listing,
flex and priority prices and /v1/batch it does not sell, and no none
reasoning effort. Add azure_ai/gpt-6-astra mirroring the
azure/gpt-6-astra Standard Global sheet the way azure_ai/gpt-5.5 mirrors
azure/gpt-5.5, and extend the cost, reasoning-effort, and wildcard
listing tests to the Foundry route.
OpenAI rejects every reasoning.effort on chat-latest except medium. With supports_reasoning set and no declared levels the entry resolved to None, so /model_group/info and the dashboard effort pickers had nothing to narrow the offered levels with
Format the model-picked id with %r so control characters in it cannot
break the log line. The regression test for the unlisted id keeps to
generic scoping wording
The sidebar and header were still keyed on legacy ?page= ids and mapped
back and forth through MIGRATED_PAGES, legacyPageHref and
legacyKeyForPathname. Leaves are now plain Next links to their path
route, the active item and breadcrumb come from usePathname, and the
setPage/defaultSelectedKey prop chain is gone.
The id-to-route table moves next to the dashboard root page as its only
consumer. That redirect now forwards the remaining query params instead
of dropping them, so deep links such as the proxy's MCP env-var setup
link (?page=mcp-servers&fill_env_vars=) no longer rely on the target page
reading the pre-redirect URL during its first render. The proxy builds
that link as /ui/mcp-servers?fill_env_vars= directly, and the Playground
warnings link to the real routes instead of relative ?page= URLs.
migratedHref is renamed uiHref, the /ui base-path helper it always was.
Every retrieve of a batch through the proxy shares one spend row, the batch id
plus the batch cost suffix, and spend log inserts skip duplicates. A poll that
landed while the batch was still validating or in progress wrote that row at
$0 and no later retrieve could overwrite it, and every completed retrieve after
the first added the cost to the key, team, and user counters again with no new
row to show for it.
The cost callback now writes nothing for a batch retrieve until the batch is
final, releasing the poll's budget reservation instead, and once it is final it
charges only when no spend row for that batch is queued for flush or already
stored. Batch cost rows are flushed to the database right away so a second
instance sees them, and the logger prices a batch only once it is final, which
also covers a failed batch that never produced an output file.
A one-character value in the provider secret bundle was masked too, which
turned every 1 in the run log into ***, including the pass numbers and the
gateway addresses, so the only public diagnostics were unreadable
Emulated file_search now warns when the model returns a vector_store_id that is
not one of the request's stores, naming the dropped id and the stores that were
searched instead. H16 asserts the warning is emitted exactly once.
A live cost map older than this release, or a proxy whose map fetch lags, could
strip `reasoning` from a model this release knows accepts it. The bundled map is
now the floor: any OpenAI entry it flags as reasoning keeps the param whatever
the live map says. Fine-tuned ids with an empty suffix (`ft:gpt-4o-2024-08-06:org::id`)
now resolve to their base entry instead of failing open, `chat-latest` carries
the flag, and the schema test keeps every codex, deep-research, and chat-latest
entry flagged. The none-effort check goes through a public wrapper so the
responses config stops importing a private helper.
_average_latency skipped integer samples in the sum while counting them in the denominator, which contradicted its own
Sequence[float | int] signature; it now averages every sample. Both success loggers in cost-based routing computed
response_ms / completion_tokens and threw the result away, so a chat response with zero completion tokens raised
ZeroDivisionError inside the handler. The proxy swallows and logs it, but the handler then skips that request's tpm and
rpm update, so cost-based routing undercounts the deployment's usage. The QA run for the latency fix hit it on real
gpt-5.5 traffic through /v1/chat/completions and /v1/messages
The regression test's recording logger overrode async_post_call_failure_hook
with untyped parameters. It now mirrors the base signature, and the
UserAPIKeyAuth import moves to module level so the annotation resolves.
The revert restored the dict annotation that #39121 had loosened to Mapping,
and the LIT001 ceiling has been ratcheted down since, so the gate rejected the
one reintroduced hit. Annotation only, no behavior change.
The emulated file_search handler searched whatever vector_store_id the
model returned, so a model steered to an id outside the request's
file_search tool reached a store the per-key vector store permission
check never saw. An id outside the request's stores now falls back to
those stores; an id that is one of them still narrows the search to it.
The changed-tests workflow overrode the suite's `--reruns 1` with `--reruns 0`, so a
transport blip failed a pass that pytest.ini already scopes to network errors and
5xx responses. Pass 2 of run 33692484803 also went red 15s after a model write with
"no healthy deployments": the barrier only polled /v1/models through nginx, which
proves one gateway converged, and the next request rolled the other. The stack now
exports LITELLM_PROXY_REPLICA_URLS, the barrier polls every replica with the full
budget before settling, and up.sh refuses to boot without DD_API_KEY, since the
gateway config enables the datadog callback on every run
Latency-based routing averaged a deployment's cached samples with total / len(samples) and raised ZeroDivisionError
once an entry held none, which the proxy answered as a 500 for every later request on that model group. Cost-based
routing writes the same {model_group}_map entry with minute counters only, so a group used by both strategies hit this
on every latency-routed request. A deployment with no samples now counts as 0 latency, the same as one the router has
never seen
Resolves LIT-7053
The Public Model Hub dialog in ModelHubTable was never opened (its open setter had no callers), but its See Page button navigated to /model_hub_table?key=<session key>. Delete the dialog, its state, the handler and the unused router import so the path cannot be revived.
A failed pass-through call logged the httpx traceback, whose message
quotes the upstream URL with the provider API key in its query string,
into the spend log's error information and into every failure callback.
The error information built for logging now redacts its traceback and
error message, and the traceback is redacted once before the failure
callbacks receive it.
OCI Cohere restates the whole assistant text on the chunk that carries
the tool calls and again on the terminal chunk that carries chatHistory.
Only the terminal restatement was dropped, so a tool-calling turn streamed
the text twice. Treat a toolCalls-bearing chunk as a restatement too and
drop its text once deltas were already emitted.
Resolves LIT-6819
The test-quality gate counts every patch of a litellm internal against a ceiling, and the three patches this test needs pushed it over. The callback imports increment_spend_counters, update_cache and proxy_logging_obj from proxy_server inside its own body, so there is no seam to inject fakes through; every other test in this file uses the same three patches for the same reason
Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
The previous cleanup only ran inside a caller's own except handler, so a
build that failed after its only caller had already been cancelled left
the failed task cached with nothing left to clear it. Move the cleanup
onto the task itself as a done-callback, which fires whether or not
anyone is still awaiting it, so the next request always gets a fresh
attempt instead of replaying the stale failure.
Adds a regression test for exactly that ordering (cancel the only
caller, let the build fail unobserved, then confirm the next request
builds successfully); it fails against the previous except-based
cleanup, which left the task cached.
test_load_state_from_db_handles_unknown_request_type compared the
WRITING cell after load against a cold-start value captured for
GENERAL. They happened to be equal for this fixture (the fast model's
empty strengths list makes every request type's prior identical), which
hid that the assertion was comparing the wrong baseline. Capture each
request type's own cold-start value instead.
The proxy cost callback zeroed response_cost whenever cache_hit was true. That rule dates from Jan 2024 when it was the only place cache hits were priced. The logging layer has priced the LLM share at 0 on a cache hit since Aug 2024, and since guardrail cost joined the standard logging payload the proxy-side zeroing has thrown away a real provider charge: a pre_call guardrail runs before the cache is consulted, so a cached response still cost whatever the guardrail billed. Drop the redundant zeroing so the payload's response_cost, which is already LLM 0 + guardrail cost, reaches spend logs, daily tables and budgets untouched
Claude-Session: https://claude.ai/code/session_01EX13mWex6RaBo9PYnkAtFW
_ensure_routelayer previously awaited asyncio.to_thread(...) directly
inside the lock. Under cancel_on_disconnect, cancelling that await
released the lock while the worker thread kept running, so a second
concurrent request would see no lock held and start a duplicate billed
build. Store the build as a task on self and have every caller await it
through asyncio.shield: cancelling one caller's wait no longer cancels
the build or lets another caller start a second one. A real build
failure (not merely a cancelled caller) clears the slot so the next
call retries fresh instead of replaying the same failure forever.
Also replaces the two tests that monkeypatched _build_routelayer (an
anti-pattern per this repo's conventions) with ones that instrument the
already-injected embedding router dependency instead, and adds a third
proving the cancellation race is actually closed.
- Shrink the fallback comments to one line each; the fuller rationale
was redundant per repo comment policy.
- Add test_pick_model_favors_the_cheaper_model_info_priced_deployment
and its hybrid-router counterpart, which exercise pick_model's actual
Thompson-sampling/scoring output instead of only asserting the
model_to_cost dict. Both are ordered so the expensive model wins
pick_best's insertion-order tie-break on the pre-fix code (proven by
reverting the production diff and rerunning), so they fail before the
fix and pass after it.
Both places that build an adaptive router's model_to_cost (the plain
auto_router/adaptive_router path in router.py, and the hybrid
adaptive-inside-complexity_router path in complexity_router.py) read
input_cost_per_token from litellm_params only. Custom pricing is
conventionally declared under model_info everywhere else in LiteLLM
(cost_calculator.py, add_deployment's litellm.model_cost registration),
so a deployment priced that way silently costs 0.0 in adaptive-router
scoring: every candidate ties on cost, the cost term contributes
nothing, and routing runs on quality alone with no warning.
Fall back to model_info at both call sites when litellm_params does not
declare a cost, matching how quality_router.py already sources cost.
litellm_params still wins when both are set.
Fixes#31481.
load_state_from_db assigned a DB row's (alpha, beta) straight into the
bandit cell, discarding the cold-start prior _init_cold_start_cells had
already put there. AdaptiveRouterUpdateQueue.flush_state_to_db only ever
persists accumulated deltas (its upsert creates a row with the raw delta
as the initial value, then increments it), never a full posterior, so a
cell whose first flush sees only one kind of signal persists a one-sided
row: e.g. alpha=1.0, beta=0.0. Loading that row as the whole cell hands
thompson_sample() a Beta(alpha, 0), and random.betavariate raises
'gammavariate: alpha and beta must be > 0.0' on every draw from that
cell from then on, surviving restarts since the bad row stays in place.
Fix: add the row on top of a freshly computed prior instead of replacing
the cell with it. Deltas are never negative, so both parameters stay
positive.
Fixes#35590. Fixes#29397.
AutoRouter's cold-start route layer construction (SemanticRouter with
auto_sync="local") ran directly on the event loop, doing at least one
synchronous embedding HTTP call inline behind a bare "if routelayer is
None" check with no lock, so it blocked the whole worker and let
concurrent cold-start requests each build a duplicate layer.
ComplexityRouter already solved this identically for its own semantic
keyword matching (_ensure_semantic_routelayer: a lock plus
asyncio.to_thread). Give AutoRouter the same treatment: extract the
build into _build_routelayer and gate it behind _ensure_routelayer's
double-checked async lock.
Fixes#33204.