Router._add_deployment called get_llm_provider without the deployment's api_base, so a config entry with a bare model plus a known OpenAI-compatible endpoint failed startup validation with LLM Provider NOT provided and the proxy returned 400 no healthy deployments for that model group. acompletion had the same gap at request time: it forwarded only base_url into its get_llm_provider call, dropping the api_base kwarg the router passes. Both now forward api_base so endpoint matching resolves the provider the same way sync completion already does
Together AI moved its canonical API host from api.together.xyz to
api.together.ai. Default the provider api_base and the rerank handler to
the new host, make rerank honor api_base and TOGETHER_AI_API_BASE like
chat already does, map both hosts to together_ai when passed as
api_base, and delete the dead models/info fetch in factory.py.
Adds live e2e coverage for the Bedrock combinations behind recent customer
incidents: llm_provider-* response-header forwarding on /chat/completions
(nonstream and stream), regional us.anthropic.* inference-profile ids over
the invoke route, and the Admin UI Test Connection probe for a
responses-mode Bedrock Mantle deployment. Registers the matching cells in
the coverage registry and publishes the provider x feature matrix table in
its README.
`requests` has no default timeout, so a host that accepts the connection and
never answers blocks the calling thread forever.
The one on the request path is the HiddenLayer guardrail's `_get_jwt`. It runs
synchronously inside `_call_hiddenlayer` whenever the hour-long JWT expires and
the API answers 401, so a stalled auth host parked the worker's whole event
loop, not just the guarded request. The other eight are the teams and users CLI
clients, which pin the operator's terminal instead.
`TeamsManagementClient` and `UsersManagementClient` now take the same
`timeout: int = 30` their `HTTPClient` sibling already had, and `Client` threads
its own timeout down to teams. `_poll_for_ready_data` already passed a timeout
through a TypedDict that ruff could not see into; passing the argument directly
retires both the TypedDict and the suppression it would have needed.
Graduate S113 into ruff.toml so the next `requests` call without a timeout fails
the lint step.
Backfill 21 serverless chat models, the multilingual-e5 embedding model, and
Llama-Guard-4-12B from the live Together catalog with per-token pricing and
capability flags. Mark 25 delisted together_ai entries with their documented
deprecation_date and point superseded models at a live successor via metadata.
Reprice Llama-3.3-70B-Instruct-Turbo to Together's current rate.
Mantle 400s ("Invalid 'input': value did not match any expected variant")
on the Codex history item types agent_message, context_compaction, and
local_shell_call, killing every Codex multi-agent session on the first
sub-agent turn. Rewrite agent_message into an assistant output_text message
(preserving encrypted_content slot payloads, which carry the plaintext task
through Mantle), context_compaction into Mantle's supported compaction
spelling, and local_shell_call into the function_call its recorded
function_call_output already pairs with.
* fix(proxy): reset a stuck team member's budget
A per-team-member budget check reads a cross-pod spend counter that
nothing ever invalidates. Once a member exceeds their per-member
budget, resetting the key's spend, raising the user's or the team's
own budget, or issuing a new key all leave the member stuck, because
none of them touch this counter or its cached membership object.
Add POST /team/{team_id}/member/{user_id}/reset_spend to reset a
member's tracked spend, and invalidate the same cached state from
/team/member_update when it raises a member's own budget, so that
path also takes effect immediately instead of waiting on the
membership cache's TTL. Name the entity in the check's error message
so a stuck member is diagnosable from the 429 body alone.
* fix(proxy): close reset-vs-floor-read race and surface double Redis write failure on member spend reset
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): broadcast spend reset as a SET so the handler's self-delivered message cannot erase the reset guard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): omit null fields from the invalidation message so plain evictions keep the old wire format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Read strict from the caller's output_format/output_config.format instead
of hardcoding true, defaulting to false to match OpenAI's API default.
Explicit true/false values are preserved and output_format still takes
precedence over output_config.format.
* fix(http_handler): dispose aiohttp session when finalized without a running loop
AsyncHTTPHandler.__del__ can only schedule an async close when a running
event loop exists at finalization time; in any other context (worker
threads whose loop has closed, sync contexts, interpreter shutdown) the
RuntimeError from get_running_loop() is swallowed and the underlying
aiohttp ClientSession is abandoned to GC, emitting 'Unclosed client
session' / 'Unclosed connector' warnings.
This is the disposal gap left after the recycle-time fix: clients created
for short-lived event loops (the loop-id-keyed LLM client cache mints one
handler per loop) are never recycled - they live and die with their loop,
and their finalization is precisely the loop-less case.
Fix:
- no running loop: fall back to the connector's synchronous teardown via
LiteLLMAiohttpTransport._mark_connector_closed - the same finalizer-safe
path used for dead-loop recycles - honoring _owns_session so a shared
session is never closed.
- running loop: keep the async close, but hold a strong reference to the
scheduled task until it completes (a bare create_task() result may be
collected before running), mirroring _background_close_tasks.
Tests: loop-less finalization closes a dead-loop session; running-loop
finalization registers and drains the close task; the sync fallback
respects session ownership. All three fail without the fix.
* lint: conform new finalizer code to the type-discipline budget
Final on the five never-rebound locals (LIT010); the class-level task
registry keeps its mutable set with the sanctioned mutable-ok reason,
mirroring the aiohttp transport's registry (LIT001).
* lint: reasoned pyright ignore on the cross-class teardown call
The handler deliberately reuses the transport's finalizer-safe connector
teardown; no public seam exists and an async close can never run at
loop-less finalization. Clears the net-new reportPrivateUsage the
basedpyright budget gate flagged once the LIT stage passed.
* fix(http_handler): retrieve exceptions from finalizer close tasks
A bare discard done-callback dropped the task without consuming its
exception, so a failing aclose() emitted "Task exception was never
retrieved" at GC, the same noise class this path exists to remove.
Mirror the transport's _on_close_task_done: discard, early-return on
cancellation, retrieve and debug-log the exception.
* fix(http_handler): dispose foreign-loop sessions instead of scheduling aclose on the live loop
GC on a live loop (e.g. the app's) of a handler whose session belongs to
another, possibly dead, loop scheduled aclose() on the current loop, the
cross-loop path the transport refuses. Route both that case and the
loop-less case through the transport's lifecycle-aware
_close_recycled_session, which picks async close on the session's own
loop, threadsafe handoff, or the synchronous connector teardown.
Regression test: a dead-loop session collected while another loop runs
is disposed without scheduling anything on that loop.
* chore: retrigger CI (test_mcp_logging payload-order flake, also failed on litellm_spendlogs_fallback_metadata minutes earlier)
* test(mcp): select the MCP tool-call payload instead of the last-delivered one
TestMCPLogger kept a single last-writer slot; an async success event from
another call (a mocked acompletion whose log task lands late) races the
MCP event for it, so the cost assertions intermittently read the wrong
payload. This PR's finalizer change shifts task interleaving on the loop
and tips that latent race over (also seen on an unrelated PR minutes
earlier). Collect call_type=call_mcp_tool payloads in their own list and
assert on those.
* test(mcp): MCPLoggerHook inherits the order-independent payload capture
It duplicated TestMCPLogger's init and success handler verbatim; the
hook test reads the same MCP payload selection, so subclass instead.
A cleared field stays mounted carrying an empty value, so it was neither listed for
deletion here nor sent in the update, since the caller drops empty values from the
payload. The old value survived: clearing a destination left the previous one still
configured and still receiving requests, which for a federated credential now carry a
minted token.
Emptiness counts as a deletion now. A masked but untouched value is a non-empty
string, so it is still preserved.
A deployment federated through a named credential holds no federation field of its
own, and the gate resolved only the name the write supplied. So a patch that cleared
litellm_credential_name, alongside an api_base or api_key of the caller's choosing,
left nothing federated to find and the write was allowed. Clearing was the way out of
a rule meant to have no way out.
Both names count now, the one already on the deployment and the one being written,
since detaching an administrator's federated credential is itself an administrator's
action.
A form-encoded body writes a space as "+", not %20, and neither side of the
comparison accounted for it. The sanitiser percent-decoded but left "+" alone, so a
passphrase echoed in its wire shape did not line up with the secret, and the wire
form the Keycloak source declared used quote where urlencode actually applies
quote_plus, so the shape being compared was not the shape that went out.
The comparison now also runs over a plus-decoded copy, keeping the percent-decoded
one alongside it: "+" is a base64 character, and decoding it away would lose a run
the undecoded copy still matches on. The declared wire form uses quote_plus, which
is what urlencode itself applies.
Assigning to a Final inside the loop that walked the rendered and percent-decoded
candidates was a type error, and it is the reason lint went red. The candidates are
built in one shot now, which is the shape the rest of this module already uses.
The federation disable sentinel also needed its two registrations following through:
the key-set test pins the registered set deliberately so a new key cannot be added
without being exercised, and declaring the field on the params model changed the
proxy's OpenAPI spec, so the dashboard's generated types are regenerated to match.
When a client redirects api_base, the handler sets a sentinel so the deployment
stops federating and no server credential is minted for a host the caller chose.
get_litellm_params then rebuilds litellm_params from kwargs and did not carry that
field, so it was dropped in transit and the deployment federated anyway. A flag the
next hop discards is worse than no flag, because the code reads as protected.
The sentinel now rides the funnel like the other federation fields, which also makes
it request-banned, correctly: a caller must not be able to set it or clear it. It is
declared on the params model for the same reason the others are, so it survives the
strict dump rather than being rebuilt away.
Every repository handed its `.table` back untyped, so a dozen modules had
each grown a private `_PrismaTableActions` Protocol to paper over it. They
had drifted: some declared `update` as returning the row, others the row or
None, and none agreed on whether `find_many` was covariant
Replace all of them with a single `TableActions[RowT_co]` in
`litellm/repositories/prisma_protocols.py`, keyed to the prisma row each
repository is bound to. Query inputs stay `Mapping[str, object]` so callers
keep passing plain dicts, and `find_many` returns `Sequence` so the row type
stays covariant
Typing the nullable returns honestly surfaced paths that were already
crashing. A team admin could never edit or delete a memory entry owned by
their team: the write-auth check fed a raw prisma row to a helper that
expects the domain model, so `members_with_roles` arrived as plain dicts and
the request died as a 500 instead of applying the edit. Non-admin members hit
the same 500 in place of the 403 they were owed, so refusal and breakage were
indistinguishable. `/v2/model/info?user_models_only=true` dereferenced a
missing user row rather than returning the 400 the route already had, three
team routes dereferenced a team deleted between the read and the write, and
the agent registry dereferenced a missing agent instead of naming it
basedpyright drops 2,132 errors, 1,454 of them reportAny and 73
reportExplicitAny. The dashboard's generated types pick up `string[]` where
they had `unknown[]` for a team's members, admins and models
Reducing only the response to credential characters, while comparing it against an
untouched secret, stopped matching any secret carrying spaces or punctuation of its
own. A hand-set passphrase echoed back whole therefore reached the caller, which the
earlier contiguous match had caught. Both sides are reduced now, and a verbatim check
runs first so the result does not depend on what the secret is made of.
Percent-escaping is reversible and applies to any field, so the comparison also runs
over a decoded copy rather than asking each caller to enumerate that shape. What a
caller still declares is an encoding the redactor cannot reverse: client_secret_basic
sends base64 of id:secret, which decodes straight back to the secret.
The docstring no longer claims more than this does. Reducing to a shared alphabet
stops an accidental or naive echo; an endpoint that deliberately re-encodes or
interleaves the credential defeats any substring match, and this was never the control
keeping the credential from an endpoint that already holds it.
client_secret_basic sends base64 of "id:secret", and an endpoint that echoes that
blob back hands over material that decodes straight to the secret. Redaction only
ever compared the raw value, so the encoded form reached the caller and the
telemetry intact.
The redactor now takes every shape the credential went out in, and the Keycloak
source declares both, so an echo of either is caught. Only client_secret_basic
carries the encoded form; client_secret_post sends the secret as a plain field.
Matching a contiguous slice of the assertion left two ways for credential material
to travel on. A token endpoint that echoed a fragment shorter than the probe shared
no long run with it, and one that echoed a fragment broken up by delimiters shared
none either, so both reached the caller and the telemetry intact.
The rendered error is now reduced to the characters a credential is made of before
being compared, which lines a fragment up with the assertion wherever it starts and
however it was split. An error that merely happens to share such a run gets redacted
too, which is the right way to be wrong here.
type the strategy-router health check params instead of a bare dict, annotate
the new interactions usage locals Final, drop a reportUnnecessaryIsInstance
suppression by narrowing the grounding tool list before iterating it, and delete
the duplicated file-id decode comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The LLM classifier capped every prior turn at 200 characters independently, so a
785 character turn was cut even when the whole block it belonged to was 353
characters. A character budget now bounds the block: turns are taken newest first
and quoted whole while they fit, older turns are dropped whole once it runs out,
and only the turn straddling the boundary is cut. The per-turn cap stays as an
optional clamp for operators who set it deliberately, defaulting to unset.
Adds a scheduled GitHub Actions lane on top of the merged record/replay
transport. A Saturday cron records the `replayable` e2e tests against the
real providers and publishes the fixture bundle as a private
`e2e-fixtures-bundle` artifact with a SHA-256 sidecar. Weekday crons pull
that artifact by its pinned digest, verify the checksum before extracting,
and replay it with provider credentials set to bogus values, so a run that
ever reached a real provider fails instead of passing.
An egress sentinel pins the provider hostnames to a local sink for the whole
replay job and counts every connection that reaches them; the job asserts
that count is zero, so hermeticity is proven by measurement. A red Saturday
publishes no bundle, so the next weekday finds nothing fresh and fails loudly
rather than replaying a week-old recording, and the transport's seven-day
freshness gate hard-fails any bundle that has drifted too far. The lane also
runs on demand from the Actions tab with a record/replay `mode` input.
Tests join the lane with `@pytest.mark.replayable`. The streaming Anthropic
test now counts to twenty so its recorded response banks several content
deltas, matching the assertion that the stream arrives incrementally.
A request carrying a session/trace header fans the header value into
litellm metadata as both trace_id and session_id. LangSmith then rejected
the whole ingest batch: a root run's trace_id must equal the run id
embedded in dotted_order (400), and a run-body session_id must reference
an existing tracer session (404/422). Override caller trace_id on runs
that post as roots and drop session_id only when it mirrors trace_id,
so deliberate child-run and valid tracer-session fields still pass through.