Making the in-memory issuer reflect a trust-on-first-use discovered value fixed
the registry/row token-identity drift, but it overloaded a single field: the
carry-forward gate keyed on issuer truthiness as a proxy for "endpoints are
anchored to a pinned issuer, fail-closed". A discovered issuer is truthy yet not
anchored, so a resource-rooted server that had learned its issuer would drop its
last-known-good endpoints on a transient discovery blip instead of carrying them
forward.
Anchoring is now a first-class property rather than a proxy. MCPServer carries
issuer_is_anchored, set at both build paths from the single _uses_issuer_anchor
definition (a pinned issuer on a discovery auth type). issuer stays the identity
value used by the token-identity tuple and the serializers; issuer_is_anchored is
the provenance value the carry-forward gate reads to decide fail-closed. The two
properties can no longer be conflated, so a discovered issuer keeps its
resource-rooted endpoints carrying forward while a pinned issuer still fails
closed.
Regression tests pin both directions: a discovered-but-not-anchored server
restores its endpoints on a discovery blip, an anchored server does not, and the
build sets issuer_is_anchored true only when the issuer is pinned
Pin a session's first-turn model for the rest of the session by default
instead of reclassifying every turn. Keeps multi-turn sessions on a single
model, preserving provider prompt caches and avoiding cross-model
conversation-history errors (e.g. Anthropic rejecting a thinking block
produced by a different model). Requests without a resolvable session_id are
unaffected. Set session_affinity: false to opt out.
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cli): add `lite up`/`lite down` to ambiently route Claude Code through the proxy
Patches ~/.claude/settings.json in place (env.ANTHROPIC_BASE_URL + apiKeyHelper
via `lite auth print-token`) so any `claude` session started afterward, from
any terminal, routes through the local LiteLLM proxy with no wrapper command
needed, unlike the existing `lite claude` subprocess-exec approach. Backs up
the original file first and restores it on Ctrl-C/SIGTERM, or via `lite down`
after an unclean exit. Cursor is not supported: no equivalent file-based config
to patch.
* feat(cli): add lite autoroute to QA complexity-based auto-routing against a real proxy (#33249)
* feat(cli): add lite autoroute to QA complexity-based auto-routing against a real proxy
Lets a customer try litellm's complexity_router against models they already
have on their existing, unmodified production proxy, with no config.yaml
edits and no new infra. lite autoroute configure discovers accessible
models via /model_group/info and walks through tier assignment (plus
optional LLM classifier / semantic matching / adaptive selection); every
referenced model becomes its own litellm_proxy/<name> deployment forwarding
back to the real proxy with the real key, so every actual call, routed
completions, classifier calls, embedding calls, still lands on their real
proxy. lite autoroute up launches that generated config as an ephemeral
local proxy, patches ~/.claude/settings.json to point Claude Code at it, and
streams routing decisions live; Ctrl-C/SIGTERM (or lite autoroute down
after an unclean exit) restores everything.
Also adds lite model-groups list (a thin CLI wrapper over the existing
ModelGroupsManagementClient), and generalizes up.py's settings-backup/restore
helpers to take explicit paths so this feature can reuse them instead of
duplicating the logic.
Depends on litellm_lite_up_down (#33231) for that generalization.
* feat(cli): allow multiple models per autoroute tier
complexity_router already supports a pool of models per tier (randomly
picked per request; adaptive mode specifically needs a pool to choose
within), but the configure wizard only ever let you assign one. Tiers are
now a tuple of model names; the wizard prompt accepts comma-separated
indices to pick more than one per tier.
* feat(cli): fuzzy model picker and auto-route Claude Code to autorouter
Numbered-index selection didn't scale past a handful of models, so switch
the tier picker to InquirerPy's fzf-style fuzzy search. Also set
ANTHROPIC_DEFAULT_{SONNET,HAIKU,OPUS}_MODEL to "autorouter" in Claude
Code's settings, since Router resolves auto-router deployments by literal
model name with no wildcard support, so a "*" catch-all model_name would
never match real traffic.
* feat(cli): allow installing lite CLI from source via LITELLM_CLI_REF
Lets testers try an unreleased branch's CLI changes with the same
curl-piped installer, instead of waiting for a PyPI release.
* fix(ci): modernize type hints to clear ruff strict-rule budget
* fix(ci): bump httplib2 and setuptools to patched versions
Clears osv-scan findings for PYSEC-2026-3444 and PYSEC-2026-3447.
* fix(cli): write autoroute's secret-bearing files with mode 0600
commands.py wrote config.yaml (embeds the real proxy key) and Claude
Code's settings.json (embeds the ephemeral proxy's master key) with
plain open(), landing at the umask-derived default (commonly 0644)
until a later chmod call caught up. That window, and the missed case
where settings.json already exists (chmod never ran at all there),
left a credential-bearing file readable by another local account.
secure_create() fixes the mode via fchmod on the fd before any
content is written, covering both the brand-new-file and
already-exists cases, and commands.py/wizard.py now route their
sensitive writes through it.
* docs(cli): warn that a stale Claude Code session can leak to a squatted port
lite autoroute up's master key is embedded statically (unlike lite up's
apiKeyHelper, resolved per request), so a Claude Code session still
running after teardown keeps sending it, along with prompt content, to
a now-unbound loopback port that another local account can bind. This
is the same one-time-patch tradeoff lite up already accepts, just with
a static secret instead of a re-resolved one -- document it in the
README's Caveats section and surface it in the teardown message itself.
* fix(cli): address greptile review feedback on autoroute PR
- terminate the ephemeral proxy child process when its health check
fails, instead of leaking an orphaned, unrecoverable process bound
to the port
- replace bare assert isinstance checks (no-ops under python -O) with
click.ClickException in the model-groups list and configure wizard
code paths
- close launch_proxy's log file handle once the child process has
inherited its fd, instead of leaking it
- add build_generated_proxy_config to config.py's __all__
* fix(cli): close TOCTOU window in lite up's settings backup write
write_backup wrote the backup (which can embed the original
apiKeyHelper/settings content) with plain open() + a chmod call after
the fact -- the same permissive-until-corrected window already fixed
for autoroute's config.yaml and Claude settings writes, and missed
entirely when the backup file already exists with broader permissions.
Moves secure_create (atomic-enough 0600 via fchmod before any content
is written) to up.py, the module both lite up and lite autoroute
share, and has autoroute/process.py import it from there instead of
keeping its own copy.
* fix(cli): refuse autoroute up when a stale backup exists from a crash
The pid-record check only catches a still-live duplicate process; a
SIGKILL'd `up` leaves no live pid but does leave AUTOROUTE_BACKUP_PATH
behind. Without this guard, a fresh `up` overwrote that backup with
the currently-patched Claude settings instead of the true originals,
so `down`/Ctrl-C would restore the wrong content permanently. up.py's
`lite up` already guards the analogous case; mirror it here.
* fix(cli): bind the ephemeral autoroute proxy to loopback only
proxy_cli.py defaults --host to 0.0.0.0 when not passed explicitly.
launch_proxy never passed it, so the ephemeral proxy -- despite every
base_url in this module being built from 127.0.0.1 -- was actually
reachable from other hosts on the network, including its
unauthenticated-until-config-lands routes before the master key is
wired in.
* docs(cli): show curl install for the autoroute QA flow
Points readers at scripts/install-cli.sh's curl one-liner instead of
assuming uv/pip is already set up, and documents the LITELLM_CLI_REF
override for trying an unreleased branch or commit.
* fix(cli): surface a clean error on an empty or corrupt autoroute config
A configure run killed between secure_create's O_TRUNC and the write
completing leaves an empty config.yaml on disk. The next up read that
via yaml.safe_load (None) into the generated-config TypeAdapter
uncaught, surfacing a raw pydantic.ValidationError instead of pointing
the user back at `lite autoroute configure`.
* fix(cli): bind lite up's apiKeyHelper to the proxy it was started against
_ensure_fresh_login only checked token freshness, not which proxy the
cached token belonged to, and resolve_api_key_helper built a bare
`lite auth print-token` command with no --base-url. A user logged into
proxy A who ran `up --base-url proxy-b` (or LITELLM_PROXY_URL=proxy-b)
would silently get proxy A's real token wired into Claude Code's
apiKeyHelper; since apiKeyHelper is invoked bare, print-token's
existing origin check never engaged, so proxy B -- attacker-controlled
or not -- received every subsequent request's Authorization header
carrying proxy A's credential.
_ensure_fresh_login now requires the cached token's base_url to match
before treating it as usable, forcing a fresh login for the selected
proxy otherwise. resolve_api_key_helper now takes that base_url and
threads it through as an explicit --base-url, so print-token's
existing (but previously unreachable in the apiKeyHelper flow)
base_url_explicit check actually enforces the match at request time
too.
* fix(cli): surface clean errors instead of raw tracebacks in lite up/down
load_json_or_empty and read_backup both delegate to pydantic's
validate_json, which raises ValidationError on invalid JSON or a
non-object root -- neither up() nor down() caught it, so a corrupt
settings or backup file surfaced an unformatted Python traceback
instead of a clean CLI error. Both now convert to UpError, and down()
(previously uncaught entirely) and up()'s teardown path now handle it.
restore_claude_settings also gained a parent.mkdir guard before
rewriting CLAUDE_SETTINGS_PATH: if ~/.claude/ was removed while `lite
up` was running, the restore would crash before deleting the backup
file, permanently stranding it and breaking every future `lite down`.
* docs(cli): call out env-var auth for autoroute commands
* fix(cli): clean up leaked proxy and surface clean errors in autoroute
Three related gaps, all following an UpError getting raised somewhere
that wasn't catching it yet:
- up() left the just-launched ephemeral proxy running with no pid
record if load_json_or_empty/write_backup/secure_create raised after
the health check passed, mirroring the existing ProcessLaunchError
cleanup for the health-check-failure branch.
- _teardown() didn't catch restore_claude_settings raising UpError
(e.g. a corrupt backup at stop time), which would otherwise escape
to Click as an unhandled error in the normal-exit path, or print
"Error in atexit" in the atexit path. up.py's own _restore_once
handles the identical case the same way.
- read_pid_record let a corrupt PID file surface a raw
pydantic.ValidationError instead of a clean message, and did so in
down(), the command specifically meant for crash recovery. down()
now clears an unreadable pid record and continues cleanup instead of
aborting, since a corrupt pid file must never block the one command
meant to recover from exactly this kind of crash.
* docs(cli): warn against running lite up and lite autoroute up together
* fix(ui): stop sending the complexity-router pseudo-model to /health/test_connection
Test Connection on a saved auto-router model sent the raw
"auto_router/complexity_router" model string to the generic
health-check endpoint, which always failed with "Unmapped LLM
provider" since it's a routing-strategy config, not a real
completion endpoint.
Reuse the per-tier connection test already built for the Add Auto
Router wizard: for complexity-router models, test each configured
tier's underlying model group instead of the router pseudo-model.
Semantic-type auto routers (auto_router_config) have no equivalent
tier-based test yet, so the button is hidden for them instead of
guaranteed to fail.
* fix(ui): address review feedback on auto-router test connection fix
Type the complexity-router config parsing instead of using `any`, use
NotificationsManager.warning instead of fromBackend for the
client-generated "no tiers configured" message, remove comments added
in the previous commit, and also test the deployment's configured
complexity_router_default_model as a fallback target when it isn't
already covered by a configured tier (matches the fallback Router
itself uses for unconfigured tiers).
Two lifecycle gaps let the issuer trust anchor drift out of sync with the
endpoints it governs. Changing or clearing a previously pinned issuer left the
authorization_url and token_url that were resolved under the old issuer in the
row, so clearing the anchor could revive stale, possibly untrusted endpoints
instead of re-discovering. And a build that discovered an issuer
trust-on-first-use persisted it to the row while the returned in-memory server
kept the issuer unset, so the registry and the row disagreed and the per-user
OAuth token identity, which includes the issuer, differed between that build and
the next rebuild and forced a spurious re-auth.
update_mcp_server now treats a change to a previously pinned issuer the same as a
url or auth_type change and clears the auth-flow-scoped endpoint fields that were
resolved under it. The trigger fires only when an issuer was already pinned and
is now changed or cleared, so establishing one for the first time, including the
trust-on-first-use discovery write-back, does not wipe the fields it just
resolved.
Both build paths, build_mcp_server_from_table and load_servers_from_config, now
construct the server with effective_issuer = manual_issuer or the discovered
issuer, skipping an origin-fallback guess exactly as the persistence does, so the
in-memory object always reflects what the row will hold.
Regression tests pin each case: clearing and re-pointing a pinned issuer clear
the stale endpoints, a first-time establish preserves the discovered fields, and
a build reflects the discovered issuer while an origin-fallback guess is not
reflected
* fix(complexity_router): return empty dict from _classifier_call_metadata when metadata is absent
The LLM classifier reads request_kwargs.get("litellm_metadata"), but the proxy stores request metadata under "metadata", so this returned None. _classifier_call_metadata then passed None straight through to the classifier acompletion call, which assumes a dict and blows up with 'NoneType' object has no attribute 'update'; the router swallowed it and silently fell back to heuristic scoring, so the configured LLM classifier never ran. Returning an empty dict keeps the classifier call well-formed.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): cover complexity-router LLM classifier routes over the proxy
Add a live e2e regression for the complexity auto-router: a lexically simple but hard prompt ("Is P equal to NP?") is routed by the LLM classifier to the higher-tier anthropic backend, read back from the spend log's model. Before the metadata fix the classifier silently crashed and the router fell back to heuristic SIMPLE scoring on the openai backend, so this test fails pre-fix and passes post-fix.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
When an admin pins an issuer, RFC 8414 section 3.3 makes that issuer the sole
authoritative source of the authorization and token endpoints, so a compromised
or misconfigured upstream cannot smuggle a token endpoint by echoing the pinned
authorize URL. The first cut enforced that only on the database build path; the
carry-forward, persistence, config-load, serialization and sanitization paths
could still restore or emit upstream-derived endpoints for an issuer-anchored
server, which is the class of gap the review flagged.
Every site now routes through one predicate. _endpoints_yield_to_issuer returns
all-None whenever the issuer is the anchor, so both build paths,
has_all_upstream_oauth_fields, needs_discovery and the endpoint merge defer to
the issuer. _carry_forward_resolved_oauth_endpoints carries only scopes for an
issuer-anchored server and fails closed on endpoints.
_persist_discovered_oauth_endpoints skips endpoint writes under the anchor. The
two table serializers round-trip the issuer and both non-admin sanitizers redact
it. Scope selection stays resource-driven per the MCP authorization spec:
_fetch_issuer_anchored_oauth_metadata takes endpoints from the issuer document
and scopes from the resource document.
The OAuth metadata resolution and corroboration gating for the database build
path move into _resolve_table_oauth_metadata so build_mcp_server_from_table
stays within the cyclomatic-complexity budget without changing behavior.
Regression tests pin the invariant at each site: the issuer overrides stored
endpoints even when they are populated, carry-forward does not restore endpoints
under the anchor, persistence does not write endpoints under the anchor, a url or
auth_type change clears stale issuer-scoped fields even when resubmitted
unchanged, the Azure heuristic stays reachable under a required issuer, and
anchored metadata takes endpoints from the issuer while scopes come from the
resource
A single ddtrace constraint now covers every supported Python version, so this collapses the version split introduced in #33438. Also aligns the build_from_pip image pin and updates the type-only Tracer import to its current module path
The issuer anchor is for the token/registration endpoints only (the RFC 9700
mix-up). Scope selection stays resource-driven per the MCP authorization spec
Scope Selection Strategy: _fetch_issuer_anchored_oauth_metadata now validates
the issuer document (RFC 8414 §3.3) for the endpoints and separately fetches the
resource's advertised scopes (WWW-Authenticate challenge, else RFC 9728
scopes_supported) for the scope value, instead of using the issuer document's
own scopes_supported. The resource can influence only the requested scope, which
the authorization server and user consent bound (RFC 6749 §3.3), never the token
endpoint.
Discover the issuer from the upstream and persist it trust-on-first-use (fill-empty-only,
frozen thereafter), so admins do not have to type it; an admin-configured issuer always wins
and is never overwritten. Re-pointing the server url now clears the discovered issuer and
endpoints (matching the existing auth_type-change clearing) so a new upstream re-discovers
instead of anchoring on the previous upstream's issuer. Adds the issuer to the per-user OAuth
token identity so re-pointing it purges stale tokens. Surfaces the issuer as an optional,
auto-discovered, overridable field in the create and edit MCP server forms.
C901 gate shows +1 vs staging; that is inherited from the #33317 stack base (delta 0 against
Adds an admin-configured issuer to MCP servers. When set, OAuth metadata is
fetched from the issuer's own origin and adopted only when the document
self-attests that same issuer (RFC 8414 §3.3), making token_endpoint,
registration_endpoint, and scopes authoritative for the pinned issuer instead
of a document the MCP resource server chose. This closes the mix-up where a
compromised resource echoes a pinned authorization_url to smuggle its own
token endpoint and inflated scopes past the corroboration gate. Discovery is
same-authority against the issuer origin, fails closed on a §3.3 mismatch, and
does not fall back to resource-rooted discovery. Rows without an issuer keep
the existing corroboration-gate behavior unchanged.
Backend + schema only; UI field and live-proxy proof follow.
Reverts the over-correction that restricted a pinned-authorization_url server's
discovered scopes to the authorization server's own scopes_supported. Per the MCP
authorization spec Scope Selection Strategy and RFC 9700 §2.3, the scopes a client
requests are resource-driven: the WWW-Authenticate 401 challenge scope, else the
RFC 9728 protected-resource scopes_supported. The authorization server's RFC 8414
scopes_supported is a non-exhaustive capability list (the server MAY omit supported
scopes) and is never the selection source; scope inflation by a compromised resource
is bounded by the authorization server and user consent (RFC 6749 §3.3), not by the
client restricting the request. The corroboration gate now rejects only the
uncorroborated token_url/registration_url (the RFC 9700 endpoint mix-up) and leaves
scopes untouched. Removes the now-unused authorization_server_scopes field.
Adds a 'passthrough' feature row to the Claude Code compat matrix that
drives the real claude CLI in each cloud's native mode against
LiteLLM's passthrough routes (the LLM-gateway setup from
code.claude.com/docs/en/gateway) instead of the /v1/messages
translation layer:
- anthropic: ANTHROPIC_BASE_URL={proxy}/anthropic, forwarded verbatim
to api.anthropic.com
- bedrock_invoke: CLAUDE_CODE_USE_BEDROCK=1 against {proxy}/bedrock;
the router resolves the alias in /model/{alias}/invoke-with-response-stream
- vertex_ai: CLAUDE_CODE_USE_VERTEX=1 against {proxy}/vertex_ai/v1;
alias, project, location and credentials resolve from the deployment,
which now sets use_in_pass_through: true (and the canonical
vertex_project/vertex_location param names) in test_config.yaml
- azure: CLAUDE_CODE_USE_FOUNDRY=1 against {proxy}/azure via the
AZURE_API_BASE/AZURE_API_KEY fallback (documented in the cron env
example)
- bedrock_converse: not_applicable; Claude Code has no Converse-wire
client
The shared cell body lives in _passthrough.py with injectable runner
and env (no monkeypatching), unit-covered in
_driver_unit_tests/test_passthrough.py including pins on the per-mode
CLI env contracts captured from a real claude CLI (2.1.210) run
against a request-logging sink.
* fix(ui): remove New badge from Projects sidebar item
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): add reusable BetaBadge and use it for Projects sidebar item
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): rename disableShowNewBadge flag to disableShowBadges and make BetaBadge respect it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): use blue for BetaBadge to match existing New badge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(ui): keep disableShowNewBadge localStorage key to preserve existing user opt-outs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore: remove non-functional md artifacts from PR
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* build: drop requires-python upper cap so Python 3.14 resolves to current releases
The <3.14 cap made pip on Python 3.14 fall back to litellm 1.83.7, a
pre-April release whose old auth flow fails with 400s. The cap was added
in d9a460277a because deps lacked 3.14 wheels and uv could not resolve
the 3.14 split; both are fixed now via the existing python_version
markers plus a ddtrace version split (2.x has no cp314 wheels, 3.16+
does). Verified on 3.14.5: uv sync --all-extras installs, litellm and
proxy_server import (rust bridge falls back to pure python), real
provider calls succeed sync/async/streaming, and the core-utils test
suite passes.
* build: cap requires-python at <3.15 and keep ddtrace on one major per python band
Reviewer preference to bound the supported window at the newest tested
minor rather than leaving it open-ended, and Greptile flagged the
ddtrace 3.14+ range spanning two majors; every ddtrace 4.x ships cp314
wheels so the band is now >=4.0,<5.0, matching the single-major
convention of the 2.x band.
Clients calling the standalone apply_guardrail endpoint had no way to pass
per-request configuration to custom guardrail implementations. This adds an
optional metadata field to ApplyGuardrailRequest and forwards it to
CustomGuardrail.apply_guardrail via request_data, only when the client sends
it. The messages guard is aligned to the same is-not-None semantics so an
explicitly-sent empty list is forwarded instead of silently dropped.
The Admin UI's Guardrail Test Playground gains an optional Metadata JSON
input (validated client-side) wired through applyGuardrail in networking.tsx,
so parameterized guardrails can be exercised from the dashboard.
Tests cover metadata alone, metadata with messages, explicit empty values,
the omitted-field passthrough, and the UI panel's parse/error behavior
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Three process-lifetime retention points kept full request payloads
(messages included) alive after the request finished. Under bursts of
large-token traffic (~73K tokens/request mean) this presented as
stepwise RSS growth that never returned to baseline, ending in OOM:
1. Logging.pre_call/post_call stored their entire locals() (messages,
the Logging object, complete_input_dict) in the module-level
litellm.error_logs dict, pinning the most recent request's payload
per worker forever. Nothing reads that dict; the writes are removed.
2. LLMCachingHandler.request_kwargs kept litellm_logging_obj inside the
stored kwargs while the handler itself hangs off
logging_obj._llm_caching_handler, closing a reference cycle
(Logging -> LLMCachingHandler -> kwargs -> Logging). Cyclic payloads
are only reclaimed by generational GC, so megabytes of dead request
data lingered until a rare gen-2 pass, and the transient copies
fragment the allocator into a permanent RSS high-water mark. The
handler now drops litellm_logging_obj from its stored kwargs; the
caching layer never reads it.
3. The router stored every request's kwargs in the ITPM/OTPM contextvar
even when no deployment configures itpm/otpm. Pooled resources
created mid-request (e.g. redis connections) capture the asyncio
context, extending that pin far past the request. The slot is now
populated only for deployments with io token limits and overwritten
with None otherwise.
Live-proxy verification (bursts of 30 x ~300KB requests, PII guardrail
+ prometheus + redis cache): unfixed grows 16-29MB per burst without
release; fixed grows under 1MB per burst after warmup and flattens.
Resolves LIT-4434
Replace list(set(...)) dedupe with dict.fromkeys so callback insertion
order is preserved deterministically instead of being randomized by set
iteration order (influenced by PYTHONHASHSEED). Applies to both the
Logging and ProxyLogging implementations.
Fixes#33003
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* feat(guardrails): add Compresr guardrail for query-aware context compression
Adds a first-class guardrail that compresses bulky message content (tool
outputs, RAG chunks, search results) through the Compresr API before the
request reaches the LLM, via the apply_guardrail / structured_messages hook
so it covers /chat/completions, /v1/messages, and /v1/responses (the latter
through the texts channel, mirrored only when the replacement is
unambiguous; anything ambiguous is left uncompressed).
Distinct from whole-conversation compressors:
- Query-aware: each message is compressed against the intent that produced
it (a tool output against its originating tool call's name + arguments,
resolved via tool_call_id; otherwise the last user message).
- Recoverable: each compressed message carries a hash marker and the request
gains a compresr_retrieve tool, so the model can pull the original content
back through the agentic loop when the compressed version is not enough.
Originals are cached in-process, scoped to the caller's virtual-key hash
plus the request's litellm_call_id, with a TTL and a per-call byte cap;
recovery is skipped when no caller scope is available so one caller can
never read another's originals. The store is per-process, so multi-worker
deployments need sticky routing (or enable_retrieval=false).
Fail-closed by default (fail_open configurable), SSRF-validated api_base
(alternate IP-literal encodings included), cross-tenant-isolated recovery
store, and upstream errors redacted from client-facing responses. The
outbound client follows redirects and re-resolves DNS per request, so the
api_base host/IP checks are defense-in-depth, not a full SSRF guarantee;
this is documented as a known limitation. Requests where nothing was
actually compressed are returned untouched (same object identity) so
handlers skip the write-back. Auto-discovered via the guardrail_hooks
registry.
* fix(guardrails): cap Compresr recovery store total memory
The recovery store bounded bytes per call and entry count, but had no
aggregate cap: 256 tracked call ids at the 10 MiB per-call default could
retain ~2.5 GiB per worker. A flood of requests with distinct
x-litellm-call-id values and large compressible tool outputs could
exhaust a shared proxy worker.
Add a global byte budget (_MAX_TOTAL_STORE_BYTES, 256 MiB) across all
entries. A running total is maintained on every insert/eviction so the
cap is enforced without re-encoding the whole store on the request path;
oldest entries are evicted once the budget is exceeded, always keeping
the most-recent entry so recovery still works for the request populating
the store. +2 regression tests.
* fix(guardrails): gate and bound Compresr recovery loop
Two hardening fixes to the compresr_retrieve agentic loop:
1. Only run the loop when a retrieve call resolves to recovery state this
guardrail actually created for the request. Previously the gate checked
only that the caller-supplied tool list contained a compresr_retrieve
function and that the model emitted a call, so a caller could define
their own same-named tool and force an extra provider round-trip with
nothing to recover. The plan now returns run_agentic_loop=False when no
requested hash resolves.
2. Bound the follow-up against retrieval amplification: each distinct hash
is expanded at most once (repeats get a short marker) and at most
_MAX_RETRIEVALS_PER_LOOP calls are honored, so prompting the model to
call compresr_retrieve many times with the same marker cannot balloon
the follow-up. _retrieve_original now returns None on miss.
+3 regression tests; two existing security tests updated to assert the
stronger veto behavior (forged/cross-tenant hashes now stop the loop
entirely instead of returning a not-found follow-up).
* fix(guardrails): warn when Compresr recovery is skipped without auth scope
When enable_retrieval is on (the default) but the proxy has no per-key
auth, the request has no caller scope, so recovery is silently disabled:
content is compressed but the compresr_retrieve tool is never injected and
the originals are dropped, with no runtime indication. Emit a one-shot
call-time warning so operators can see recovery is being suppressed and
configure virtual-key auth. +1 regression test.
* style(guardrails): tighten Compresr guardrail comments
Condense the verbose multi-line inline comments and the api_base docstring
to concise form. No behavior change.
* fix(guardrails): keep injected tool on Responses API + bound recovery markers by byte cap
Two fixes for reviewer-flagged defects in the Compresr guardrail:
- Responses API: _merge_tools_after_guardrail iterated only over the
request's original tools, dropping any tool a guardrail appended (the
compresr_retrieve recovery tool) whenever the request already had tools.
Keep the appended tools so recovery works on /v1/responses.
- Recovery markers: markers + originals were built for every compressed
target before the per-call byte cap trimmed the store, so an evicted
original left a marker the model could never retrieve. Attach recovery
only while the store (existing entries under the same key + this call's
originals) stays within the cap, so a shipped marker is always retrievable
-- including on a later turn that reuses the store key.
Adds regression tests for both paths.
* refactor(guardrails): extract _existing_originals to keep apply_guardrail under the complexity gate
The byte-cap fix added a branch to apply_guardrail, tipping it past the
C901 complexity ceiling. Move the store lookup into a small helper; no
behavior change.
* fix(guardrails): harden Compresr SSRF blocklist, re-arm no-scope warning, tolerate odd tool shapes
* fix(guardrails): rerun input guardrails on Compresr retrieval follow-up
* chore: remove unrelated deepkeep files committed by mistake
---------
Co-authored-by: charafkamel <charafkamel@live.com>
CHAT_ROUTES was built once at module load via migratedHref, capturing an
empty server root path before the UI-config bootstrap sets it. Under
SERVER_ROOT_PATH every chat route came out unprefixed, so router.push
hard-navigated to a 404; during a send that page unload also aborted the
streaming request, so the first message of a new conversation never
rendered and only a second send appeared to work. Compute the routes at
render time via getChatRoutes(), update the send URL with a shallow
history.pushState instead of a router navigation, and source the active
conversation id from the hook's local state so it propagates without a
router round-trip