BerriAI/litellm CI's type-discipline budget check flagged the
subclass __init__ shim. Fixes:
- Drop the __init__ override entirely — the subclass now inherits
__init__ from _BaseConductGuard (when the standalone package is
installed) or from CustomGuardrail (fallback). Removes both the
banned **kwargs (LIT008) and all four inert # type: ignore markers
(LIT009 x 4).
- Move the missing-package check into a dedicated
raise_if_missing_package() helper called by
initialize_guardrail before construction. Preserves the
cursor[bot] fix (silent-drop-on-import-failure) without needing
a custom __init__.
- Fallback branch aliases _BaseConductGuard = CustomGuardrail
directly, no type-ignore comment needed.
- Test updated to exercise the helper instead of the removed
__init__ path; new companion test asserts the helper is a no-op
when the package IS installed.
Local: ruff --select ANN,TID passes clean. ruff format applied.
Same behavioral surface — user-visible error message unchanged.
Rename fail_mode → unreachable_fallback (typed field)
─────────────────────────────────────────────────────
The shim was reading a free-form ``fail_mode`` field; a typo silently
defaulted the plugin to fail-open behavior. Switch to the typed
``LitellmParams.unreachable_fallback`` field so Pydantic validates the
value at config load. The plugin's constructor kwarg stays as
``fail_mode`` — the initializer maps the typed field onto it.
(yucheng-berri, devin-ai-integration)
Fix timeout default (was silently discarded)
────────────────────────────────────────────
``getattr(litellm_params, "timeout", 8.0)`` only applied the default
when the attribute was missing; ``LitellmParams.timeout`` always
exists and defaults to ``None``, so the intended 8-second budget was
never used. Change to ``getattr(..., None) or 8.0`` so ``None`` (and
``0``) fall through to the default.
(cursor[bot])
Move ImportError from module-load to __init__
─────────────────────────────────────────────
Raising ImportError at module load caused the guardrail-hook
auto-discovery loop to silently drop the registration when
``conduct-litellm-guard`` was missing. Users saw configs load with
no guardrail active and no error. Import lazily; raise the friendly
``pip install`` error at ``ConductGuardrail.__init__`` when
actionable.
(cursor[bot])
Advertise only supported event hooks
────────────────────────────────────
``during_call`` mode was advertised in the guardrail config but the
class never overrode ``async_moderation_hook`` — every request in that
mode silently bypassed policy. Override ``get_supported_event_hooks``
to return only ``pre_call`` so LiteLLM validates configs against
supported modes at load time. ``during_call`` / ``post_call`` support
lands with plugin 0.3.x once the underlying response-gate is wired
through ``guard_check_response``.
(veria-ai)
Text-completion + full-turn prompt scanning
───────────────────────────────────────────
Fixed in the standalone package: ``conduct-litellm-guard 0.2.2``
(BerriAI/litellm PR #38143 companion, shipping to PyPI shortly).
Pinned in the docstring here as the minimum supported version.
(veria-ai — text_completion bypass + 4KB truncation)
Tests
─────
* ``test_only_pre_call_event_hook_advertised`` — regression for
``during_call`` silent-bypass finding
* ``test_initialize_prefers_typed_unreachable_fallback`` — regression
for typo silent-fail-open finding
* ``test_initialize_applies_timeout_default_when_field_is_none`` —
regression for silently-discarded 8.0 default
* ``test_missing_standalone_package_raises_at_construction`` —
regression for silent-drop-on-import-failure finding (previous
module-load raise replaced with lazy import + init-time raise)
Codecov flagged the __init__.initialize_guardrail body as uncovered
(30% patch coverage on that file). Added a test that mocks
litellm.logging_callback_manager and calls initialize_guardrail with
a SimpleNamespace stand-in for LitellmParams — exercises the full
function body and confirms the callback is registered.
The wrapper module imports its runtime from the conduct-litellm-guard
PyPI package. When the package is not installed in the CI environment,
the smoke tests can't verify wiring (the import raises before any test
runs). Use pytest.importorskip so BerriAI's default CI env doesn't
fail on this integration, while environments that do install the
package (via 'pip install conduct-litellm-guard[dev]' or similar)
still get the smoke coverage.
Full behavioural test coverage lives in the conduct-litellm-guard
package's own CI.
The full adapter (response parser, session-ID chain, fail-mode logic,
HTTP client) lives in the conduct-litellm-guard package on PyPI. The
upstream tree hosts a thin re-export + the LiteLLM registration wiring.
Matches the Aporia / Lakera pattern — vendor SDK on PyPI, upstream
integration is a tiny adapter.
Benefits:
- Passes ruff-strict-budget and type-discipline-budget without new
violations.
- Users get the same install experience as any other guardrail vendor:
pip install conduct-litellm-guard
- Vendor keeps ownership of the parser + fail-mode semantics; upstream
keeps a stable interface.
Tests slimmed to smoke coverage (imports work, class is a
CustomGuardrail, enum + registries wired, missing-package error path).
Full behavioural coverage stays in the PyPI package.
Local runs of both scripts/ruff_strict_gate.py and
scripts/type_discipline_gate.py against upstream/litellm_internal_staging:
both pass.
Adds Conduct Guard as a first-class LiteLLM guardrail. Point any
LiteLLM proxy at Conduct and every LLM call routed through it is
policy-checked before the upstream request goes out — block, warn,
audit, or trigger a human-in-the-loop approval, with the same signed
configuration + hash-chained audit log Conduct exposes on its native
enforcement surfaces.
- litellm/types/guardrails.py: add CONDUCT to SupportedGuardrailIntegrations.
- litellm/proxy/guardrails/guardrail_hooks/conduct/__init__.py: registration
via guardrail_initializer_registry and guardrail_class_registry, picked
up by the auto-discovery in guardrail_registry.py.
- litellm/proxy/guardrails/guardrail_hooks/conduct/conduct.py: the adapter.
CustomGuardrail subclass, async_pre_call_hook, response envelope parser
for the five Conduct verdicts (ok / advisory / WARNING / BLOCKED /
PENDING approval), fail-mode logic, session-ID resolution chain
(litellm_metadata.trace_id → X-Conduct-Session-Id → hash fallback).
- tests/test_litellm/proxy/guardrails/test_conduct_guardrail.py: envelope
parsing, pre-call allow/block/approval, config precedence, missing-token
construction error.
```yaml
guardrails:
- guardrail_name: conduct-guard
litellm_params:
guardrail: conduct
mode: pre_call
api_base: https://api.conductai.ai # optional, default
api_key: os.environ/CONDUCT_AGENT_TOKEN # cond_agt_* token
fail_mode: fail_closed # or fail_open
tool_name: llm_call # scoped tool_name
```
A standalone PyPI package `conduct-litellm-guard` shipped ahead of this
PR for teams pinned to older LiteLLM versions. Once this integration
merges, the standalone README will point at the native support as the
preferred path.
- PyPI: https://pypi.org/project/conduct-litellm-guard/
- Product: https://conductai.ai/guard
Contact: sudhi@b2bsphere.com
A team or key priority that is not Latin-1 encodable (for example CJK text) was
attached as a response header by the dynamic rate limiter v3 post-call hook, and
Starlette then raised UnicodeEncodeError while writing headers, turning a
successful /v1/messages call into HTTP 500. The header is now omitted for such
values while x-litellm-rate-limiter-version and the v3 rate limit headers are
still attached.
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): offload spend tracking to a pod-local spend worker sidecar
py-spy on the gateway showed the post-response _PROXY_track_cost_callback,
spend-log and DBSpendUpdateWriter work running on the inference workers'
event loop, so a DB or Redis stall backed up the request path.
When LITELLM_SPEND_WORKER_ENABLED=true, _ProxyDBLogger serializes one compact
typed SpendEvent per success and hands it to a SpendEventProducer that ships
it over a unix socket (default) or loopback-only TCP to a sidecar started as
`python -m gateway.spend_worker`. The sidecar runs the unchanged
_ProxyDBLogger pipeline against the pod's PgBouncer (pooled_database_url).
When the sidecar is unreachable, the buffer is full, or the gateway shuts
down with events still queued or in flight, the producer applies
LITELLM_SPEND_WORKER_ON_UNAVAILABLE (fallback in-process, or drop). The
sidecar half-closes producers on SIGTERM and drains, the producer treats
EOF as unavailable, and the gateway flushes buffered spend counters on
shutdown. The sidecar honors LITELLM_LOG so its writes are visible in its
own process log.
Helm: both charts gain an opt-in spend-worker sidecar container sharing an
emptyDir socket dir, and the componentized chart's HPA uses a
ContainerResource CPU metric scoped to the gateway container so sidecar
CPU does not drive inference scaling.
* feat(terraform): opt-in spend-worker sidecar for the AWS and GCP gateway stacks
Adds spend_worker_* inputs to both modules. On ECS Fargate the sidecar is a second, non-essential container in the gateway task; on Cloud Run it is a second container in the gateway service. Both listen on loopback TCP, share the gateway's DB/Redis/secret env, and set LITELLM_JOB_ROLE=spend_worker. Disabled by default. Plan-only tests cover both, and the terraform CI workflow now runs the gcp module too
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): retrieve a completed batch in the in-process spend path test
The base now defers cost tracking for batches that are still in flight, so an in_progress batch never reaches update_database
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): rename the spend worker sidecar to collector
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): run the collector from the installed litellm package and finish in-flight fallbacks on shutdown
The sidecar command becomes python -m litellm.proxy.collector so the classic image, whose runtime
stage copies only the installed package, can run it. The module now assembles DATABASE_URL and the
pod-local pgbouncer URL itself, replacing gateway/collector.py
The componentized collector sidecar inherits gateway.volumeMounts so custom CA mounts reach it.
SpendEventProducer shields an in-progress fallback from the writer task cancellation so close()
no longer loses an event already handed to the in-process pipeline
Helpers used across modules (address_argument, should_store_prompts_and_responses_in_spend_logs,
flush_spend_counters_on_shutdown) become public so the change adds no reportPrivateUsage errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(terraform): drop the gcp job duplicated by the aws/gcp matrix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(collector): keep metrics env off the classic sidecar and reject shared loopback ports
The classic chart no longer hands PROMETHEUS_METRICS_PORT and the billing metrics env to the collector container, and gives it the same /.npm scratch mount as the proxy on a read-only root. AWS and GCP now refuse a plan where the spend collector and the metrics sidecar bind the same loopback port. A regression test drives a sidecar crash mid-stream on asyncio and uvloop and checks no event is billed by both the sidecar and the in-process fallback; the producer docstring spells out why a failed drain() cannot double count
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(proxy): format pooled_database_url after the pgbouncer rebase
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep the cache-hit preset key and survive dead producers on collector drain
Cache hits updated the logging object after the early return, so the offloaded spend event carried
preset_cache_key=None and the collector re-hashed reconstructed kwargs. Also guard write_eof() against
producer transports uvloop already closed so one dead connection cannot abort the drain
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(terraform): keep the gcp collector port off the metrics sidecar health port
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): collector connects to Postgres directly under IAM or Entra token auth
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): mark the collector's DATABASE_URL as pooled when it uses the pod's pgbouncer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The cascade zeroed end-user spend with a single update_many whose where
clause enumerated every dependent user id. Prisma compiles that IN-list
into one prepared statement carrying one bind variable per customer, and
PostgreSQL caps a statement at 32,767 of them. Once a shared budget had
more dependents than that the statement could not be parsed at all, so
the atomic cascade rolled back, budget_reset_at never advanced, and the
tier stayed due on every later tick forever. Customers sitting at their
cap were blocked indefinitely with only a recurring log line to show for
it.
End users now match on budget_id like every other gated table, plus a
NULL-budget_id branch for the implicitly created rows that carry no link
and ride the default tier. The statement's bind count now tracks the
number of expiring tiers rather than the customer population, so a reset
costs the same whether a budget has ten dependents or a million.
Fixes#40564
Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
* feat(proxy): share database connections across workers with an in-container pgbouncer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): parse pgbouncer options iteratively to satisfy the recursion gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): refuse pgbouncer with token db auth and retry failed pooler restarts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): build pgbouncer 1.25.2 from a pinned source archive and verify pooler replacements
The public Wolfi repository only carries pgbouncer 1.24.1-r3, which the image
scan rejects (CVE-2026-6664, CVE-2026-6665, CVE-2026-6666, CVE-2025-12819).
All three images now compile the checksummed 1.25.2 release in a builder stage.
The supervisor now waits for a replacement pooler to listen before treating it
as recovered, ends and retries one that never does, and takes the same lock for
stop() and spawn so no replacement can be started after shutdown began.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): refuse to start pgbouncer on a loopback port another process already owns
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): count pgbouncer ready only once its own unix socket answers, not any listener on the port
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): refuse pgbouncer older than 1.19, whose unix socket cannot vouch for the tcp port
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(helm,terraform): expose the in-container pgbouncer pool for the componentized gateway
Add database.connectionPool to helm/litellm and gateway_connection_pool_* to
terraform/litellm/aws so the componentized gateway can receive the
LITELLM_PGBOUNCER_* env the classic image already honours. Both reject the
pool under IAM or Entra token auth at render/plan time: the pooler holds one
static database password for the life of the pod or task.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(gateway): launch the componentized gateway image through a pgbouncer-aware supervisor (#40592)
The componentized gateway image started uvicorn directly, so the in-container
PgBouncer never ran for it: every worker opened its own Prisma pool to the
database. It also passed no keep-alive timeout, so behind a load balancer with
a 60s idle timeout uvicorn's 5s default closed idle connections first and the
balancer returned 502s on scale-out
gateway.launch assembles DATABASE_URL, starts PgBouncer once per pod when
LITELLM_PGBOUNCER_ENABLED is set, hands the workers the loopback URL and then
runs uvicorn on gateway.main:app with KEEPALIVE_TIMEOUT as --timeout-keep-alive.
The image builds PgBouncer 1.25.2 from a checksummed tarball, copies the
compiled Rust extension into the /app source tree it imports from (it was only
in site-packages, which PYTHONPATH=/app shadows) and asserts the native bridge
loads. The app user is added to stats_users so operators can read the PgBouncer
console with the application credentials
The supervisor returns the pooled URL instead of writing into the mapping it
was handed, a database user whose name PgBouncer would split into several
stats_users entries is refused before the config is written, and the launcher
tests drive main() with an injected serve callable
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(terraform): describe the gateway.launch pooler entrypoint in the aws module README
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): run pgbouncer exit hooks only in the parent and copy the CA into the runtime dir
Gunicorn workers inherit the parent's atexit table, so a recycled worker (max_requests) stopped the shared pooler and removed its runtime dir, then hung in the inherited Popen lock. The hooks now no-op unless os.getpid() is the process that started PgBouncer
A verified TLS upstream named the operator's CA bundle directly, which is often a 0600 root-owned file that nobody (the user PgBouncer drops to) cannot read, so every server connection failed with "failed to load CA". The bundle is copied into the runtime dir next to the ini and chowned with it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(terraform): run the gateway through gateway.launch and add the gcp connection-pool variables
Cloud Run and ECS overrode the image command with uvicorn gateway.main:app, which skips the supervisor that starts the in-container PgBouncer, so LITELLM_PGBOUNCER_ENABLED was inert on both stacks. Both now exec python -m gateway.launch (under ddtrace-run when USE_DDTRACE is set), and the gcp module gains gateway_connection_pool_enabled / gateway_pool_max_db_connections / gateway_pool_max_client_conn wired to the gateway service only
The test_launch password_env fixture now restores DATABASE_URL even when it was unset: monkeypatch.delenv records nothing for an absent var, so main() left postgresql://...@db.internal in the xdist worker's environ and the key-rotation e2e test in the same proxy-infra shard stopped skipping and tried to reach db.internal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(pgbouncer): keep channel_binding and gssencmode off the loopback URL
Prisma would demand TLS channel binding from a pooler that only speaks
plain TCP on 127.0.0.1. Also pass the request the marketplace test
started needing after #40518 landed on top of #40496
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The Logs drawer gets messages/response from GET /spend/logs/ui/{request_id};
the list endpoint omits those heavy columns for every caller, admins included.
That detail route was missing from LiteLLMRoutes.spend_tracking_routes, and
check_route_access anchors patterns, so /spend/logs/ui never matched it. Every
internal_user got a 403 before the handler ran and the UI fell back to the
"Request/Response Data Not Available" banner, even on their own requests
Adds the route to spend_tracking_routes so internal_user, internal_user_view_only,
admin_viewer and org_admin all inherit it, and drops the now-redundant explicit
entry from admin_viewer_routes. The handler already authorizes non-admins per row
via _assert_user_can_view_request_id, so no handler-side scoping change is needed
That helper returned silently when no spend-log row existed, which the detail
handler treats as authorized before asking every custom logger for the payload by
raw request_id. With retention pruning the row can be gone while the payload is
still in cold storage, so opening the route would have let a non-admin read
another tenant's prompt out of S3/GCS. A missing row now falls through to the
same 403 as a foreign row, which also removes the exists-but-not-yours oracle
Fixes#34099
Adds object_permission.skills to keys and teams, enforces it on
/claude-code/marketplace.json?key=, /claude-code/plugins and
/claude-code/plugins/{name}, and exposes an Allowed Skills selector in
the key and team create/edit forms of the Admin UI
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): reject initialize with 403 when the key grants no MCP servers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): e2e expects 403 initialize for a key with no MCP servers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): mention IP filtering in the no-servers initialize denial and keep zero-grant tool coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A request whose model field matched no configured model was rejected with 400 but its failure row still persisted the raw client string as the model, so a client that concatenated its prompt into the model field wrote that prompt into LiteLLM_SpendLogs and the daily spend tables, where /user/daily/activity/aggregated returned it as a breakdown.models key. The spend log payload now records such rejections under the constant unknown-model, keeping the failed request counted without persisting client input as a model name.
* feat(rust_bridge): count budget-check input tokens in Rust on all LLM routes
Rust counts input tokens from the raw JSON body with the GIL released inside the existing budget reservation, covering every LLM route the auth dependency guards. It only fires for models on the Anthropic tokenizer when a budget is set, and Python counts whenever Rust is off, missing, or declines a body shape.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(rust): count byte-level BPE tokens without the GPT-2 split regex (#40594)
The oniguruma run of the ByteLevel pre-tokenizer regex is about 90% of
encode_fast on a 100k token body (100 ms of the ~110 ms Rust admission
count in the gateway pod). A hand-written scanner that yields the same
pieces, then feeds the model directly, counts the same text in 10 ms.
It only engages for tokenizers with the Anthropic shape (optional NFKC,
ByteLevel without prefix space, no post-processor) and falls back to the
full encoder when the text contains an added token. Parity with
encode_fast is tested on random texts, the pieces are compared with the
real pre-tokenizer, and the \p{L}/\p{N}/\s tables are checked against
oniguruma for every code point.
NFKC runs through unicode-normalization-alignments, the crate and
Unicode tables NormalizedString::nfkc already uses, so the fast path
normalizes exactly what the full encoder would. Using the newer
unicode-normalization crate changed the count for 171 code points that
gained compatibility decompositions after Unicode 9 (U+32FF, U+A7F1..).
The fast normalizer is compared with the tokenizer's for every scalar
value and on random texts.
The scanner is built without mutable state: byte_char and mapped_len replace the const table builders and the reusable mapped buffer, and iter::successors replaces the stateful piece iterator. byte_chars_match_the_byte_level_alphabet checks the byte mapping against ByteLevel for every scalar value.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust_bridge): bound concurrent token-count encodes and share the Anthropic tokenizer predicate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The writer health probe only ran SELECT 1, which a read-only Postgres
session answers fine, so a pooled connection left pointing at a demoted
primary kept failing every write with SQLSTATE 25006 until the pod was
restarted. Probe transaction_read_only instead, treat a 25006 on the
request path as a signal to recreate the client, and back off
exponentially while the database as a whole stays read-only so a replica
or an in-progress failover does not get its engine killed every cycle.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
While the Redis circuit breaker is open every guarded call was refused with a bare
Exception that each swallowing catch site logged as an ERROR traceback, so a
sub-second latency blip turned into thousands of tracebacks per minute and pinned
every replica's CPU. The sync guard also recorded socket timeouts as hard
connectivity failures, so with least-busy routing the breaker opened on the first
slow replies and the timeout-only min-duration guard never applied.
Refusals now raise RedisCircuitBreakerOpenError, and the catch sites route it
through log_redis_failure, which logs a refusal at DEBUG and everything else at
the caller's level. The sync guard passes is_timeout like the async one.
The Admin UI and the docs already present Organizations as an enterprise
feature, but every /organization route served unlicensed proxies. A router
level dependency now enforces the license on all of them, and it resolves
the auth dependency first so a bad key still gets 401 rather than 403.
Claude-Session: https://claude.ai/code/session_01Se8ERtsqMQ3eVzWiLVMyNS
Profiling the sidecar-enabled gateway at 700 rps showed ~2.4% of all samples
in get_model_group_info called per request from budget reservation, plus
get_deployment_model_info for tiered pricing tables. Both are read-only lookups
over the model list, so serve them from the Router's lru caches and clear the
deployment cache alongside the group cache when the model list changes.
The deployment-info cache is a per-router lru_cache built in __init__ rather
than a class-level decorated method, so it does not pin Router instances in a
process-wide cache and is dropped with the router.
A price data reload replaces litellm.model_cost without touching model_list, so
the reload replay also clears both caches; otherwise reservation would keep
pricing against the old catalog until an unrelated model-list change.
Co-authored-by: yassin <yassin@berri.ai>