Commit graph

36203 commits

Author SHA1 Message Date
Sameer Kankute
6b8b99b4db
docs(router): add health check driven routing guide
New standalone page covering the full health check routing feature:
allowed_fails_policy integration, health_check_ignore_transient_errors,
architecture SVG, step-by-step setup, and gotchas (TTL, AllowedFails semantics).

Replaces the inline section in health.md with a link to the new page.
Added to the Routing & Load Balancing sidebar.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 14:27:09 -07:00
Sameer Kankute
9a605e2bc5
fix(router): also exclude 429/408 from health state cache when ignore_transient_errors set
The previous fix only skipped cooldown counter increments. The health state
cache was still marking 429/408 endpoints as is_healthy=False, causing the
binary health check filter to exclude them from routing.

Now, when health_check_ignore_transient_errors=True, 429/408 endpoints are
also excluded from the unhealthy list passed to build_deployment_health_states(),
so the binary filter treats them as unaffected (not unhealthy).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 14:26:53 -07:00
Sameer Kankute
4fb8932e36
feat(router): add health_check_ignore_transient_errors flag
When enabled, health check failures with 429 (rate limit) or 408 (timeout)
status codes are skipped from the cooldown pipeline. These are transient
load issues, not broken deployments. Auth errors (401), 404, and 5xx errors
still increment counters and trigger cooldown as before.

Config (general_settings):
  health_check_ignore_transient_errors: true

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 14:26:53 -07:00
Sameer Kankute
23c31560e2
feat(router): integrate allowed_fails_policy into health check failures
Health check failures now increment the same per-deployment failure
counters used by allowed_fails_policy, so users can control how many
health check failures of each error type are required before a
deployment enters cooldown.

- ahealth_check() preserves the original exception in its return dict
- run_with_timeout() returns a litellm.Timeout on health check timeout
- _perform_health_check() propagates exceptions to unhealthy endpoints
- _write_health_state_to_router_cache() calls _set_cooldown_deployments
  for each unhealthy endpoint that has an exception
- When allowed_fails_policy is set, the binary health check filter is
  bypassed so cooldown is the sole routing exclusion mechanism
- Safety net: if all deployments are in cooldown with
  enable_health_check_routing=True, the cooldown filter is bypassed

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 14:26:21 -07:00
Sameer Kankute
e9f2065be1
fix(docker): include enterprise bridge in non-root runtime image
Copy the /app/enterprise bridge package into the non-root runtime image so enterprise proxy hooks register correctly (including managed_files).
2026-04-03 14:26:21 -07:00
Sameer Kankute
45bc2df360
fix: revert accidental _litellm_uuid import back to _uuid
The isort hook picked up a stale rename from the working directory.
Both router.py and proxy_server.py need litellm._uuid, not _litellm_uuid.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:26:21 -07:00
Sameer Kankute
d32a282be9
fix: re-attach model_id after endpoint cleaning, bump log level
- model_id is now added after _clean_endpoint_data() so it survives
  health_check_details: False (MINIMAL_DISPLAY_PARAMS filtering)
- Health state write failures logged at warning instead of debug

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:26:21 -07:00
Sameer Kankute
9a442e6e31
feat(router): add health-check-driven routing behind opt-in flag
Background health checks now feed deployment health state into the
router candidate-filtering pipeline. Unhealthy deployments are excluded
proactively instead of waiting for request failures to trigger cooldown.

Gated by `enable_health_check_routing: true` in general_settings.
Off by default — zero behavior change for existing users.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:26:21 -07:00
Sameer Kankute
af4ac2b3ce
Fix codeql 2026-04-03 14:26:21 -07:00
Sameer Kankute
7ca560aba7
Fix test 2026-04-03 14:26:21 -07:00
Sameer Kankute
758cd7d8aa
Fix test 2026-04-03 14:26:21 -07:00
Sameer Kankute
e23c8a9589
fix: revert accidental _litellm_uuid import back to _uuid
The isort hook picked up a stale rename from the working directory.
Both router.py and proxy_server.py need litellm._uuid, not _litellm_uuid.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:25:56 -07:00
Sameer Kankute
b06019dfea
fix: re-attach model_id after endpoint cleaning, bump log level
- model_id is now added after _clean_endpoint_data() so it survives
  health_check_details: False (MINIMAL_DISPLAY_PARAMS filtering)
- Health state write failures logged at warning instead of debug

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:25:56 -07:00
Sameer Kankute
94816c7d83
feat(router): add health-check-driven routing behind opt-in flag
Background health checks now feed deployment health state into the
router candidate-filtering pipeline. Unhealthy deployments are excluded
proactively instead of waiting for request failures to trigger cooldown.

Gated by `enable_health_check_routing: true` in general_settings.
Off by default — zero behavior change for existing users.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:25:56 -07:00
yuneng-jiang
397f9ab0fc
Merge pull request #24611 from Sameerlite/Sameerlite/order-fallback2
feat(router): order-based fallback across deployment priority levels
2026-04-03 14:25:03 -07:00
Sameer Kankute
40d245ce02
Fix tests 2026-04-03 14:25:03 -07:00
Sameer Kankute
5badcaa7f3
fix(tests): reset module-level cache in stale alias bypass tests
Reset _ENABLE_TEAM_STALE_ALIAS_BYPASS to None in both test functions
to ensure test isolation and prevent ordering-dependent failures

Made-with: Cursor
2026-04-03 14:25:03 -07:00
Sameer Kankute
b10daef606
fix(router): address final Greptile P1/P2 comments
- Reorder team_public_model_name assignment to happen before model_name mutation for clarity
- Add comment explaining no-rename fast-exit case in _update_existing_team_model_assignment
- Add comment explaining final patch_data.model_name = None applies to all code paths

Made-with: Cursor
2026-04-03 14:25:03 -07:00
Sameer Kankute
6a3d2553a4
fix(router): address remaining Greptile review comments
- Cache LITELLM_ENABLE_TEAM_STALE_ALIAS_BYPASS at module level to avoid hot-path secret lookups
- Add clarifying comments for should_include_deployment team isolation logic
- Add negative assertion for update_team.assert_not_called() in test
- Add docstring clarification for _get_team_deployments helper pattern
- Add explicit assertion message in test_get_model_list_alias_optimization

Made-with: Cursor
2026-04-03 14:25:03 -07:00
Sameer Kankute
d92bb1b4d8
fix(router): address Greptile P1/P2 review comments
- Add deduplication guard in _update_team_model_index to prevent duplicate indices
- Add wildcard comment in map_team_model for clarity
- Add monkeypatch to test_team_alias_stale_bypass_disabled_by_default for determinism
- Extract _get_team_deployments helper to centralize DB access pattern
- Add clarifying comments for team_public_model_name assignment ordering

Made-with: Cursor
2026-04-03 14:25:03 -07:00
Sameer Kankute
a174dc6db2
Fix greptile reviews and mock test 2026-04-03 14:25:03 -07:00
Sameer Kankute
2169ac43b3
Fix greptile reviews and mock test 2026-04-03 14:25:03 -07:00
Sameer Kankute
7af677af97
Fix greptile reviews and mock test 2026-04-03 14:25:03 -07:00
Sameer Kankute
55e8fa777d
Fix code qa issues 2026-04-03 14:25:03 -07:00
Sameer Kankute
b904e78899
Fix greptile comments 2026-04-03 14:25:03 -07:00
Sameer Kankute
acbe06b8f2
Fix greptile comments 2026-04-03 14:24:45 -07:00
Sameer Kankute
d00d7939cf
Fix greptile comments 2026-04-03 14:24:45 -07:00
Sameer Kankute
5f15ab864e
Fix greptile comments 2026-04-03 14:24:45 -07:00
Sameer Kankute
dc60e8e6f7
Fix greptile comments 2026-04-03 14:24:44 -07:00
Sameer Kankute
b1c97dcad9
fix(routing): address state consistency and type safety issues
- Check alias target pattern to detect stale team aliases
- Fix PrismaClient type annotation to Optional
- Eliminate in-place mutation in index update logic

Made-with: Cursor
2026-04-03 14:24:21 -07:00
Sameer Kankute
ef5591b561
perf(routing): optimize team model checks and improve test coverage
- Use O(1) team index lookup instead of map_team_model in alias guard
- Fix MockPrismaClient to validate where clause filters
- Add comment explaining DB query trade-off for team deployments

Made-with: Cursor
2026-04-03 14:24:21 -07:00
Sameer Kankute
b1800a6d38
fix(routing): prevent stale model_aliases from interfering with team routing
- Skip model_aliases rewrite if model resolves to team deployments
- Add test coverage for sibling-preservation branch
- Update MockPrismaClient to support sibling deployment scenarios

Made-with: Cursor
2026-04-03 14:24:21 -07:00
Sameer Kankute
023b5a52f0
fix(router): guard None model_info and deduplicate team index logic
- Guard against None model_info in sibling deployment check
- Extract _update_team_model_index helper to eliminate duplication

Made-with: Cursor
2026-04-03 14:24:05 -07:00
Sameer Kankute
d4969cfb1b
fix(management): query DB directly for sibling deployments on rename
- Add clarifying comments to test assertions
- Query prisma DB instead of in-memory router to avoid stale state
- Prevents incorrect deletion of old public name when siblings exist

Made-with: Cursor
2026-04-03 14:24:05 -07:00
Sameer Kankute
ea00646d5c
fix(router): prevent cross-team deployment leakage in fallback path
Guard should_include_deployment fallback to only return deployments
matching the requested team_id, preventing public-name collisions
from leaking deployments across teams

Made-with: Cursor
2026-04-03 14:24:05 -07:00
Sameer Kankute
b932aac33a
fix(router): address Greptile P1/P2 performance issues
- Guard against llm_router=None to prevent silent deletion
- Add O(1) team_model index to avoid O(n) scan on every team request

Made-with: Cursor
2026-04-03 14:24:05 -07:00
Sameer Kankute
da29da3bc3
fix(router): address remaining Greptile P0/P1 issues
- Update map_team_model test to expect public name return
- Only remove old public name if no sibling deployments use it

Made-with: Cursor
2026-04-03 14:24:04 -07:00
Sameer Kankute
114afc3a2d
fix(router): address Greptile review comments
- Add None guard for original_model_name in _add_team_model_to_db
- Remove stale old public name when renaming team model
- Add comment clarifying team deployment early-return priority

Made-with: Cursor
2026-04-03 14:23:49 -07:00
Sameer Kankute
2ddb56e00d
chore(team-routing): remove temporary candidate pool logs
Remove temporary fire-emoji router logs used for local verification while keeping team sibling deployment routing behavior unchanged.

Made-with: Cursor
2026-04-03 14:23:49 -07:00
Sameer Kankute
1245ab49bb
fix(team-routing): keep team model routing on public names
Remove team model_alias rewrites and resolve team deployments by team_public_model_name with team_id so sibling deployments stay in the routing candidate pool, with explicit logs showing candidate selection before load balancing.

Made-with: Cursor
2026-04-03 14:23:49 -07:00
Sameer Kankute
fbc4baebc4
docs: remove enable_pre_call_checks requirement from order docs
Order-based routing and fallback work without enable_pre_call_checks
in the current code. Remove the stale requirement from both doc files.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:22:28 -07:00
Sameer Kankute
6fdc29a292
fix(router): handle non-standard fallback formats with order-based fallback
When fallbacks use non-standard formats (e.g. ["claude-3-haiku"] or
[{"model": "...", "messages": [...]}]), detect them with
_check_non_standard_fallback_format and pass them through directly
instead of trying to parse with get_fallback_model_group which only
handles the standard dict-keyed format.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:22:28 -07:00
Sameer Kankute
08d1303c3c
feat(router): add order-based fallback so higher order deployments are tried on failure
When a request to an order=1 deployment fails, the router now
automatically tries order=2, order=3, etc. before falling through to
external fallbacks. Works for all error types (429, 404, connection
errors). Requires enable_pre_call_checks=True.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:22:28 -07:00
Sameer Kankute
040c7b53d3
feat(router): add order-based fallback so higher order deployments are tried on failure
When order=1 deployments fail, the router now automatically tries order=2,
then order=3, etc. before falling through to external fallbacks. This removes
the need for enable_pre_call_checks and makes order work as a true priority-based
fallback chain within a model group.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 14:22:04 -07:00
Harshit28j
6300a92aed
fix: address Greptile review feedback on key rotation lock 2026-04-03 14:21:35 -07:00
Krrish Dholakia
40879e7066 fix: remove dead migration sql 2026-03-23 20:14:37 -07:00
Ishaan Jaffer
659d4013c5 fix: address Greptile review issues
- GCPIAMCredentialProvider now inherits from redis.credentials.CredentialProvider
  so redis-py's async path calls get_credentials_async() properly
- move _redis_credential_provider import to top of _redis.py (PEP 8)
- remove dead else-branch that silently no-oped (gcp_service_account from
  redis_kwargs.get() was always None since it's popped by _get_redis_client_logic)
- remove mid-function 'from litellm import get_secret_str' inline import
- remove unused 'call' import from test_redis.py
2026-03-23 10:17:43 -07:00
Ishaan Jaffer
a1494b59a3 refactor(redis): move GCPIAMCredentialProvider to its own file
Extract GCPIAMCredentialProvider and _generate_gcp_iam_access_token
into litellm/_redis_credential_provider.py. _redis.py imports them
from there, keeping the public API unchanged.
2026-03-23 10:17:43 -07:00
Ishaan Jaffer
67b752ca8f fix(redis): regenerate GCP IAM token per connection for async cluster clients
Async RedisCluster was generating the IAM token once at startup and
storing it as a static password. After the 1-hour GCP token TTL, any
new connection (including to newly-discovered cluster nodes) would fail
to authenticate.

Fix: introduce GCPIAMCredentialProvider that implements redis-py's
CredentialProvider protocol. It calls _generate_gcp_iam_access_token()
on every new connection, matching what the sync redis_connect_func
already does. async_redis.RedisCluster accepts a credential_provider
kwarg which is invoked per-connection.
2026-03-23 10:17:43 -07:00
yuneng-jiang
9e98a7d7b3 Merge branch 'main' of github.com:BerriAI/litellm into litellm_rc_branch 2026-03-21 23:50:45 -07:00