litellm/litellm
mateo-berri caf9bbbd5a fix(least-busy): keep the shared count readable, counted once, and off the loop
A Redis outage read as "every deployment is idle", because batch_get_cache
swallows the failure and answers with an empty dict. batch_get_counts and its
async twin raise instead, so a worker that cannot reach Redis falls back to its
own numbers rather than routing on zeros.

The counter's TTL is now set only on a key that has none, so a +1 left behind by
a worker that died mid-request ages out an hour after the key was created. It
used to be refreshed on every touch, which kept that stuck count alive for as
long as the group took traffic.

Two least-busy groups counted the same request twice, since the pre-call list
kept a selector per group while the success list deduped by class. The selector
now goes on through add_litellm_input_callback, which dedupes the same way.

A prompt-management model picked its deployment on the synchronous path, so the
new Redis read landed on the event loop and configured routing plugins never
ran. It awaits the async selector now.
2026-09-06 00:20:05 -07:00
..
a2a_protocol Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_r4 2026-09-02 23:32:10 +00:00
anthropic_interface fix(anthropic_endpoints): return Anthropic type:error envelope for /v1/messages errors 2026-08-31 16:23:36 -07:00
assistants refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
batch_completion
batches refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
caching fix(least-busy): keep the shared count readable, counted once, and off the loop 2026-09-06 00:20:05 -07:00
completion_extras Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_headroom_ccr_streaming_responses 2026-09-01 22:25:59 -07:00
compression refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
containers Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5_r4 2026-09-02 23:32:10 +00:00
endpoints/speech/speech_to_completion_bridge refactor(speech): freeze httpx response header dicts (LIT002) 2026-08-31 21:12:53 -07:00
evals
experimental_mcp_client chore: merge origin/litellm_internal_staging into litellm_decrease_anys_opus5_r4 2026-09-04 18:59:35 -07:00
files refactor(typing): replace Any with proven types in 42 more backend files 2026-09-02 15:35:01 +00:00
fine_tuning refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
google_genai refactor(types): replace Any with precise types across 73 modules 2026-09-01 11:05:02 +00:00
images refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
integrations fix(azure_sentinel): split batches under the 1MB ingestion cap (#39880) 2026-09-05 17:15:36 -07:00
interactions refactor(typing): replace Any with proven types in 42 more backend files 2026-09-02 15:35:01 +00:00
litellm_core_utils fix(proxy): redact provider keys from pass-through failure tracebacks 2026-09-05 15:58:33 -07:00
llms fix(anthropic): keep provider_specific_fields off the native /v1/messages wire (#39967) 2026-09-05 17:50:54 -07:00
models feat(mcp): add opt-in per-server oauth relay discovery (#39936) 2026-09-05 18:04:35 -07:00
ocr feat(ocr): add Cohere Parse support for cohere and azure_ai 2026-09-04 21:25:46 -07:00
passthrough Merge origin/litellm_internal_staging into litellm_techdebt_20260901 2026-09-01 19:38:19 +00:00
proxy feat(router): gate heuristic v1 tuning (#39952) 2026-09-05 19:24:00 -07:00
proxy_auth
rag chore: merge origin/litellm_internal_staging into litellm_decrease_anys_opus5_r4 2026-09-04 18:59:35 -07:00
realtime_api Merge pull request #39851 from BerriAI/litellm_fix_realtime_backend_close_hang 2026-09-05 09:51:15 -07:00
repositories feat(router): gate heuristic v1 tuning (#39952) 2026-09-05 19:24:00 -07:00
rerank_api fix(rerank): adopt declared authenticating providers in arerank instead of resolving them 2026-09-01 14:47:59 -07:00
responses fix(file_search): escape the dropped vector_store_id in the warning 2026-09-05 16:44:54 -07:00
router_strategy fix(least-busy): keep the shared count readable, counted once, and off the loop 2026-09-06 00:20:05 -07:00
router_utils feat(router): gate heuristic v1 tuning (#39952) 2026-09-05 19:24:00 -07:00
rust_bridge chore: merge litellm_internal_staging into litellm_techdebt_20260903 2026-09-04 16:23:52 +00:00
sandbox
search fix(search): forward search-tool params through the router, complete Parallel AI v1 param mapping (#37883) 2026-09-01 21:46:46 -07:00
secret_managers chore: merge origin/litellm_internal_staging into litellm_decrease_anys_opus5_r4 2026-09-04 18:59:35 -07:00
skills refactor(typing): replace Any with proven types in 42 more backend files 2026-09-02 15:35:01 +00:00
types feat(mcp): add opt-in per-server oauth relay discovery (#39936) 2026-09-05 18:04:35 -07:00
vector_store_files refactor(typing): replace Any with proven types in 42 more backend files 2026-09-02 15:35:01 +00:00
vector_stores Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci 2026-09-03 00:16:19 -07:00
videos refactor(videos): make video upload param keyword-only on public edit fns 2026-08-24 15:55:07 -07:00
__init__.py Revert "perf: lazy-load SDK symbols so import litellm stays under 60 MB RSS (…" 2026-09-05 16:07:09 -07:00
_internal_context.py
_lazy_imports.py fix(lazy_imports): type import_map as Mapping to stay under the LIT001 budget 2026-09-05 16:12:45 -07:00
_lazy_imports_registry.py Revert "perf: lazy-load SDK symbols so import litellm stays under 60 MB RSS (…" 2026-09-05 16:07:09 -07:00
_logging.py fix(proxy): stop leaking internal exception details to clients (#39380) 2026-09-02 17:32:00 -07:00
_redis.py fix(redis): coerce env var string types and fix param discovery through decorator wrappers (#30644) 2026-08-31 20:51:31 -07:00
_redis_credential_provider.py refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
_service_logger.py chore: merge origin/litellm_internal_staging into litellm_decrease_anys_opus5_r4 2026-09-04 18:59:35 -07:00
_uuid.py
_version.py
anthropic_beta_headers_config.json
anthropic_beta_headers_manager.py
blog_posts.json
budget_manager.py
constants.py perf(logging): scan large base64 payloads for log truncation off the event loop (#39890) 2026-09-05 11:51:15 -07:00
cost.json
cost_calculator.py fix(xai): merge litellm_internal_staging and keep the stream builder from short-circuiting xAI's reported cost 2026-09-02 16:46:35 -07:00
exceptions.py refactor(typing): replace Any with proven types in 89 more backend files 2026-09-02 23:08:42 +00:00
main.py Merge pull request #38975 from BerriAI/litellm_fix_azure_ai_reclassify 2026-09-03 13:15:40 -07:00
model_prices_and_context_window_backup.json Merge pull request #39764 from BerriAI/litellm_govcloud_profiles_lit6421 2026-09-05 17:15:22 -07:00
policy_templates_backup.json
provider_endpoints_support_backup.json fix(ocr): send each provider a health-check document it accepts 2026-09-04 22:52:12 -07:00
py.typed
router.py fix(least-busy): keep the shared count readable, counted once, and off the loop 2026-09-06 00:20:05 -07:00
scheduler.py
setup_wizard.py feat(models): add Claude Fable 5.1 across Anthropic, Bedrock, Vertex AI, and Azure AI 2026-09-01 18:07:06 +00:00
timeout.py
utils.py feat(ocr): add Cohere Parse support for cohere and azure_ai 2026-09-04 21:25:46 -07:00