litellm/litellm/llms
devin-ai-integration[bot] 96f58fac53
fix(router): don't cool down parent deployment on advisor sub-call failure (#33792)
* fix(router): don't cool down parent deployment on advisor sub-call failure

Advisor orchestration issues a sub-call to a different provider/credentials than the selected deployment. When that sub-call fails (e.g. a 401 because no advisor API key is configured), the exception propagates up and the router's deployment_callback_on_failure attributes it to the healthy parent deployment's model_info.id, cooling it down and rejecting unrelated callers to the same model group.

Tag advisor sub-call failures on the exception and skip cooldown for them in deployment_callback_on_failure. The exception is tagged rather than wrapped so its type is preserved and retry/fallback classification and the client-facing error are unchanged. Genuine executor/deployment failures are untagged and still cool down as before.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(router): tag advisor orchestration failures via provider-neutral util

Address review on LIT-4565: move the cooldown-exemption marker into
litellm/router_utils/cooldown_handlers.py so the router imports it at
module top instead of an in-function anthropic import, and extend the
exemption to AdvisorMaxIterationsError so a max-iterations orchestration
failure no longer cools down the healthy executor deployment.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-25 10:17:13 -07:00
..
a2a fix(a2a): populate response usage in a2a chat transformation (#31980) 2026-07-03 09:28:36 +05:30
ai21/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
aiml style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
aiohttp_openai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
amazon_nova style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
anthropic fix(router): don't cool down parent deployment on advisor sub-call failure (#33792) 2026-07-25 10:17:13 -07:00
apiserpent style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
aws_polly style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
azure fix(azure): build responses input_items url with path before query string (#32270) 2026-07-06 14:02:44 -07:00
azure_ai fix(vertex,azure): model-aware mid-conversation system for Claude /v1/messages 2026-07-20 16:48:29 -07:00
base_llm feat(guardrails): streaming text transformation in generic_guardrail_api (#33110) 2026-07-14 17:38:11 -07:00
baseten style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
bedrock fix(bedrock-mantle): backfill usage on non-streaming /v1/messages responses 2026-07-23 23:37:40 +00:00
bedrock_mantle Merge origin/litellm_internal_staging into litellm_mantle_codex_additional_tools (resolve overlap with #33228 hoist) 2026-07-20 20:30:34 -07:00
black_forest_labs style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
brave/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
bytez style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cerebras fix: add reasoning param support for GPT OSS cerebras 2026-02-02 17:20:04 +05:30
chatgpt style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
clarifai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cloudflare/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
codestral/completion style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cohere style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cometapi style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
compactifai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
custom_httpx fix(rust): route agentic-completion-hook /messages requests to Python for all stream modes (#34126) 2026-07-22 00:36:47 +00:00
dashscope fix(cost): coerce string tiered-pricing costs and share tier helper 2026-07-11 14:05:26 -07:00
databricks fix(anthropic): override custom_llm_provider in provider config subclasses so capability probes use the right namespace 2026-07-11 12:15:59 -07:00
dataforseo/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
datarobot/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
deepgram style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
deepinfra style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
deepseek test: litellm fix failing tests (#32577) 2026-07-09 13:54:45 -07:00
deprecated_providers style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
docker_model_runner/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
duckduckgo/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
e2b style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
elevenlabs style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
empower/chat LiteLLM Common Base LLM Config (pt.3): Move all OAI compatible providers to base llm config (#7148) 2024-12-10 17:12:42 -08:00
exa_ai/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
fal_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
fastcrw style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
featherless_ai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
firecrawl style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
fireworks_ai fix(fireworks_ai): restore Content-Type application/json header (fixes 415) (#33929) 2026-07-20 09:52:35 -07:00
friendliai/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
galadriel/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
gdc feat(gdc): implement Google Distributed Cloud (GDC) Gemini provider (#31895) 2026-07-01 17:31:07 -07:00
gemini fix(realtime): stop second Gemini Live setup, retry hung handshake, close guardrail bypass (#31519) 2026-06-28 08:52:20 +05:30
gigachat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
github/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
github_copilot fix(anthropic): override custom_llm_provider in provider config subclasses so capability probes use the right namespace 2026-07-11 12:15:59 -07:00
google_pse/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
gradient_ai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
groq style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
heroku/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
hosted_vllm style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
huggingface chore(proxy): clean up request parameter validation and provider destination handling (#34189) 2026-07-22 00:57:58 +00:00
hyperbolic Revert "Litellm dev 07 21 2025 p1 (#12848)" 2025-07-22 18:28:36 -07:00
inception style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
infinity style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
jina_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
lambda_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
langflow style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
langgraph style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
lemonade style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
linkup style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
litellm_proxy style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
llamafile/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
lm_studio style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
manus style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
meta_llama/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
milvus/vector_stores style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
minimax style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
mistral style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
modelscope style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
moonshot/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
morph style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
nebius Integration with Nebius AI Studio added (#11143) 2025-05-27 11:05:22 -07:00
nlp_cloud style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
novita/chat build(deps-dev): bump black to 26.3.1 and apply formatting (#28525) 2026-05-21 17:24:18 -07:00
nscale/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
nvidia_nim style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
nvidia_riva style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
oci chore(lint): zero out crash-class pyright rules and ban new type: ignore comments (#32152) 2026-07-04 16:56:12 -07:00
ollama chore(lint): zero out crash-class pyright rules and ban new type: ignore comments (#32152) 2026-07-04 16:56:12 -07:00
oobabooga chore(proxy): clean up request parameter validation and provider destination handling (#34189) 2026-07-22 00:57:58 +00:00
openai feat(guardrails): add Compresr guardrail for query-aware context compression (#33295) 2026-07-15 13:53:41 -07:00
openai_like fix(anthropic): override custom_llm_provider in provider config subclasses so capability probes use the right namespace 2026-07-11 12:15:59 -07:00
openrouter style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
opensandbox style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
ovhcloud style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
parallel_ai/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
pass_through style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
perplexity style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
petals style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
pg_vector/vector_stores style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
predibase style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
ragflow style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
recraft style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
reducto style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
replicate style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
runwayml style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
s3_vectors style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
sagemaker test(sagemaker): cover sync native streaming path via injectable make_sync_call 2026-07-23 04:21:12 +00:00
sambanova style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
sap chore(lint): zero out crash-class pyright rules and ban new type: ignore comments (#32152) 2026-07-04 16:56:12 -07:00
scaleway/audio_transcription style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
searchapi style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
searxng style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
serper/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
snowflake style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
soniox style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
stability style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
tavily/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
tencent feat(tencent): add Tencent TokenHub as a provider (#31903) 2026-07-02 18:31:59 -07:00
tinyfish/search feat(tinyfish): make search provider permissive, attribute errors (#31997) 2026-07-03 10:17:11 -07:00
together_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
topaz style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
triton style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
v0 style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
vercel_ai_gateway style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
vertex_ai fix: handle explicit outputInfo: null in Vertex AI batch response (#34473) 2026-07-24 15:10:16 -07:00
vllm style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
volcengine style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
voyage style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
wandb (feat): Add W&B Inference to LiteLLM 2025-09-11 00:07:30 +05:30
watsonx style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
xai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
xinference/image_generation style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
you_com style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
zai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
__init__.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
base.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
custom_llm.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
maritalk.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
README.md LiteLLM Minor Fixes and Improvements (09/13/2024) (#5689) 2024-09-14 10:02:55 -07:00

File Structure

August 27th, 2024

To make it easy to see how calls are transformed for each model/provider:

we are working on moving all supported litellm providers to a folder structure, where folder name is the supported litellm provider name.

Each folder will contain a *_transformation.py file, which has all the request/response transformation logic, making it easy to see how calls are modified.

E.g. cohere/, bedrock/.