litellm/litellm/passthrough
Sameer Kankute e8e5b47fc2
feat(passthrough): add configurable pass-through request timeouts (#30266)
* feat(passthrough): add configurable pass-through request timeouts

Allow operators to set general_settings.pass_through_request_timeout and per-endpoint timeout values, and apply them to native HTTP passthrough routes and SDK passthrough paths such as Bedrock /converse.

* fix(passthrough): address CI lint and regenerate dashboard API types

* refactor(passthrough): extract timeout utils to proxy-free module, fix router_timeout drop

- Move resolve_llm_passthrough_timeout + resolve_pass_through_request_timeout to
  litellm/passthrough/timeout_utils.py (no fastapi/proxy imports at module scope)
- router.py and passthrough/main.py now import from timeout_utils directly,
  avoiding the fastapi transitive import in pure SDK contexts
- pass_through_endpoints.py re-imports from timeout_utils for backward compat
- resolve_llm_passthrough_timeout now accepts router_timeout so Router(timeout=X)
  is respected for passthrough calls instead of being silently dropped
- Use _get_httpx_client (cached) instead of bare HTTPHandler(...) in sync path
  to avoid creating an unclosed client per call

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(router): use _explicit_timeout for passthrough to not shadow general_settings

self.timeout defaults to litellm.request_timeout (6000s) when the user
doesn't pass timeout= to Router(). Using it as router_timeout caused
general_settings.pass_through_request_timeout to be silently ignored.

Only pass router_timeout when the user explicitly set Router(timeout=X).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(timeout_utils): avoid fastapi transitive import by using sys.modules

resolve_pass_through_request_timeout previously did a lazy
`from litellm.proxy.proxy_server import general_settings` which loads
the proxy module (and transitively fastapi) even in pure SDK contexts.

Replace with a sys.modules lookup: if the proxy module is already loaded
(i.e. we're inside the proxy), read general_settings from it; otherwise
skip and fall back to the 600s default. No import is triggered.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(lint): remove unused imports from pass_through_endpoints.py

DEFAULT_PASS_THROUGH_REQUEST_TIMEOUT_SECONDS and resolve_llm_passthrough_timeout
are not used in this file; only resolve_pass_through_request_timeout is.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pass_through_endpoints): re-export DEFAULT_PASS_THROUGH_REQUEST_TIMEOUT_SECONDS

Tests import this constant directly from pass_through_endpoints.py;
re-add it to the import from timeout_utils for backward compatibility.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pass_through_endpoints): re-export resolve_llm_passthrough_timeout for backward compat

Tests import both DEFAULT_PASS_THROUGH_REQUEST_TIMEOUT_SECONDS and
resolve_llm_passthrough_timeout from pass_through_endpoints.py; use
noqa comments to suppress the unused-import lint warning on re-exports.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-12 07:40:02 -07:00
..
__init__.py build(VLLM-Passthrough-with-loadbalancing-support-(enables-using-model-list-for-VLLM-/classify-endpoint)): Closes #11205 2025-05-31 09:00:04 -07:00
main.py feat(passthrough): add configurable pass-through request timeouts (#30266) 2026-06-12 07:40:02 -07:00
README.md build(VLLM-Passthrough-with-loadbalancing-support-(enables-using-model-list-for-VLLM-/classify-endpoint)): Closes #11205 2025-05-31 09:00:04 -07:00
timeout_utils.py feat(passthrough): add configurable pass-through request timeouts (#30266) 2026-06-12 07:40:02 -07:00
utils.py fix(proxy): Bedrock Knowledge Base pass-through: preserve SigV4 headers and signed request body (#27526) 2026-05-25 19:21:55 +05:30

This makes it easier to pass through requests to the LLM APIs.

E.g. Route to VLLM's /classify endpoint:

SDK (Basic)

import litellm


response = litellm.llm_passthrough_route(
    model="hosted_vllm/papluca/xlm-roberta-base-language-detection",
    method="POST",
    endpoint="classify",
    api_base="http://localhost:8090",
    api_key=None,
    json={
        "model": "swapped-for-litellm-model",
        "input": "Hello, world!",
    }
)

print(response)

SDK (Router)

import asyncio
from litellm import Router

router = Router(
    model_list=[
        {
            "model_name": "roberta-base-language-detection",
            "litellm_params": {
                "model": "hosted_vllm/papluca/xlm-roberta-base-language-detection",
                "api_base": "http://localhost:8090", 
            }
        }
    ]
)

request_data = {
    "model": "roberta-base-language-detection",
    "method": "POST",
    "endpoint": "classify",
    "api_base": "http://localhost:8090",
    "api_key": None,
    "json": {
        "model": "roberta-base-language-detection",
        "input": "Hello, world!",
    }
}

async def main():
    response = await router.allm_passthrough_route(**request_data)
    print(response)

if __name__ == "__main__":
    asyncio.run(main())

PROXY

  1. Setup config.yaml
model_list:
  - model_name: roberta-base-language-detection
    litellm_params:
      model: hosted_vllm/papluca/xlm-roberta-base-language-detection
      api_base: http://localhost:8090
  1. Run the proxy
litellm proxy --config config.yaml

# RUNNING on http://localhost:4000
  1. Use the proxy
curl -X POST http://localhost:4000/vllm/classify \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <your-api-key>" \
-d '{"model": "roberta-base-language-detection", "input": "Hello, world!"}' \

How to add a provider for passthrough

See VLLMModelInfo for an example.

  1. Inherit from BaseModelInfo
from litellm.llms.base_llm.base_utils import BaseLLMModelInfo

class VLLMModelInfo(BaseLLMModelInfo):
    pass
  1. Register the provider in the ProviderConfigManager.get_provider_model_info
from litellm.utils import ProviderConfigManager
from litellm.types.utils import LlmProviders

provider_config = ProviderConfigManager.get_provider_model_info(
    model="my-test-model", provider=LlmProviders.VLLM
)

print(provider_config)