mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
Ports the shunt technique for Claude Code (https://engineering.atspotify.com/2026/9/portal-by-spotify-cut-my-claude-code-token-usage-by-90) into the auto router, server-side, so it works for any client with no plugin to install. shunt intercepts large file reads and boilerplate generation at the client with PreToolUse hooks and hands them to a cheap worker model. This makes the same decision on the proxy: a new always-on ShuntGuardrail injects bulk_read and code_write tool definitions pre-call, then post-call rewrites large Read/Bash/bulk_read/code_write tool_use blocks into a Bash command carrying shunt's own wc -l conditional, so small files still get read directly and only large ones are delegated. Two new endpoints, /v1/bulk_read and /v1/code_write, run the worker call through llm_router.acompletion with shunt's verbatim system prompts, so worker spend is tracked against the calling key and team like any other request. The guardrail arms per request off an auto-router marker's own litellm_params (auto_router_shunt_min_lines and the two worker-model fields), the same shape auto_router_compression already uses, so no guardrails: config entry is needed. It stays inert when the resolved model has no shunt config. In the UI, "Shunt" is one entry in the existing preset dropdown. Its tiers resolve through fallback chains against whatever models the proxy actually has, so it greys out only when there are no chat models at all, never merely because the named models are absent, and all three settings stay editable under Advanced. Streaming buffers the whole response before rewriting, matching tool_permission.py, because input_json_delta fragments split mid-token. An unparseable stream passes through untouched: shunt is an optimization, not a safety control, so a request should never fail because the rewrite could not run. The CLI's Bash allow rules only reduce prompts for the generated commands. Claude Code matches rule text before the first wildcard with no host-aware matching, so they cannot pin the destination, and the docs recommend a PreToolUse hook where a real boundary is needed. |
||
|---|---|---|
| .. | ||
| agent_tests | ||
| audio_tests | ||
| base_sdk_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| enterprise | ||
| guardrails_tests | ||
| image_gen_tests | ||
| integration | ||
| litellm-proxy-extras | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| rust-python-harness | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| unified_google_tests | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _wait_helpers.py | ||
| _ws_vcr.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_organizations.py | ||
| test_otel_thread_leak.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_rust_python_harness.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/test_litellm
This folder can only run mock tests.