litellm/tests
moe-berri 98a0cf306f fix(shadow_eval): size the judge output cap for a judge that reasons
The cap covers reasoning tokens as well as the verdict, and the models people
pick as judges reason before answering whether the call asks them to or not:
Anthropic's 5 family thinks adaptively and cannot be told not to, so the
reasoning bills against max_tokens with nothing in the request to opt out.

At 1500 the reasoning consumed the budget and the reply arrived empty or cut
off mid-object, which the attempt recorded as an unparseable judge verdict
rather than a result. Headroom costs nothing: max_tokens is a ceiling and only
generated tokens bill, so the only movement is that judge calls which used to
bill their full budget and return nothing now return a verdict.

Deliberately not passing reasoning_effort to bound the reasoning instead:
is_thinking_enabled treats any reasoning_effort as thinking-enabled, which
drops the forced tool_choice that json_mode relies on and turns thinking on
with a 1024-token floor for judges that were not reasoning at all.
2026-09-04 15:18:14 -07:00
..
agent_tests
audio_tests
base_sdk_tests
basic_proxy_startup_tests
batches_tests
benchmarks
code_coverage_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci 2026-09-04 09:01:13 -07:00
documentation_tests Merge remote-tracking branch 'origin/main' into litellm_bedrock_messages_disconnect_billing 2026-08-31 08:58:37 -07:00
e2e Merge pull request #39684 from BerriAI/litellm_lit_4738_table_scrolling 2026-09-03 17:30:53 -07:00
enterprise Merge pull request #39661 from BerriAI/litellm_lit_4741_copy_id_search 2026-09-03 16:17:59 -07:00
guardrails_tests fix(logging): blocked requests no longer report guardrail_status=success in multi-guardrail configs (#39596) 2026-09-03 17:33:31 -07:00
image_gen_tests
integration
litellm-proxy-extras test(proxy-extras): fake run_prisma instead of subprocess.run in the migrate deploy harness 2026-09-03 16:10:44 -07:00
litellm_utils_tests test(aiohttp): pin NO_PROXY so proxy env cannot hijack the refused-port probe 2026-08-28 12:45:46 -07:00
llm_responses_api_testing fix(guardrails): deliver modify_response block as valid SSE on streaming chat and Responses 2026-08-31 16:01:39 -07:00
llm_translation test: exempt MockTransport request-shape embedding tests from VCR replay 2026-09-01 13:29:12 -07:00
load_tests feat(proxy): per-worker admission control that rejects excess requests with 503 (#39352) 2026-09-03 18:19:04 -07:00
local_testing fix(cache): use sync Redis batch reads (#39358) 2026-09-03 14:37:48 -07:00
logging_callback_tests Merge branch 'litellm_internal_staging' into litellm_fix_failing_request_slowdown 2026-09-03 00:20:51 -07:00
mcp_tests fix(mcp): scope allow-all servers to virtual keys (#39531) 2026-09-03 10:32:03 -07:00
multi_instance_e2e_tests
ocr_tests fix(vertex): avoid duplicate DeepSeek OCR model namespace 2026-09-01 14:39:15 -07:00
openai_endpoints_tests refactor(tests): assign the streamed id and lock poll once instead of rebinding 2026-09-03 13:49:43 -07:00
otel_tests
pass_through_tests fix(logging): key bridged /v1/messages rows on the id the caller received 2026-09-03 01:15:06 -07:00
pass_through_unit_tests test(websearch): carry a reasoned test-quality suppression on the router patch 2026-08-31 22:32:05 -07:00
proxy_admin_ui_tests refactor(tests): assign the streamed id and lock poll once instead of rebinding 2026-09-03 13:49:43 -07:00
proxy_behavior feat(team): report per-user spend within a team for JWT traffic (#39771) 2026-09-04 12:00:47 -07:00
proxy_e2e_anthropic_messages_tests
proxy_migration_tests fix(proxy-extras): kill the whole Prisma process group when a command times out 2026-09-02 18:29:55 -07:00
proxy_security_tests
proxy_unit_tests Merge pull request #39568 from BerriAI/litellm_fix-batch-spend-key-double-hash-bcae 2026-09-04 10:47:34 -07:00
router_unit_tests test: repair four chronically failing CI tests 2026-09-04 10:10:47 -07:00
rust-python-harness test: address review notes on the chronic-test repairs 2026-09-04 10:18:15 -07:00
search_tests
spend_tracking_tests
store_model_in_db_tests test: repair four chronically failing CI tests 2026-09-04 10:10:47 -07:00
test_litellm fix(shadow_eval): size the judge output cap for a judge that reasons 2026-09-04 15:18:14 -07:00
unified_google_tests
vector_store_tests fix(vector-store): carry request metadata into the Router executor built from the router kwarg 2026-09-02 16:53:05 -07:00
windows_tests
__init__.py
_fake_openai_endpoint_server.py test(timeout): time out against the local fake endpoint instead of api.openai.com 2026-09-03 09:53:30 -07:00
_flush_vcr_cache.py
_live_test_helpers.py
_openai_record_replay_proxy.py
_vcr_conftest_common.py
_vcr_redis_persister.py
_wait_helpers.py
_ws_vcr.py
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py test: address review notes on the chronic-test repairs 2026-09-04 10:18:15 -07:00
test_new_vector_store_endpoints.py
test_openai_endpoints.py
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_rust_python_harness.py test: address review notes on the chronic-test repairs 2026-09-04 10:18:15 -07:00
test_service_logger_otel.py
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.