litellm/tests
mateo-berri 415cf3d6b1 RALPH: compat matrix slice 2 - add 4 provider columns for basic_messaging_non_streaming (#26478, PRD #26476)
Slice 2 of the Claude Code Compatibility Matrix: extend the tracer-bullet
cell from slice 1 across all four remaining provider columns for
basic_messaging_non_streaming. Proves the multi-provider, multi-model,
all-must-pass aggregation logic against a 1x5 grid that exercises every
status state.

What landed:

- tests/claude_code/basic_messaging_non_streaming/test_bedrock_invoke.py
- tests/claude_code/basic_messaging_non_streaming/test_bedrock_converse.py
- tests/claude_code/basic_messaging_non_streaming/test_vertex_ai.py
  Per-provider files modeled on test_anthropic.py: each parametrizes
  over Haiku 4.5 / Sonnet 4.6 / Opus 4.7 (the three Claude tiers
  required by the PRD), drives the real `claude` CLI through the
  driver, and reports pass/fail via `compat_result`. Per-cell error
  strings always include `[<model>]` so the docs tooltip can name the
  failing model when a cell goes red.

- tests/claude_code/basic_messaging_non_streaming/test_azure.py
  All three (Azure, Claude) cells report `not_applicable` with a
  reason: Azure OpenAI Service does not host Anthropic models. The
  test still parametrizes over the same three model ids so the test
  count per cell is uniform across columns, and a future "Azure adds
  Anthropic" announcement only requires flipping the body, not the
  parametrization.

- tests/claude_code/sample_compatibility-matrix.json
  Hand-authored 1x5 sample updated to reflect the slice 2 outcome:
  anthropic / bedrock_invoke / bedrock_converse / vertex_ai = pass,
  azure = not_applicable.

- tests/claude_code/_builder_unit_tests/test_matrix_builder.py
  Two new golden-file tests:
  1. 1x5 grid: feed the per-model results the four new test files
     produce on a real run; assert the builder output equals the
     hand-authored sample byte-for-byte.
  2. fail-with-model-named: feed pass/fail/pass for one cell and assert
     the cell aggregates to fail with the failing model id surfaced
     in the error string (acceptance criterion: "the error string
     identifies which model broke").

Key decisions:

- Duplication across the four per-provider files is accepted (per the
  PRD) rather than extracted into a helper. Each file is self-contained
  so a test author touching one provider doesn't accidentally regress
  the others.
- Per-provider model alias names: `claude-<tier>-<provider-suffix>`
  (e.g. `claude-haiku-4-5-bedrock-invoke`). These are the alias names
  the proxy operator wires up in the routing config; the test only
  knows the alias, the proxy knows the upstream model id and region.
- Azure is `not_applicable` rather than `not_tested` because the
  cell will never apply, not "we haven't gotten to it yet" - the two
  states are visually and semantically distinct in the rendered grid.
- Sample shows the realistic best-case outcome (4 pass + 1 NA). The
  React renderer's coverage of the `fail` and `not_tested` states is
  exercised by other cells in v1+, not the v0 sample.

Tests: 31 -> 34 passing (added 2 builder golden tests + 3 Azure
not_applicable parametrizations that pass without env vars).

Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The companion update to compatibility-matrix.json in the docs repo.
  The hand-authored sample in this repo is the artifact the docs PR
  copies; opening that doc PR is the next step in this slice.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-06 23:27:05 +00:00
..
agent_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
audio_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
basic_proxy_startup_tests build: migrate packaging, CI, and Docker from Poetry to uv (#25007) 2026-04-09 11:46:23 -07:00
batches_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
benchmarks Add CodSpeed performance benchmarks (#23676) 2026-03-14 18:44:36 -07:00
claude_code RALPH: compat matrix slice 2 - add 4 provider columns for basic_messaging_non_streaming (#26478, PRD #26476) 2026-05-06 23:27:05 +00:00
code_coverage_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
documentation_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
enterprise [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
guardrails_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
image_gen_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
litellm [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
litellm-proxy-extras style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
litellm_core_utils Merge branch 'litellm_internal_staging' into litellm_staging_03_22_2026 2026-04-20 19:56:00 +05:30
litellm_utils_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
llm_responses_api_testing [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
llm_translation [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
load_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
local_testing [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
logging_callback_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
mcp_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_yj_apr17 2026-04-17 17:36:40 -07:00
multi_instance_e2e_tests
ocr_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
old_proxy_tests/tests fix: cleanup tests 2026-03-30 16:24:35 -07:00
openai_endpoints_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
otel_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
pass_through_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
pass_through_unit_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
proxy_admin_ui_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
proxy_e2e_anthropic_messages_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
proxy_security_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
proxy_unit_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
router_unit_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
scim_tests
search_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
spend_tracking_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
store_model_in_db_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_litellm [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
unified_google_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
vector_store_tests fix: drop milvus dbName and partitionNames from MILVUS_OPTIONAL_PARAMS 2026-04-30 11:51:32 -07:00
windows_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
__init__.py
_flush_vcr_cache.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
_vcr_conftest_common.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
_vcr_redis_persister.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
eval_swe_bench.py Prompt Compression - add it to the proxy (#25729) 2026-04-20 15:08:00 -07:00
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
README.MD
test_budget_management.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_callbacks_on_proxy.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_config.py
test_debug_warning.py
test_default_encoding_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_end_users.py
test_entrypoint.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_health.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
test_keys.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_litellm_proxy_responses_config.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_logging.conf
test_models.py test: replace test_add_and_delete_models integration test with mock 2026-03-30 21:30:57 -07:00
test_new_vector_store_endpoints.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_openai_endpoints.py
test_organizations.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_otel_thread_leak.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_passthrough_endpoints.py
test_presidio_latency.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_proxy_server_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_ratelimit.py [Fix] test_ratelimit: skip over-limit cases that race with background RPM tracking 2026-04-11 13:08:59 -07:00
test_resource_cleanup.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_service_logger_otel.py
test_spend_logs.py
test_team.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_team_logging.py test: cleanup dead tests 2026-03-28 20:49:02 -07:00
test_team_members.py
test_users.py Litellm fix update bedrock models (#24947) 2026-04-01 19:22:54 -07:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.