litellm/tests
mateo-berri e87e5c1199 RALPH: compat matrix slice 4 - daily cron VM publishes matrix to docs (#26480, PRD #26476)
Slice 4 of the Claude Code Compatibility Matrix: stand up the daily-cron
pipeline that publishes `compatibility-matrix.json` to the docs repo. After
this slice lands, the hand-authored matrix in the docs repo is replaced by
auto-generated output, and the docs page begins reflecting real test runs
against the latest stable LiteLLM release.

What landed:

- tests/claude_code/resolver.py
  Latest Stable LiteLLM Resolver. Calls the GitHub Releases API and
  returns the newest tag matching `v*-stable`. Sort is numeric on
  (major, minor, patch) so v1.10.0-stable correctly outranks
  v1.9.5-stable. Injectable `http_get` so tests run offline.

- tests/claude_code/publisher.py
  Daily-cron orchestrator. Resolves the latest stable tag, pulls
  `ghcr.io/berriai/litellm:<tag>`, starts it as the proxy, installs
  `@anthropic-ai/claude-code@latest`, runs `pytest tests/claude_code/`,
  invokes the Matrix JSON Builder, and direct-pushes
  `compatibility-matrix.json` to the docs repo's main branch using a
  GitHub App installation token (`DOCS_REPO_TOKEN`). Idempotent: a no-op
  if the JSON is byte-identical to what's already on main.

- tests/claude_code/_publisher_unit_tests/test_resolver.py
  test_publisher.py
  14 unit tests covering the small pure helpers — version sort,
  non-stable filtering, http-get injection, commit message determinism,
  Docker image-name builder, and the file allowlist that enforces the
  "only `compatibility-matrix.json` ever ships" guarantee. Per the PRD's
  "Testing Decisions" section, the publisher's full subprocess
  orchestration intentionally ships without a unit-test harness; the
  daily-cron failure surface is itself the test.

- .github/workflows/claude_code_compat_matrix.yml
  GitHub Actions workflow with three triggers (daily cron at 06:00 UTC,
  `release: published` filtered to `*-stable` tags, and
  `workflow_dispatch`). Mints a docs-repo installation token from a
  GitHub App scoped to `BerriAI/litellm-docs` only with `contents:
  write`, then runs the publisher.

- .gitignore
  Add `compatibility-matrix.json` (cron VM output).

Key decisions:

- "Isolated VM" is realized as a GitHub-hosted ubuntu-latest runner —
  every run gets a fresh ephemeral VM, and the always-latest Claude
  Code CLI is only ever installed inside that ephemeral environment,
  so a malicious or broken Claude Code release cannot affect the
  trusted PR-gate CI in CircleCI.
- File-level restriction on the GitHub App's broad `contents: write`
  scope is enforced by `select_files_to_commit` (script correctness),
  per the PRD's explicit acknowledgement that GitHub does not support
  file-path-scoped tokens.
- `release` runs are filtered to tags ending in `-stable` at the
  workflow level, so a `v1.84.0-rc1` release does not republish the
  matrix.
- Resolver and publisher live under `tests/claude_code/` alongside
  `matrix_builder.py` and `cli_driver.py` — production code that
  supports the test suite, kept colocated with it to match the slice
  1+2 layout.

Out of scope / blockers for next iteration:

- Provisioning the GitHub App itself (creating it under BerriAI's
  org, installing it on litellm-docs only, generating the private key
  and registering `COMPAT_MATRIX_APP_ID` / `COMPAT_MATRIX_APP_PRIVATE_KEY`
  as repo secrets) is an operator/infra step that cannot land via a
  code change in this repo.
- The first successful cron run is what removes the hand-authored
  `compatibility-matrix.json` from the docs repo and replaces it with
  generated output — that happens after this PR merges and the App is
  installed; not a code change here.

Tests: 34 -> 45 passing (added 7 resolver tests + 7 publisher helper
tests, all unit-only and offline). The 12 per-cell failures under
`tests/claude_code/basic_messaging_non_streaming/` remain by design —
they require a running proxy which the cron VM provides.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-25 05:03:16 +00:00
..
agent_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
audio_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
basic_proxy_startup_tests build: migrate packaging, CI, and Docker from Poetry to uv (#25007) 2026-04-09 11:46:23 -07:00
batches_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
benchmarks Add CodSpeed performance benchmarks (#23676) 2026-03-14 18:44:36 -07:00
claude_code RALPH: compat matrix slice 4 - daily cron VM publishes matrix to docs (#26480, PRD #26476) 2026-04-25 05:03:16 +00:00
code_coverage_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
documentation_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
enterprise Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_migration_projects 2026-04-24 12:52:10 -07:00
guardrails_tests fix(proxy): guardrail header dedupe, mypy during_call, test mock kwargs 2026-04-22 23:22:35 +03:00
image_gen_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
litellm Merge branch 'litellm_internal_staging' into litellm_staging_03_22_2026 2026-04-20 19:56:00 +05:30
litellm-proxy-extras style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
litellm_core_utils Merge branch 'litellm_internal_staging' into litellm_staging_03_22_2026 2026-04-20 19:56:00 +05:30
litellm_utils_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
llm_responses_api_testing style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
llm_translation Merge pull request #26283 from BerriAI/litellm_internal_staging 2026-04-22 19:55:27 -03:00
load_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
local_testing Merge branch 'litellm_internal_staging' into litellm_staging_03_22_2026 2026-04-24 13:12:19 -03:00
logging_callback_tests feat(guardrails): LLM-as-a-Judge guardrail (#26360) 2026-04-24 17:15:32 -07:00
mcp_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_yj_apr17 2026-04-17 17:36:40 -07:00
multi_instance_e2e_tests
ocr_tests Merge branch 'litellm_internal_staging' into litellm_pages_support_for_ocr 2026-04-17 19:29:45 -07:00
old_proxy_tests/tests fix: cleanup tests 2026-03-30 16:24:35 -07:00
openai_endpoints_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
otel_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
pass_through_tests [Infra] CCI: pin Ruby and Node.js installs in proxy_pass_through_endpoint_tests 2026-04-22 21:28:42 -07:00
pass_through_unit_tests replace retired claude-3-haiku-20240307 with claude-haiku-4-5-20251001 in anthropic messages passthrough test 2026-04-20 15:44:11 -07:00
proxy_admin_ui_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
proxy_e2e_anthropic_messages_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
proxy_security_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
proxy_unit_tests Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_migration_projects 2026-04-24 12:52:10 -07:00
router_unit_tests Merge branch 'litellm_internal_staging' into litellm_adaptive_routing 2026-04-20 15:28:08 -07:00
scim_tests
search_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
spend_tracking_tests [Fix] Harden spend accuracy test against transient aiohttp connection errors 2026-04-22 17:40:42 -07:00
store_model_in_db_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_litellm feat(proxy): add /v1/memory CRUD endpoints (#26218) 2026-04-24 18:38:07 -07:00
unified_google_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
vector_store_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
windows_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
__init__.py
eval_swe_bench.py Prompt Compression - add it to the proxy (#25729) 2026-04-20 15:08:00 -07:00
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
README.MD
test_budget_management.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_callbacks_on_proxy.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_config.py
test_debug_warning.py
test_default_encoding_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_end_users.py
test_entrypoint.py
test_fallbacks.py Revert "fix: prevent error when max_fallbacks exceeds available models (#20071)" 2026-02-03 15:15:30 +05:30
test_gpt5_azure_temperature_support.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_health.py
test_keys.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_litellm_proxy_responses_config.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_logging.conf
test_models.py test: replace test_add_and_delete_models integration test with mock 2026-03-30 21:30:57 -07:00
test_new_vector_store_endpoints.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_openai_endpoints.py
test_organizations.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_otel_thread_leak.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_passthrough_endpoints.py
test_presidio_latency.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_proxy_server_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_ratelimit.py [Fix] test_ratelimit: skip over-limit cases that race with background RPM tracking 2026-04-11 13:08:59 -07:00
test_resource_cleanup.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_service_logger_otel.py fix(langfuse_otel): prevent empty proxy request spans from being sent to Langfuse 2026-01-28 15:35:35 +01:00
test_spend_logs.py Fix CI: Revert security scan changes and add GitGuardian ignore rules (#18358) 2025-12-22 17:03:53 -08:00
test_team.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_team_logging.py test: cleanup dead tests 2026-03-28 20:49:02 -07:00
test_team_members.py
test_users.py Litellm fix update bedrock models (#24947) 2026-04-01 19:22:54 -07:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.