Strip VCR wiring from the batches test conftest. Drops:
- import of `_vcr_conftest_common` helpers
- the `vcr_config` fixture, `pytest_recording_configure`,
`_vcr_outcome_gate`, `pytest_runtest_makereport`
- the `apply_vcr_auto_marker_to_items` call in
`pytest_collection_modifyitems`
- `VerboseReporterState` / its `pytest_configure` /
`pytest_runtest_logreport` hooks (purely VCR-verdict plumbing)
Why: every test in this directory creates ephemeral OpenAI / Bedrock /
vLLM resources whose IDs change per run (file-XXX, batch-XXX,
ft-XXX, ...). VCR's path/query/body matchers don't match across runs,
so `record_mode="new_episodes"` was silently passing through to the
live API and recording many new cassette entries every run. Cassette
bloat without replay benefit.
Behaviour after this change is identical to running the directory
without `CASSETTE_REDIS_URL` set: tests that have keys hit live APIs,
tests that don't continue to skip via their existing skipif markers.
Conftest now keeps only path setup and the session-scoped `event_loop`
fixture.
OpenAI announced gpt-3.5-turbo-0125 (and fine-tuning of gpt-3.5-turbo
in general) for shutdown on 2026-10-23, with the announcement landing
2026-04-22. The hard-fail date is ~5 months out, but timing fits the
recent uptick in this test flaking and OpenAI may already be running
the deprecated model's pipeline with deprioritized infra.
Bump to gpt-4o-mini-2024-07-18 — currently supported for fine-tuning,
no announced shutdown. Updates the live test plus the mocked test for
consistency. Belt-and-suspenders with the existing propagation-retry
helper.
Previous fix polled `litellm.afile_retrieve` for `status == "processed"`
before calling the fine-tuning endpoint. That doesn't actually solve
the race:
- OpenAI's `FileObject.status` field is deprecated per the SDK type and
not authoritative — it can read "processed" before the file is usable.
- The retrieve and fine-tuning endpoints don't share a consistency
model, so retrieve succeeding tells you nothing about FT visibility.
Replace with a retry around the actual `acreate_fine_tuning_job` call
that catches the OpenAI 400 `'file-... does not exist'` and backs off
exponentially (1s → cap 8s, 12 attempts, ~70s total budget). The
operation succeeding is the only reliable signal that propagation
finished.
OpenAI file uploads are eventually consistent — a freshly uploaded file
may briefly 404 from `retrieve` and is rejected by the fine-tuning
endpoint with `'file-... does not exist'` until processing finishes.
The async fine-tuning test called `acreate_fine_tuning_job` immediately
after `acreate_file` and flaked on this race.
Add a polling helper that waits up to ~30s for `status=processed` (and
short-circuits on `error`), called between upload and FT job creation.
Mirrors the same propagation lag covered by the `await asyncio.sleep(1)`
in the sister batches test, but more robust against longer delays.
OpenAI deprecated the gpt-4o-realtime-preview-2024-10-01 snapshot,
which caused these E2E tests to fail consistently in CI. Bump to the
unversioned gpt-4o-realtime-preview alias to match the sibling
test_openai_realtime_simple.py and stay current as OpenAI rolls the
alias forward.
- Add `_strip_image_b64_payloads` filter: rewrites `data[*].b64_json` in
image-gen responses to a 4-byte placeholder before the cassette is saved.
Image-edit and image-gen cassettes (193 MB / 184 MB / 104 MB / ...) will
shrink to <100 KB on next record. Tests assert response shape only, so
coverage is preserved.
- Add `_normalize_multipart_boundary` filter: replaces httpx's per-request
random multipart boundary with a fixed string in both Content-Type header
and body bytes. Audio-transcription / Whisper tests have been effectively
unmocked — every CI run hit live providers and was silently capped at
MAX_EPISODES_PER_CASSETTE=50. Both record and replay now see identical
bytes; the safe_body matcher works.
- Fix test_evals_api.py body poisoning: replace `int(time.time())` in eval
names with `hashlib.sha1(test_node_name)[:12]`, add a function-scoped
`managed_eval` fixture that creates and deletes the eval, and switch
`get_eval` / `update_eval` from `list_evals().data[0].id` (which made
the URL vary by run) to `managed_eval.id`. Net coverage gain: delete is
now actually exercised.
- Swap arxiv PDF URL in BaseOCRTest for the in-repo `dummy.pdf` (589 B)
served via sha-pinned jsdelivr.
- Swap etsystatic image URL in BaseLLMChatTest.test_image_url for the
in-repo LiteLLM logo (9.2 KB) served via the same jsdelivr pin.
- Add `tests/llm_translation/test_vcr_filters.py` with 14 unit tests
covering both new filters: replacement, idempotency, nesting, content-
length update, two-distinct-boundaries-converge-after-normalize, etc.
Cassettes recorded with the prior patterns will mismatch on the first CI
run after merge; recommend flushing the cassette Redis once (post-merge)
so re-records save under the new format from the start.
* feat(xai): add grok-4.3 and grok-4.3-latest to model_prices_and_context_window.json
xAI's docs page now lists grok-4.3 as the recommended chat / coding model:
"We strongly recommend all API callers use grok-4.3. It is the most
intelligent and fastest model we've built." (https://docs.x.ai/docs/models)
Pricing/specs sourced from xAI's published model metadata:
- input: $1.25 / 1M tokens (<=200k), $2.50 / 1M tokens (>200k)
- output: $2.50 / 1M tokens (<=200k), $5.00 / 1M tokens (>200k)
- cached: $0.20 / 1M tokens (<=200k), $0.40 / 1M tokens (>200k)
- context: 1,000,000 tokens
- capabilities: vision, reasoning, function calling, structured outputs,
prompt caching, web search
Adds two entries: `xai/grok-4.3` (canonical) and `xai/grok-4.3-latest` (alias),
mirroring the pattern used for the rest of the xAI/Grok-4 family.
* test(xai): add model_info test for grok-4.3 + sync backup cost map
- Mirror xai/grok-4.3 and xai/grok-4.3-latest entries into
litellm/model_prices_and_context_window_backup.json so the bundled
model cost map matches the canonical model_prices_and_context_window.json.
- Add tests/test_litellm/test_xai_grok_4_3_model_metadata.py covering
pricing tiers, capability flags, context window, provider routing,
and parity between the main and backup cost maps.
- Point 'source' at the live xAI models page (the per-model URL
https://docs.x.ai/docs/models/grok-4.3 currently 404s).
---------
Co-authored-by: ishaan-berri <155045088+ishaan-berri@users.noreply.github.com>
Co-authored-by: shin-watcher <shin-watcher@berri.ai>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* feat(xai): add grok-4.3 and grok-4.3-latest to model_prices_and_context_window.json
xAI's docs page now lists grok-4.3 as the recommended chat / coding model:
"We strongly recommend all API callers use grok-4.3. It is the most
intelligent and fastest model we've built." (https://docs.x.ai/docs/models)
Pricing/specs sourced from xAI's published model metadata:
- input: $1.25 / 1M tokens (<=200k), $2.50 / 1M tokens (>200k)
- output: $2.50 / 1M tokens (<=200k), $5.00 / 1M tokens (>200k)
- cached: $0.20 / 1M tokens (<=200k), $0.40 / 1M tokens (>200k)
- context: 1,000,000 tokens
- capabilities: vision, reasoning, function calling, structured outputs,
prompt caching, web search
Adds two entries: `xai/grok-4.3` (canonical) and `xai/grok-4.3-latest` (alias),
mirroring the pattern used for the rest of the xAI/Grok-4 family.
* test(xai): add model_info test for grok-4.3 + sync backup cost map
- Mirror xai/grok-4.3 and xai/grok-4.3-latest entries into
litellm/model_prices_and_context_window_backup.json so the bundled
model cost map matches the canonical model_prices_and_context_window.json.
- Add tests/test_litellm/test_xai_grok_4_3_model_metadata.py covering
pricing tiers, capability flags, context window, provider routing,
and parity between the main and backup cost maps.
- Point 'source' at the live xAI models page (the per-model URL
https://docs.x.ai/docs/models/grok-4.3 currently 404s).
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
---------
Co-authored-by: shin-watcher <shin-watcher@berri.ai>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
The get_project endpoint's team-membership check only iterated
team.members_with_roles, so users present in team.admins but not
members_with_roles were denied access. Restore the admin check to
match _check_user_permission_for_project.
Changing the default to True silently broke existing Prometheus/Grafana
scrape jobs that hit /metrics without an API key. Restore the prior
opt-in behavior so monitoring keeps working on upgrade.
These three cells were failing for reasons unrelated to LiteLLM
translation:
- vision: tests passed `--image <path>`, a flag that no longer exists
in Claude Code 2.x (image attachment is now via the Files API or via
`--input-format stream-json` with inline content blocks). Rewrite
the cells to feed an Anthropic-shaped user message containing both
text and a base64 `image` content block through stdin in stream-json
mode. Hermetic — no temp file or Files API upload needed.
- extended_thinking: tests set `MAX_THINKING_TOKENS=4096` as an env
var, which Claude Code 2.x ignores. Switch to `--effort max` (the
current CLI knob) and use a non-trivial prompt (3-gallon / 5-gallon
jug puzzle). With trivial arithmetic the modern Sonnet/Opus tiers
optimize away the thinking step and arrive without a thinking block,
which made the test silently false-fail.
- web_search: assertion looked for `server_tool_use` /
`web_search_tool_result` blocks, but Claude Code's `WebSearch` is
a *client-side* tool: the CLI executes the search itself and feeds
the result back as a regular `tool_result` block. The Anthropic
server-side `web_search_20250305` tool only fires when injected
into the request directly (which the CLI does not do). Update the
assertion to look for a `tool_use` block whose name is
`WebSearch` — that's the right signal that the proxy preserved
both the request-side tool definition and the response-side tool_use
block end-to-end.
Driver change required to support stream-json input + variadic flags:
- cli_driver: insert `--` before the prompt positional. Variadic
flags like `--allowed-tools <tools...>` (commander.js) greedily
consume every following token, so the prompt was being eaten as a
tool name and the CLI would error out with "Input must be provided
either through stdin or as a prompt argument when using --print".
- cli_driver: thread a `stdin_input` parameter through `run_claude`
and `run_claude_models_parallel` so the vision rewrite can pipe
stream-json events to the CLI on stdin.
Validated end-to-end against a live LiteLLM proxy: all three Anthropic
cells now pass on Haiku 4.5, Sonnet 4.6, and Opus 4.7. Driver unit
tests (120) still green.
* Refactor Bedrock response stream shape handling
- Introduced a module-level constant `BEDROCK_RESPONSE_STREAM_SHAPE` to cache the response stream shape, eliminating the need for per-instance caching in `BedrockEventStreamDecoderBase`.
- Updated relevant methods to utilize the new constant, improving performance by avoiding redundant loading of the shape.
- Added tests to ensure the shape is loaded correctly at import time and is consistent across different modules.
- Added a new mock server script for testing Bedrock pass-through functionality.
* Refactor response parsing for Bedrock and SageMaker
- Improved code readability by formatting the parsing method calls in `AWSEventStreamDecoder` for both Bedrock and SageMaker response stream shapes.
- Added blank lines for better separation of code blocks in `invoke_handler.py` and `common_utils.py` to enhance maintainability.
* Enhance error handling for Bedrock and SageMaker response stream shape loading
- Wrapped the loading logic in `_load_bedrock_response_stream_shape` and `_load_sagemaker_response_stream_shape` with try-except blocks to gracefully handle exceptions.
- Added logging to warn when the response stream shape cannot be pre-loaded, ensuring the module imports cleanly.
- Updated tests to verify that loading failures return `None` instead of propagating exceptions.
* Implement error handling for missing response stream shapes in Bedrock and SageMaker
- Added checks in `_parse_message_from_event` methods to raise appropriate errors when `BEDROCK_RESPONSE_STREAM_SHAPE` or `SAGEMAKER_RESPONSE_STREAM_SHAPE` is None, ensuring clearer error reporting.
- Updated logging messages to reflect the unavailability of event-stream decoding for both Bedrock and SageMaker.
- Enhanced unit tests to verify that the correct exceptions are raised when the response stream shapes are not loaded.
* [Chore] CI: Block PRs that drop overall code coverage
Tighten Codecov project status threshold from 1% to 0% so any drop in
overall project coverage relative to the base commit fails the
codecov/project check. target: auto keeps the bar floating with the
codebase, no manual maintenance needed as coverage moves up over time.
* [Chore] CI: Always post Codecov status regardless of CI outcome
Set codecov.require_ci_to_pass: false and codecov.notify.wait_for_ci:
false so Codecov posts the codecov/project and codecov/patch checks as
soon as the expected uploads arrive, instead of withholding them when
unrelated CI jobs fail. The coverage-regression check is independent
of test pass/fail, and CI failures are already enforced by their own
required-status checks.
The assert-shard-coverage guard in test-unit-proxy-db.yml failed because
test_request_size_limit_middleware.py was added under tests/proxy_unit_tests/
but not referenced by any matrix entry. Assigning it to the proxy-runtime
shard, which already covers other server-runtime tests (proxy_routes,
proxy_gunicorn, server_root_path).
Cursor security review flagged that run_claude() forwarded the entire
parent environment to the externally installed claude CLI binary. In
the PR gate flow the binary is dynamically installed from npm, and
the surrounding job loads every upstream provider credential
(ANTHROPIC_API_KEY, AWS_*, AZURE_FOUNDRY_*, VERTEXAI_CREDENTIALS,
GITHUB_TOKEN, ...) into its env so the proxy can route requests. A
compromised CLI release would have read access to all of them — even
though the CLI itself only ever talks to the proxy via the explicit
ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN we set.
Build the subprocess env from a small allowlist of process-runtime
vars (PATH, HOME, NVM_DIR, locale) rather than inheriting all of
os.environ. Caller-supplied extra_env still rides on top, which is
the sanctioned way for tests to opt-in to passing additional vars
(e.g. extended_thinking sets MAX_THINKING_TOKENS).
Add unit tests pinning the contract: PATH/HOME flow through, secrets
do not, and extra_env can still override anything.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Bugbot flagged that the awk extracting PINNED_UV_VERSION required the
value to start with a specifier (`==`, `>=`, etc.) and silently
fell through to the system uv if pyproject.toml later switched to a
bare `required-version = "0.10.9"` form. Strip any leading
specifier prefix from the quoted value rather than requiring one to
be present, so both prefixed and bare forms resolve to the same
download URL.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Cursor security review flagged that the cron VM proxy binds to
0.0.0.0 (litellm's default when --host is omitted) while running
under the predictable default `LITELLM_MASTER_KEY=sk-cron-matrix`.
On a VM where :PROXY_PORT is reachable, anything that can hit the
port can authenticate with the predictable fallback secret and burn
upstream provider credentials.
The populator proxy is only ever talked to by the same-host pytest
run (the health check and tests both use http://127.0.0.1:${PORT}),
so there's no reason for it to listen on external interfaces. Pass
`--host 127.0.0.1` explicitly.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
- Switch all per-cell tests from @pytest.mark.parametrize("model", ...)
(3 sequential invocations) to a single test that fans out to all 3
Claude tiers via run_claude_models_parallel. Per-cell wall time is now
bounded by the slowest model rather than the sum.
- Add 5 new v0 feature dirs (5 providers each, 25 new test files):
web_search, pdf_input, prompt_caching_1h,
tool_use_streaming, thinking_with_tool_use
Manifest expanded to match.
- Add cross-process token-bucket rate limiter (rate_limiter.py + tests)
so xdist workers stay under per-provider req/s limits during full-grid
runs. New env knobs: LITELLM_COMPAT_RATE_{ANTHROPIC,AZURE,VERTEX_AI,
BEDROCK_CONVERSE,BEDROCK_INVOKE}.
- conftest.py: write per-worker shards under <artifact>.shards/, merge
in the controller; preserve the "don't write empty artifact" guard so
unit-test runs don't clobber a real compat-results.json.
- Vertex test_config.yaml: route project/location through env so the
cron VM can target a different GCP project than the upstream default.
- Add run_compat.sh wrapper for binary-searching ideal req/s per
provider against compat-rate-limit-summary.json output.
Cursor security review flagged that the cron VM downloads the uv
release tarball and pipes it straight through `tar -xzO ... > file ;
chmod +x`, with no integrity check. A tampered release artifact would
execute in a credentialed cron context with access to the docs-repo
push token.
Download the tarball + the official .sha256 sidecar Astral publishes
alongside every uv release to a temp dir, run `sha256sum -c` against
the sidecar, and only extract+install on success. On mismatch we wipe
the tempdir and `die` with a clear refusal.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Bugbot flagged that the resolver fetches only the first page (default
30 entries) of the releases endpoint. LiteLLM ships multiple non-stable
releases per day, so 30+ non-stable releases between consecutive
v*-stable tags is routinely the case — when it happens, the jq filter
matches nothing, LITELLM_VERSION is empty, and the daily matrix update
silently dies.
Walk pages 1..5 (100 per page = 500 releases max) and short-circuit as
soon as a page contains at least one v*-stable tag. Same final jq
filter, same sort order.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Bugbot flagged that `run_claude` placed the prompt as cmd[7] and then
appended extra_args after it. `claude --print` takes the prompt as the
final positional argument; flags appearing after it (e.g.
`--allowed-tools Bash`, `--image <path>`) are swallowed by the prompt
parser, which silently breaks the tool_use and vision cells.
Build the flag list first, then append the prompt last. Add a unit
test that pins the ordering.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
The pytest_sessionfinish hook in tests/claude_code/conftest.py is loaded
for every pytest run under tests/claude_code/, including sibling unit-
test trees (_driver_unit_tests/, _builder_unit_tests/, ...). Without a
guard, those runs wrote an empty results artifact and silently
overwrote any real artifact from a prior compat-test run.
- Add pytest_sessionstart hook in tests/claude_code/conftest.py that
clears the module-level _COLLECTOR so results don't leak across
pytest.main() invocations within the same process.
- Export LITELLM_MASTER_KEY=${PROXY_API_KEY} when launching the cron
VM proxy so the auth token the tests send actually matches what
the proxy expects.
The Claude Code Compatibility Matrix PR gate booted the proxy with only
Anthropic + Bedrock credentials, so every Vertex AI and Azure cell (12
out of 30) failed at startup because test_config.yaml resolves
VERTEXAI_CREDENTIALS, AZURE_FOUNDRY_API_KEY, and AZURE_FOUNDRY_API_BASE
from the environment. Forward those three env vars from the CircleCI
context so the proxy has working routes for all five providers.
Also drop the leftover debug prints in
basic_messaging_non_streaming/test_anthropic.py that wrote the proxy
api_key to stdout (and from there to JUnit XML / CI logs).
The resolver picks the latest `v*-stable` tag of BerriAI/litellm. Until
the compat matrix work itself lands in a stable release, that tag's
tree won't contain `tests/claude_code/` at all — the proxy config and
test files only exist on the work-in-progress stack. Without a shim,
the cron dies with 'proxy config not found at .../test_config.yaml'
on every run.
Fix: after `git checkout <tag>`, if
`<worktree>/tests/claude_code/test_config.yaml` is missing, copy the
directory from ${LITELLM_REPO} (the dev checkout, which has the
in-flight matrix work). The `git clean -e tests/claude_code` line
preserves the shim across runs.
Once the matrix work is in the resolved tag, the if-branch is a no-op
and the shim is never used. No code path needs to be removed later;
the bridge self-disables.
Tier 1 (single anthropic cell) and tier 2 (full
basic_messaging_non_streaming row across 5 providers) now run cleanly
from `run_daily.sh` on the GCP VM. Five issues showed up during
validation; each is fixed in this commit.
1. uv version pin
----------------
The litellm worktree pins an exact uv version in
`pyproject.toml`'s `[tool.uv] required-version` field. The cron
VM's system uv (currently 0.11.8) refused to sync against the
v1.83.10-stable lockfile (which pins ==0.10.9). Fix: parse the
pinned version out of the worktree's pyproject.toml, download the
matching standalone binary into `<worktree>/.uv-bin/uv-<version>`,
and use it for sync/run/proxy. Cached across runs.
2. Missing proxy + dev extras
---------------------------
`uv sync --frozen` only installed the base dependency set, so
`uv run litellm` died at startup with
`ModuleNotFoundError: No module named 'websockets'`. Per the
repo's AGENTS.md the canonical incantation is
`uv sync --frozen --group proxy-dev --extra proxy`.
3. Wrong env vars for the test driver
-----------------------------------
The script was setting `ANTHROPIC_BASE_URL` and
`ANTHROPIC_AUTH_TOKEN`, which is what Claude Code itself reads,
but the test files read `LITELLM_PROXY_BASE_URL` and
`LITELLM_PROXY_API_KEY` (search the test_config-driven test files
for `PROXY_BASE_URL_ENV = "LITELLM_PROXY_BASE_URL"`). Tests were
marking themselves `fail` with
"missing required env: set LITELLM_PROXY_BASE_URL...". Fix: rename
the two env vars in the pytest invocation. The driver still
propagates them onward as ANTHROPIC_BASE_URL/AUTH_TOKEN to Claude.
4. Cleanup couldn't find the proxy
--------------------------------
The previous setup did
`( ... && setsid uv run litellm ... ) &; PROXY_PID=$!`. `setsid`
detaches the inner uv into its own session, but `$!` records the
PID of the outer subshell, not the long-lived python proxy. So
`kill -TERM "-${PROXY_PID}"` in the EXIT trap targeted the wrong
pgid and the proxy survived as an orphan whenever the script was
killed externally. Fix: replace the subshell with
`setsid bash -c '...'` that writes $$ to a known pid file before
exec'ing the proxy. The cleanup trap reads that file and uses it
as the pgid. Belt-and-braces: the trap also `pgrep -f`s by port
number and SIGKILLs survivors. Trap now fires on `INT TERM` too,
not just `EXIT`.
5. .uv-bin cache survives git clean
---------------------------------
The original `git clean -fdx -e .venv` wiped `.uv-bin/` between
runs, forcing re-download of the pinned uv binary on every
invocation. Now excluded.
Things that worked first try
----------------------------
* Worktree clone + checkout to the resolved tag.
* gh auth on this VM (mateo-berri account, collaborator on
BerriAI/litellm-docs).
* The matrix builder produced a well-formed
`compatibility-matrix.json` with the right per-cell aggregation
even when 4 of 5 cells failed (tier 2 was: anthropic=pass,
bedrock_invoke=fail, bedrock_converse=fail, vertex_ai=fail with a
real 403 from GCP for insufficient scopes, azure=fail with
timeout).
The previous iteration of this PR ported the populator to a Python
module (`publisher.py`) with 8 unit-tested pure helpers for branch
naming and PR-body rendering. After the docker code came out, the
GitHub App auth came out, and the worktree-vs-tempdir decision was
made, what was left was: 'git fetch + git checkout + uv sync + start
a subprocess + run pytest + clone docs + git commit + gh pr create'.
That's a bash script.
This commit replaces 800 lines of Python (publisher.py + resolver.py +
their unit tests) with a 297-line run_daily.sh and a 48-line
build_matrix.py whose only job is to be importable Python that can
call into the matrix_builder we already have. Net deletion: -822 lines.
Removed
-------
* tests/claude_code/publisher.py — the full Python orchestrator.
Every code path it had is now in run_daily.sh.
* tests/claude_code/resolver.py — the GitHub Releases v*-stable
resolver. Replaced by ~10 lines of jq inside run_daily.sh.
* tests/claude_code/_publisher_unit_tests/ — both test files. The
pure helpers they covered (commit message, file allowlist, branch
name, PR title/body) only existed because publisher.py was Python.
The bash equivalents are short heredoc strings.
Added
-----
* tests/claude_code/cron_vm/run_daily.sh — the actual cron job,
structured as numbered phases (resolve / worktree / proxy /
pytest / build / publish) so journalctl output is readable.
* tests/claude_code/cron_vm/build_matrix.py — a 48-line CLI that
calls the existing matrix_builder.build_from_paths. Kept in
Python because the builder itself is Python and well-tested.
Modified
--------
* tests/claude_code/cron_vm/litellm-compat-matrix.service —
ExecStart now invokes run_daily.sh instead of
'python -m tests.claude_code.publisher'.
* tests/claude_code/cron_vm/README.md — updated layout table,
file roles, and operating commands to match.
Why this is the right shape
---------------------------
* The failure mode at 06:00 UTC is 'read journalctl, see the literal
failing command with its + prefix, copy-paste it into a shell to
reproduce'. Bash makes that immediate; Python's subprocess.run
output looks similar but the surrounding orchestration is harder
to step through interactively.
* Every operation the script does is already a shell command (git,
uv, gh, jq, curl, pytest). The Python wrapper was translating
between argv arrays and back.
* The two pieces that genuinely benefit from being in a typed
language are matrix_builder (already Python) and the resolver's
semver sort (now done in jq, with the version_key tuple sort
inline). 'Already Python' wins, 'tiny jq pipeline' wins.
What's preserved
----------------
* Idempotency: same (litellm, claude, UTC date) -> same branch ->
same PR. force-with-lease push, gh-pr-create no-op-on-exists.
* Byte-identical-JSON early return (git diff --cached --quiet).
* Per-feature status table in the PR body (jq pipeline mirroring
the Python pr_body_for_matrix logic).
* Persistent worktree approach so disk doesn't grow unboundedly.
* Proxy bound to :4100 to avoid colliding with a developer's :4000.
* SKIP_PUBLISH=1 and PYTEST_K=... operator escape hatches.
Operator escape hatch for first-time validation after a Claude Code CLI
upgrade or a proxy-config change. Setting `PYTEST_K=anthropic and
basic_messaging_non_streaming` (for example) narrows the matrix run to
one cell, which is enough to confirm the worktree+uv+proxy+pytest
plumbing without burning a full $2 in provider tokens.
Cells the narrowed run doesn't touch are filled with `not_tested` by
the matrix builder, so a PYTEST_K-narrowed run is structurally safe to
publish — though in practice operators pair it with `--skip-publish`
during validation.
The daily compat-matrix runs from the dedicated GCP VM
`litellm-compatibility-matrix-populator` rather than from a GitHub
Actions runner. The VM has no docker daemon, has `gh` already
authenticated against an account with `pull-requests: write` on
`BerriAI/litellm-docs`, and runs a long-lived litellm checkout we can
reuse across runs. That makes Docker, the GitHub App auth flow, and the
GHA workflow itself dead code.
Removed
-------
* `.github/workflows/claude_code_compat_matrix.yml` — no longer
triggers anything; the systemd timer in this PR owns the daily fire.
* `docker_image_for_tag` + `DOCKER_IMAGE_BASE` constants and the two
unit tests that covered them.
* `_start_proxy(image, port)` / `_stop_proxy(container_id)` /
`docker run` flow, replaced by direct `uv run litellm` subprocess
management with a sigterm-the-process-group teardown.
* `--skip-proxy` CLI flag (was only useful when the GHA workflow
split docker-bringup from publish into separate jobs).
* `docs_token` parameter and `DOCS_REPO_TOKEN` env var; `gh` on
the VM is already authenticated, so we don't pass an explicit
token through the publisher.
Added
-----
* Persistent worktree flow in `publisher.py`. First run clones
`BerriAI/litellm` into `~/litellm-cron-worktree/`; subsequent runs
do `git fetch --tags && git checkout --force <stable-tag> &&
uv sync --frozen`. Disk footprint is bounded because uv sync
removes packages no longer pinned and `git clean -fdx -e .venv`
wipes per-run cruft while keeping the venv around.
* `tests/claude_code/cron_vm/` containing systemd units and a
setup README:
- `litellm-compat-matrix.service` (`Type=oneshot`, runs as the
`mateo` user, sources `/etc/litellm-compat-matrix.env` for
provider creds, hardened with `NoNewPrivileges` /
`ProtectSystem=strict` / `PrivateTmp`);
- `litellm-compat-matrix.timer` (`OnCalendar=*-*-* 06:00:00 UTC`,
`Persistent=true` so a missed run fires when the VM is back up,
`RandomizedDelaySec=10min`);
- `.env.example` documenting the provider-credential surface;
- `README.md` covering one-time install, daily operation,
`journalctl` debugging, and the gotchas (proxy port `4100` to
avoid colliding with a developer's `:4000`, `gh` token
rotation, what to do after a Claude Code CLI upgrade).
Operator notes
--------------
* The proxy now binds `:4100` by default so a developer SSH'd into
the VM with their own `:4000` proxy isn't preempted by the cron.
* The Claude Code CLI is exercised as-is from the system install;
the populator does NOT `npm install` it. Operators upgrade the
CLI by running `npm install -g @anthropic-ai/claude-code@latest`
out of band, typically after watching a `--skip-publish` run to
verify the matrix doesn't suddenly turn red.
* 20 publisher unit tests pass (`pytest
tests/claude_code/_publisher_unit_tests/`).
* End-to-end validation on the VM happens after this PR lands as
follow-up commits on the same branch — the systemd unit is
`Type=oneshot` so a manual `systemctl start` reproduces the cron.
The daily Claude Code compatibility-matrix cron has been direct-pushing
`compatibility-matrix.json` to litellm-docs's main branch. Switch to
opening (or updating) a pull request so docs maintainers can review each
matrix update before it ships to readers.
Behavioural changes
-------------------
publisher.publish() now:
* checks out a deterministic head branch
(`compat-matrix/<litellm>-<claude>-<UTC-date>`) before staging the
JSON, instead of committing on top of the docs branch directly;
* `git push --force-with-lease` so a same-day rerun updates the
existing branch (and therefore the existing PR), without
overwriting any docs-maintainer fixup commit on the same branch;
* shells out to `gh pr create` against `docs_repo` with a
title/body that surfaces the resolved versions and a per-feature
status summary, so reviewers can triage from the inbox;
* treats 'a pull request for branch ... already exists' as success,
so two cron runs on the same day produce one PR, not two.
Idempotency contract
--------------------
* Same (litellm_version, claude_code_version, UTC date) -> same
branch -> same PR. Verified by the new
`test_pr_branch_name_is_deterministic_per_inputs` /
`...changes_when_any_component_changes` tests.
* Byte-identical JSON to the docs branch -> early-return before
push, same as the previous direct-push path.
* Empty version inputs are rejected up front so two distinct PRs
can never silently collapse onto one branch.
Tests
-----
* 8 new tests in `_publisher_unit_tests/test_publisher.py` cover
`pr_branch_name`, `pr_title_for_matrix`, and `pr_body_for_matrix`
(determinism, content, ordering, missing-provider rectangularity,
empty-input rejection).
* Existing 7 `commit_message_for_matrix` /
`docker_image_for_tag` / `select_files_to_commit` tests are
unchanged and still pass.
Workflow
--------
`.github/workflows/claude_code_compat_matrix.yml` updates only the
header doc comment to reflect that the GitHub App now needs
`pull-requests: write` in addition to `contents: write`. `gh` is
preinstalled on `ubuntu-latest` (also used by
`auto_update_price_and_context_window.yml`), so no install step is
needed.
Operator action required (one-time)
-----------------------------------
The compat-matrix GitHub App installation on `BerriAI/litellm-docs`
needs `pull-requests: write` added to its installation permissions
before the next cron run. Without it, the new `gh pr create` call
will fail with a 403; `compat-results.json` and
`compatibility-matrix.json` will still upload as workflow artifacts
for debugging.
Anthropic and Microsoft announced Claude Haiku 4.5, Sonnet 4.5/4.6, and
Opus 4.1/4.6/4.7 in Microsoft Foundry on 2025-11-18, so the matrix's
Azure column should exercise a real route through the LiteLLM proxy
rather than report not_applicable.
Foundry serves Claude on an Anthropic-shape /anthropic/v1/messages
endpoint (not the Azure OpenAI chat-completions route), and LiteLLM
already supports it via the azure_ai/claude-* provider prefix
(litellm/llms/azure_ai/anthropic/{handler,transformation,messages_transformation}.py).
- test_config.yaml: add 3 azure aliases pointing at azure_ai/claude-*
with AZURE_FOUNDRY_API_BASE / AZURE_FOUNDRY_API_KEY env
- 6x test_azure.py: replace not_applicable stubs with real run_claude
drivers, mirroring the existing test_vertex_ai.py shape exactly
- sample_compatibility-matrix.json: Azure cells flip to pass
- _builder_unit_tests: pin the new invariant (run_claude is used,
not_applicable is gone) and feed pass results across all 5 providers
in the 6x5 golden test
Slice 5 of the Claude Code Compatibility Matrix: extend the published
matrix from the 1x5 grid that landed in slice 2 to the full v0 6x5
grid described in the PRD's "Features in v0" section. After this
slice merges and the daily cron runs, the docs page reflects all six
v0 features against all five providers.
What landed:
- tests/claude_code/manifest.yaml
Five new entries appended in PRD row order:
basic_messaging_streaming, tool_use, prompt_caching_5m, vision,
extended_thinking. The manifest is the row-order source of truth
the matrix builder respects.
- tests/claude_code/<feature>/test_<provider>.py (25 new files)
For each of the five new features, five per-provider test files
modeled on slice 2's basic_messaging_non_streaming/. Each non-Azure
file parametrizes over Haiku 4.5 / Sonnet 4.6 / Opus 4.7 and drives
the real `claude` CLI through the driver with feature-specific
options:
* basic_messaging_streaming — count-1-to-5 prompt; asserts the
stream-json wire actually emitted events plus a non-empty reply.
* tool_use — `--allowed-tools Bash` plus an `echo pong` prompt;
asserts a `tool_use` content block was emitted.
* prompt_caching_5m — same baseline prompt as non-streaming, but
asserts the upstream usage block reports
cache_creation_input_tokens or cache_read_input_tokens > 0
(Claude Code stamps cache_control on its system prompt by default,
so a single live call surfaces it).
* vision — decodes a checked-in 1x1 PNG (base64 const) into
`tmp_path` and attaches it via `--image`; asserts a non-empty
reply.
* extended_thinking — sets `MAX_THINKING_TOKENS=4096`; asserts a
`thinking` content block was emitted.
All five Azure files report `not_applicable` with the standard
reason: Azure OpenAI Service does not host Anthropic models.
- tests/claude_code/sample_compatibility-matrix.json
Hand-authored 6x5 sample showing the realistic best-case outcome:
4 pass + 1 not_applicable (Azure) per row.
- tests/claude_code/_builder_unit_tests/test_v0_layout.py
New structural unit tests pinning the on-disk shape so future edits
can't silently flip the matrix shape:
* manifest lists all six v0 feature ids in PRD order
* manifest lists all five v0 provider columns in PRD order
* every (feature, provider) has a test file at the inferred path
* every test file references all three required Claude tiers
* every Azure test file is a `not_applicable` declaration
- tests/claude_code/_builder_unit_tests/test_matrix_builder.py
Renamed the slice-2 1x5 golden test to
test_build_matrix_6x5_grid_matches_published_sample and rebuilt
its inputs to feed all six features. The golden file is now the
6x5 sample.
Key decisions:
- Per-feature per-provider test bodies are deliberately duplicated
(per the PRD: "Duplication across per-provider files is accepted").
Each file is self-contained so a contributor touching one cell
doesn't accidentally regress neighbors.
- Only Azure cells are marked `not_applicable` in this slice. Other
combinations that turn out to genuinely not apply on the live cron
run (e.g. a provider that doesn't support `thinking` for a tier)
will be tightened to `not_applicable` reasons in a follow-up; for
now they fail honestly, which the matrix renderer paints red.
- prompt_caching_5m's assertion (cache tokens > 0 in the usage block)
exercises the path Claude Code customers care about: that the proxy
preserves `cache_control` annotations end-to-end. It does not try
to differentiate cache_creation vs cache_read across runs.
- The vision PNG fixture is generated at test time from a base64
const rather than checked into git as a binary — keeps the diff
text-only and avoids needing PIL or any image-generation library.
Tests: 31 -> 142 unit tests passing (no proxy / no `claude` CLI
required). Test counts:
* 12 builder tests (was 11; +1 for 6x5 golden, the slice-2 1x5
test was renamed in place)
* 100 v0_layout structural tests (new)
* 10 driver tests (unchanged)
* 9 compat_result tests (unchanged)
* 14 publisher unit tests (unchanged)
* 8 PR-gate version-resolver tests (unchanged)
* 6 CircleCI structural tests (unchanged)
The 90 per-cell tests under tests/claude_code/<feature>/ continue to
require a running proxy + `claude` CLI; they only run inside the
CircleCI PR gate or the daily-cron VM (both established in slices
3 and 4).
Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The companion update to compatibility-matrix.json in the docs repo.
After slice 4's daily-cron lands the App credentials, the cron run
will replace the docs-side hand-authored JSON automatically; until
then the slice-2 1x5 sample remains in the docs repo.
Notes for next iteration:
- The exact `claude` CLI flags for tool-allowlist (`--allowed-tools`),
vision (`--image`), and extended thinking (`MAX_THINKING_TOKENS`)
are best-guess from the current Claude Code surface; if the live
PR-gate run reveals different flag names, tighten in place.
- Several non-Azure cells will likely need `not_applicable`
declarations once the cron VM produces real outcomes (e.g.
Bedrock Invoke + extended_thinking is uncertain). That refinement
is an iteration-2 follow-up driven by data, not a blocker for this
slice.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Slice 4 of the Claude Code Compatibility Matrix: stand up the daily-cron
pipeline that publishes `compatibility-matrix.json` to the docs repo. After
this slice lands, the hand-authored matrix in the docs repo is replaced by
auto-generated output, and the docs page begins reflecting real test runs
against the latest stable LiteLLM release.
What landed:
- tests/claude_code/resolver.py
Latest Stable LiteLLM Resolver. Calls the GitHub Releases API and
returns the newest tag matching `v*-stable`. Sort is numeric on
(major, minor, patch) so v1.10.0-stable correctly outranks
v1.9.5-stable. Injectable `http_get` so tests run offline.
- tests/claude_code/publisher.py
Daily-cron orchestrator. Resolves the latest stable tag, pulls
`ghcr.io/berriai/litellm:<tag>`, starts it as the proxy, installs
`@anthropic-ai/claude-code@latest`, runs `pytest tests/claude_code/`,
invokes the Matrix JSON Builder, and direct-pushes
`compatibility-matrix.json` to the docs repo's main branch using a
GitHub App installation token (`DOCS_REPO_TOKEN`). Idempotent: a no-op
if the JSON is byte-identical to what's already on main.
- tests/claude_code/_publisher_unit_tests/test_resolver.py
test_publisher.py
14 unit tests covering the small pure helpers — version sort,
non-stable filtering, http-get injection, commit message determinism,
Docker image-name builder, and the file allowlist that enforces the
"only `compatibility-matrix.json` ever ships" guarantee. Per the PRD's
"Testing Decisions" section, the publisher's full subprocess
orchestration intentionally ships without a unit-test harness; the
daily-cron failure surface is itself the test.
- .github/workflows/claude_code_compat_matrix.yml
GitHub Actions workflow with three triggers (daily cron at 06:00 UTC,
`release: published` filtered to `*-stable` tags, and
`workflow_dispatch`). Mints a docs-repo installation token from a
GitHub App scoped to `BerriAI/litellm-docs` only with `contents:
write`, then runs the publisher.
- .gitignore
Add `compatibility-matrix.json` (cron VM output).
Key decisions:
- "Isolated VM" is realized as a GitHub-hosted ubuntu-latest runner —
every run gets a fresh ephemeral VM, and the always-latest Claude
Code CLI is only ever installed inside that ephemeral environment,
so a malicious or broken Claude Code release cannot affect the
trusted PR-gate CI in CircleCI.
- File-level restriction on the GitHub App's broad `contents: write`
scope is enforced by `select_files_to_commit` (script correctness),
per the PRD's explicit acknowledgement that GitHub does not support
file-path-scoped tokens.
- `release` runs are filtered to tags ending in `-stable` at the
workflow level, so a `v1.84.0-rc1` release does not republish the
matrix.
- Resolver and publisher live under `tests/claude_code/` alongside
`matrix_builder.py` and `cli_driver.py` — production code that
supports the test suite, kept colocated with it to match the slice
1+2 layout.
Out of scope / blockers for next iteration:
- Provisioning the GitHub App itself (creating it under BerriAI's
org, installing it on litellm-docs only, generating the private key
and registering `COMPAT_MATRIX_APP_ID` / `COMPAT_MATRIX_APP_PRIVATE_KEY`
as repo secrets) is an operator/infra step that cannot land via a
code change in this repo.
- The first successful cron run is what removes the hand-authored
`compatibility-matrix.json` from the docs repo and replaces it with
generated output — that happens after this PR merges and the App is
installed; not a code change here.
Tests: 34 -> 45 passing (added 7 resolver tests + 7 publisher helper
tests, all unit-only and offline). The 12 per-cell failures under
`tests/claude_code/basic_messaging_non_streaming/` remain by design —
they require a running proxy which the cron VM provides.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Slice 3 of the Claude Code Compatibility Matrix: wire the
`tests/claude_code/` suite into CircleCI as a pre-merge gate. A red
status on the new `claude_code_compat_pr_gate` job blocks merge into
the staging branch.
What landed:
- tests/claude_code/pr_gate_version_resolver.py
The Claude Code PR-Gate Version Resolver described in the PRD's
"Version resolvers" section. Queries the npm registry for
`@anthropic-ai/claude-code` and returns the newest version whose
publish timestamp is at least 3 days old. The 3-day window is a
security review buffer: a malicious or broken Claude Code release
has at least 72 hours to be detected before it can land in our PR
gate. Importable function (with `metadata=` / `fetcher=` / `as_of=`
injection seams for tests) and a `python -m ...` CLI for the CI step.
- tests/claude_code/test_config.yaml
Proxy routing config that maps the per-cell aliases the tests use
(`claude-haiku-4-5`, `claude-haiku-4-5-bedrock-invoke`, ...,
`claude-opus-4-7-vertex`) to real upstream model ids on Anthropic /
Bedrock (Invoke + Converse) / Vertex AI. Azure intentionally has no
entries here because every Azure × claude-code cell is
`not_applicable` (Azure OpenAI doesn't host Claude).
- .circleci/config.yml
New `claude_code_compat_pr_gate` job. Pattern modeled on
`proxy_e2e_anthropic_messages_tests` (load PR-built docker image,
start postgres, mount config.yaml). New step in the middle:
resolve the Claude Code version from the resolver, install Node 20
via the machine image's preinstalled nvm, and `npm install -g
@anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}` (pinned, never
`latest`). Wired into `workflows.build_and_test` with a
`requires: [build_docker_database_image]` gate and the same
`*main_branches` filter the other proxy e2e job uses.
- tests/claude_code/_pr_gate_unit_tests/
16 new unit tests:
* 8 against the version resolver: boundary (>= 3d inclusive),
empty / all-too-new metadata, semver-vs-publish-time tiebreak,
custom min_age, fetcher injection, npm `time.created` /
`time.modified` skipping.
* 8 structural tests against `.circleci/config.yml`: job exists,
is in the workflow, requires the docker image, invokes the
resolver, install command is pinned (rejects unpinned `latest`),
runs `tests/claude_code/`, mounts `test_config.yaml`, exports
`LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY`. Plus one
regression test: the existing `proxy_e2e_anthropic_messages_tests`
job is unchanged in shape (acceptance criterion).
Key decisions:
- "Newest version" in the resolver is by **publish time**, not by
semver string ordering — if a patch lands on an older major after a
newer release, the patched line is the eligible one. (Tested.)
- The resolver's CLI prints the announcement to stderr and the bare
version to stdout, so the CI step can do
`CLAUDE_CODE_VERSION=$(uv run python -m ...)` cleanly while still
surfacing the selected version in the job log (acceptance criterion:
"the selected Claude Code version is logged").
- The structural CircleCI tests live under `_pr_gate_unit_tests/` so
the conftest path-inference hook skips them (the leading underscore
is the existing convention from `_driver_unit_tests/` /
`_builder_unit_tests/`); they don't pollute the matrix artifact.
- No `--no-verify` style supply-chain safety relaxation. Per the PRD,
Claude Code's pinning is the 3-day publish-age window, not a fixed
hash — by design, since the daily cron also pulls newer versions.
Tests: 47 -> 47 passing for the unit suite (16 new + 31 from slices
1 and 2). The end-to-end cells under `basic_messaging_non_streaming/`
require `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` and a
running proxy + `claude` CLI; they only run inside the new CircleCI
job.
Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- No docs PR is needed for this slice — the gate produces a status
check, not a published artifact. The compat matrix JSON the docs
page consumes is published by the daily-cron job (a future slice),
not by the PR gate.
Notes for next iteration:
- The daily cron / matrix publisher is the next slice. Several
pieces this slice introduces (the `tests/claude_code/test_config.yaml`
proxy config, the structure of the compat-results.json artifact)
will be reused by it.
- The bedrock-converse / vertex_ai aliases in `test_config.yaml` use
best-guess upstream model ids (`us.anthropic.claude-{tier}` and
`vertex_ai/claude-{tier}`); the real ids may need to be tightened
once the gate runs against live AWS / GCP credentials and we see
what resolves.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>