mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-17 23:51:30 +00:00
7 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c1955a15ca |
RALPH: fix compat matrix - Azure now hosts Claude via Microsoft Foundry
Anthropic and Microsoft announced Claude Haiku 4.5, Sonnet 4.5/4.6, and
Opus 4.1/4.6/4.7 in Microsoft Foundry on 2025-11-18, so the matrix's
Azure column should exercise a real route through the LiteLLM proxy
rather than report not_applicable.
Foundry serves Claude on an Anthropic-shape /anthropic/v1/messages
endpoint (not the Azure OpenAI chat-completions route), and LiteLLM
already supports it via the azure_ai/claude-* provider prefix
(litellm/llms/azure_ai/anthropic/{handler,transformation,messages_transformation}.py).
- test_config.yaml: add 3 azure aliases pointing at azure_ai/claude-*
with AZURE_FOUNDRY_API_BASE / AZURE_FOUNDRY_API_KEY env
- 6x test_azure.py: replace not_applicable stubs with real run_claude
drivers, mirroring the existing test_vertex_ai.py shape exactly
- sample_compatibility-matrix.json: Azure cells flip to pass
- _builder_unit_tests: pin the new invariant (run_claude is used,
not_applicable is gone) and feed pass results across all 5 providers
in the 6x5 golden test
|
||
|
|
28cdbb4485 |
RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476)
Slice 5 of the Claude Code Compatibility Matrix: extend the published
matrix from the 1x5 grid that landed in slice 2 to the full v0 6x5
grid described in the PRD's "Features in v0" section. After this
slice merges and the daily cron runs, the docs page reflects all six
v0 features against all five providers.
What landed:
- tests/claude_code/manifest.yaml
Five new entries appended in PRD row order:
basic_messaging_streaming, tool_use, prompt_caching_5m, vision,
extended_thinking. The manifest is the row-order source of truth
the matrix builder respects.
- tests/claude_code/<feature>/test_<provider>.py (25 new files)
For each of the five new features, five per-provider test files
modeled on slice 2's basic_messaging_non_streaming/. Each non-Azure
file parametrizes over Haiku 4.5 / Sonnet 4.6 / Opus 4.7 and drives
the real `claude` CLI through the driver with feature-specific
options:
* basic_messaging_streaming — count-1-to-5 prompt; asserts the
stream-json wire actually emitted events plus a non-empty reply.
* tool_use — `--allowed-tools Bash` plus an `echo pong` prompt;
asserts a `tool_use` content block was emitted.
* prompt_caching_5m — same baseline prompt as non-streaming, but
asserts the upstream usage block reports
cache_creation_input_tokens or cache_read_input_tokens > 0
(Claude Code stamps cache_control on its system prompt by default,
so a single live call surfaces it).
* vision — decodes a checked-in 1x1 PNG (base64 const) into
`tmp_path` and attaches it via `--image`; asserts a non-empty
reply.
* extended_thinking — sets `MAX_THINKING_TOKENS=4096`; asserts a
`thinking` content block was emitted.
All five Azure files report `not_applicable` with the standard
reason: Azure OpenAI Service does not host Anthropic models.
- tests/claude_code/sample_compatibility-matrix.json
Hand-authored 6x5 sample showing the realistic best-case outcome:
4 pass + 1 not_applicable (Azure) per row.
- tests/claude_code/_builder_unit_tests/test_v0_layout.py
New structural unit tests pinning the on-disk shape so future edits
can't silently flip the matrix shape:
* manifest lists all six v0 feature ids in PRD order
* manifest lists all five v0 provider columns in PRD order
* every (feature, provider) has a test file at the inferred path
* every test file references all three required Claude tiers
* every Azure test file is a `not_applicable` declaration
- tests/claude_code/_builder_unit_tests/test_matrix_builder.py
Renamed the slice-2 1x5 golden test to
test_build_matrix_6x5_grid_matches_published_sample and rebuilt
its inputs to feed all six features. The golden file is now the
6x5 sample.
Key decisions:
- Per-feature per-provider test bodies are deliberately duplicated
(per the PRD: "Duplication across per-provider files is accepted").
Each file is self-contained so a contributor touching one cell
doesn't accidentally regress neighbors.
- Only Azure cells are marked `not_applicable` in this slice. Other
combinations that turn out to genuinely not apply on the live cron
run (e.g. a provider that doesn't support `thinking` for a tier)
will be tightened to `not_applicable` reasons in a follow-up; for
now they fail honestly, which the matrix renderer paints red.
- prompt_caching_5m's assertion (cache tokens > 0 in the usage block)
exercises the path Claude Code customers care about: that the proxy
preserves `cache_control` annotations end-to-end. It does not try
to differentiate cache_creation vs cache_read across runs.
- The vision PNG fixture is generated at test time from a base64
const rather than checked into git as a binary — keeps the diff
text-only and avoids needing PIL or any image-generation library.
Tests: 31 -> 142 unit tests passing (no proxy / no `claude` CLI
required). Test counts:
* 12 builder tests (was 11; +1 for 6x5 golden, the slice-2 1x5
test was renamed in place)
* 100 v0_layout structural tests (new)
* 10 driver tests (unchanged)
* 9 compat_result tests (unchanged)
* 14 publisher unit tests (unchanged)
* 8 PR-gate version-resolver tests (unchanged)
* 6 CircleCI structural tests (unchanged)
The 90 per-cell tests under tests/claude_code/<feature>/ continue to
require a running proxy + `claude` CLI; they only run inside the
CircleCI PR gate or the daily-cron VM (both established in slices
3 and 4).
Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The companion update to compatibility-matrix.json in the docs repo.
After slice 4's daily-cron lands the App credentials, the cron run
will replace the docs-side hand-authored JSON automatically; until
then the slice-2 1x5 sample remains in the docs repo.
Notes for next iteration:
- The exact `claude` CLI flags for tool-allowlist (`--allowed-tools`),
vision (`--image`), and extended thinking (`MAX_THINKING_TOKENS`)
are best-guess from the current Claude Code surface; if the live
PR-gate run reveals different flag names, tighten in place.
- Several non-Azure cells will likely need `not_applicable`
declarations once the cron VM produces real outcomes (e.g.
Bedrock Invoke + extended_thinking is uncertain). That refinement
is an iteration-2 follow-up driven by data, not a blocker for this
slice.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
94d323fa56 | Merge branch 'sandcastle/issue-26480-cron-vm-publish' into sandcastle/compat-matrix-stack | ||
|
|
069dba6e1d |
RALPH: compat matrix slice 3 - wire PR gate in CircleCI (#26479, PRD #26476)
Slice 3 of the Claude Code Compatibility Matrix: wire the
`tests/claude_code/` suite into CircleCI as a pre-merge gate. A red
status on the new `claude_code_compat_pr_gate` job blocks merge into
the staging branch.
What landed:
- tests/claude_code/pr_gate_version_resolver.py
The Claude Code PR-Gate Version Resolver described in the PRD's
"Version resolvers" section. Queries the npm registry for
`@anthropic-ai/claude-code` and returns the newest version whose
publish timestamp is at least 3 days old. The 3-day window is a
security review buffer: a malicious or broken Claude Code release
has at least 72 hours to be detected before it can land in our PR
gate. Importable function (with `metadata=` / `fetcher=` / `as_of=`
injection seams for tests) and a `python -m ...` CLI for the CI step.
- tests/claude_code/test_config.yaml
Proxy routing config that maps the per-cell aliases the tests use
(`claude-haiku-4-5`, `claude-haiku-4-5-bedrock-invoke`, ...,
`claude-opus-4-7-vertex`) to real upstream model ids on Anthropic /
Bedrock (Invoke + Converse) / Vertex AI. Azure intentionally has no
entries here because every Azure × claude-code cell is
`not_applicable` (Azure OpenAI doesn't host Claude).
- .circleci/config.yml
New `claude_code_compat_pr_gate` job. Pattern modeled on
`proxy_e2e_anthropic_messages_tests` (load PR-built docker image,
start postgres, mount config.yaml). New step in the middle:
resolve the Claude Code version from the resolver, install Node 20
via the machine image's preinstalled nvm, and `npm install -g
@anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}` (pinned, never
`latest`). Wired into `workflows.build_and_test` with a
`requires: [build_docker_database_image]` gate and the same
`*main_branches` filter the other proxy e2e job uses.
- tests/claude_code/_pr_gate_unit_tests/
16 new unit tests:
* 8 against the version resolver: boundary (>= 3d inclusive),
empty / all-too-new metadata, semver-vs-publish-time tiebreak,
custom min_age, fetcher injection, npm `time.created` /
`time.modified` skipping.
* 8 structural tests against `.circleci/config.yml`: job exists,
is in the workflow, requires the docker image, invokes the
resolver, install command is pinned (rejects unpinned `latest`),
runs `tests/claude_code/`, mounts `test_config.yaml`, exports
`LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY`. Plus one
regression test: the existing `proxy_e2e_anthropic_messages_tests`
job is unchanged in shape (acceptance criterion).
Key decisions:
- "Newest version" in the resolver is by **publish time**, not by
semver string ordering — if a patch lands on an older major after a
newer release, the patched line is the eligible one. (Tested.)
- The resolver's CLI prints the announcement to stderr and the bare
version to stdout, so the CI step can do
`CLAUDE_CODE_VERSION=$(uv run python -m ...)` cleanly while still
surfacing the selected version in the job log (acceptance criterion:
"the selected Claude Code version is logged").
- The structural CircleCI tests live under `_pr_gate_unit_tests/` so
the conftest path-inference hook skips them (the leading underscore
is the existing convention from `_driver_unit_tests/` /
`_builder_unit_tests/`); they don't pollute the matrix artifact.
- No `--no-verify` style supply-chain safety relaxation. Per the PRD,
Claude Code's pinning is the 3-day publish-age window, not a fixed
hash — by design, since the daily cron also pulls newer versions.
Tests: 47 -> 47 passing for the unit suite (16 new + 31 from slices
1 and 2). The end-to-end cells under `basic_messaging_non_streaming/`
require `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` and a
running proxy + `claude` CLI; they only run inside the new CircleCI
job.
Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- No docs PR is needed for this slice — the gate produces a status
check, not a published artifact. The compat matrix JSON the docs
page consumes is published by the daily-cron job (a future slice),
not by the PR gate.
Notes for next iteration:
- The daily cron / matrix publisher is the next slice. Several
pieces this slice introduces (the `tests/claude_code/test_config.yaml`
proxy config, the structure of the compat-results.json artifact)
will be reused by it.
- The bedrock-converse / vertex_ai aliases in `test_config.yaml` use
best-guess upstream model ids (`us.anthropic.claude-{tier}` and
`vertex_ai/claude-{tier}`); the real ids may need to be tightened
once the gate runs against live AWS / GCP credentials and we see
what resolves.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
e87e5c1199 |
RALPH: compat matrix slice 4 - daily cron VM publishes matrix to docs (#26480, PRD #26476)
Slice 4 of the Claude Code Compatibility Matrix: stand up the daily-cron pipeline that publishes `compatibility-matrix.json` to the docs repo. After this slice lands, the hand-authored matrix in the docs repo is replaced by auto-generated output, and the docs page begins reflecting real test runs against the latest stable LiteLLM release. What landed: - tests/claude_code/resolver.py Latest Stable LiteLLM Resolver. Calls the GitHub Releases API and returns the newest tag matching `v*-stable`. Sort is numeric on (major, minor, patch) so v1.10.0-stable correctly outranks v1.9.5-stable. Injectable `http_get` so tests run offline. - tests/claude_code/publisher.py Daily-cron orchestrator. Resolves the latest stable tag, pulls `ghcr.io/berriai/litellm:<tag>`, starts it as the proxy, installs `@anthropic-ai/claude-code@latest`, runs `pytest tests/claude_code/`, invokes the Matrix JSON Builder, and direct-pushes `compatibility-matrix.json` to the docs repo's main branch using a GitHub App installation token (`DOCS_REPO_TOKEN`). Idempotent: a no-op if the JSON is byte-identical to what's already on main. - tests/claude_code/_publisher_unit_tests/test_resolver.py test_publisher.py 14 unit tests covering the small pure helpers — version sort, non-stable filtering, http-get injection, commit message determinism, Docker image-name builder, and the file allowlist that enforces the "only `compatibility-matrix.json` ever ships" guarantee. Per the PRD's "Testing Decisions" section, the publisher's full subprocess orchestration intentionally ships without a unit-test harness; the daily-cron failure surface is itself the test. - .github/workflows/claude_code_compat_matrix.yml GitHub Actions workflow with three triggers (daily cron at 06:00 UTC, `release: published` filtered to `*-stable` tags, and `workflow_dispatch`). Mints a docs-repo installation token from a GitHub App scoped to `BerriAI/litellm-docs` only with `contents: write`, then runs the publisher. - .gitignore Add `compatibility-matrix.json` (cron VM output). Key decisions: - "Isolated VM" is realized as a GitHub-hosted ubuntu-latest runner — every run gets a fresh ephemeral VM, and the always-latest Claude Code CLI is only ever installed inside that ephemeral environment, so a malicious or broken Claude Code release cannot affect the trusted PR-gate CI in CircleCI. - File-level restriction on the GitHub App's broad `contents: write` scope is enforced by `select_files_to_commit` (script correctness), per the PRD's explicit acknowledgement that GitHub does not support file-path-scoped tokens. - `release` runs are filtered to tags ending in `-stable` at the workflow level, so a `v1.84.0-rc1` release does not republish the matrix. - Resolver and publisher live under `tests/claude_code/` alongside `matrix_builder.py` and `cli_driver.py` — production code that supports the test suite, kept colocated with it to match the slice 1+2 layout. Out of scope / blockers for next iteration: - Provisioning the GitHub App itself (creating it under BerriAI's org, installing it on litellm-docs only, generating the private key and registering `COMPAT_MATRIX_APP_ID` / `COMPAT_MATRIX_APP_PRIVATE_KEY` as repo secrets) is an operator/infra step that cannot land via a code change in this repo. - The first successful cron run is what removes the hand-authored `compatibility-matrix.json` from the docs repo and replaces it with generated output — that happens after this PR merges and the App is installed; not a code change here. Tests: 34 -> 45 passing (added 7 resolver tests + 7 publisher helper tests, all unit-only and offline). The 12 per-cell failures under `tests/claude_code/basic_messaging_non_streaming/` remain by design — they require a running proxy which the cron VM provides. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
658c41244c |
RALPH: compat matrix slice 2 - add 4 provider columns for basic_messaging_non_streaming (#26478, PRD #26476)
Slice 2 of the Claude Code Compatibility Matrix: extend the tracer-bullet
cell from slice 1 across all four remaining provider columns for
basic_messaging_non_streaming. Proves the multi-provider, multi-model,
all-must-pass aggregation logic against a 1x5 grid that exercises every
status state.
What landed:
- tests/claude_code/basic_messaging_non_streaming/test_bedrock_invoke.py
- tests/claude_code/basic_messaging_non_streaming/test_bedrock_converse.py
- tests/claude_code/basic_messaging_non_streaming/test_vertex_ai.py
Per-provider files modeled on test_anthropic.py: each parametrizes
over Haiku 4.5 / Sonnet 4.6 / Opus 4.7 (the three Claude tiers
required by the PRD), drives the real `claude` CLI through the
driver, and reports pass/fail via `compat_result`. Per-cell error
strings always include `[<model>]` so the docs tooltip can name the
failing model when a cell goes red.
- tests/claude_code/basic_messaging_non_streaming/test_azure.py
All three (Azure, Claude) cells report `not_applicable` with a
reason: Azure OpenAI Service does not host Anthropic models. The
test still parametrizes over the same three model ids so the test
count per cell is uniform across columns, and a future "Azure adds
Anthropic" announcement only requires flipping the body, not the
parametrization.
- tests/claude_code/sample_compatibility-matrix.json
Hand-authored 1x5 sample updated to reflect the slice 2 outcome:
anthropic / bedrock_invoke / bedrock_converse / vertex_ai = pass,
azure = not_applicable.
- tests/claude_code/_builder_unit_tests/test_matrix_builder.py
Two new golden-file tests:
1. 1x5 grid: feed the per-model results the four new test files
produce on a real run; assert the builder output equals the
hand-authored sample byte-for-byte.
2. fail-with-model-named: feed pass/fail/pass for one cell and assert
the cell aggregates to fail with the failing model id surfaced
in the error string (acceptance criterion: "the error string
identifies which model broke").
Key decisions:
- Duplication across the four per-provider files is accepted (per the
PRD) rather than extracted into a helper. Each file is self-contained
so a test author touching one provider doesn't accidentally regress
the others.
- Per-provider model alias names: `claude-<tier>-<provider-suffix>`
(e.g. `claude-haiku-4-5-bedrock-invoke`). These are the alias names
the proxy operator wires up in the routing config; the test only
knows the alias, the proxy knows the upstream model id and region.
- Azure is `not_applicable` rather than `not_tested` because the
cell will never apply, not "we haven't gotten to it yet" - the two
states are visually and semantically distinct in the rendered grid.
- Sample shows the realistic best-case outcome (4 pass + 1 NA). The
React renderer's coverage of the `fail` and `not_tested` states is
exercised by other cells in v1+, not the v0 sample.
Tests: 31 -> 34 passing (added 2 builder golden tests + 3 Azure
not_applicable parametrizations that pass without env vars).
Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The companion update to compatibility-matrix.json in the docs repo.
The hand-authored sample in this repo is the artifact the docs PR
copies; opening that doc PR is the next step in this slice.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
0bf013f620 |
RALPH: tracer-bullet for Claude Code compatibility matrix (#26477, PRD #26476)
Slice 1 of the Claude Code Compatibility Matrix: the thinnest end-to-end path through every layer for a single (feature, provider) cell, so a future docs page can render a real green cell sourced from a real test. What landed in this repo: - tests/claude_code/manifest.yaml — feature manifest with one entry (basic_messaging_non_streaming) plus the v0 provider column order. - tests/claude_code/cli_driver.py — Claude Code CLI Driver. One entry point (run_claude); handles subprocess assembly, env overlay, stream-JSON parsing, and structured failure modes. `runner=` is a unit-test seam. - tests/claude_code/conftest.py — `compat_result` fixture (tagged-union recorder) + pytest_runtest_makereport hook that infers (feature, provider) from the file path and writes a structured compat-results.json artifact. - tests/claude_code/basic_messaging_non_streaming/test_anthropic.py — the one cell, parametrized over Haiku/Sonnet/Opus per the PRD's per-cell model rule. - tests/claude_code/matrix_builder.py — pure-function builder from (manifest, results, run-metadata) to the v1 JSON schema. Aggregates per- model results into one cell (pass iff all pass). build_from_paths is the thin I/O wrapper for the publisher. - tests/claude_code/sample_compatibility-matrix.json — hand-authored sample of the v1 JSON; copied to the docs repo by hand as part of this slice. - Unit tests: 10 driver tests (mocked subprocess), 9 compat_result tests, 10 matrix-builder golden-file tests. 29/29 pass. Key decisions: - (feature, provider) is inferred from file path, not declared in metadata — mirrors the PRD's "no drift" goal. - Driver injects subprocess via a `runner` kwarg so unit tests don't need the real `claude` CLI; production callers leave it default. - Builder is a pure function on Mappings/Sequences; load/write live in a thin `build_from_paths` wrapper. Golden-file tests pin the schema. - `_driver_unit_tests/` and `_builder_unit_tests/` are prefixed with `_` so the conftest's path-inference hook skips them and they don't pollute the matrix artifact. - `compat-results.json` added to .gitignore (CI-only output). Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs): - The MDX page `docs/tutorials/claude-code-compatibility` and the `<CompatibilityMatrix />` React component. The hand-authored compatibility-matrix.json (`sample_compatibility-matrix.json` in this repo) is the artifact those docs files will consume; opening that doc PR is the next step in this slice. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |