* chore(e2e): move the compat-matrix populator from a GCE VM to a Render cron job The daily Claude Code compatibility-matrix job ran as a systemd timer on the litellm-compatibility-matrix-populator VM in the vertex-check GCP project. Replace that with a Render Docker cron job built from a new Dockerfile in tests/e2e/claude_code/cron_vm: pinned and checksummed debian base, gh, uv, and Claude Code CLI, a non-root populator user, and run_daily.sh as the entrypoint. run_daily.sh now clones a fresh blobless checkout per run (Render cron disks are ephemeral), reads the publish PAT from the github-token secret file under CREDENTIALS_DIRECTORY, and its comments no longer describe systemd. The .service and .timer units are gone; README.md and the env example describe the Render service, its secret files, and the local docker build instead. * docs(e2e): name the plan and trigger route Render's cron-job API accepts Render answers a bare 404 for the legacy pro_max plan name on a cron job (4c-16g is the same 4 CPU / 16 GB size) and the manual trigger route is /v1/cron-jobs, not /v1/cronjobs. * fix(e2e): install the published litellm wheel instead of building the tag from source The tag builds a Rust extension through maturin, which needs a C and Rust toolchain the cron image does not carry, so the first Render run failed at uv sync with "linker cc not found". Sync the locked dependencies with --no-install-project, install the PyPI wheel (what users run) with --no-build, and pass --no-sync to every uv run so uv never puts the source build back. * fix(e2e): keep the SKIP_PUBLISH matrix where a Render run can read it The validation run wrote the matrix into the image checkout, which nobody can read once the container exits. Save it under HOME and print it at the end of the log instead. * fix(e2e): let the stale compat-matrix PR sweep see past the newest 100 docs PRs The docs repo has a few hundred open PRs, so a 100-item list never reached the week-old compat-matrix PR and the sweep left it open on every run. * docs(e2e): say the Render cron needs a manual deploy after each merge Pushes never started a deploy during setup because Render only hears about them through its GitHub app, which the org does not have, so the README now carries the deploy command and the wait-for-live rule * ci: build the compat-matrix cron image on pull requests The CI coverage gate requires every Dockerfile to be built by a job, and building this one on each PR that touches it also catches a broken pin or checksum before Render does * fix(e2e): shim the whole tests/e2e tree into the compat-matrix worktree The five-file helper allowlist missed fixture_mode, which e2e_config now imports, so the first Render run died at conftest load with ModuleNotFoundError. Copy the image's whole tests/e2e tree instead and keep pytest from loading the EKS-harness conftest with --confcutdir * fix(e2e): scope the compat-matrix sweep to the publishing account's own PRs The stale-PR sweep selected every open docs PR whose head branch starts with compat-matrix/, so a contributor's fork PR under that name would have been closed once a newer matrix PR existed. The sweep now resolves the publishing login from the token and only closes same-repo PRs that account opened --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| build_matrix.py | ||
| check_regressions.py | ||
| Dockerfile | ||
| litellm-compat-matrix.env.example | ||
| README.md | ||
| run_daily.sh | ||
Render cron job for the Claude Code compatibility-matrix populator
The populator runs daily as the Render cron job litellm-compat-matrix
(Docker runtime, built from the Dockerfile in this directory) rather
than as a GitHub Action or on a dedicated VM. Trade-offs:
- ✅ No machine to keep on or patch. Render builds the image from this
directory on every push to
mainthat touchestests/e2e/**and runs it on the schedule. - ✅ Credentials live in Render env vars and secret files, scoped to this one service, instead of on a VM filesystem.
- ✅ The publish token still uses the
mateo-berriaccount, which is a collaborator onBerriAI/litellm-docs, so no GitHub App withpull-requests: writehas to be provisioned. - ⚠️ The disk is ephemeral, so every run starts from a fresh (blobless)
clone of litellm plus a cold
uv sync. That adds a few minutes on top of the ~10 minute test run; the job's 12 hour ceiling is nowhere near. - ⚠️ The Claude Code CLI version under test is pinned in the
Dockerfile(CLAUDE_CODE_VERSION+ its checksum). Bumping it is a PR, see the gotchas below.
Layout
| File | Purpose |
|---|---|
Dockerfile |
The image Render builds: Debian bookworm-slim plus pinned, checksum-verified gh, uv, and the Claude Code CLI, with this tests/e2e/ tree copied to /opt/litellm/tests/e2e/. Runs as the non-root user populator (uid/gid 1000, which is what Render's secret files are readable by). |
run_daily.sh |
The actual cron job. Resolves versions, clones the worktree, boots the proxy, runs pytest, builds the JSON, opens (or updates) a docs PR, sweeps stale compat-matrix PRs. |
build_matrix.py |
Tiny Python CLI that wraps claude_code.matrix_builder.build_from_paths. Exists only because the bash script needs some way to render the per-cell aggregation, and the builder is already Python. |
check_regressions.py |
Tiny Python CLI that wraps claude_code.matrix_builder.find_regressions. Diffs the freshly built matrix against the currently-published one and exits 3 if any cell flipped green→red, which gates auto-merge. |
litellm-compat-matrix.env.example |
The service's env vars, one per line, with what each is for. |
What run_daily.sh does
- Resolves the latest LiteLLM final release tag (newest bare
vX.Y.Z, skipping-rc.N/-dev.Npre-releases) by paging the GitHub Releases API (curl | jq). - Reads the Claude Code CLI version via
claude --version. That is whatever theDockerfilepins; the job never upgrades it on its own. - Clones the worktree at
~/litellm-cron-worktree/(a--filter=blob:noneclone, so only the checked-out tag's blobs are fetched),git checkout --force <tag>, thenuv sync --frozen --no-install-projectagainst a uv-managed CPython 3.12 followed byuv pip install --no-build litellm==<version>, so the proxy under test is the published PyPI wheel (what users install) rather than a source build: the tag builds a Rust extension through maturin, and the image ships no C or Rust toolchain. Then shims the test suite:tests/e2e/in the worktree is replaced by the image's copy of this whole tree, so the cron always runs today's tests against the latest stable proxy, and pytest runs with--confcutdirpointed atclaude_code/so the tree's EKS-harnessconftest.py(whose imports the stable venv doesn't install) is never loaded. The tag's owntests/e2e/is deliberately not used. - Boots the proxy as a
setsidbackground process on port4100bound to loopback, then polls/health/livelinessuntil it's up. - Runs pytest on
tests/e2e/claude_code/withLITELLM_PROXY_URLpointed at the proxy andCOMPAT_RESULTS_PATHset so the conftest hook writes the per-test results artifact. Test failures becomefailcells in the JSON, not script errors. - Builds
compatibility-matrix.jsonby handing the artifact + manifest tobuild_matrix.py. - Opens or updates a docs PR:
gh repo cloneoflitellm-docsinto a tempdir, deterministic head branch (compat-matrix/<litellm-version>-<claude-code-version>-<UTC-date>),--forcepush directly toBerriAI/litellm-docs(themateo-berritoken has write access, so this is a same-repo branch, not a fork),gh pr create. A re-run on the same day fast-forwards the existing branch andgh pr createno-ops ("a pull request for branch ... already exists" is treated as success). If the JSON is byte-identical to whatmainalready publishes, the push is skipped entirely. These PRs are not gated on a second human review. - Gates auto-merge on a regression check: before enabling
auto-merge,
check_regressions.pydiffs the new matrix against the one currently onmain. Auto-merge (gh pr merge --auto --squash) is only enabled when no cell flipped green→red — i.e. every transition is red→green, green→green, or red→red. A pre-existing red cell (e.g. a provider that's out of API credits) isred→redand does not block; only apass→failflip does. When a regression is detected the PR is still opened/updated (with a warning banner naming the offending cells) but auto-merge is left off — and any auto-merge a prior same-day run enabled is explicitly disabled — so a human reviews before it lands on the public table. The check fails closed: if it errors, auto-merge is withheld. - Sweeps stale compat-matrix PRs: once today's PR exists, every
other open
compat-matrix/*PR that the publishing account opened from a branch on the docs repo itself is closed (and its bot-owned branch deleted), so at most one compat-matrix PR is ever open — the newest. A contributor's PR under that prefix is never touched.
The Render service
Everything below is what the live service is set to; recreate it with the same values if it ever has to be rebuilt.
| Setting | Value |
|---|---|
| Workspace | Litellm (the one that already builds the other litellm services) |
| Type | Cron job, Docker runtime |
| Repo / branch | BerriAI/litellm @ main |
| Dockerfile path | tests/e2e/claude_code/cron_vm/Dockerfile |
| Docker build context | tests/e2e (the repo root .dockerignore excludes tests, so the context has to start below it) |
| Build filter | included paths tests/e2e/** |
| Schedule | 0 6 * * * (06:00 UTC daily) |
| Plan / region | 4c-16g (4 CPU, 16 GB, what the dashboard calls Pro Max; the suite fans out to ~75 concurrent CLI calls) / Oregon |
| Env vars | every key in litellm-compat-matrix.env.example |
| Secret files | github-token (the publish PAT, one line) and vertex-service-account.json (the Vertex service-account key) |
Render mounts secret files at /etc/secrets/<name>, which is where
CREDENTIALS_DIRECTORY and GOOGLE_APPLICATION_CREDENTIALS in the env
example point. Render also passes env vars to docker build as build
args, which is why the Dockerfile declares no ARG that could ever
be given a secret's name.
Creating it through the API looks like this (fill envVars and
secretFiles from the env example and the two secrets; ownerId is
the workspace id from GET /v1/owners):
curl -fsS https://api.render.com/v1/services \
-H "Authorization: Bearer ${RENDER_API_KEY}" \
-H 'Content-Type: application/json' \
-d '{
"type": "cron_job",
"name": "litellm-compat-matrix",
"ownerId": "<workspace id>",
"repo": "https://github.com/BerriAI/litellm",
"branch": "main",
"autoDeploy": "yes",
"buildFilter": {"paths": ["tests/e2e/**"], "ignoredPaths": []},
"envVars": [{"key": "ANTHROPIC_API_KEY", "value": "..."}],
"secretFiles": [{"name": "github-token", "content": "..."},
{"name": "vertex-service-account.json", "content": "..."}],
"serviceDetails": {
"runtime": "docker",
"schedule": "0 6 * * *",
"plan": "4c-16g",
"region": "oregon",
"envSpecificDetails": {
"dockerfilePath": "tests/e2e/claude_code/cron_vm/Dockerfile",
"dockerContext": "tests/e2e"
}
}
}'
Operating it
# Trigger a real run right now (PRs to litellm-docs). The id is the
# service id (`crn-...`) from the dashboard URL or `GET /v1/services`.
curl -fsS -X POST "https://api.render.com/v1/cron-jobs/${CRON_ID}/runs" \
-H "Authorization: Bearer ${RENDER_API_KEY}"
# Follow a run: the Logs tab on the service, or the API.
curl -fsS "https://api.render.com/v1/logs?ownerId=${OWNER_ID}&resource=${CRON_ID}&limit=100" \
-H "Authorization: Bearer ${RENDER_API_KEY}"
# Rebuild the image after a merge that touches tests/e2e/** (see the
# auto-deploy gotcha below). The deploy is done once its status is
# `live`; a run triggered before that still uses the previous image.
curl -fsS -X POST "https://api.render.com/v1/services/${CRON_ID}/deploys" \
-H "Authorization: Bearer ${RENDER_API_KEY}" \
-H 'Content-Type: application/json' -d '{"clearCache": "do_not_clear"}'
curl -fsS "https://api.render.com/v1/services/${CRON_ID}/deploys?limit=1" \
-H "Authorization: Bearer ${RENDER_API_KEY}"
# A run that does NOT open a PR (first-time validation, CLI bumps):
# set SKIP_PUBLISH=1 on the service, trigger a run, then remove it.
# The matrix JSON is printed at the end of the run's log (nothing on
# the container's disk outlives the run) and saved to
# ~/compatibility-matrix.json for a local docker run.
# PYTEST_K='basic_messaging_non_streaming and anthropic' narrows the
# run to one cell the same way.
# Build and run the image locally (docker on Apple silicon needs the
# platform flag; the context is tests/e2e, see the table above).
docker build --platform linux/amd64 \
-f tests/e2e/claude_code/cron_vm/Dockerfile -t compat-matrix tests/e2e
docker run --rm --platform linux/amd64 \
--env-file litellm-compat-matrix.env -e SKIP_PUBLISH=1 \
-v "$PWD/secrets:/etc/secrets:ro" compat-matrix
Gotchas
- The venv is pinned to Python 3.12 (
CRON_PYTHON_VERSION). The e2e suite uses PEP 695typealiases, which the image's Debian Python can't parse;run_daily.shhas uv fetch a managed CPython into~/litellm-cron-worktree/.uv-python/and syncs the venv against it. - The proxy port is
4100, not4000. Kept from the VM days so a developer running the script locally next to their own:4000proxy doesn't collide. Override withPROXY_PORT=.... uv sync --frozenrequires the resolved tag to be tagged on GitHub, and the wheel install requires it on PyPI. If the latest stable release was made but not pushed as a git tag, thegit checkoutstep fails; push the tag, then rerun. PyPI has had every stable version days before its GitHub release so far (1.102.0 was uploaded 2026-09-20, released on GitHub 2026-09-22), so the--no-buildinstall failing means the wheel is genuinely missing, not late.- Pushes do not redeploy the service; deploy by hand.
autoDeployisyeson the service, but Render only hears about pushes through its GitHub app, which is not installed on theBerriAIorg (an org admin step), so no push to the branch has ever started a deploy. After a merge that changes anything undertests/e2e/**, run the deploy command from the operating section (or "Manual Deploy" on the dashboard) and wait forlivebefore triggering a run, otherwise the next scheduled run still executes the old image. - Publish-token rotation is your problem. The cron does not
refresh the token; if
mateo-berri's PAT in thegithub-tokensecret file expires, the run fails at thegit push/gh pr createstep with a 401 ("Bad credentials" / "Authentication failed"). Mint a fresh PAT and replace the secret file on the service. The token needs write access toBerriAI/litellm-docs(classicreposcope, or fine-grained Contents:RW + Pull requests:RW). It is delivered as a file, not an env var, so pytest, the proxy, and the claude CLI never inherit it; manual runs exportGITHUB_TOKENinstead. - Bumping the Claude Code CLI is a PR. Change
CLAUDE_CODE_VERSIONin theDockerfileand setCLAUDE_CODE_SHA256to thelinux-x64checksum fromhttps://downloads.claude.ai/claude-code-releases/<version>/manifest.json. The first run on a new CLI is the riskiest one: if the new CLI changes its wire format the matrix run can produce systematic failures, so trigger aSKIP_PUBLISH=1run before the next scheduled fire.ghanduvbump the same way, with the checksum from the release'sgh_<version>_checksums.txtand the tarball's.sha256sidecar respectively. - A local build on Apple silicon only proves the image assembles.
Under QEMU the Claude Code binary (a Bun executable) dies with
CPU lacks AVX supportandghpanics in the Go runtime, soclaude --versionand a full run are verified with aSKIP_PUBLISH=1run on Render, not locally. - Nothing persists between runs. A failed run leaves no half-installed venv behind, but also no cache: don't expect a rerun to be faster than the first one.