litellm/tests/e2e/coverage_registry
devin-ai-integration[bot] b21e44cbf9
feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375)
* test(e2e): jwt auto_register map-existing-key repro

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(jwt): route existing-key lookup through VerificationTokenRepository

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key

Skip Admin UI session keys and keys whose allowed_routes restrict them to
anything other than llm_api_routes (management, read_only, password-reset
sessions). Mapping a JWT to one of those left the user with 401s or 403s on
every LLM call, since the mapping persists.

* fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team

Only reuse a key whose team_id matches the team auth_builder resolved for
the JWT (no team matches no team), so a personal key can no longer bypass
the resolved team's model and budget limits.

With the flag on, the first JWT request now falls through to the same
virtual-key checks later mapped requests get, instead of returning early,
so a reused key's own limits apply from request one rather than 200 then
403. Flag off keeps the early return unchanged.

* fix(jwt): keep the early return when no master key is set

Without a master key the generic virtual-key path returns a bare
INTERNAL_USER object, so falling through on the first auto-registered
request dropped the key's team, models and budgets. Only fall through when
a master key is configured.

Tests now assert the reused key per team rather than the query shape, and
cover the flag-off early return and the no-master-key case.

* test(jwt): assert on race-loser's returned key, not only mocks (TQ002)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes

A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key)

Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget

Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401

Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter

* test(e2e): create the reused key in the team the JWT resolves to

The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass

* test(integration): match the held statement across TCP reads

The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read

* fix(jwt): gate key reuse on the claim field, not on the claim value

Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path

* fix(jwt): let an issuer's own user field replace the global one when gating key reuse

An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key

* test(jwt): make the flag-off test fail when the flag no longer gates key reuse

The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings

* chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 12:11:33 -07:00
..
__init__.py chore: consolidate CLAUDE.md into AGENTS.md 2026-09-19 02:30:35 +00:00
collector.py fix(e2e-changed): keep the gate off suites the stack cannot run 2026-09-05 21:03:50 -07:00
guardrail.yaml test(e2e): move live-provider legacy tests into tests/e2e (#44120) 2026-10-02 00:02:18 -07:00
llm_claude_code_compat.yaml test(e2e): replay a real tool-search assistant turn back to Bedrock Invoke (#36856) 2026-08-17 11:59:26 -07:00
llm_conversational.yaml test(e2e): bill Sail windows that synchronous calls can still use (#44058) 2026-10-01 12:54:26 -07:00
llm_nonconversational.yaml test(e2e): assert only litellm-owned batch behavior and move the blank S3 env pin to an integration test (#43321) 2026-09-26 11:17:03 -07:00
logging.yaml feat(s3_v2): add s3_partition_granularity option for hourly S3 folders (#43748) 2026-10-01 11:00:21 -07:00
management_cases.py test: enforce isolated actors and stop OIDC process groups 2026-09-12 13:49:49 -07:00
mcp.yaml test(e2e): restore LIT-3467 implementation for rework 2026-09-19 16:21:53 -07:00
mgmt.yaml fix(proxy): delete large teams without per-member transaction fan-out (#42998) 2026-09-30 13:49:18 -07:00
other.yaml feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375) 2026-10-02 12:11:33 -07:00
quota_management.yaml fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870) 2026-09-29 12:54:17 -07:00
README.md chore: consolidate CLAUDE.md into AGENTS.md 2026-09-19 02:30:35 +00:00
registry.py ci: gate tests/e2e on zero basedpyright errors in pre-commit and lint CI 2026-07-11 10:25:22 -07:00
reliability.yaml test(e2e): hold every worker under an idle RSS budget before any traffic (#42552) 2026-09-22 14:47:14 -07:00
schema.py feat(sail): add Sail as a provider with service_tier mapped to its completion window (#42840) 2026-09-26 12:57:48 -07:00
test_collector.py fix(e2e): exclude skipped tests from coverage-registry numerator 2026-07-30 22:19:30 -07:00

e2e coverage registry

This directory is the denominator for e2e test coverage: the set of behaviors we want covered, one row per behavior, checked into the repo so coverage is a number we can track instead of a guess. It implements the plan in the "E2E Coverage Tracking" note; the naming grammar lives in tests/e2e/AGENTS.md.

The model

A cell is one customer-noticeable behavior a single e2e test can assert pass/fail on, for example llm.chat_completions.bedrock_converse.tool_use.stream.works. Cells are grouped module > feature > test, with LLM cells split into Core LLMs and Non-Core LLMs for dashboarding. Each cell carries a tier (P0/P1/P2), a source, and a fail_before_fix flag.

The rows live in per-prefix YAML files (llm_*.yaml, mgmt.yaml, mcp.yaml, reliability.yaml, quota_management.yaml, logging.yaml, guardrail.yaml, other.yaml) and validate against the discriminated union in schema.py, so an LLM row cannot carry a guardrail field and vice versa. llm rows with subject_endpoint of chat_completions, messages, or responses roll up to Core LLMs; all other LLM endpoints roll up to Non-Core LLMs. LLM endpoint, route, and capability values are typed in schema.py, so new taxonomy values require an explicit schema change. logging and guardrail are two id-prefixes that roll up into the single Logging & Guardrails dashboard module.

A test declares what it covers with a marker:

@pytest.mark.covers("llm.chat_completions.openai.tool_use.stream.works")
def test_openai_streaming_tool_calls(self) -> None:
    ...

The number

collector.py diffs the registry against those markers and reports coverage per module. It is static: a collect-only pass reads the markers, so it runs no test and needs no live proxy. Whether a covered cell currently passes or fails is a separate, live concern.

A skipped test asserts nothing, so its markers do not count. A cell is covered only when at least one test pytest would actually run declares it; a cell claimed by both a live test and a skipped one stays covered. Skip state comes from pytest's own evaluator, so skip and skipif resolve exactly as they do in the e2e run, which also means a skipif on an absent credential makes that cell uncovered in the environments where the test cannot run. Cells left uncovered this way are listed under the headline (and counted by litellm_e2e_coverage_skipped_markers) so an unskipped-pending gap is visible rather than inflating the number. The one skip the collector cannot see is pytest.skip() called from inside a test body, since it does not exist until the test runs.

cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector

Use --format loki after the e2e pytest run in the same Kubernetes job/pod to print structured stdout lines for Loki:

cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector --format loki --strict

This emits exactly one COVERAGE_TOTAL line and one COVERAGE_MODULE line per module in MODULE_ORDER, in that order. Loki uses log-safe module= labels from LOKI_MODULE_LABELS (core_llms, management_ui, etc.) so existing JSON and Prometheus consumers keep their human-readable module names unchanged.

The headline is overall coverage. The collector also lists markers that point at ids not in the registry, so a typo or an unenumerated behavior surfaces instead of being silently dropped.

Use strict mode in CI once existing draft markers are reconciled:

cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector --strict

Strict mode exits non-zero on @pytest.mark.covers(...) ids that are not checked into the registry. Add --fail-on-collection-errors when the job should also fail on pytest collection errors.

Provider x feature matrix: customer-run Bedrock combinations

The provider and feature combinations customers actually run get explicit cells, expanded here as incidents surface new ones. The current Bedrock set, seeded from a customer's production shape (regional us.anthropic.* inference-profile ids over both chat routes, provider response headers for AWS-side correlation, and the Test Connection probe for a responses-mode Bedrock Mantle deployment):

Cell Feature Covering test
llm.chat_completions.bedrock_converse.basic.nonstream.works regional us. id, Converse llm_translation/test_chat_completions_regression_e2e.py
llm.chat_completions.bedrock_converse.basic.stream.works regional us. id, Converse stream llm_translation/test_chat_completions_regression_e2e.py
llm.chat_completions.bedrock_invoke.basic.nonstream.works regional us. id, Invoke llm_translation/test_bedrock_provider_matrix_e2e.py
llm.chat_completions.bedrock_invoke.basic.stream.works regional us. id, Invoke stream llm_translation/test_bedrock_provider_matrix_e2e.py
llm.chat_completions.bedrock_converse.response_headers.nonstream.works llm_provider-* headers llm_translation/test_bedrock_provider_matrix_e2e.py
llm.chat_completions.bedrock_converse.response_headers.stream.works llm_provider-* headers, stream llm_translation/test_bedrock_provider_matrix_e2e.py
mgmt.model.test_connection.happy_path Test Connection, Bedrock Mantle management/test_model_test_connection_e2e.py

Status: this is a draft for review

The cells were enumerated from the codebase and the tiers are a first proposal. Known things to settle before treating the set as final:

  • tiers are proposed, not signed off; 125 P0 is a lot to prove fail-before-fix, so P0 may want tightening
  • a few cells need a support check or a prune (for example llm.embeddings.anthropic.* and reliability.perf.throughput.under_slo)
  • auth is covered in two places (other.auth.* and the mgmt authz assertions); the boundary needs a decision, and the auth cluster may deserve promotion to its own module
  • the P2 "niche" cells each stand in for a large tail of integrations/providers by design, so the denominator is deliberately P0-weighted rather than a full inventory