docs(agents): ban semicolons and colons in human-facing text (#45684)

* docs(agents): ban semicolons and colons in human-facing text

Tighten the writing rule so ";" and ":" are only allowed where the sentence would be nonsensical without them, rewrite the existing semicolons in AGENTS.md to follow it, and note that docs (litellm-docs) keep their trailing "."

* docs(agents): rewrite semicolons in nested AGENTS.md files

* docs(agents): hand-maintain the dashboard Next.js warning and sync e2e CONTRIBUTING lines

* docs(agents): point at the PR template path that exists

* docs(agents): restore the generated Next.js block in the dashboard AGENTS.md
This commit is contained in:
Mateo Wang 2026-10-09 16:36:32 -07:00 • committed by GitHub
parent 4cb47edd2a
commit 3069a467d0
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
38 changed files with 216 additions and 216 deletions

View file

@ -25,7 +25,7 @@ Same thing for bug fixes. The tests should make it so that this specific bug can
Never test structure of code only function of it
A test must only fail when litellm code changes. Never pin facts we don't own (a vendor's price, a third party's field, an upstream default, today's date) as literals or as "X must be absent"; assert the invariant our code guarantees instead, e.g. two rows agree, a value is within range, a field is derived from another. If an outside fact is truly load-bearing, cite its source and date next to the assertion so a reader can tell stale from broken
A test must only fail when litellm code changes. Never pin facts we don't own (a vendor's price, a third party's field, an upstream default, today's date) as literals or as "X must be absent". Assert the invariant our code guarantees instead, e.g. two rows agree, a value is within range, a field is derived from another. If an outside fact is truly load-bearing, cite its source and date next to the assertion so a reader can tell stale from broken
`tests/unit/` mirrors `litellm/` in a parallel path (see `tests/unit/AGENTS.md`). Name tests `test_<filename>.py`, but always match the existing test file in the directory you touch — many provider dirs use longer descriptive names (e.g. `test_anthropic_chat_transformation.py`) to avoid ambiguity across sibling folders. For bug fixes, extend the existing mapped test file rather than creating a new one. Only create a new test file for a new feature (provider, endpoint, or transformation module) that has no mapped test yet, following that directory's naming convention (or `test_<filename>.py` if you're the first test there). One focused regression test beats many shallow ones
@ -33,20 +33,20 @@ End-to-end tests belong in `tests/e2e/` and must follow the harness conventions
When creating PRs, target the repository's current default branch for both internal and external / OSS contributions. Check it with `python3 scripts/default_branch.py --branch` instead of assuming a branch name or relying on cached `origin/HEAD`
When writing a PR body, treat the comments and imperative instructions inside .github/pull_request_template.md as rules to follow, not just layout. Agent harnesses may strip HTML comments from copies of that file injected into context, so read .github/pull_request_template.md from disk before writing a PR body to make sure you see every comment rule. A section you have nothing to put in (Relevant issues, Affected release, Linear ticket, Caveats, QA runbook, and so on) is removed entirely, heading included, never left as an empty title
When writing a PR body, treat the comments and imperative instructions inside .github/PULL_REQUEST_TEMPLATE/general.md as rules to follow, not just layout. Agent harnesses may strip HTML comments from copies of that file injected into context, so read .github/PULL_REQUEST_TEMPLATE/general.md from disk before writing a PR body to make sure you see every comment rule. A section you have nothing to put in (Relevant issues, Affected release, Linear ticket, Caveats, QA runbook, and so on) is removed entirely, heading included, never left as an empty title
Same applies for filing bug reports and feature requests, with .github/ISSUE_TEMPLATE/bug_report.yml and .github/ISSUE_TEMPLATE/feature_request.yml, respectively
If you're resolving a linear ticket, in the "## Linear ticket" section of the PR, say "Resolves LIT-1234", replacing "LIT-1234" with the actual ticket id that you're resolving. If you don't have the ticket id, don't make one up or search for it. Just drop the section
Never use `pytest` commands or the like as "Screenshots / Proof of Fix". We prefer curl'ing a live proxy instance running on localhost:4000 (I like to run it with `python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --detailed_debug --reload --use_v2_migration_resolver 2>&1 | tee litellm.log`; the Admin UI dev server is `npm run dev` in `ui/litellm-dashboard`, served on port 3000) and showing both the command run and the output. Also, it should hit real LLM provider APIs, not mocks, and cost real $$$ because that is the most realistic test. The proof of fix should be exactly what the end user / customer would see / do. The run logs in PR #27703 is a prime example of how to do it (not a huge fan of using a python test script that future me and the team will have no visibility into; I prefer just curl commands or a short list of bash commands (e.g., using `for`)). If it's a UI thing, or the main use case runs through a headful agentic coding tool like Claude Code or Codex, drive that surface yourself and embed your own before and after screenshots of it in the PR (the Admin UI page, or what the coding tool shows), next to an ordered list of the URLs to go to (e.g., http://localhost:4000/ui/?page=logs), where to click, and what fields to fill out so a reviewer can reproduce it
Never use `pytest` commands or the like as "Screenshots / Proof of Fix". We prefer curl'ing a live proxy instance running on localhost:4000 (I like to run it with `python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --detailed_debug --reload --use_v2_migration_resolver 2>&1 | tee litellm.log`. The Admin UI dev server is `npm run dev` in `ui/litellm-dashboard`, served on port 3000) and showing both the command run and the output. Also, it should hit real LLM provider APIs, not mocks, and cost real $$$ because that is the most realistic test. The proof of fix should be exactly what the end user / customer would see / do. The run logs in PR #27703 is a prime example of how to do it (not a huge fan of using a python test script that future me and the team will have no visibility into. I prefer just curl commands or a short list of bash commands (e.g., using `for`)). If it's a UI thing, or the main use case runs through a headful agentic coding tool like Claude Code or Codex, drive that surface yourself and embed your own before and after screenshots of it in the PR (the Admin UI page, or what the coding tool shows), next to an ordered list of the URLs to go to (e.g., http://localhost:4000/ui/?page=logs), where to click, and what fields to fill out so a reviewer can reproduce it
If you ever write any human-facing text (pull requests, issues, commit messages, discussion posts, github comments, release notes, docs, etc.), always follow these guidelines to sound less AI-y:
- don't use emojis
- don't use "—". Instead, reach for ",", ".", conjunction words, ":", ";", etc. in descending order of preference: vary among them, weighted toward the front of the list, and skip "," where it would cause a comma splice or the sentence is getting long. Overusing any one of them, ";" especially, also feels AI-y. A word cap does not penalize you for adding more sentences: when writing under tight word budgets, prefer a period split or a conjunction over ";", and keep to at most one ";" per message
- don't use "—". Don't use ";" or ":" unless it would be nonsensical not to. Instead, reach for ",", ".", conjunction words, etc. in descending order of preference: vary among them, weighted toward the front of the list, and skip "," where it would cause a comma splice or the sentence is getting long. Overusing any one of them also feels AI-y. A word cap does not penalize you for adding more sentences
- don't use the pattern "It's not X, it's Y", "You're not X, you're Y", etc.
- unless explicitly asked, don't use bulleted or numbered lists unless it would be nonsensical not to. Instead, prefer prose
- don't add a trailing "." at the end of paragraphs (just like this file). That means every paragraph, not just the last one (of the markdown file, PR description, GitHub comment, etc.). Rule of thumb: if you're adding new line(s) before the next sentence, don't add a "."
- don't add a trailing "." at the end of paragraphs (just like this file). That means every paragraph, not just the last one (of the markdown file, PR description, GitHub comment, etc.). Rule of thumb: if you're adding new line(s) before the next sentence, don't add a ".". Exception is that docs (i.e., litellm-docs) should have trailing "."
- don't use →. Instead, prefer not to use arrows, and if need be, use -> instead
- use plain, simple, everyday engineering language: the common phrase engineers actually say over rare compact phrasing, in grammatically complete sentences. When explicitly asked to use bullets or ordered lists and structure legitimately helps the reader, prefer nested bullets (any depth is fine) over dense lines in a flat structure
@ -58,17 +58,17 @@ The four lint gates (`scripts/ruff_strict_gate.py`, `scripts/type_discipline_gat
`make check` (f.k.a. `make pre-commit`, which still works identically as an alias) saves its complete output to a log file in .git (overwriting previous logs) and prints that path as its first and last output lines. To inspect a run, read or grep that log instead of re-running the multi-minute checks just to see a different slice
`make check`, `make lint`, `scripts/pre_commit_lint.sh`, and the standalone lint gates (`scripts/ruff_strict_gate.py`, `scripts/type_discipline_gate.py`, `scripts/type_check_gate.py`) each hold one of 2 machine-wide slots, so when other sessions or worktrees on the same box are already running heavy work, yours prints "all N machine-wide slots are busy; queueing" and then stays quiet until a slot frees. Give the command a long timeout and let it wait rather than killing it, retrying it, or assuming it hung. Don't change the # of machine-wide slots or make it unlimited by setting `LITELLM_GATE_SLOTS=0`
`make check`, `make lint`, `scripts/pre_commit_lint.sh`, and the standalone lint gates (`scripts/ruff_strict_gate.py`, `scripts/type_discipline_gate.py`, `scripts/type_check_gate.py`) each hold one of 2 machine-wide slots, so when other sessions or worktrees on the same box are already running heavy work, yours prints "all N machine-wide slots are busy" followed by "queueing" and then stays quiet until a slot frees. Give the command a long timeout and let it wait rather than killing it, retrying it, or assuming it hung. Don't change the # of machine-wide slots or make it unlimited by setting `LITELLM_GATE_SLOTS=0`
If you're trying to create a new function that relies on untyped stuff, instead of adding more Any's and pushing `reportAny` / `reportExplicitAny` closer to their cap in `ANY_CAPS`, just validate it in the caller with Pydantic (a model or `TypeAdapter` that returns the typed thing or raises will do) and then pass the now typed variable in
If you get an LIT001 fail, refactor the code to follow functional programming best practices rather than introducing mutable data structures. For example, build values in one shot with comprehensions or generators wrapped in `tuple()` / `MappingProxyType()` / `frozenset()` instead of seeding an empty `list`/`dict`/`set` and mutating it over time. Ideally, `# mutable-ok` is never used; reach for it only as a genuine last resort when an immutable rewrite is truly impossible, and always pair it with a real reason
If you get an LIT001 fail, refactor the code to follow functional programming best practices rather than introducing mutable data structures. For example, build values in one shot with comprehensions or generators wrapped in `tuple()` / `MappingProxyType()` / `frozenset()` instead of seeding an empty `list`/`dict`/`set` and mutating it over time. Ideally, `# mutable-ok` is never used. Reach for it only as a genuine last resort when an immutable rewrite is truly impossible, and always pair it with a real reason
Every lint or type suppression must name the exact rule inside brackets and carry a reason comment, e.g. `# pyright: ignore[reportArgumentType] # stubs lack async overload` or `# noqa: TID251 # <reason>`. `# type: ignore` is banned (LIT009): pyrightconfig.json sets `enableTypeIgnoreComments` to false, so it silently does nothing
Commit and push your work when you're done without asking
When referencing or running models (coding, QA'ing, writing docs, writing tests, etc.), use the latest model in that model family unless otherwise specified; treat your training knowledge, memories, configs, and tests as stale, and determine the family's latest with model_prices_and_context_window.json or the web
When referencing or running models (coding, QA'ing, writing docs, writing tests, etc.), use the latest model in that model family unless otherwise specified. Treat your training knowledge, memories, configs, and tests as stale, and determine the family's latest with model_prices_and_context_window.json or the web
Always pull before starting any work. The checkout or worktree may be sitting on a stale branch
@ -84,7 +84,7 @@ Monkeypatching attributes of a class to do testing is an anti-pattern. Prefer de
Do not put names of customers or customer company names in code, PR descriptions, issue bodies, etc. This means never mention literally any company name. Especially if you're about to say a sentence mentioning that the reason the PR exists was a feature/model/bug fix/etc. requested by a company. That's the indication that you should replace that company name with "the customer". e.g. not "Model request from Acme (Pylon #1234)" but "Model request from a customer (Pylon #1234)". This is because the codebase is public. The only exception is for publicly known providers or vendors such as OpenAI, Anthropic, AWS Bedrock, etc. only IF we're adding support for that provider/vendor in general and NOT if that PR or whatnot was a request by one of them, and they're actually one of our customers.
CI supply-chain safety: Never pipe a remote script into a shell (`curl ... | bash`, `wget ... | sh`); download the artifact to a file, verify its SHA-256 checksum, then install. Pin every external tool to a specific version with a full URL (not `latest` or `stable`). Verify checksums for all downloaded binaries, using the provider's official `.sha256` / `.sha256sum` sidecar when available. These rules apply to every download in CI
CI supply-chain safety: Never pipe a remote script into a shell (`curl ... | bash`, `wget ... | sh`). Download the artifact to a file, verify its SHA-256 checksum, then install. Pin every external tool to a specific version with a full URL (not `latest` or `stable`). Verify checksums for all downloaded binaries, using the provider's official `.sha256` / `.sha256sum` sidecar when available. These rules apply to every download in CI
Prisma migrations apply synchronously at proxy boot, before it serves traffic, so a migration must only change schema, never rewrite rows. No `UPDATE`, `DELETE` or `MERGE`, and no `INSERT ... SELECT`: on a spend-log-sized table any of those is minutes of downtime plus a doubled heap that plain autovacuum won't give back. `tests/code_coverage_tests/check_migrations_no_data_rewrites.py` enforces this. When a rewrite is genuinely bounded and has to ship inside the migration, mark the statement `-- data-migration-ok: <what bounds it>`
@ -92,19 +92,19 @@ Follow these coding conventions for new/updated code (a three-line fix in a lega
- Composition over inheritance
- Never-nester: early returns over deep nesting
- Don't throw; model failures as values (One function (e.g., raise_public) maps error union to existing public exception contracts via exhaustive match + assert_never)
- No mutation; don't reassign variables, global or local. Instead of mutable lists and dicts, prefer tuples, frozen dataclasses (with slots=True), `MappingProxyType`, etc.
- Annotate every variable with `: Final` (LIT010). Unpacking and walrus targets cannot carry the annotation, so they are implicitly final. Don't rebind them. Never rebind or mutate function parameters (LIT011); `self`/`cls` attribute stores are the exception. If rebinding or in-place mutation is truly unavoidable, suppress with `# rebind-ok: <reason>`
- Don't throw. Model failures as values (One function (e.g., raise_public) maps error union to existing public exception contracts via exhaustive match + assert_never)
- No mutation. Don't reassign variables, global or local. Instead of mutable lists and dicts, prefer tuples, frozen dataclasses (with slots=True), `MappingProxyType`, etc.
- Annotate every variable with `: Final` (LIT010). Unpacking and walrus targets cannot carry the annotation, so they are implicitly final. Don't rebind them. Never rebind or mutate function parameters (LIT011). `self`/`cls` attribute stores are the exception. If rebinding or in-place mutation is truly unavoidable, suppress with `# rebind-ok: <reason>`
- Qualify every TypedDict field with `ReadOnly[...]` (LIT012), which nests freely with `Required` / `NotRequired` / `Annotated` in any order. If making the key writable is truly unavoidable, suppress with `# writable-ok: <reason>`
- Comprehensions take at most one `for` clause and one `if` clause (LIT014); split stacked clauses into a helper generator, a named intermediate, or a plain loop. Suppress with `# comprehension-ok: <reason>` only when unavoidable
- Comprehensions take at most one `for` clause and one `if` clause (LIT014). Split stacked clauses into a helper generator, a named intermediate, or a plain loop. Suppress with `# comprehension-ok: <reason>` only when unavoidable
- Every pydantic model must be frozen (LIT015), set via `model_config = ConfigDict(frozen=True)`, a dict-literal `model_config`, an inner `class Config`, or the class keywords. Subclasses inherit it unless they override it. Replace in-place field writes with `model_copy(update=...)`. If making the model mutable is truly unavoidable, suppress with `# frozen-ok: <reason>` on the `class` line
- Use dependency injection
- Fully typed; no `Any` or coarse types like `dict[str, Any]` or just `dict`. Every function parameter must be strongly typed
- Fully typed. No `Any` or coarse types like `dict[str, Any]` or just `dict`. Every function parameter must be strongly typed
- Use tagged unions + match
- No monster files or god objects
- No file sprawl: deliberate file and folder structure
- Standard over hand-rolled: use the official SDK or a library where one exists; where none does, follow industry standards instead of inventing local conventions
- API-fragmentation-aware: when logic must branch on which API surface produced or consumes data (e.g. chat completions vs Anthropic Messages vs Responses API shapes), proactively look for an existing shared helper (e.g. `litellm_core_utils/prompt_templates/factory.py`) before writing per-surface parsing in the new module; if none exists, add one there instead of duplicating the same format-detection logic in every new guardrail/integration
- Standard over hand-rolled: use the official SDK or a library where one exists. Where none does, follow industry standards instead of inventing local conventions
- API-fragmentation-aware: when logic must branch on which API surface produced or consumes data (e.g. chat completions vs Anthropic Messages vs Responses API shapes), proactively look for an existing shared helper (e.g. `litellm_core_utils/prompt_templates/factory.py`) before writing per-surface parsing in the new module. If none exists, add one there instead of duplicating the same format-detection logic in every new guardrail/integration
Follow conventional commits for commit names and PR titles
@ -130,4 +130,4 @@ Before implementing:
Ask yourself: "Would a senior engineer say this is overcomplicated?" If yes, simplify
Before requesting maintainer review, verify the current PR tip passes required CI and code coverage, meets Greptile confidence of at least 4/5, and has acceptable Veria and Bugbot reviews. Inspect warnings and findings, fix actionable issues, and rerun the affected checks and reviewers after changes. Record evidence for any false positive or unavailable review; never treat a pending or missing bot result as a pass. Do not lower coverage thresholds or raise `ANY_CAPS` to satisfy a check
Before requesting maintainer review, verify the current PR tip passes required CI and code coverage, meets Greptile confidence of at least 4/5, and has acceptable Veria and Bugbot reviews. Inspect warnings and findings, fix actionable issues, and rerun the affected checks and reviewers after changes. Record evidence for any false positive or unavailable review. Never treat a pending or missing bot result as a pass. Do not lower coverage thresholds or raise `ANY_CAPS` to satisfy a check

View file

@ -4,7 +4,7 @@ For diagnostic tracing changes, follow [.agents/skills/rust-tracing/SKILL.md](.a
For string-valued enums and their Serde conversions, follow [.agents/skills/rust-string-enums/SKILL.md](.agents/skills/rust-string-enums/SKILL.md)
For fieldless enums, derive `strum::VariantArray` and use `VARIANTS` instead of a hand-listed `ALL` array; derive Strum string conversions instead of hand-written variant-to-string matches
For fieldless enums, derive `strum::VariantArray` and use `VARIANTS` instead of a hand-listed `ALL` array. Derive Strum string conversions instead of hand-written variant-to-string matches
Use the derived conversions directly (`<&'static str>::from(x)` / `.into()`, `str::parse`) with no `as_str`/`parse` wrapper that only delegates, and spell each variant with explicit `#[strum(serialize = "...")]` instead of `serialize_all`
@ -15,20 +15,20 @@ Use the derived conversions directly (`<&'static str>::from(x)` / `.into()`, `st
- A test that only uses the crate's public API lives in `crates/<crate>/tests/<subject>.rs`, next to `src/`
- Split a mixed test file along that line instead of widening visibility to move it
- A test for another crate's item belongs in that crate, not in a downstream one
- Never set `autotests = false` or hand-list `[[test]]` targets; every file directly under `tests/` is discovered by cargo, and a shared helper goes in `tests/<name>/mod.rs` or `tests/<subject>/support.rs` so it is not picked up as a test crate of its own
- Never set `autotests = false` or hand-list `[[test]]` targets. Every file directly under `tests/` is discovered by cargo, and a shared helper goes in `tests/<name>/mod.rs` or `tests/<subject>/support.rs` so it is not picked up as a test crate of its own
## Test fixtures and cases
Use [`#[rstest]`](https://docs.rs/rstest/latest/rstest/attr.rstest.html) for new and updated tests and [`#[fixture]`](https://docs.rs/rstest/latest/rstest/attr.fixture.html) for reusable setup, injected through typed test arguments. Express input variations as named `#[case::name(...)]` cases instead of loops or duplicated tests so each failure identifies its case. Keep behavior assertions in the test body and fixtures focused on setup. Use the workspace `rstest` dependency
- Never loop over inputs (`for`, `.iter().for_each`, `.all`) inside a test body; give each input its own `#[case::name(...)]`, or use `#[values(...)]` for a cross product
- Never loop over inputs (`for`, `.iter().for_each`, `.all`) inside a test body. Give each input its own `#[case::name(...)]`, or use `#[values(...)]` for a cross product
- Exception: a test pinning a Rust table against a repo-owned data file (for example `include_str!` of a JSON config) may iterate that file's entries
## Error definitions
- A crate's errors live in `src/error.rs`, defined with `thiserror`, and re-exported from `lib.rs`
- Put message templates in the variant's `#[error(...)]` declaration. Callers pass only the small typed arguments needed to fill them, never `Error::Variant(format!(...))` or a preformatted message. Keep the smallest set of neutral variants that callers need to distinguish; different wording or providers do not justify new variants
- Default to one top-level `Error` enum per crate, with one variant per failure mode and a `#[error(...)]` message on each. A failure mode is something a caller handles differently (phase, status code, retry, a message Python parity pins exactly); failures no caller tells apart share one variant and differ only in its message
- Put message templates in the variant's `#[error(...)]` declaration. Callers pass only the small typed arguments needed to fill them, never `Error::Variant(format!(...))` or a preformatted message. Keep the smallest set of neutral variants that callers need to distinguish. Different wording or providers do not justify new variants
- Default to one top-level `Error` enum per crate, with one variant per failure mode and a `#[error(...)]` message on each. A failure mode is something a caller handles differently (phase, status code, retry, a message Python parity pins exactly). Failures no caller tells apart share one variant and differ only in its message
- Keep shared error enums minimal and provider-neutral. Provider names, credential types, configuration fields, and setup guidance belong in caller-supplied data, not dedicated variants or hardcoded shared messages. Reuse a variant for the same failure mode across providers, such as `MissingApiBase { provider: "Azure", guidance: "..." }`. An exact parity message does not justify a provider-specific variant when caller-supplied context can preserve it
- Wrap a lower-level error as a variant with `#[from]` or `#[source]` instead of flattening it to a string
- Exception: split into separate types when different functions fail in disjoint ways, especially when different callers see them. A shared enum would force every caller to match variants its function can never return

View file

@ -24,6 +24,6 @@ Keep unary caching independent of stream-only methods. Store streams only after
Test each contract in its owner: storage capabilities in backend tests, envelopes and freshness here, reuse and replay in core, Python callback and fallback behavior at the bridge, and HTTP behavior at the gateway. Run backend contract checks and Python response-codec fixtures before exposing a new backend
`ScopedCache` requires an explicit shared or isolated scope at construction. Per-call `CachePolicy` controls reads, writes, expiry, and freshness without replacing the attached scope or service. `CacheOptions` binds that policy to an explicit scope for storage requests and has no default sharing policy. Versioned native envelopes reject incompatible API surfaces and versions as misses; this envelope is distinct from the legacy Python response codec
`ScopedCache` requires an explicit shared or isolated scope at construction. Per-call `CachePolicy` controls reads, writes, expiry, and freshness without replacing the attached scope or service. `CacheOptions` binds that policy to an explicit scope for storage requests and has no default sharing policy. Versioned native envelopes reject incompatible API surfaces and versions as misses. This envelope is distinct from the legacy Python response codec
Response storage is not the source of budget or rate-limit coordination dependencies. Keep counters, reservations, and atomic admission operations out of `ResponseCacheService`, including when both services happen to use Redis

View file

@ -1,20 +1,20 @@
- Target invariants, not completion claims
- This crate owns compatibility for all existing Python callbacks and loggers, including `CustomLogger`. `mapping.rs` owns the executable call bindings and the inventory of Python-owned hooks. A Python-owned entry records an existing path, never permission to invoke it a second time. The native call adapter preserves the `Logging` contract (`function_setup`, the deployment hooks, `pre_call`/`post_call`, the sync and async success and failure fan-out, the deferred proxy release, the argument sharing those callbacks rely on)
- Smell test: if a future callback host (`callbacks-v1-python`, WASM, in-process Rust) could share a piece of this crate, it does not belong here
- SDK request policy (credential inheritance, the budget and retry-count limits) is a separate hook supplied by `python-bridge`; compose it after this adapter so logging adopts the final keyword view before policy mutates or rejects it
- The driver in `litellm-host-python`, the routes and core see one `PythonCallHooks` using the shared `CallEvent`; they never learn which Python objects consume a call
- SDK request policy (credential inheritance, the budget and retry-count limits) is a separate hook supplied by `python-bridge`. Compose it after this adapter so logging adopts the final keyword view before policy mutates or rejects it
- The driver in `litellm-host-python`, the routes and core see one `PythonCallHooks` using the shared `CallEvent`. They never learn which Python objects consume a call
- Every litellm Python internal Rust still borrows is a variant of `LegacyPython`, grouped by subsystem, with its signature pinned in `python_contract.json`
- The enum only shrinks: when Rust owns a subsystem, delete its group rather than adding a Rust path beside it
- Calling a user's own callback directly is permanent Python surface and gets its own type outside `LegacyPython`
- `PublicCall` is the caller's call as `Logging` sees it: the positional arguments, the keyword view as the call rewrites it (setup, deployment hook, preflight) and the bound request object backing omitted keywords; shared bridge composition hands it to `LegacyLogging`; routes use the neutral call boundary
- `PublicCall` is the caller's call as `Logging` sees it: the positional arguments, the keyword view as the call rewrites it (setup, deployment hook, preflight) and the bound request object backing omitted keywords. Shared bridge composition hands it to `LegacyLogging`. Routes use the neutral call boundary
- `LoggingOperation` selects legacy logging entrypoints and response handling. It belongs here rather than in shared inference data contracts
- `setup` reuses a `Logging` passed as `litellm_logging_obj` (the proxy and Router) and otherwise builds one through `function_setup`; which callbacks run is `Logging`'s decision, never this crate's
- Callbacks receive the caller's own objects and may mutate them; this crate alone carries that obligation
- Retain complete boundary arguments, opaque values, aliases, omitted/default distinctions and deliberate copies; preserve the deployment-hook kwargs view
- `setup` reuses a `Logging` passed as `litellm_logging_obj` (the proxy and Router) and otherwise builds one through `function_setup`. Which callbacks run is `Logging`'s decision, never this crate's
- Callbacks receive the caller's own objects and may mutate them. This crate alone carries that obligation
- Retain complete boundary arguments, opaque values, aliases, omitted/default distinctions and deliberate copies. Preserve the deployment-hook kwargs view
- Before `pre_call`, re-alias every body key whose value equals the caller's argument to the caller's own object, resolved through `litellm_host_python::lookup`
- Retain body/header roots from `pre_call` to `post_call`; in-place mutation reaches the wire, envelope field replacement is visible to later callbacks only
- Retain body/header roots from `pre_call` to `post_call`. In-place mutation reaches the wire, envelope field replacement is visible to later callbacks only
- Success and failure handlers receive the exact selected public response or exception
- A failure-handler error cannot suppress the other eligible family or replace the mapped provider error; a cancellation ends the call with no further dispatch
- Dispatch errors never replay provider work or trigger the opposite outcome; the proxy releases deferred success at most once
- A failure-handler error cannot suppress the other eligible family or replace the mapped provider error. A cancellation ends the call with no further dispatch
- Dispatch errors never replay provider work or trigger the opposite outcome. The proxy releases deferred success at most once
- Delivery follows the registry, not the callable's type: direct, awaited, executor-submitted, logging-worker and deferred paths stay distinct
- Traverse every retained Python edge; `close` is idempotent and restores the correlation context once
- Traverse every retained Python edge. `close` is idempotent and restores the correlation context once

View file

@ -1,9 +1,9 @@
# Requirements
Core must pause mid-call to ask the host for things it cannot do itself (Python callbacks, secret and token reads, `before_send` rewrites, stream demand), then continue where it stopped. Any change to this crate must keep every requirement below; the alternatives section says which one each rejected design breaks
Core must pause mid-call to ask the host for things it cannot do itself (Python callbacks, secret and token reads, `before_send` rewrites, stream demand), then continue where it stopped. Any change to this crate must keep every requirement below. The alternatives section says which one each rejected design breaks
- R1 Core never calls the host: it names an op and waits for the answer, so it stays free of PyO3 and of any other host runtime
- R2 Async host work is awaited by the host's own driver in the caller's asyncio task (`litellm/rust_bridge/lifecycle.py`), so `contextvars` writes reach the caller; a Rust-side `into_future` would run it in a copied context
- R2 Async host work is awaited by the host's own driver in the caller's asyncio task (`litellm/rust_bridge/lifecycle.py`), so `contextvars` writes reach the caller. A Rust-side `into_future` would run it in a copied context
- R3 The body awaits real I/O (HTTP, `spawn_blocking`, timers) between yields, so `resume` is itself a future driven by the caller's runtime
- R4 Route code stays straight-line async (`host.route(OcrOp::ReadDocument).await?`) instead of hand-written states
- R5 Each op fixes its answer type at compile time: a host cannot answer `ReadDocument` with a token, and core never matches a result variant it did not ask for
@ -20,7 +20,7 @@ Core must pause mid-call to ask the host for things it cannot do itself (Python
- `corosensei` and other stackful coroutines: sync bodies on their own stack, no async I/O inside (R3)
- A hand-written phase enum with an `advance` match (the old `HostPhase`): every await point becomes a state (R4)
- An injected host trait with `async fn`s: core would call the host itself (R1, R2)
- Sans-IO, where core does no I/O and HTTP becomes one more host op: keeps every requirement and makes `resume` a pure step function, but HTTP, streaming, retries and timeouts would move out of core into every bridge; the one real alternative, not taken
- Sans-IO, where core does no I/O and HTTP becomes one more host op: keeps every requirement and makes `resume` a pure step function, but HTTP, streaming, retries and timeouts would move out of core into every bridge. The one real alternative, not taken
- Temporal's Rust workflow SDK (`WorkflowFuture`, `WfContext`) is the closest precedent: an `async fn` polled in place, commands sent over a channel with a oneshot to unblock them. Roles are inverted there (the language SDK owns the program, core answers), and its workflow body may not do real I/O
# Tradeoffs accepted

View file

@ -2,9 +2,9 @@
Inbound authentication uses a verifier, identity resolver, and authorizer injected into `Auth`. `Auth::from_config` composes the current master-key implementation. `Auth::new` accepts alternate implementations and an injected clock
`CredentialExtractor` selects a token or transport credential from HTTP request parts. The default `Bearer` extractor is strict; a host can supply another extractor with `Auth::with_extractor` for a separate route profile. A transport credential is only a selection marker, never proof of identity
`CredentialExtractor` selects a token or transport credential from HTTP request parts. The default `Bearer` extractor is strict. A host can supply another extractor with `Auth::with_extractor` for a separate route profile. A transport credential is only a selection marker, never proof of identity
`Authenticator` receives that selected credential and HTTP request parts. Transport verifiers must obtain proof from trusted server extensions, not client-supplied identity headers. It returns `VerifiedIdentity` only after verification. Request parts provide headers and trusted transport extensions for integrations; the verifier must validate any client-controlled claims before treating them as identity. There is no fallback chain after verification failure
`Authenticator` receives that selected credential and HTTP request parts. Transport verifiers must obtain proof from trusted server extensions, not client-supplied identity headers. It returns `VerifiedIdentity` only after verification. Request parts provide headers and trusted transport extensions for integrations. The verifier must validate any client-controlled claims before treating them as identity. There is no fallback chain after verification failure
`IdentityResolver` maps verified identity into an internal principal and permissions. The caller retains the original authentication evidence and credential restrictions independently of that mapping. Principals include an authority, subject, and kind. External subjects must remain authority-scoped unless the resolver explicitly maps them to a shared internal identity
@ -14,13 +14,13 @@ Inbound authentication uses a verifier, identity resolver, and authorizer inject
The Axum `authenticate` middleware currently accepts one Authorization header using the existing Bearer format. Missing, duplicate, empty, or malformed credentials fail. Authentication replaces any preexisting caller extension and checks the method and matched route before dispatch. Handlers extract `AuthenticatedRequest` and authorize their parsed operation before calling a provider. Missing authenticated context fails closed
The gateway also inserts an MCP `Authorization` extension and a `SessionOwner` derived from the verified principal and credential reference. The MCP host checks the injected policy at initialization and before sending resolved upstream operations. Standalone MCP library hosts retain their existing explicitly trusted registry behavior when no policy is installed; production mounts must install authentication and the policy extension together
The gateway also inserts an MCP `Authorization` extension and a `SessionOwner` derived from the verified principal and credential reference. The MCP host checks the injected policy at initialization and before sending resolved upstream operations. Standalone MCP library hosts retain their existing explicitly trusted registry behavior when no policy is installed. Production mounts must install authentication and the policy extension together
Local UI login continues using axum-login and tower-sessions. Only after session and CSRF validation does `UiSession` expose an authenticated caller. Its credentials are restricted to session-info and logout operations, so the UI CSRF bearer cannot authorize inference. Future UI management routes must extend that explicit scope
Authentication evidence and principals contain no raw token, password, request body, or mutable accounting state. Session ownership is scoped by principal authority, subject, verifier, and credential ID, so separate credentials do not silently share an MCP session. Credential rotation may retain ownership when the verifier preserves a stable credential ID. Scope and expiry checks still run for each operation
Failures distinguish invalid or expired credentials, forbidden operations, unavailable authentication services, and missing server configuration/context. HTTP adapters map those outcomes to status codes; inference keeps its API-specific error envelopes and MCP translates policy failures to its own protocol
Failures distinguish invalid or expired credentials, forbidden operations, unavailable authentication services, and missing server configuration/context. HTTP adapters map those outcomes to status codes. Inference keeps its API-specific error envelopes and MCP translates policy failures to its own protocol
Virtual-key storage, JWT verification, OAuth2 introspection, trusted-proxy validation, SSO, and custom Python hook adapters are not implemented here yet. They should supply these contracts rather than bypassing the shared authorization boundary. Request-body-dependent custom hooks will need a bounded endpoint adapter after parsing

View file

@ -1,6 +1,6 @@
- Expose a mountable Axum router; listener binding, server lifecycle, and shared inbound middleware belong to `gateway`
- Own endpoint paths, request parsing, model alias resolution, and API-specific response and SSE error formats; delegate hosted call execution and HTTP body delivery to host-http
- Delegate inference execution to the `inference-<fmt>` crates and provider transformations and authentication to `llms` and the auth crates; do not duplicate them in handlers
- Let the `inference-<fmt>` crates validate inference fields and supported features, then map their errors to HTTP responses; do not add gateway checks for temporary inference limitations
- Use injected deployments, HTTP pools, settings, and secret sources; do not load process configuration or construct independent clients in handlers
- Test HTTP contracts here, including status codes, forwarded headers, error envelopes, and streaming behavior; keep inference and provider tests in their owning crates
- Expose a mountable Axum router. Listener binding, server lifecycle, and shared inbound middleware belong to `gateway`
- Own endpoint paths, request parsing, model alias resolution, and API-specific response and SSE error formats. Delegate hosted call execution and HTTP body delivery to host-http
- Delegate inference execution to the `inference-<fmt>` crates and provider transformations and authentication to `llms` and the auth crates. Do not duplicate them in handlers
- Let the `inference-<fmt>` crates validate inference fields and supported features, then map their errors to HTTP responses. Do not add gateway checks for temporary inference limitations
- Use injected deployments, HTTP pools, settings, and secret sources. Do not load process configuration or construct independent clients in handlers
- Test HTTP contracts here, including status codes, forwarded headers, error envelopes, and streaming behavior. Keep inference and provider tests in their owning crates

View file

@ -1,5 +1,5 @@
- Keep this crate a thin composition layer: mount endpoint routers and serve the supplied listener
- Server lifecycle and shared inbound middleware belong here, including client authentication, rate limiting, and request logging
- Endpoint paths, request handling, model resolution, and response encoding belong to the mounted crates; provider execution belongs to the `inference-<fmt>` crates and `llms`
- Inject shared state and infrastructure; avoid global runtimes, duplicate client pools, and abstractions for hypothetical endpoint groups
- Test mounting and server lifecycle through public HTTP behavior; test endpoint semantics in the owning crate
- Endpoint paths, request handling, model resolution, and response encoding belong to the mounted crates. Provider execution belongs to the `inference-<fmt>` crates and `llms`
- Inject shared state and infrastructure. Avoid global runtimes, duplicate client pools, and abstractions for hypothetical endpoint groups
- Test mounting and server lifecycle through public HTTP behavior. Test endpoint semantics in the owning crate

View file

@ -1,7 +1,7 @@
`litellm-host-native` is the Rust driver for hosted calls. `Driver` owns the machine, a `HostCallHandler` and a `Interceptors`; `advance()` answers services and hooks inline and returns at completion or at the next stream boundary, holding the `Reply<ControlFlow<()>>` until the consumer calls `advance()` or `detach()` again. Dropping the driver drops the machine and so cancels the call
`litellm-host-native` is the Rust driver for hosted calls. `Driver` owns the machine, a `HostCallHandler` and a `Interceptors`. `advance()` answers services and hooks inline and returns at completion or at the next stream boundary, holding the `Reply<ControlFlow<()>>` until the consumer calls `advance()` or `detach()` again. Dropping the driver drops the machine and so cancels the call
The consumer decides demand, so the driver never spawns a producer task and never buffers chunks ahead of demand. `litellm-host-http` polls it from the response body; `in_process::run_hosted` polls it on behalf of a `StreamConsumer`. Both observe lifecycle terminals themselves, the driver reports none
The consumer decides demand, so the driver never spawns a producer task and never buffers chunks ahead of demand. `litellm-host-http` polls it from the response body. `in_process::run_hosted` polls it on behalf of a `StreamConsumer`. Both observe lifecycle terminals themselves, the driver reports none
Depend on `litellm-host` only. HTTP encoding stays in `litellm-host-http`; `litellm-host-python` drives the machine directly so Python callbacks stay in the caller's asyncio task
Depend on `litellm-host` only. HTTP encoding stays in `litellm-host-http`. `litellm-host-python` drives the machine directly so Python callbacks stay in the caller's asyncio task
`services.rs` owns `HostCallHandler` and its borrowed and no-service implementations. This is the Rust driver's handler contract; the shared host crate owns the service request protocol
`services.rs` owns `HostCallHandler` and its borrowed and no-service implementations. This is the Rust driver's handler contract. The shared host crate owns the service request protocol

View file

@ -1,4 +1,4 @@
- Target invariants; implementation and runtime validation may lag these rules
- Target invariants. Implementation and runtime validation may lag these rules
## Boundary with Python consumers
@ -6,15 +6,15 @@ This crate owns CPython execution mechanics for generic `litellm-host` machines
`src/native.rs` belongs here: it runs a generic machine through the Python runtime and owns its pending execution and abort handle. Keep provider selection, request projection and public exception policy out of it. A rename to `machine_runner.rs` is optional and must not change behavior
The execution handle receives its Python lifecycle binding from its consumer through `PythonLifecycle` rather than import a fixed `litellm.rust_bridge` module. Generic suspension and execution state validation belong here; public stream wrappers and `_hidden_params` conventions belong to the consumer
The execution handle receives its Python lifecycle binding from its consumer through `PythonLifecycle` rather than import a fixed `litellm.rust_bridge` module. Generic suspension and execution state validation belong here. Public stream wrappers and `_hidden_params` conventions belong to the consumer
Creating a resolved asyncio Future from an already constructed Python value belongs here, alongside runtime waiting, interpreter detachment and panic containment. Choosing which callable exceptions become a public `RuntimeError` belongs to the consumer; `python-bridge::callable::wrap_failure` owns that policy
Creating a resolved asyncio Future from an already constructed Python value belongs here, alongside runtime waiting, interpreter detachment and panic containment. Choosing which callable exceptions become a public `RuntimeError` belongs to the consumer. `python-bridge::callable::wrap_failure` owns that policy
The driver owns ordering: start, argument preparation, prepared-argument hooks, binding decode and machine start. Fallible per-call resource setup supplied by the consumer runs after all argument hooks, using the prepared argument view, and before provider work. Setup failure follows the existing terminal failure path. Creating or discarding an unstarted coroutine must not initialize clients, acquire credentials or capture execution context
Boundary tests exercise behavior with a supplied lifecycle binding without importing the LiteLLM Python package. Pin inline awaiting, awaitable final values, exception identity, cancellation, re-entry and release of retained objects, rather than module names or source layout
`HookChain` composes Python runtime hooks in order. Each argument, wire-request and response transformation feeds its result to the next hook. After all argument transformations, the driver calls `arguments_prepared` on every hook in order. Retained callback views must adopt that dictionary before later policy hooks can mutate or reject it. SDK policy is supplied by bridge composition as a hook, never a separate driver phase or parameter. Hooks implement only the stages they need; default stages preserve the supplied values
`HookChain` composes Python runtime hooks in order. Each argument, wire-request and response transformation feeds its result to the next hook. After all argument transformations, the driver calls `arguments_prepared` on every hook in order. Retained callback views must adopt that dictionary before later policy hooks can mutate or reject it. SDK policy is supplied by bridge composition as a hook, never a separate driver phase or parameter. Hooks implement only the stages they need. Default stages preserve the supplied values
Terminal notifications share the selected response or exception. An ordinary notification error is reported as unraisable and does not skip the next hook or replace the selected outcome. Preparation, interception and transformation errors stop the chain. Cancellation stops all further hook dispatch. Suspensions stay inline in the existing driver, and the chain traverses retained event values for GC
@ -23,21 +23,21 @@ Terminal notifications share the selected response or exception. An ordinary not
- Keep this crate the CPython runtime adapter and nothing more: Serde marshalling, interpreter detachment, tokio/asyncio glue, the `Execution` handle, the call driver and the `PythonBinding`, `PythonHostCalls` and `PythonOwned` traits, and the `PythonRuntime` specialization of `host::hooks::CallHooks`
- No LiteLLM domain dependencies beyond `litellm-host`: no route types, no `Logging` policy, no public API registration, no cdylib build features
- `PythonCallHooks` only constrains the shared call-stage interface to `PythonRuntime` and Python ownership. It must not redeclare the stages
- `PythonCallEvent` is a specialization of the shared `CallEvent`, never a separately defined lifecycle. The driver emits `Succeeded` or `Failed` exactly once and never dispatches callbacks after a cancellation; which Python objects consume those events is the legacy adapter's business
- `PythonCallEvent` is a specialization of the shared `CallEvent`, never a separately defined lifecycle. The driver emits `Succeeded` or `Failed` exactly once and never dispatches callbacks after a cancellation. Which Python objects consume those events is the legacy adapter's business
- `CallOptions` can publish snapshots independently of callback delivery. Terminal snapshots follow completed hook dispatch, and cancellation never calls a Python callback. For Python-driven calls, leave the machine's observation publisher unset so the driver is the sole publisher of intercepted provider-response snapshots
- `PythonBinding::decode_request` receives the keyword view returned by `prepare_arguments` and updated by `arguments_prepared`, not the caller's dict; a binding that decodes from it inherits all composed argument rewrites
- A native failure, including one a host op returns as `InvokeError::Native`, is classified exactly once through the binding's `map_error`; a Python exception raised inside the call, and a failure in `prepare_arguments` or `transform_response`, is raised as is
- `PythonBinding::decode_request` receives the keyword view returned by `prepare_arguments` and updated by `arguments_prepared`, not the caller's dict. A binding that decodes from it inherits all composed argument rewrites
- A native failure, including one a host op returns as `InvokeError::Native`, is classified exactly once through the binding's `map_error`. A Python exception raised inside the call, and a failure in `prepare_arguments` or `transform_response`, is raised as is
- A failing `map_error` is raised with the native error's text as its `__context__`, never swallowed
- Use standard PyO3 ownership and conversion APIs
- Prefer `Bound<'py, T>` for attached operations/results, `Py<T>` for retention; binding/unbinding does not copy payloads
- Use `pythonize` for selected Serde data, never a JSON-text round trip; share conversion with `Pythonized<T>`
- Preserve `PythonizeError`'s standard conversion into `PyErr`; do not stringify original Python exceptions into new `ValueError`s
- Prefer `Bound<'py, T>` for attached operations/results, `Py<T>` for retention. Binding/unbinding does not copy payloads
- Use `pythonize` for selected Serde data, never a JSON-text round trip. Share conversion with `Pythonized<T>`
- Preserve `PythonizeError`'s standard conversion into `PyErr`. Do not stringify original Python exceptions into new `ValueError`s
- Keep serializer-panic containment in `Pythonized<T>`: async output conversion can run in an unjoined blocking task and otherwise strand delivery
- Use `Python::detach` for Rust-only work; Python operations require attachment
- Keep diagnostic counters in the consumer; wrapper invocations do not measure every interpreter release
- Release exclusive class borrows/locks before Python calls or decrements that can invoke finalizers; expose retained Python edges to GC without calling Python during traversal
- Use `Python::detach` for Rust-only work. Python operations require attachment
- Keep diagnostic counters in the consumer. Wrapper invocations do not measure every interpreter release
- Release exclusive class borrows/locks before Python calls or decrements that can invoke finalizers. Expose retained Python edges to GC without calling Python during traversal
- Keep coroutine driving in the shared Python driver and the native handle
- Shared driver implementation: `litellm/rust_bridge/lifecycle.py`; handle: `src/handle.rs`; call driver: `src/driver.rs`; native-backed behavior tests: `tests/lifecycle.py`. The consumer supplies the lifecycle binding
- Every lifecycle suspension is awaited inline in the caller's task; `into_future` creates a separate task and cannot satisfy this contract
- Shared driver implementation: `litellm/rust_bridge/lifecycle.py`, handle: `src/handle.rs`, call driver: `src/driver.rs`, native-backed behavior tests: `tests/lifecycle.py`. The consumer supplies the lifecycle binding
- Every lifecycle suspension is awaited inline in the caller's task. `into_future` creates a separate task and cannot satisfy this contract
- References: [ownership](https://pyo3.rs/v0.29.2/types.html), [conversions](https://pyo3.rs/v0.29.2/conversions/traits.html), [pythonize errors](https://docs.rs/pythonize/0.29.0/src/pythonize/error.rs.html)
- [GC](https://pyo3.rs/v0.29.2/class/protocols.html#garbage-collector-integration), [re-entry](https://pyo3.rs/v0.29.2/class/call.html), [parallelism](https://pyo3.rs/v0.29.2/parallelism.html), [async delivery source](https://docs.rs/pyo3-async-runtimes/0.29.0/src/pyo3_async_runtimes/generic.rs.html)

View file

@ -2,28 +2,28 @@
| Responsibility | Shared contract | HTTP | Python |
| --- | --- | --- | --- |
| Input and output conversion | `Protocol::{Request, Response, StreamHead, Chunk, Error}` | Typed input; `ResponseEncoder` and `StreamEncoder` produce HTTP values | `PythonBinding` decodes prepared arguments, encodes public values and maps native errors |
| Host services | `Protocol::HostCall`, `HostServices::call` | `host-native::services::HostCallHandler` answers typed calls; `()` handles protocols without host calls | `PythonHostCalls` invokes retained Python objects; it may share an owner with the binding |
| Input and output conversion | `Protocol::{Request, Response, StreamHead, Chunk, Error}` | Typed input. `ResponseEncoder` and `StreamEncoder` produce HTTP values | `PythonBinding` decodes prepared arguments, encodes public values and maps native errors |
| Host services | `Protocol::HostCall`, `HostServices::call` | `host-native::services::HostCallHandler` answers typed calls. `()` handles protocols without host calls | `PythonHostCalls` invokes retained Python objects. It may share an owner with the binding |
| Active hooks | `Interceptors::{before_provider_request, after_provider_response}` and `hooks::CallHooks<Runtime>` | Request interception and fallible execution callbacks | `PythonCallHooks` also prepares arguments, transforms public responses and receives stream callbacks |
| Passive observation | `observation::ObservationSender` | Queued execution and lifecycle snapshots, retained by the response body | Public Python callbacks remain active hooks with their existing failure policy |
| Runtime driving | `Machine`, `HostRequest::{HostCall, Intercept, Stream}` | Body polling controls demand | Native polling and inline caller-task Python awaits control progress |
`hosted_call(request, observers, execute)` starts with a typed request. Its route closure receives separate `HostServices`, `ChannelInterceptors` and optional observation publisher; it returns `CallOutput`. Hosted-call plumbing alone forwards the returned stream through demand replies. Lower-level `CallMachine` users receive a `CallContext` containing separately named services, interceptors, observers and stream delivery
`hosted_call(request, observers, execute)` starts with a typed request. Its route closure receives separate `HostServices`, `ChannelInterceptors` and optional observation publisher. It returns `CallOutput`. Hosted-call plumbing alone forwards the returned stream through demand replies. Lower-level `CallMachine` users receive a `CallContext` containing separately named services, interceptors, observers and stream delivery
Core route constructors prepare their dependencies and return a closure accepting the typed request. Python starts that closure only after argument preparation, preflight and decoding succeed. These steps remain inside the driver's terminal and error handling. Decoding may retain objects for subsequent host service calls
Each driver owns terminal dispatch. Hooks can change or fail execution; passive observers return no result. HTTP observes success after response conversion or stream exhaustion, failure on errors, and cancellation on body drop. Python preserves exception identity and maps native failures once. Explicit Python stream close reports success for delivered chunks; cancellation stops further callback dispatch
Each driver owns terminal dispatch. Hooks can change or fail execution. Passive observers return no result. HTTP observes success after response conversion or stream exhaustion, failure on errors, and cancellation on body drop. Python preserves exception identity and maps native failures once. Explicit Python stream close reports success for delivered chunks. Cancellation stops further callback dispatch
`interceptors.rs` owns `Interceptors` and its request/response payload types. `lifecycle.rs` owns `CallObserver`, `CallEvent`, `ExecutionEvent`, timing, failure origin, and the observation wrappers. Event payloads are generic so a runtime can retain its own response, exception and raw-response references without introducing a language dependency. `snapshot()` projects them into the owned observation contract without retaining runtime objects. Pass interceptors and observers separately at direct route and HTTP entrypoints. Routes publish execution events independently of interception
`hooks.rs` owns the call-stage interface and its runtime-associated types. It contains no Python types or legacy callback policy. A runtime supplies its context and continuation representation through `HookRuntime`
`protocol.rs` owns `Protocol` and suspension messages, including `InterceptRequest` and `StreamDelivery`. `call.rs` owns route outputs and their adaptation into a hosted machine. Rust service handling belongs in `host-native::services`; coroutine channel handles stay in `machine/context.rs`
`protocol.rs` owns `Protocol` and suspension messages, including `InterceptRequest` and `StreamDelivery`. `call.rs` owns route outputs and their adaptation into a hosted machine. Rust service handling belongs in `host-native::services`. Coroutine channel handles stay in `machine/context.rs`
Rust handlers answer suspensions through `litellm-host-native::Driver`, which `litellm-host-http` and `litellm_host_native::in_process` share. `in_process::Host` is an assembly of services, interceptors, stream consumer and optional observation publisher. It is not a trait mirroring every suspension. Use `run_hosted` to preserve the distinction between stream completion and detachment
Keep API policy in gateway-inference and python-bridge, and legacy callback policy in callbacks-legacy-python. Python bindings and hooks expose retained references through `PythonOwned`, with idempotent close and GC traversal. Runtime machinery stays in driver, native, handle and runtime modules
Interceptors run inline and can rewrite values or fail execution. Observers consume owned `CallEvent` snapshots from `observation_channel`; its bounded `ObservationSender` never waits for delivery and counts events dropped when the queue is full or closed. The host owns receiver processing and draining. Pass the same publisher to machine construction and the driver when one receiver should collect execution and lifecycle events. Legacy Python callbacks retain their existing awaited, fallible behavior through the Python adapter
Interceptors run inline and can rewrite values or fail execution. Observers consume owned `CallEvent` snapshots from `observation_channel`. Its bounded `ObservationSender` never waits for delivery and counts events dropped when the queue is full or closed. The host owns receiver processing and draining. Pass the same publisher to machine construction and the driver when one receiver should collect execution and lifecycle events. Legacy Python callbacks retain their existing awaited, fallible behavior through the Python adapter
`ExecutionFacts` and `ResultSource` describe execution without pricing or budget policy. `Interceptors::result_ready` delivers these facts through an awaited `InterceptRequest::ResultReady`; hosts receive them before response transformation or stream delivery. `ExecutionEvent::ResultReady` is the matching lifecycle event and can also be published as a passive snapshot. Accounting must consume the awaited path rather than a lossy observation queue
`ExecutionFacts` and `ResultSource` describe execution without pricing or budget policy. `Interceptors::result_ready` delivers these facts through an awaited `InterceptRequest::ResultReady`. Hosts receive them before response transformation or stream delivery. `ExecutionEvent::ResultReady` is the matching lifecycle event and can also be published as a passive snapshot. Accounting must consume the awaited path rather than a lossy observation queue

View file

@ -19,7 +19,7 @@ Credential acquisition contracts and reusable adapters belong in `litellm-auth-t
The coroutine polls the route future until it completes or yields a `HostRequest`. Each request carries a typed `Reply` that its driver must answer before resuming, or abandon when interrupting or dropping the execution. A pending network future is an ordinary async wait, not a host suspension
`CallContext` gives the route separate capabilities: `HostServices` requests host operations, `ChannelInterceptors` requests interception, `ObservationSender` publishes events, and `StreamSender` delivers stream values. Keep their yield-and-reply mechanics in `context.rs`. The actual service and hook implementations belong to the host. This follows the effect-handler pattern: the route requests an operation, the driver handles it, and the route continues with the reply. The suspended computation stays in the coroutine; `Reply` only supplies its result
`CallContext` gives the route separate capabilities: `HostServices` requests host operations, `ChannelInterceptors` requests interception, `ObservationSender` publishes events, and `StreamSender` delivers stream values. Keep their yield-and-reply mechanics in `context.rs`. The actual service and hook implementations belong to the host. This follows the effect-handler pattern: the route requests an operation, the driver handles it, and the route continues with the reply. The suspended computation stays in the coroutine. `Reply` only supplies its result
Stream replies use `std::ops::ControlFlow<()>`. `ControlFlow::Continue(())` permits stream execution to continue. `ControlFlow::Break(())` tells it that the consumer stopped reading. Holding the reply applies backpressure until the consumer advances. Keep stream forwarding and the distinction between stream exhaustion and detachment in `crate::call::hosted_call`
@ -29,4 +29,4 @@ Stream replies use `std::ops::ControlFlow<()>`. `ControlFlow::Continue(())` perm
Rust drivers live in `litellm-host-native` and `litellm-host-http`. The Python driver lives in `litellm-host-python` and awaits Python hooks in the caller's task. Keep runtime scheduling, encoding, terminal observation, and callback policy in those layers and their adapters
Public behavior tests belong in `crates/host/tests`; private behavior tests stay inline with their owning implementation. Test suspension answers, interruption, resource release, and stream demand through behavior, rather than asserting file layout
Public behavior tests belong in `crates/host/tests`. Private behavior tests stay inline with their owning implementation. Test suspension answers, interruption, resource release, and stream demand through behavior, rather than asserting file layout

View file

@ -1,3 +1,3 @@
- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md)
- This crate owns Chat Completions call orchestration; payloads belong in `litellm-llms-types`, transformations in `llms/src/base_llm/chat` and `llms/src/<provider>/chat`
- This crate owns Chat Completions call orchestration. Payloads belong in `litellm-llms-types`, transformations in `llms/src/base_llm/chat` and `llms/src/<provider>/chat`
- Chat Completions is currently non-streaming: the route returns its completed response, cached through `litellm_inference::caching::execute_unary`

View file

@ -7,8 +7,8 @@
- delegate authentication policy, beta selection, payload rewriting, and response interpretation to those adapters
- keep provider policy out of request preparation and transport handlers
- calling a concrete provider helper for every provider is still a policy dependency
- Route types such as `MessagesCall`, prepared requests, and response wrappers containing live streams describe execution; reuse the shared Messages payload types inside them instead of defining another request or response schema here
- Route types such as `MessagesCall`, prepared requests, and response wrappers containing live streams describe execution. Reuse the shared Messages payload types inside them instead of defining another request or response schema here
- The route returns `litellm_host::call::CallOutput`: a completed response, or a stream head and chunks
- Per-call dependencies are grouped in `litellm_inference::context::CallContext`; `src/lib.rs` explicitly sequences cache lookup, provider execution, result acceptance (`CallContext::result_ready`), and cache storage
- Per-call dependencies are grouped in `litellm_inference::context::CallContext`. `src/lib.rs` explicitly sequences cache lookup, provider execution, result acceptance (`CallContext::result_ready`), and cache storage
- Preserve the order of validation, normalization, caller-requested parameter removal, and provider transformation when that order affects observable behavior
- Test provider dispatch, auth precedence, header handling, transformations, and responses through behavior, not source structure

View file

@ -1,5 +1,5 @@
- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md)
- This crate owns OCR call orchestration; payloads belong in `litellm-llms-types`, transformations and the request handler in `llms/src/base_llm/ocr` and `llms/src/<provider>/ocr`
- This crate owns OCR call orchestration. Payloads belong in `litellm-llms-types`, transformations and the request handler in `llms/src/base_llm/ocr` and `llms/src/<provider>/ocr`
- The route returns its completed response directly
- Route closures supply OCR-specific host capabilities, such as the caller's Azure AD token provider (`src/route.rs`)
- Provider code reaches the caller's hooks mid-call only through `litellm_llms::base_llm::ocr::handler::CallHooks`, which this crate implements over its host until it folds into `litellm_host::interceptors::Interceptors`

View file

@ -1,4 +1,4 @@
- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md)
- This crate owns Responses API call orchestration; payloads belong in `litellm-llms-types`, transformations in `llms/src/base_llm/responses` and `llms/src/<provider>/responses`
- This crate owns Responses API call orchestration. Payloads belong in `litellm-llms-types`, transformations in `llms/src/base_llm/responses` and `llms/src/<provider>/responses`
- The HTTP route returns `litellm_host::call::CallOutput`: a completed response, or a stream head and chunks, cached through `litellm_inference::caching::execute_streaming`
- WebSocket sessions (`src/websocket.rs`) stay separate from the HTTP call driver because a connection can accept multiple requests while receiving events

View file

@ -1,3 +1,3 @@
- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md)
- This crate owns audio transcription call orchestration; transformations belong in `llms/src/base_llm/audio_transcription` and `llms/src/<provider>/audio_transcription`
- `AudioTranscriptionRoute::execute` returns the completed response directly; there is no hosted route machine or response cache for transcription yet
- This crate owns audio transcription call orchestration. Transformations belong in `llms/src/base_llm/audio_transcription` and `llms/src/<provider>/audio_transcription`
- `AudioTranscriptionRoute::execute` returns the completed response directly. There is no hosted route machine or response cache for transcription yet

View file

@ -6,7 +6,7 @@
- diagnostic spans (`src/diagnostic.rs`), outbound send and signing (`src/outbound.rs`), provider resolution (`src/provider.rs`), `CoreResources` (`src/resources.rs`)
- response caching (`src/caching.rs`): `Cachable`, `StreamCachable`, `CacheRequest`, `CallCache`, `execute_unary`, `execute_streaming`, stream capture
- Shared test helpers belong in litellm-inference-testing, used only as a dev dependency
- Nothing here names an API format; format-specific code, constants, tests and test builders live in their `inference-<fmt>` crate
- Nothing here names an API format. Format-specific code, constants, tests and test builders live in their `inference-<fmt>` crate
- Never depend on an `inference-*` crate from here
## Format crates
@ -26,22 +26,22 @@
- Hosts assemble route objects from shared `CoreResources`, HTTP settings, and secret sources
- each route owns its provider client and authentication dependencies
- gateway routes live for the gateway lifetime; Python assembles routes per call from its settings snapshot
- gateway routes live for the gateway lifetime. Python assembles routes per call from its settings snapshot
- Calls pass `Interceptors` and an optional `ObservationSender` separately
- use `&()` for no hooks and `None` for no observer
- handlers accept `Interceptors`, never a concrete `ChannelInterceptors`
- Construction does no work; preparation and lifecycle observation begin when the future is polled
- Native observers receive start and terminal events through the shared call runner; a stream retains its lifecycle until exhaustion, error, or drop. Hosted routes leave terminal observation to their driver
- A hosted format's `route.rs` declares the concrete `Protocol` and a route method that takes a typed request and builds a `litellm_host::call::HostedMachine` with `hosted_call`; audio transcription has no hosted machine and returns its response directly
- Construction does no work. Preparation and lifecycle observation begin when the future is polled
- Native observers receive start and terminal events through the shared call runner. A stream retains its lifecycle until exhaustion, error, or drop. Hosted routes leave terminal observation to their driver
- A hosted format's `route.rs` declares the concrete `Protocol` and a route method that takes a typed request and builds a `litellm_host::call::HostedMachine` with `hosted_call`. Audio transcription has no hosted machine and returns its response directly
- the shared call plumbing owns stream opening, delivery, backpressure, and detachment
- request decoding belongs to the boundary before the machine starts
- route closures only supply execution dependencies and route-specific host capabilities
- Python uses its own shared driver and preserves caller-task callback execution
- Not here: serving HTTP (axum routes, extractors), config file reading, rollout state, databases, or callback execution of any kind. Routes run as machines that yield host operations and call events; which integrations consume those events is the host's business
- Not here: serving HTTP (axum routes, extractors), config file reading, rollout state, databases, or callback execution of any kind. Routes run as machines that yield host operations and call events. Which integrations consume those events is the host's business
## Crate boundaries
- Dependencies only point down; Python package names identify counterparts, not ownership
- Dependencies only point down. Python package names identify counterparts, not ownership
- `litellm-llms-types` owns shared inference API contracts, grouped by format: pure serde data and shape validation, no I/O
- `litellm-core-utils` mirrors `litellm/litellm_core_utils/`: pure helpers (provider resolution, prompt factory, call arguments, settings lookup and layer merge), no network I/O
- `litellm-http` is Rust-only and route-neutral: settings resolution, pooled `reqwest` clients, TLS, proxies, the SSRF-safe media fetcher, request and header helpers, and transport errors
@ -49,35 +49,35 @@
- `litellm-inference-<fmt>` mirrors the route packages (`litellm/ocr/`, `litellm/messages/`, ...)
- Provider code never imports from `inference` or `inference-*`
- Handlers belong in `inference-*` or `llms`, never in a host crate
- Import every item from its canonical path. Never re-export another crate's items or give an item a second public path; the only allowed re-exports are a private submodule surfacing its item at its module root (`mod error; pub use error::Error;`) and each format crate's `pub use litellm_inference::RouteError as Error;`, which keeps the pre-split `<fmt>::Error` name
- Import every item from its canonical path. Never re-export another crate's items or give an item a second public path. The only allowed re-exports are a private submodule surfacing its item at its module root (`mod error; pub use error::Error;`) and each format crate's `pub use litellm_inference::RouteError as Error;`, which keeps the pre-split `<fmt>::Error` name
## Error placement
- The workspace `Error definitions` rules shape each crate's error; this section decides which crate and module a failure belongs to
- The workspace `Error definitions` rules shape each crate's error. This section decides which crate and module a failure belongs to
- A failure is declared once, by the lowest crate that raises it
- every crate above nests it unchanged (`#[error(transparent)] Auth(#[from] litellm_auth::Error)`) or maps it once at its boundary, as `src/error.rs` does for `litellm_llms::Error`
- `RouteError` collects route failures for every format crate and never re-declares a variant a lower crate raises
- Scope follows the concept, not the first caller
- an error type under `litellm-llms`'s `<provider>/` is private to that provider: no other provider and nothing in `base_llm` may import it
- a failure two providers or two routes can hit (wire framing, stream event decoding, a malformed provider response) belongs to the crate that owns the concept: `litellm-framer` for framing, `litellm_llms::Error` for the transformation layer
- `litellm_llms::Error` (`crates/llms/src/error.rs`) is the one transformation error for every provider and API; `base_llm/ocr/error.rs` is the recorded exception until OCR folds into it
- `litellm_llms::Error` (`crates/llms/src/error.rs`) is the one transformation error for every provider and API. `base_llm/ocr/error.rs` is the recorded exception until OCR folds into it
## Response caching and accounting
- Attach a `litellm_cache_response::ScopedCache` with `route.with_cache(cache)`; cached and uncached routes use the same `execute` and `machine` methods
- `CallOptions` carries a scope-free `CachePolicy` and observation; per-call policy never replaces the attached scope or service
- Provider transport does not own cache orchestration; stream capture stays in `src/caching.rs`
- Attach a `litellm_cache_response::ScopedCache` with `route.with_cache(cache)`. Cached and uncached routes use the same `execute` and `machine` methods
- `CallOptions` carries a scope-free `CachePolicy` and observation. Per-call policy never replaces the attached scope or service
- Provider transport does not own cache orchestration. Stream capture stays in `src/caching.rs`
- This crate and the format crates own request identity, typed response reconstruction, and stream capture and replay
- `cache-response` owns cache policy, namespacing, scope encoding, versioned envelopes, and freshness
- The SDK explicitly chooses shared scope; the gateway derives isolated scope from authenticated identity before attaching its service
- The SDK explicitly chooses shared scope. The gateway derives isolated scope from authenticated identity before attaching its service
- `ExecutionFacts` are delivered through the awaited `ResultReady` host operation for both provider and cached results, before public response processing or stream opening
- facts carry resolved model/provider and result source, including the hit key
- usage remains in the typed response or delivered stream, where completion and cancellation determine what was actually reported
- passive observation is not an accounting delivery mechanism
- Inference does not calculate prices, charge budgets, or update rate-limit counters
- the legacy Python callback adapter translates execution facts into the existing Python logging contract; Python remains the accounting owner on that path
- the legacy Python callback adapter translates execution facts into the existing Python logging contract. Python remains the accounting owner on that path
- native gateway accounting belongs to gateway dependencies, independently of `host-python`
- response-cache services expose no coordination counters or reservation APIs; a shared Redis deployment does not make response storage and accounting coordination the same dependency
- response-cache services expose no coordination counters or reservation APIs. A shared Redis deployment does not make response storage and accounting coordination the same dependency
- Cache lookup follows provider preparation, credential resolution, and the request interceptor
- keys describe the effective provider URL, authenticated headers, and rewritten body
- signed requests bypass caching until the signing identity has a stable cache representation

View file

@ -10,7 +10,7 @@ The same ownership rule applies to Messages, Responses, Chat Completions, OCR, a
- Keep one canonical definition and import path when moving a contract, updating consumers together instead of adding duplicate models or compatibility re-exports
- Keep shared provider-specific wire types and extensions under `providers`
- Provider types may reuse format types; format types must not depend on provider types
- Provider types may reuse format types. Format types must not depend on provider types
- A field belonging to an API format stays under `formats` even when provider support varies. Including it in a type does not promise provider support
- Add a typed provider extension when a consumer needs to interpret or construct it. Keep adapter-only projections in `llms` until a shared public data contract is needed
- Keep one authoritative representation of each field, preserving unknown fields without duplicating typed values in an extension map

View file

@ -14,11 +14,11 @@ Use named `#[rstest]` cases for independent input/output scenarios instead of lo
Shared OCR document and response contracts live in `litellm-llms-types::formats::ocr`. `BaseOcrConfig` and decoding into adapter errors remain in `src/base_llm/ocr/transformation.rs`. Rust context/environment types support the runtime. `BaseOcrConfig::prepare_request` corresponds to Python's HTTP-handler preparation rather than a `BaseOCRConfig` method, and `validate_request_body` is a Rust-only hook. `src/base_llm/ocr/error.rs` and `src/base_llm/ocr/document.rs` are Rust-only: the OCR error taxonomy shared with the route, and inline-document helpers shared by several providers
For Mistral, `async_transform_ocr_request` uses the base default in both languages. `resolve_headers` and `build_ocr_url` implement the respective environment and URL operations, and `normalize_response` implements the typed part of response transformation. Existing auth key/header handling and top-level response-extra preservation differ between languages; layout refactors must preserve those behaviors and verify them with the existing tests
For Mistral, `async_transform_ocr_request` uses the base default in both languages. `resolve_headers` and `build_ocr_url` implement the respective environment and URL operations, and `normalize_response` implements the typed part of response transformation. Existing auth key/header handling and top-level response-extra preservation differ between languages. Layout refactors must preserve those behaviors and verify them with the existing tests
For non-OCR pairs, order corresponding methods as parameter support/mapping, environment validation, URL construction, request transformation, and response transformation, followed by Rust-only runtime hooks. Auth resolution remains split between configs and route preparation in the `litellm-inference-<fmt>` crates. Chat `supported_openai_param_mappings` describes accepted OpenAI/provider name pairs, unlike Python's `get_supported_openai_params` name list. Audio `map_transcription_params` remains a Rust filtering helper
Azure Messages maps to `llms/azure_ai/anthropic/messages_transformation.py`; Bedrock Converse maps to `llms/bedrock/chat/converse_transformation.py`. `AnthropicConfig`, `AmazonConverseConfig`, and the non-OCR base traits are partial ports. `OpenAiResponsesApiConfig` implements WebSocket transformations and a direct HTTP Responses path. Its HTTP path does not implement Python model-specific parameter rewriting or Responses-to-Chat emulation. Preserve their acceptance gates, passthrough behavior, and host fallback contracts when aligning layout
Azure Messages maps to `llms/azure_ai/anthropic/messages_transformation.py`. Bedrock Converse maps to `llms/bedrock/chat/converse_transformation.py`. `AnthropicConfig`, `AmazonConverseConfig`, and the non-OCR base traits are partial ports. `OpenAiResponsesApiConfig` implements WebSocket transformations and a direct HTTP Responses path. Its HTTP path does not implement Python model-specific parameter rewriting or Responses-to-Chat emulation. Preserve their acceptance gates, passthrough behavior, and host fallback contracts when aligning layout
## Provider and format boundaries

View file

@ -1,10 +1,10 @@
- Target invariants, not completion claims; these supersede the crate guidance below where they conflict
- Target invariants, not completion claims. These supersede the crate guidance below where they conflict
## Boundary migration
Keep domain composition here and execution mechanics in `litellm-host-python`. A helper does not belong in the runtime adapter merely because it uses PyO3. LiteLLM argument rules, provider defaults, public responses, public exception policy and cache or secret-manager compatibility remain product responsibilities
The bridge supplies the Python lifecycle binding and public stream construction to the runtime adapter. Preserve the single inline coroutine driver; removing the host's hardcoded import must not introduce another driver or a separate asyncio task for caller hooks
The bridge supplies the Python lifecycle binding and public stream construction to the runtime adapter. Preserve the single inline coroutine driver. Removing the host's hardcoded import must not introduce another driver or a separate asyncio task for caller hooks
`src/callable.rs::wrap_failure` owns callable exception policy here. Resolved-Future construction uses `litellm-host-python::ready_future`, passing an already constructed Python value. Keep cache-specific serialization and disabled-cache results here
@ -16,18 +16,18 @@ Each step needs focused regression tests in the owning crate and Python integrat
- Keep this crate the product-specific PyO3 consumer of `litellm-host-python`
- Own registration, input projection, the route host and the caller callables it answers operations with (file readers, token providers), public response/error construction and the per-call composition of machine, route host and callback contract
- Legacy callback sharing (the caller's args, kwargs and request object, body/header roots, re-aliasing unchanged body keys) lives in `litellm-callbacks-legacy-python` behind `PublicCall` and `LegacyLogging`; the bridge hands the public call over and keeps no copy
- Value-oriented execution, sync waiting, nested-runtime checks, signal polling and panic containment live in `litellm-host-python`; native async work uses `pyo3-async-runtimes`, Serde output uses `Pythonized<T>`
- Core owns typed native state, the route machine, provider preparation/I/O and normalization; the host driver owns terminal events; the legacy adapter in `litellm-callbacks-legacy-python` owns `Logging` dispatch policy
- Python, Rust SDK and gateway use one lifecycle-bearing core route entrypoint; provider helpers stay private, never bridge-accessible transport drivers
- Built-in provider/config/secret/auth/document preparation stays in Rust; caller-authored callbacks and focused Python-file reads run only at core-selected points
- Target GIL-enabled CPython explicitly with `#[pymodule(gil_used = true)]`; detach Rust-only work
- Legacy callback sharing (the caller's args, kwargs and request object, body/header roots, re-aliasing unchanged body keys) lives in `litellm-callbacks-legacy-python` behind `PublicCall` and `LegacyLogging`. The bridge hands the public call over and keeps no copy
- Value-oriented execution, sync waiting, nested-runtime checks, signal polling and panic containment live in `litellm-host-python`. Native async work uses `pyo3-async-runtimes`, Serde output uses `Pythonized<T>`
- Core owns typed native state, the route machine, provider preparation/I/O and normalization. The host driver owns terminal events. The legacy adapter in `litellm-callbacks-legacy-python` owns `Logging` dispatch policy
- Python, Rust SDK and gateway use one lifecycle-bearing core route entrypoint. Provider helpers stay private, never bridge-accessible transport drivers
- Built-in provider/config/secret/auth/document preparation stays in Rust. Caller-authored callbacks and focused Python-file reads run only at core-selected points
- Target GIL-enabled CPython explicitly with `#[pymodule(gil_used = true)]`. Detach Rust-only work
- GIL and tokio invariants, each pinned by a test in `host-python` (`runtime.rs`,
`gil.rs`) so a regression fails there before it deadlocks a proxy:
- Never hold the GIL while waiting on the runtime. A sync entrypoint releases it with
`release_gil` around `block_on`, because every task that attaches would otherwise wait
on the thread that is waiting on them ([pyo3 parallelism](https://pyo3.rs/v0.29.2/parallelism.html))
- Never `block_on` from a tokio worker; the sync entrypoints refuse with "cannot run from
- Never `block_on` from a tokio worker. The sync entrypoints refuse with "cannot run from
a Tokio context" instead of panicking inside the runtime ([tokio `Runtime::block_on`](https://docs.rs/tokio/latest/tokio/runtime/struct.Runtime.html#method.block_on))
- Inside a future, `Python::attach` only for GIL-cheap work: cloning a `Py<T>`, building
a small value, reading a settings snapshot. Anything that can block (a secret manager
@ -39,33 +39,33 @@ Each step needs focused regression tests in the owning crate and Python integrat
`threading.get_ident()` test). Like any foreign-thread attach it therefore has no running
asyncio loop and a fresh `contextvars` context: do not hand it a coroutine or anything
bound to the caller's loop
- Dropping the await (an asyncio cancel) does not interrupt the Python call; it runs to
- Dropping the await (an asyncio cancel) does not interrupt the Python call. It runs to
completion and its result is discarded. A panic in it reaches the awaiting task as a panic
- Add a case to `gil.rs` when a new seam changes any of these; the tests are the spec
- Free-threading requires separate runtime/concurrency validation; omitting the attribute does not opt out on PyO3 0.28+
- Add a case to `gil.rs` when a new seam changes any of these. The tests are the spec
- Free-threading requires separate runtime/concurrency validation. Omitting the attribute does not opt out on PyO3 0.28+
- Preserve public argument binding and Python object provenance
- Project only consumed fields at reference read points; no eager whole-graph serialization or equality-based alias reconstruction
- Preserve provider-specific upload/submission/poll observation and encoding boundaries; signed/build-captured bytes must not be silently reserialized
- Project only consumed fields at reference read points. No eager whole-graph serialization or equality-based alias reconstruction
- Preserve provider-specific upload/submission/poll observation and encoding boundaries. Signed/build-captured bytes must not be silently reserialized
- Conversion errors and every failure after the call starts are terminal
- Disabled/unavailable native execution may select legacy once; callback exceptions never authorize fallback or replay
- Disabled/unavailable native execution may select legacy once. Callback exceptions never authorize fallback or replay
- Use one ordinary inline `async def` driver in `litellm/rust_bridge/lifecycle.py`, with the native handle and call driver in `litellm-host-python`
- Contract: `start`, `resume_value`, `resume_error`, idempotent `close`; explicitly tagged `Await`/`Complete` preserve awaitable final values
- Validate Created/Running/Suspended/Closed protocol states; the machine yields ops, the driver emits one terminal event, the adapter chooses dispatch policy
- Defer effectful setup/context reads/timestamps until start; unstarted-handle destruction releases inputs independently of Python `finally`
- Catch only the selected await's errors; start/resume errors propagate, `GeneratorExit` closes without further awaits
- Inline hooks preserve caller task/thread/loop and context writes; `into_future` creates a separate task and cannot satisfy this contract
- Contract: `start`, `resume_value`, `resume_error`, idempotent `close`. Explicitly tagged `Await`/`Complete` preserve awaitable final values
- Validate Created/Running/Suspended/Closed protocol states. The machine yields ops, the driver emits one terminal event, the adapter chooses dispatch policy
- Defer effectful setup/context reads/timestamps until start. Unstarted-handle destruction releases inputs independently of Python `finally`
- Catch only the selected await's errors. Start/resume errors propagate, `GeneratorExit` closes without further awaits
- Inline hooks preserve caller task/thread/loop and context writes. `into_future` creates a separate task and cannot satisfy this contract
- Finalize fallible public response/error construction, replacements and metadata before terminal dispatch
- Make ownership safe across suspension, re-entry, cancellation and GC
- Keep native provider state typed in core; do not shuttle it through opaque Python transport/response classes
- Prefer one retained `Py<PyBaseException>` via `PyErr::into_value(py)`; reconstruct transient `PyErr`s, preserving identity, traceback, cause and context
- Traverse every owned Python edge, including duplicate references; traversal cannot call Python
- Keep native provider state typed in core. Do not shuttle it through opaque Python transport/response classes
- Prefer one retained `Py<PyBaseException>` via `PyErr::into_value(py)`. Reconstruct transient `PyErr`s, preserving identity, traceback, cause and context
- Traverse every owned Python edge, including duplicate references. Traversal cannot call Python
- Take state out and mark Running under a short borrow, release borrows/locks before Python invocation, publish terminal state before finalizer-capable drops
- Close/GC/deferred release are idempotent and re-entry-safe, including during Rust unwinding; release only owned references, never clear caller containers or mask the selected error
- The machine owns its in-flight provider future; `interrupt` drops it synchronously, so provider captures are released before the driver returns and no task outlives the call
- Close/GC/deferred release are idempotent and re-entry-safe, including during Rust unwinding. Release only owned references, never clear caller containers or mask the selected error
- The machine owns its in-flight provider future. `interrupt` drops it synchronously, so provider captures are released before the driver returns and no task outlives the call
- Verify behavior through a fresh, provenance-checked installed extension and positive native execution evidence before replacing the custom coroutine
- Cover admitted provider workflows, binding/read-point/identity behavior, failure continuation, finalization, no replay, deferred gates, re-entry, GC and cancellation termination
- Measure real conversion/copy costs before optimizing; preserve input contracts and capture lifetimes with `PyBackedBytes`, and lookup timing when interning names
- Ship accurate `_native.pyi` declarations and typing markers; distinguish Future-returning bindings from coroutine-returning bindings
- Measure real conversion/copy costs before optimizing. Preserve input contracts and capture lifetimes with `PyBackedBytes`, and lookup timing when interning names
- Ship accurate `_native.pyi` declarations and typing markers. Distinguish Future-returning bindings from coroutine-returning bindings
- References: [ownership](https://pyo3.rs/v0.29.2/types.html), [GC](https://pyo3.rs/v0.29.2/class/protocols.html#garbage-collector-integration), [exception transfer](https://docs.rs/pyo3/0.29.2/pyo3/struct.PyErr.html#method.into_value), [re-entry](https://pyo3.rs/v0.29.2/class/call.html)
- [GIL policy](https://pyo3.rs/v0.29.2/free-threading.html), [experimental async limits](https://pyo3.rs/v0.29.2/async-await.html), [task conversion](https://docs.rs/pyo3-async-runtimes/0.29.0/pyo3_async_runtimes/fn.into_future_with_locals.html), [native cancellation/delivery](https://docs.rs/pyo3-async-runtimes/0.29.0/pyo3_async_runtimes/tokio/fn.future_into_py.html)
- [performance](https://pyo3.rs/v0.29.2/performance.html), [PyBackedBytes](https://docs.rs/pyo3/0.29.2/pyo3/pybacked/struct.PyBackedBytes.html), [typing](https://pyo3.rs/v0.29.2/python-typing-hints.html)
@ -87,7 +87,7 @@ GIL handling to `litellm-host-python`.
measured reason.
- Provider dispatch belongs in the `litellm-inference-*` route crate (e.g.
`litellm_inference_messages`), not in this PyO3 crate.
- Python owns rollout state and fallback. Rust should return errors; Python
- Python owns rollout state and fallback. Rust should return errors. Python
decides whether to raise or fall back. For a rust-only provider/route (no
Python reference), the Python side is a thin dispatch that calls Rust and
raises when the bridge is unavailable, with no fallback.
@ -101,7 +101,7 @@ GIL handling to `litellm-host-python`.
Python implementation, switch its dispatch to `NO_PYTHON` in the same change
- Keep the Python interface minimal (well under 100 lines per route): it only
marshals inputs and calls Rust. Do not add per-route feature flags, and do
not put provider dispatch in `litellm/main.py`; it lives in a thin dispatch
not put provider dispatch in `litellm/main.py`. It lives in a thin dispatch
class under `litellm/llms/<provider>/<route>/`.
## Data Handling

View file

@ -2,6 +2,6 @@
This folder owns how Rust inference reaches the selected cache: global cache selection, route admission, delegation to a Python cache and the experimental V2 native handles. Cache algorithms, storage protocols and response-cache semantics belong to their cache crates
`mod.rs` exposes the cache boundary to routes and module registration; adapter directories remain private. `selection.rs` owns global cache selection, route admission and inference protocol composition for both adapters
`mod.rs` exposes the cache boundary to routes and module registration. Adapter directories remain private. `selection.rs` owns global cache selection, route admission and inference protocol composition for both adapters
`python/` delegates operations to the selected Python cache without discovering configuration. `native/` owns the experimental V2 native cache handles. Neither adapter depends on shared selection or the other adapter. Shared composition depends on the adapters, and routes use only the parent module's exports

View file

@ -1,12 +1,12 @@
# Route boundary
These invariants apply to lifecycle-bearing public calls; value-oriented APIs that return an already running Future retain their explicit contract
These invariants apply to lifecycle-bearing public calls. Value-oriented APIs that return an already running Future retain their explicit contract
Route modules own public argument projection, route-specific host operations, response and error construction, and composition of the core route with the Python host. Provider dispatch, transport execution and normalization belong to core and provider crates. Runtime waiting, cancellation mechanics and execution state validation belong to `litellm-host-python`
Before execution starts, perform only admission checks needed to select native execution or legacy fallback. Do not fully project a request just to decide admission. Keep effectful settings reads, HTTP client acquisition, secret-source construction and caller context capture inside execution, after argument preparation and SDK preflight. Configure resources from that prepared view, never the original entrypoint kwargs
The host driver owns sequencing and terminal events; the bridge supplies fallible resource composition without exposing route types to the driver. An unstarted async call performs no resource setup. Setup errors after start follow the terminal failure contract and never authorize fallback or provider replay
The host driver owns sequencing and terminal events. The bridge supplies fallible resource composition without exposing route types to the driver. An unstarted async call performs no resource setup. Setup errors after start follow the terminal failure contract and never authorize fallback or provider replay
Use the shared `run_public_call` boundary with hooks supplied by bridge composition. `callbacks-legacy-python` owns legacy argument sharing and `Logging` dispatch behind `PublicCall` and `LegacyLogging`. Route bindings supply `callbacks-legacy-python::LoggingOperation` when composing legacy logging and may retain the request needed for projection, but must not duplicate the legacy callback contract

View file

@ -7,17 +7,17 @@
- Cache-key derivation: sha256 over `str(value)` must match Python byte for byte (`repr::to_str`)
- Byte-identical writes where Python compares raw values (`json::dumps`, `repr`)
- Relation to [`py_literal`](https://docs.rs/py_literal/latest/py_literal/): replaced, do not reintroduce
- Its pest grammar backtracks: parse time doubles per nested `[`/`{` (105 ms at depth 16); ours is linear (19 µs at depth 128)
- Its formatter is not `repr` (`2e-1`, always single quotes, escapes non-ASCII); it also lost `-0.0`, `(1+2j)`, `set()`
- `cache-response` and `cache-disk` still depend on it; migrate them here
- Its pest grammar backtracks: parse time doubles per nested `[`/`{` (105 ms at depth 16). Ours is linear (19 µs at depth 128)
- Its formatter is not `repr` (`2e-1`, always single quotes, escapes non-ASCII). It also lost `-0.0`, `(1+2j)`, `set()`
- `cache-response` and `cache-disk` still depend on it. Migrate them here
- Relation to [`serde-pickle`](https://docs.rs/serde-pickle/latest/serde_pickle/): the pickle codec, used only through its serde interface
- Never `serde_pickle::Value`: its `BTreeMap` dicts reorder keys
- Accepted limits: ints beyond i64, `tuple`/`set`/`frozenset` decode as lists, class references (`GLOBAL`/`REDUCE`) fail; writes protocol 3
- Every behavior is pinned by CPython output, not by reasoning; everything under `generated/` is script output, never hand-edited
- Regenerate `generated/values.json` with `scripts/generate_fixtures.py`; add a corpus row before changing behavior
- Divergences go in `KNOWN` in `tests/fixtures.rs` with a reason; an entry that starts matching fails until deleted
- Accepted limits: ints beyond i64, `tuple`/`set`/`frozenset` decode as lists, class references (`GLOBAL`/`REDUCE`) fail. Writes protocol 3
- Every behavior is pinned by CPython output, not by reasoning. Everything under `generated/` is script output, never hand-edited
- Regenerate `generated/values.json` with `scripts/generate_fixtures.py`. Add a corpus row before changing behavior
- Divergences go in `KNOWN` in `tests/fixtures.rs` with a reason. An entry that starts matching fails until deleted
- Regenerate `generated/nonprintable.rs` with `scripts/generate_nonprintable.py` when the target Python's Unicode version changes
- `scripts/verify_rust_pickles.py` checks CPython reads Rust pickles, with class resolution disabled; CI does not run Python
- `scripts/verify_rust_pickles.py` checks CPython reads Rust pickles, with class resolution disabled. CI does not run Python
- Decoders recurse, so they reject nesting beyond `MAX_DEPTH` for stack safety: deliberately stricter than CPython, whose parser takes ~200 levels and whose unpickler has no limit (pinned as `nested_150`)
- Formatters (`repr`, `json`) are unbounded; values from the decoders are already capped, a hand-built `Value` is the caller's responsibility
- Formatters (`repr`, `json`) are unbounded. Values from the decoders are already capped, a hand-built `Value` is the caller's responsibility
- `literal_eval` must stay linear in depth: `tests/limits.rs` times the deepest parse, `benches/formats.rs` measures the curve but is manual, since CI runs no Rust bench

View file

@ -6,4 +6,4 @@ The crate has no product tables, OTLP types, or named trace queries. `litellm-sp
It also applies embedded SQLx migrations through the `_sqlx_migrations` ledger
Migration files are append-only, and changed applied files are rejected by their checksums. Startup migrations must be replay-safe schema changes because the runner records success after execution without dirty states or locks. Backfills belong in coordinated jobs outside proxy startup. The replay policy lives in `ClickHouseMigrate::apply`, `dirty_version`, and `lock`; a Keeper-backed or deploy-time runner changes only those methods
Migration files are append-only, and changed applied files are rejected by their checksums. Startup migrations must be replay-safe schema changes because the runner records success after execution without dirty states or locks. Backfills belong in coordinated jobs outside proxy startup. The replay policy lives in `ClickHouseMigrate::apply`, `dirty_version`, and `lock`. A Keeper-backed or deploy-time runner changes only those methods

View file

@ -18,9 +18,9 @@ litellm/proxy/_experimental/mcp_server/
mcp_server_manager.py # upstream server registry, clients, tool routing [PR7: _create_mcp_client swaps resolve_mcp_auth -> resolve_credentials]
auth/
user_api_key_auth_mcp.py # LiteLLM admission auth and MCP request headers
token_exchange.py # OAuth token exchange handling [unchanged; V1TokenExchangeAdapter delegates here]
token_exchange.py # OAuth token exchange handling [unchanged, V1TokenExchangeAdapter delegates here]
litellm_auth_handler.py # authenticated-user adapter for MCP sessions
client_allowlist.py # gateway-level client application allowlist (mcp_allowed_clients); leaf module, no litellm.proxy imports
client_allowlist.py # gateway-level client application allowlist (mcp_allowed_clients), a leaf module with no litellm.proxy imports
outbound_credentials/ # NEW — typed upstream-credential resolution (resolve_credentials + arms)
__init__.py # public surface: resolve_credentials, the configs, CredError
result.py # Ok | Error union (pure stdlib)
@ -28,13 +28,13 @@ litellm/proxy/_experimental/mcp_server/
httpx_auth.py # NoOpAuth, StaticHeaderAuth (every mode -> one httpx.Auth)
resolver.py # resolve_credentials(): exhaustive per-mode match + assert_never
seams.py # injected Protocols (one per cache-touching mode)
v1_adapters.py # v1-backed seam bodies; delegate to auth/oauth2/db owners
v1_adapters.py # v1-backed seam bodies that delegate to auth/oauth2/db owners
adapter.py # to_subject / to_server_spec / raise_public (v1 <-> v2 boundary)
discoverable_endpoints.py # MCP OAuth metadata, authorize, token, callback
byok_oauth_endpoints.py # BYOK OAuth UI/API flow
oauth_utils.py # redirect URI and proxy base URL validation
oauth2_token_cache.py # OAuth2 and per-user token resolution/cache [PR7: resolve_mcp_auth removed; cache class stays, V1OAuth2CacheAdapter delegates to async_get_token]
db.py # MCP server, credential, env var, submission DB access [unchanged; V1ByokStore delegates to _get_byok_credential / get_user_credential]
oauth2_token_cache.py # OAuth2 and per-user token resolution/cache [PR7: resolve_mcp_auth removed, cache class stays, V1OAuth2CacheAdapter delegates to async_get_token]
db.py # MCP server, credential, env var, submission DB access [unchanged, V1ByokStore delegates to _get_byok_credential / get_user_credential]
toolset_db.py # MCP toolset DB access
rest_endpoints.py # proxy REST facade for listing/calling MCP tools [PR7: 7-arm only — pass identity + inbound token down instead of mcp_auth_header]
openapi_to_mcp_generator.py# OpenAPI spec to MCP tool generation
@ -60,7 +60,7 @@ module materially harder to understand.
## Implementation Rules
- Preserve the boundary between LiteLLM admission auth and upstream MCP auth.
Admission belongs in `auth/user_api_key_auth_mcp.py`; upstream token exchange,
Admission belongs in `auth/user_api_key_auth_mcp.py`. Upstream token exchange,
delegated auth, per-user OAuth, BYOK, and raw header forwarding belong in the
dedicated OAuth/header modules.
- Treat `none`, bearer/API key, OAuth, OAuth token exchange, delegated upstream

View file

@ -1,8 +1,8 @@
# Python boundary
This package owns native rollout and fallback selection, Python public API compatibility, settings projection and the Python bindings supplied to the Rust bridge. Rust core owns provider execution; `litellm-host-python` owns CPython runtime mechanics; `callbacks-legacy-python` owns legacy callback sharing and dispatch policy
This package owns native rollout and fallback selection, Python public API compatibility, settings projection and the Python bindings supplied to the Rust bridge. Rust core owns provider execution. `litellm-host-python` owns CPython runtime mechanics. `callbacks-legacy-python` owns legacy callback sharing and dispatch policy
`lifecycle.py` owns generic inline execution and stream iteration. `streams.py` supplies the product binding and public stream wrappers through the bridge rather than let the generic Rust host import this package by name. Keep one driver implementing `start`, `resume_value`, `resume_error` and idempotent `close`; do not create a second implementation
`lifecycle.py` owns generic inline execution and stream iteration. `streams.py` supplies the product binding and public stream wrappers through the bridge rather than let the generic Rust host import this package by name. Keep one driver implementing `start`, `resume_value`, `resume_error` and idempotent `close`. Do not create a second implementation
Generic execution steps and inline suspension handling must not depend on LiteLLM response metadata. Public stream construction, `_hidden_params` and header compatibility remain product responsibilities. Keep generic stream heads opaque and preserve public header behavior in the product wrappers

View file

@ -3,7 +3,7 @@
- Lens owns trace normalization, schemas, storage, graph assembly and querying in `BerriAI/lens`
- This package owns the HTTP client, gateway response validation and asynchronous gateway spend export. Keep the trace reader remote-only
- Gateway endpoints preserve authenticated user/team scope, trace references, pagination and public error categories
- `generated/` comes from `scripts/generate_trace_types.py` using the Lens schemas pinned in `scripts/lens_assets/source.json`. Update the canonical Lens Rust contracts and import their schema outputs before regenerating; never edit generated Python
- `generated/` comes from `scripts/generate_trace_types.py` using the Lens schemas pinned in `scripts/lens_assets/source.json`. Update the canonical Lens Rust contracts and import their schema outputs before regenerating. Never edit generated Python
- Keep the pinned schema and fixture checksums current. CI verifies assets and regenerated Python without a Rust trace implementation in this repository
- Independent gateway ClickHouse spend logging lives in `litellm/integrations/clickhouse`, using `litellm/rust_bridge/clickhouse.py` and the native spend writer
- OTLP upload routes only return setup guidance. Keep their JSON/protobuf error encoding without restoring ingestion or a native trace fallback

View file

@ -6,7 +6,7 @@ itself, no proxy at all: `tests/e2e_harness`. Two fit, split it
## What good looks like
Red when the claim in the name is broken. Prove it: mutate the behaviour, red; restore, green. Put the
Red when the claim in the name is broken. Prove it: mutate the behaviour, red. Restore, green. Put the
mutation in the PR body
```python
@ -21,7 +21,7 @@ def test_custom_price_is_reported_and_charged(gateway: Gateway) -> None:
Rates in the test, expected computed by hand, one call, `response.text` in the assert
Assert the whole value. Iterating `expected_body.items()` (`test_responses_api_request_body.py`) cannot
see an extra key; that is the shape of `stream_options.include_usage` (#19777, #28553)
see an extra key. That is the shape of `stream_options.include_usage` (#19777, #28553)
The linter catches no-assert, mock-echo and credential skips. It cannot see an assert
behind an `if` (a poll that ends in `pytest.fail` is fine), `except Exception` around the call
@ -29,7 +29,7 @@ behind an `if` (a poll that ends in `pytest.fail` is fine), `except Exception` a
## Where it goes
What the assertion depends on goes in the test; everything else in conftest. A rate in a fixture three
What the assertion depends on goes in the test. Everything else in conftest. A rate in a fixture three
directories up makes a failed assertion unreadable. Extend the file that already covers the behaviour
## Writing it so a human can read it

View file

@ -1,6 +1,6 @@
# e2e harness conventions
Code-style rules for writing tests under `tests/e2e/`. The harness already encodes the plumbing; your job is the feature-specific behavior, not reinventing it. For what a complete test must do (the lifecycle contract, asserting both recorded state and enforced behavior) and how to run a suite, see `CONTRIBUTING.md` in this directory. Repo-wide conventions live in the root `AGENTS.md`
Code-style rules for writing tests under `tests/e2e/`. The harness already encodes the plumbing. Your job is the feature-specific behavior, not reinventing it. For what a complete test must do (the lifecycle contract, asserting both recorded state and enforced behavior) and how to run a suite, see `CONTRIBUTING.md` in this directory. Repo-wide conventions live in the root `AGENTS.md`
## What good looks like
@ -24,13 +24,13 @@ assert still tears down. Assert what the caller receives
## Where it goes
By the surface a customer would name: `guardrails`, `llm_translation`, `management`. Mutation check
deferred; it needs credentials
deferred. It needs credentials
## Suite folders
Each subdirectory under `tests/e2e/` is one suite, scoped to an endpoint family or behavior area. If you add a new folder, you must add a line here describing what kind of tests belong in it, so the layout stays self-describing. `gateway/` is the exception: it holds proxy configuration only and never tests
- `migrations/` - isolated Docker startup, concurrent migration, crash recovery, and legacy database compatibility. The CircleCI migration workflow enables `LITELLM_MIGRATION_TESTS=1`; these tests own their proxy containers and databases, so they do not use the shared proxy preflight or shared database cleanup
- `migrations/` - isolated Docker startup, concurrent migration, crash recovery, and legacy database compatibility. The CircleCI migration workflow enables `LITELLM_MIGRATION_TESTS=1`. These tests own their proxy containers and databases, so they do not use the shared proxy preflight or shared database cleanup
- `llm_translation/` - LLM endpoint and provider-translation behavior: passthrough, custom pricing, OCR, and the non-chat inference endpoints (`/v1/responses`, `/v1/messages`, `/embeddings`, `/v1/rerank`, `/v1/audio/speech`, `/v1/images/generations`), each against a deployment the test creates via `/model/new` and deletes on teardown
- `access_control/` - the gateway's authorization and error-shape contract: per-key model allow-lists, route-group permissions (`allowed_routes`), and unknown-model validation
@ -38,33 +38,33 @@ Each subdirectory under `tests/e2e/` is one suite, scoped to an endpoint family
- `batches/` - the `/batches` endpoint (placeholder until the first test lands)
- `realtime/` - realtime websocket sessions, including the pipecat audio path
- `quota_management/` - quota enforcement and accounting, one subfolder per behavior: `ratelimit/` (rpm/tpm blocks, window reset, pacing headers on live traffic), `budgets/` (budget definition, enforcement, and reset windows: key, team, tag, soft, multi-window), and `spend_tracking/` (spend logging and cost attribution on `/spend/*`)
- `management/` - key/team/user/organization management routes: create/update/delete persistence via the info routes, team membership, and llm-only-key route denials (API surface; not Playwright)
- `management/` - key/team/user/organization management routes: create/update/delete persistence via the info routes, team membership, and llm-only-key route denials (API surface, not Playwright)
- `a2a/` - the A2A (agent-to-agent) surface: admin registration via `/v1/agents`, proxy-fronted card discovery at `/.well-known/agent-card.json`, and JSON-RPC `message/send` invocation, driving agents backed by the litellm completion bridge (a real provider) and asserting protocol-version normalization (0.3 vs 1.0)
- `mcp/` - the MCP server surface over api_key auth against the real Datadog remote MCP server (see "MCP suite: real Datadog only" below); plus the gateway-managed OAuth (authorization_code) path exercised through `/chat/completions` in `test_mcp_chat_completion_oauth_e2e.py` and direct MCP protocol operations in `test_mcp_oauth_happy_path_e2e.py`, the one behavior Datadog's static-header auth cannot reach, seeding the per-user upstream token via the interactive authorize dance driven with the mcp SDK's own OAuth client (headless-browser consent from a saved session) and asserting the completion or protocol call lists and executes the server's tools with the stored per-user token
- `mcp/` - the MCP server surface over api_key auth against the real Datadog remote MCP server (see "MCP suite: real Datadog only" below), plus the gateway-managed OAuth (authorization_code) path exercised through `/chat/completions` in `test_mcp_chat_completion_oauth_e2e.py` and direct MCP protocol operations in `test_mcp_oauth_happy_path_e2e.py`, the one behavior Datadog's static-header auth cannot reach, seeding the per-user upstream token via the interactive authorize dance driven with the mcp SDK's own OAuth client (headless-browser consent from a saved session) and asserting the completion or protocol call lists and executes the server's tools with the stored per-user token
- `logging/` - logging-integration delivery (datadog and friends)
- `security/` - secret handling and log-leak protection
- `router/` - routing and reliability behavior (fallbacks, cooldowns) plus the memory tests (`test_reliability_memory_e2e.py`: every worker's RSS as read at collection time, before any test traffic, must sit under a fixed idle budget, the release-gate check for a DB-backed boot that idles near the pod limit the way v1.100.x did; and a few hundred failing requests with retries and fallbacks must not grow proxy RSS past a fixed budget nor store a request snapshot past a fixed size, the release-gate check for the v1.100.0 retry-breadcrumb leak)
- `load/` - performance-category tests, kept OUT of the main suite: throughput/load SLO tests are a different testing category from functional e2e (variance-driven, historically flaky) and live outside this suite until re-implemented as their own pipeline (LIT-5163); do not add a live load test that runs in the default collection. What lives here: the weekly session-anomaly test (`test_weekly_session_anomaly_e2e.py`, Claude Code-shaped multi-turn sessions against real providers with ceilings on error rate, cache read/write, turn time, and spend; marked `weekly` and deselected unless `E2E_WEEKLY_ANOMALY` is set, driven by `.github/workflows/weekly_load_anomaly.yml`), the Redis chaos test (`test_redis_chaos_e2e.py`, locust load against mock deployments split round robin over `/chat/completions` and `/v1/messages`, one endpoint per simulated user, with `CLIENT PAUSE ALL` on the proxy's Redis mid-run to simulate it being down outright, asserting zero failed requests on every endpoint, budgeting RSS and CPU-per-request as ratios against the same run's healthy phase, and holding p50/p90/p99 latency and log-bytes-per-request to flat ceilings (a ratio cannot bound those two: an open breaker skips Redis instead of waiting on it, so the chaos phase can measure cheaper than baseline while still being far slower than a user should see); needs a proxy booted from `gateway/redis_chaos_ci_config.yml` on the same host with `E2E_PROXY_PID` and `E2E_PROXY_LOG` set, marked `redis_chaos`, deselected unless `E2E_REDIS_CHAOS` is set and excluded from the per-PR selector like the rest of `load/`, driven by `.github/workflows/test-e2e-redis-chaos.yml` and by the Buildkite `e2e-redis-chaos` step in project-releaser, which runs the proxy, Postgres and Valkey co-located with pytest in one pod and sets the opt-in). Its aggregation logic (locust, process usage, session anomaly) is covered by `tests/e2e_harness/load/`
- `other/` - the holding-pen suite for the `other.*` registry cluster with no home of its own yet: the master-key auth gate, JWT auth (access tokens issued by a real Keycloak realm, `idp.py` plus `idp_realm.json`, whose JWKS the proxy's `JWT_PUBLIC_KEY_URL` points at; see CONTRIBUTING.md for the start command and config block), and the process-lifecycle health probes (liveness, public readiness, authenticated readiness diagnostics). Promote a cluster out once it is large/stable enough for its own suite
- `router/` - routing and reliability behavior (fallbacks, cooldowns) plus the memory tests (`test_reliability_memory_e2e.py`: every worker's RSS as read at collection time, before any test traffic, must sit under a fixed idle budget, the release-gate check for a DB-backed boot that idles near the pod limit the way v1.100.x did, and a few hundred failing requests with retries and fallbacks must not grow proxy RSS past a fixed budget nor store a request snapshot past a fixed size, the release-gate check for the v1.100.0 retry-breadcrumb leak)
- `load/` - performance-category tests, kept OUT of the main suite: throughput/load SLO tests are a different testing category from functional e2e (variance-driven, historically flaky) and live outside this suite until re-implemented as their own pipeline (LIT-5163). Do not add a live load test that runs in the default collection. What lives here: the weekly session-anomaly test (`test_weekly_session_anomaly_e2e.py`, Claude Code-shaped multi-turn sessions against real providers with ceilings on error rate, cache read/write, turn time, and spend, marked `weekly` and deselected unless `E2E_WEEKLY_ANOMALY` is set, driven by `.github/workflows/weekly_load_anomaly.yml`), the Redis chaos test (`test_redis_chaos_e2e.py`, locust load against mock deployments split round robin over `/chat/completions` and `/v1/messages`, one endpoint per simulated user, with `CLIENT PAUSE ALL` on the proxy's Redis mid-run to simulate it being down outright, asserting zero failed requests on every endpoint, budgeting RSS and CPU-per-request as ratios against the same run's healthy phase, and holding p50/p90/p99 latency and log-bytes-per-request to flat ceilings (a ratio cannot bound those two: an open breaker skips Redis instead of waiting on it, so the chaos phase can measure cheaper than baseline while still being far slower than a user should see), needs a proxy booted from `gateway/redis_chaos_ci_config.yml` on the same host with `E2E_PROXY_PID` and `E2E_PROXY_LOG` set, marked `redis_chaos`, deselected unless `E2E_REDIS_CHAOS` is set and excluded from the per-PR selector like the rest of `load/`, driven by `.github/workflows/test-e2e-redis-chaos.yml` and by the Buildkite `e2e-redis-chaos` step in project-releaser, which runs the proxy, Postgres and Valkey co-located with pytest in one pod and sets the opt-in). Its aggregation logic (locust, process usage, session anomaly) is covered by `tests/e2e_harness/load/`
- `other/` - the holding-pen suite for the `other.*` registry cluster with no home of its own yet: the master-key auth gate, JWT auth (access tokens issued by a real Keycloak realm, `idp.py` plus `idp_realm.json`, whose JWKS the proxy's `JWT_PUBLIC_KEY_URL` points at, see CONTRIBUTING.md for the start command and config block), and the process-lifecycle health probes (liveness, public readiness, authenticated readiness diagnostics). Promote a cluster out once it is large/stable enough for its own suite
- `secret_manager/` - the gateway's `key_management_system` against a real secret manager: deployment keys resolved from it (`os.environ/<name>` where the name exists only in the manager) and virtual keys written to and deleted from it. The tests are backend-agnostic and each backend is its own lane, because the setting is global to the proxy: `E2E_SECRET_MANAGER=<system>` opts in and picks the backend from `secret_backends.BACKENDS`, the proxy is booted from `gateway/secret_manager_<system>_ci_config.yml` against the live manager, and the tests reach that manager through the backend's `SecretStore` (`secret_store_<system>.py`). A test needing something not every backend does carries `requires_capability(...)` and is deselected on lanes that lack it. `secret_manager/backend.sh up <system>` runs a backend in Docker and writes the proxy's and the tests' env. Marked `secret_manager`, deselected unless `E2E_SECRET_MANAGER` is set, and kept out of the per-PR selector. Backends today: `hashicorp_vault` and `cyberark` (CyberArk Conjur, which cannot delete, so the delete test is Vault-only)
- `gateway/` - proxy configuration only (`litellm-config.yml`); no tests
- `claude_code/` - the Claude Code compatibility matrix: drives the real `claude` CLI (and HTTP probes) against a proxy for each feature x provider cell, reporting tagged-union outcomes via the `compat_result` fixture; ships its own driver/builder/publisher, covered by `tests/e2e_harness/claude_code/`. The HTTP probes ride the shared transport (`ProxyClient.count_tokens` / `ProxyClient.messages`); the CLI-driving path stays bespoke
- `ui/` - the Admin UI browser suite: Playwright in TypeScript, driving the dashboard served by a live proxy on port 4000 (seeded postgres + mock LLM upstream; see its `run_e2e.sh`). It is a self-contained npm package with its own lockfile and does not use the Python harness, pytest markers, or the shared transport; the Python rules in this file (typed models, `Result` unions, basedpyright zero-error gate) do not apply inside it. Its only Python file, `fixtures/mock_llm_server/server.py`, is excluded from the e2e basedpyright gate via the root `pyrightconfig.json`
- `gateway/` - proxy configuration only (`litellm-config.yml`). No tests
- `claude_code/` - the Claude Code compatibility matrix: drives the real `claude` CLI (and HTTP probes) against a proxy for each feature x provider cell, reporting tagged-union outcomes via the `compat_result` fixture. Ships its own driver/builder/publisher, covered by `tests/e2e_harness/claude_code/`. The HTTP probes ride the shared transport (`ProxyClient.count_tokens` / `ProxyClient.messages`). The CLI-driving path stays bespoke
- `ui/` - the Admin UI browser suite: Playwright in TypeScript, driving the dashboard served by a live proxy on port 4000 (seeded postgres + mock LLM upstream, see its `run_e2e.sh`). It is a self-contained npm package with its own lockfile and does not use the Python harness, pytest markers, or the shared transport. The Python rules in this file (typed models, `Result` unions, basedpyright zero-error gate) do not apply inside it. Its only Python file, `fixtures/mock_llm_server/server.py`, is excluded from the e2e basedpyright gate via the root `pyrightconfig.json`
## MCP suite: real Datadog only
Every test under `tests/e2e/mcp/` must exercise the proxy against the real Datadog remote MCP server, except the two Linear OAuth tests `test_mcp_chat_completion_oauth_e2e.py` and `test_mcp_oauth_happy_path_e2e.py`. Do not add a compose service, FastMCP fixture, mock upstream, or any other fake MCP host for this suite
- Register via `register_datadog_mcp` in `tests/e2e/mcp/datadog_mcp.py` (or extend that helper if you need a different `toolsets=` / `allowed_tools` slice of the same Datadog endpoint). That posts `/v1/mcp/server` with `url=datadog_mcp_url(...)` and static headers `DD-API-KEY` / `DD-APPLICATION-KEY` from the process env
- Auth is Datadog's documented CI/header path, not a browser OAuth authorize/token dance. Hard-fail when `DD_API_KEY` or `DD_APP_KEY` is missing (`assert_dd_mcp_creds`); never skip for a missing fake upstream
- Prefer calling real Datadog tools that prove the product path (e.g. `search_datadog_logs` for list/call and permission denials). Seed a unique marker (`e2e-datadog-mcp-*`) in a chat completion when you need a log the tool can find; dual-read with `dd_logs` from conftest when delivery matters
- Auth is Datadog's documented CI/header path, not a browser OAuth authorize/token dance. Hard-fail when `DD_API_KEY` or `DD_APP_KEY` is missing (`assert_dd_mcp_creds`). Never skip for a missing fake upstream
- Prefer calling real Datadog tools that prove the product path (e.g. `search_datadog_logs` for list/call and permission denials). Seed a unique marker (`e2e-datadog-mcp-*`) in a chat completion when you need a log the tool can find. Dual-read with `dd_logs` from conftest when delivery matters
- Delete the MCP server (and any keys) through `resources.defer` the same way every other suite tears down
- If a new MCP behavior cannot be covered with Datadog's tool surface, say so in the PR and get agreement before inventing another upstream; the default is always Datadog
- The two standing exceptions are `test_mcp_chat_completion_oauth_e2e.py` and `test_mcp_oauth_happy_path_e2e.py`. Datadog authenticates with the static `DD-API-KEY` / `DD-APPLICATION-KEY` headers and exposes no authorize/token dance at all, so these tests drive a real Linear MCP server instead; they are still real remote upstreams, so the no-mock, no-fixture rule above holds unchanged. The direct OAuth test also uses the existing live provider edge to inspect forwarded headers without replay, and owns a separate source-built gateway for cold restarts
- If a new MCP behavior cannot be covered with Datadog's tool surface, say so in the PR and get agreement before inventing another upstream. The default is always Datadog
- The two standing exceptions are `test_mcp_chat_completion_oauth_e2e.py` and `test_mcp_oauth_happy_path_e2e.py`. Datadog authenticates with the static `DD-API-KEY` / `DD-APPLICATION-KEY` headers and exposes no authorize/token dance at all, so these tests drive a real Linear MCP server instead. They are still real remote upstreams, so the no-mock, no-fixture rule above holds unchanged. The direct OAuth test also uses the existing live provider edge to inspect forwarded headers without replay, and owns a separate source-built gateway for cold restarts
## Lay the pattern down in a class
Keep the cases for one feature inside a class so the file reads as a spec for how that feature behaves in production. The class name says what is under test; each method is one behavior. Think of it as documenting the contract, with the rough intent being
Keep the cases for one feature inside a class so the file reads as a spec for how that feature behaves in production. The class name says what is under test. Each method is one behavior. Think of it as documenting the contract, with the rough intent being
```python
# pseudo-code to convey intent, not the real API
@ -80,13 +80,13 @@ class TestPromptCompression:
assert response.cost == compressed_value # the cost was actually reduced
```
That snippet only conveys intent. What you actually write uses the real harness: the `client` fixture for your suite, the `scoped_key` fixture for an auto-deleted key, typed pydantic bodies from `models.py`, and `unwrap(...)` on the tagged-union result. `tests/e2e/llm_translation/test_custom_pricing_e2e.py` is the reference to copy from; it creates a scoped key, drives a real gemini call, polls `/spend/logs` to a deadline for the cost-breakdown row, then asserts the input and output costs match the configured custom rates and that a sibling deployment kept its own price. Read it before writing yours
That snippet only conveys intent. What you actually write uses the real harness: the `client` fixture for your suite, the `scoped_key` fixture for an auto-deleted key, typed pydantic bodies from `models.py`, and `unwrap(...)` on the tagged-union result. `tests/e2e/llm_translation/test_custom_pricing_e2e.py` is the reference to copy from. It creates a scoped key, drives a real gemini call, polls `/spend/logs` to a deadline for the cost-breakdown row, then asserts the input and output costs match the configured custom rates and that a sibling deployment kept its own price. Read it before writing yours
## Use the shared transport; never touch requests directly
## Use the shared transport, never touch requests directly
Every HTTP call goes through the shared transport, never through `requests.*` in a test. `e2e_http.py` is the only module permitted to call `requests.*`, and that is enforced in CI by `tests/code_coverage_tests/check_e2e_no_raw_requests.py`. A test that imports requests will fail the check
One deliberate exception: LLM-endpoint calls in `llm_translation/` go through the real provider SDKs (OpenAI, Anthropic) via the suite's `sdk` fixture (`llm_translation/sdk_clients.py`), because that is what customers actually run against the proxy (LIT-4577). The SDKs raise their own typed exceptions on failure, which is exactly the customer-observable contract; management routes (model/key CRUD, spend read-back) and endpoints no official SDK covers (e.g. `/v1/rerank`, `/v1/ocr`, custom passthrough paths) stay on the shared transport. Raw HTTP client imports remain banned either way
One deliberate exception: LLM-endpoint calls in `llm_translation/` go through the real provider SDKs (OpenAI, Anthropic) via the suite's `sdk` fixture (`llm_translation/sdk_clients.py`), because that is what customers actually run against the proxy (LIT-4577). The SDKs raise their own typed exceptions on failure, which is exactly the customer-observable contract. Management routes (model/key CRUD, spend read-back) and endpoints no official SDK covers (e.g. `/v1/rerank`, `/v1/ocr`, custom passthrough paths) stay on the shared transport. Raw HTTP client imports remain banned either way
The shape is layered so tests stay declarative
@ -96,7 +96,7 @@ The shape is layered so tests stay declarative
Each suite provides its own `client` fixture (see `llm_translation/passthrough_client.py`), a frozen dataclass that holds the shared `ProxyClient` (as `.proxy`) and adds suite-specific routes. Cleanup runs through that same `ProxyClient`, so whatever keys or customers your test creates get torn down by the `resources` fixture
Request and response bodies are typed pydantic models in `models.py`; only the fields a test reads are modelled, and nothing passes raw dicts. Outcomes come back as a `Result[R]` tagged union (`Success`, `NetworkError`, `UnauthorizedError`, `RateLimitedError`, `ValidationError`, `UnknownApiError`). Handle them with `match`, or call `unwrap(...)` when a non-success should fail the test. The harness hard-fails and never skips: a test marked `e2e` fails when no proxy answers its liveness probe, and once a request reaches the proxy any wrong behavior is likewise a hard failure, so a missing proxy turns the run red instead of being mistaken for a pass
Request and response bodies are typed pydantic models in `models.py`. Only the fields a test reads are modelled, and nothing passes raw dicts. Outcomes come back as a `Result[R]` tagged union (`Success`, `NetworkError`, `UnauthorizedError`, `RateLimitedError`, `ValidationError`, `UnknownApiError`). Handle them with `match`, or call `unwrap(...)` when a non-success should fail the test. The harness hard-fails and never skips: a test marked `e2e` fails when no proxy answers its liveness probe, and once a request reaches the proxy any wrong behavior is likewise a hard failure, so a missing proxy turns the run red instead of being mistaken for a pass
Mark live tests with `@pytest.mark.e2e` (on the class or the module). Coverage of the harness itself lives outside the suite in `tests/e2e_harness/` (see its `AGENTS.md`) and runs without a proxy, so nothing under `tests/e2e/` is a markerless test. Add `@pytest.mark.quiet_stack` to a test that measures the proxy itself (RSS, latency): the shared stack lock in `stack_lock.py` then runs it while no other test on the host is hitting the stack, marked or not, so the reading depends only on the test's own traffic. Use `scoped_key` for a fresh all-models key that auto-deletes, `resources` when you need to create and tear down more than a key, and `unique_marker()` from `e2e_config` to keep prompts, tags, and customer ids from colliding across concurrent runs and the shared response cache
@ -104,13 +104,13 @@ Mark live tests with `@pytest.mark.e2e` (on the class or the module). Coverage o
`E2E_FIXTURE_MODE` scopes the proxy's provider-bound traffic: `live` (the default, and what an unset variable means: nothing changes), `record` (the proxy's provider calls are forwarded to the real provider through a local edge server and written to a fixture bundle), or `replay` (the edge answers those calls from the bundle, so the run makes zero provider calls and spends nothing). Test-to-proxy traffic always goes over the wire in every mode: record and replay both need the live proxy and database, because the point is that key auth, routing, cost calculation, and spend-log writes execute for real while only the provider is swapped out. Breaking any of those in the proxy turns a replay run red
The seam is `provider_edge.py`: `start_provider_edge` boots an in-process HTTP server (one shared instance per pytest process, `e2e_config.provider_edge_base` is the accessor) that mounts each supported provider under a path prefix (`EDGE_MOUNTS`: `/openai` -> `https://api.openai.com`, `/anthropic` -> `https://api.anthropic.com`). A test participates by registering its deployment with `api_base=provider_edge_base("openai")` plus the provider's path suffix; `quota_management/spend_tracking/test_provider_edge_spend_e2e.py` is the reference. In live mode the accessor returns None and the deployment defaults to the real provider, so an edge-wired test runs in all three modes unchanged. Non-wired tests hit their providers live in every mode. The edge binds `E2E_PROVIDER_EDGE_BIND_HOST` (default 127.0.0.1) and advertises `E2E_PROVIDER_EDGE_ADVERTISE_HOST` in the api_base it hands out, for proxies running in containers
The seam is `provider_edge.py`: `start_provider_edge` boots an in-process HTTP server (one shared instance per pytest process, `e2e_config.provider_edge_base` is the accessor) that mounts each supported provider under a path prefix (`EDGE_MOUNTS`: `/openai` -> `https://api.openai.com`, `/anthropic` -> `https://api.anthropic.com`). A test participates by registering its deployment with `api_base=provider_edge_base("openai")` plus the provider's path suffix. `quota_management/spend_tracking/test_provider_edge_spend_e2e.py` is the reference. In live mode the accessor returns None and the deployment defaults to the real provider, so an edge-wired test runs in all three modes unchanged. Non-wired tests hit their providers live in every mode. The edge binds `E2E_PROVIDER_EDGE_BIND_HOST` (default 127.0.0.1) and advertises `E2E_PROVIDER_EDGE_ADVERTISE_HOST` in the api_base it hands out, for proxies running in containers
A bundle (default `tests/e2e/.fixtures`, override with `E2E_FIXTURE_DIR`) is a directory: `manifest.json` carries the record timestamp, harness git version, and format version, and each test gets a subdirectory holding one JSON file per provider call in call order (`0000-post-openai-v1-chat-completions.json`). Request headers are never stored (provider credentials never touch disk), non-JSON request bodies store a canonicalized sha256 digest instead of the bytes, `multipart/form-data` bodies store their ordinary fields plus a JSON list of the uploaded parts' `[field, filename, content-type]` triples and a digest of their content, so the per-request random boundary and the envelope never reach the key, and responses store status, filtered headers, and the verbatim body base64-encoded, which is part of why bundles are gitignored. Responses come in two shapes told apart by a `kind` tag: an ordinary one holding a single base64 body, and, for a response the provider streamed (`content-type: text/event-stream`), one holding its transfer chunks in order plus why the stream ended early if it did, so replay reproduces the split points the provider chose instead of one coalesced body. `fixture_bundle.py` owns the format, and `BUNDLE_FORMAT_VERSION` is checked on load, so a bundle recorded under older rules is refused by name rather than partially read. Record serves the proxy the same filtered stored response replay will serve later, chunk for chunk on a stream, so the two modes are byte-identical from the proxy's side of the socket. Replay does not reproduce the provider's inter-chunk timing (chunks go out as fast as the socket takes them), so a test that judges streaming on the clock, such as the `stream_event_arrivals` lead between the first content delta and `message_stop`, gates that assertion on `provider_paces_stream()` and proves only the event grammar in replay
Multipart identity is the fiddly corner, and the rules exist because each one had a collision behind it. A part counts as an upload when it carries a filename or declares its own content type, and everything else is an ordinary field. Field names get a `name[n]` suffix on repeats, with a literal `[` doubled first, so a form that repeats `purpose` never keys the same as one that literally sends `purpose[1]`. A field whose name reads as a credential is stored as `<secret>`, which stays key-preserving because the key is recomputed from the stored request rather than saved alongside it, so the live request carrying the real value still matches its redacted fixture. A field value that is not UTF-8 is stored as a base64 sha256 digest, base64 and not hex because the canonicalizer rewrites any 64-character hex run to `<sha256>` and would fold every binary value onto one key. The uploaded parts contribute a JSON list rather than a `field:filename` string, so a separator inside a filename cannot impersonate a field boundary, and their byte length is stored for a reader's benefit but deliberately left out of the key, since the canonicalizer absorbs timestamp and id drift inside a file that changes its length
Replay matches calls per test by canonical key: `fixture_canonical.py` canonicalizes the recorded request (volatile headers and credential fields out, unique markers, generated ids, uuids, and timestamps replaced with fixed placeholders, object keys sorted) and the key is the method, edge path, and a content hash, so identity survives re-records and machine changes while any real content drift comes back as an HTTP 599 naming the computed key, the closest recorded key with its file, and a content diff, and never falls through to a live call. Matching is order-independent across distinct keys (concurrent calls may interleave) and FIFO within one key (a retry loop replays its responses in recorded order); a passed test must also consume its whole recording, or teardown fails it naming a leftover key. Either way the fix is always to re-record with `E2E_FIXTURE_MODE=record`. Every rewrite rule lives in `fixture_canonical.py`, so a new volatile header, credential field name, or generated-id shape is one edit there. Record starts fresh every time: it wipes the previous bundle (refusing to wipe a directory that is not a bundle) and never reads it. A replay bundle whose manifest is older than seven days hard-fails at collection time naming the bundle's age, so replay can never certify against fixtures that have drifted more than a week from the live providers
Replay matches calls per test by canonical key: `fixture_canonical.py` canonicalizes the recorded request (volatile headers and credential fields out, unique markers, generated ids, uuids, and timestamps replaced with fixed placeholders, object keys sorted) and the key is the method, edge path, and a content hash, so identity survives re-records and machine changes while any real content drift comes back as an HTTP 599 naming the computed key, the closest recorded key with its file, and a content diff, and never falls through to a live call. Matching is order-independent across distinct keys (concurrent calls may interleave) and FIFO within one key (a retry loop replays its responses in recorded order). A passed test must also consume its whole recording, or teardown fails it naming a leftover key. Either way the fix is always to re-record with `E2E_FIXTURE_MODE=record`. Every rewrite rule lives in `fixture_canonical.py`, so a new volatile header, credential field name, or generated-id shape is one edit there. Record starts fresh every time: it wipes the previous bundle (refusing to wipe a directory that is not a bundle) and never reads it. A replay bundle whose manifest is older than seven days hard-fails at collection time naming the bundle's age, so replay can never certify against fixtures that have drifted more than a week from the live providers
A replayed response carries the recorded provider response id, and `LiteLLM_SpendLogs.request_id` (the table's primary key) is that id, so a replay against a database that still holds the record run's rows silently dedupes its spend inserts and any spend assertion goes red with zero matching rows and nothing in the proxy log. Run both modes with `E2E_RESET_SPEND_LOGS=1` (plus `DATABASE_URL` in the runner env) so each session truncates the table after itself, or replay against a fresh database, which is the CI shape
@ -125,7 +125,7 @@ E2E_FIXTURE_MODE=replay E2E_FIXTURE_DIR=/tmp/e2e-fixtures E2E_RESET_SPEND_LOGS=1
Point the proxy at bogus provider credentials for the replay run and it still has to pass: that is the whole proof that nothing left the process. Bundles are never committed. `tests/e2e/.fixtures` is gitignored because a bundle holds verbatim provider response bodies and hard-fails after seven days. CI records and replays this lane on a schedule in `.github/workflows/e2e_record_replay.yml`, publishing the bundle as a private `e2e-fixtures-bundle` artifact instead of committing it, selecting the tests with the `@pytest.mark.replayable` marker, and proving the bogus-credentials replay hermetic by counting provider egress with `.github/scripts/e2e_egress_sentinel.py`
Current limits: Bedrock cannot be mounted in record or replay (SigV4 signs the Host header, so a rewritten api_base fails signature verification); a test that needs to observe the Converse body registers its own `LiveEdge` with `provider_edge_bedrock.bedrock_signer` re-signing the forwarded request, and carries the `provider_edge_host` opt-in marker because the gateway must reach the pytest host, which the Buildkite ephemeral stack cannot, so those tests are deselected unless `E2E_PROVIDER_EDGE_HOST_REACHABLE` is set, as on a local run whose gateway can reach the pytest host. Deployments baked into the proxy's config file cannot be edge-wired (only `/model/new` registrations can carry the edge api_base), and a file upload routed by `custom_llm_provider` through the proxy's `files_settings` block never passes a deployment at all, so the batches `model_param` and `provider_fallback` scenarios keep uploading live in every mode
Current limits: Bedrock cannot be mounted in record or replay (SigV4 signs the Host header, so a rewritten api_base fails signature verification). A test that needs to observe the Converse body registers its own `LiveEdge` with `provider_edge_bedrock.bedrock_signer` re-signing the forwarded request, and carries the `provider_edge_host` opt-in marker because the gateway must reach the pytest host, which the Buildkite ephemeral stack cannot, so those tests are deselected unless `E2E_PROVIDER_EDGE_HOST_REACHABLE` is set, as on a local run whose gateway can reach the pytest host. Deployments baked into the proxy's config file cannot be edge-wired (only `/model/new` registrations can carry the edge api_base), and a file upload routed by `custom_llm_provider` through the proxy's `files_settings` block never passes a deployment at all, so the batches `model_param` and `provider_fallback` scenarios keep uploading live in every mode
## Typing
@ -133,7 +133,7 @@ The harness is fully typed with no error budget: `make lint-e2e-basedpyright` mu
## Typed test metadata
Separate from the coverage registry and additive to it: `@meta(Subject(...))` from `e2e_metadata.py` says what a test DRIVES, as closed enums rather than a string id. `@pytest.mark.covers("cell.id")` is untouched and keeps working exactly as before; the two markers coexist on the same test, and `@meta` always goes BELOW `@covers` so `Item.location` still anchors at the first decorator and every `source` deep link stays put
Separate from the coverage registry and additive to it: `@meta(Subject(...))` from `e2e_metadata.py` says what a test DRIVES, as closed enums rather than a string id. `@pytest.mark.covers("cell.id")` is untouched and keeps working exactly as before. The two markers coexist on the same test, and `@meta` always goes BELOW `@covers` so `Item.location` still anchors at the first decorator and every `source` deep link stays put
```python
@pytest.mark.covers("quota_management.budget.key.blocks_over_limit")
@ -150,7 +150,7 @@ def test_bare_key_blocks_over_its_own_budget(...) -> None: ...
`route` is the endpoint the test is checking: `TEAM_MANAGEMENT` for a `/team/update` test, `SPEND_REPORTING` for a `/spend/logs` test, `MESSAGES` for a test of spend on `/v1/messages`. A test whose chat call only triggers the behavior under test, like the budget block above, leaves it unset, since its steps already name the call
Every pytest test under `tests/e2e/` declares a `Subject` with at least its `domain`, except the `load/` suite, which is kept out of the default collection. The harness's own tests in `tests/e2e_harness/` carry none, since they drive nothing. The fields themselves stay optional, since a test that makes no LLM call has no provider, model or mode to name. Every field is a closed enum, so a typo is a basedpyright error at the call site rather than a property that silently never appears. `providers`, `models` and `capabilities` are tuples even with one member, because one test node routinely drives several: the claude_code matrix runs haiku, sonnet and opus in a single body, and a spend test calls two providers on one key. Declare every provider and every model the test drives, fallbacks included. The three are independent sets with no positional pairing between them (one provider x three models is the common case), and each is deduped and sorted at declaration so the committed run artifacts diff cleanly. `models=("gpt-5.5")` is a str and not a tuple, so anything but a tuple raises a `TypeError` where the decorator runs and shows up as a collection error naming the file. `Subject` is serialized with `dataclasses.asdict`, so a new scalar field needs no serializer edit; empty fields emit no `<property>` at all. A declared model names the constant the test drives (`CHEAP_ANTHROPIC_MODEL`, the file's own `BACKEND`), never a copy of its value, so the property cannot claim one model while an env override runs another. `e2e_metadata` and its call sites never import litellm, only the stdlib, pytest and pydantic: `Provider` mirrors litellm's `LlmProviders` values instead of importing them, because tests/e2e is shipped to the runner image on its own and a `from litellm...` at module scope would make the litellm package a hard dependency of COLLECTING the suite. `TestProviderMirrorsLitellm` in `tests/code_coverage_tests/test_e2e_metadata.py` fails on drift wherever litellm is importable and skips where it is not, so adding a provider is one line in `e2e_metadata`
Every pytest test under `tests/e2e/` declares a `Subject` with at least its `domain`, except the `load/` suite, which is kept out of the default collection. The harness's own tests in `tests/e2e_harness/` carry none, since they drive nothing. The fields themselves stay optional, since a test that makes no LLM call has no provider, model or mode to name. Every field is a closed enum, so a typo is a basedpyright error at the call site rather than a property that silently never appears. `providers`, `models` and `capabilities` are tuples even with one member, because one test node routinely drives several: the claude_code matrix runs haiku, sonnet and opus in a single body, and a spend test calls two providers on one key. Declare every provider and every model the test drives, fallbacks included. The three are independent sets with no positional pairing between them (one provider x three models is the common case), and each is deduped and sorted at declaration so the committed run artifacts diff cleanly. `models=("gpt-5.5")` is a str and not a tuple, so anything but a tuple raises a `TypeError` where the decorator runs and shows up as a collection error naming the file. `Subject` is serialized with `dataclasses.asdict`, so a new scalar field needs no serializer edit. Empty fields emit no `<property>` at all. A declared model names the constant the test drives (`CHEAP_ANTHROPIC_MODEL`, the file's own `BACKEND`), never a copy of its value, so the property cannot claim one model while an env override runs another. `e2e_metadata` and its call sites never import litellm, only the stdlib, pytest and pydantic: `Provider` mirrors litellm's `LlmProviders` values instead of importing them, because tests/e2e is shipped to the runner image on its own and a `from litellm...` at module scope would make the litellm package a hard dependency of COLLECTING the suite. `TestProviderMirrorsLitellm` in `tests/code_coverage_tests/test_e2e_metadata.py` fails on drift wherever litellm is importable and skips where it is not, so adding a provider is one line in `e2e_metadata`
Declared fields ride out as JUnit `<property>` entries behind the fixed prefix, the same way steps do: each scalar under its field name, and each plural value as a repeated property under its SINGULAR name (`provider`, `model`, `capability`). The results JSON downstream regroups them under the plural key, so `providers`, `models` and `capabilities` are arrays there, `[]` when empty
@ -195,13 +195,13 @@ Each step is its own `<property name="step">` in the JUnit XML (`junit_propertie
## Coverage registry
The set of tests we want is a registry checked into this repo, one row per behavior; that file is the definition of done and the denominator. Each e2e test declares what it covers with `@pytest.mark.covers("...")`, and a small collector diffs the registry against the tests and ships coverage to the existing Grafana. No Allure, no new dependencies
The set of tests we want is a registry checked into this repo, one row per behavior. That file is the definition of done and the denominator. Each e2e test declares what it covers with `@pytest.mark.covers("...")`, and a small collector diffs the registry against the tests and ships coverage to the existing Grafana. No Allure, no new dependencies
Coverage is organized as module > feature > test. Dashboard modules are `Core LLMs`, `Non-Core LLMs`, `MCPs`, `Management/UI`, `Reliability & Performance`, `Quota Management`, `Logging & Guardrails`, and `Other`. The Loki stdout formatter maps those display modules to log-safe labels (`core_llms`, `non_core_llms`, `mcp`, `management_ui`, `reliability_performance`, `quota_management`, `logging_guardrails`, and `other`) without changing JSON or Prometheus labels. A feature is either an endpoint (`/chat/completions`) or a behavior (fallbacks, rate limits; config-driven, with no route of its own). A cell reads like `llm.chat_completions.bedrock_converse.tool_use.stream.works`
Coverage is organized as module > feature > test. Dashboard modules are `Core LLMs`, `Non-Core LLMs`, `MCPs`, `Management/UI`, `Reliability & Performance`, `Quota Management`, `Logging & Guardrails`, and `Other`. The Loki stdout formatter maps those display modules to log-safe labels (`core_llms`, `non_core_llms`, `mcp`, `management_ui`, `reliability_performance`, `quota_management`, `logging_guardrails`, and `other`) without changing JSON or Prometheus labels. A feature is either an endpoint (`/chat/completions`) or a behavior (fallbacks, rate limits) that is config-driven, with no route of its own. A cell reads like `llm.chat_completions.bedrock_converse.tool_use.stream.works`
The metric is coverage: the share of registry rows that have a passing covering test, reported to Grafana per module so a gap surfaces as an uncovered row rather than a silent absence
Tests do not declare a dashboard module directly. They only declare the registry cell id with `@pytest.mark.covers("...")`; the registry row decides the module, tier, endpoint, and dashboard rollup. Run `python -m coverage_registry.collector --strict` when you want CI to reject unknown marker ids. Add `--fail-on-collection-errors` when the job should also fail on pytest collection errors.
Tests do not declare a dashboard module directly. They only declare the registry cell id with `@pytest.mark.covers("...")`. The registry row decides the module, tier, endpoint, and dashboard rollup. Run `python -m coverage_registry.collector --strict` when you want CI to reject unknown marker ids. Add `--fail-on-collection-errors` when the job should also fail on pytest collection errors.
Skipping a test gives its cell back to the gap list: the collector counts a cell as covered only when a test pytest would actually run declares it, and prints the cells left claimed only by skipped tests. So a `@pytest.mark.skip` on a red cell is honest bookkeeping, not a way to keep the number up.
@ -216,7 +216,7 @@ llm.<endpoint>.<route>.<capability>.<streaming>.<assertion>
| realtime
route : openai | azure_openai | anthropic | bedrock_converse | bedrock_invoke | vertex
| azure_foundry | cohere | together_ai | ollama | ollama_chat
(vocab varies per endpoint; messages is anthropic-format only)
(vocab varies per endpoint, messages is anthropic-format only)
capability : basic | tool_use | prompt_cache_5m | vision | thinking | structured_output
| service_tier | mid_conversation_system
streaming : stream | nonstream (omit where n/a)
@ -247,21 +247,21 @@ mcp.<operation>.<auth_family>.<assertion>
e.g. mcp.call_tool.oauth.succeeds
```
Reliability & Performance - behavior features (no route; endpoint is exercised_on)
Reliability & Performance - behavior features (no route, endpoint is exercised_on)
```
reliability.<behavior>.<variant>.<assertion>
behavior : fallback | retry | cooldown | timeout | routing | cache | circuit_breaker | perf
variant : <trigger> 5xx | context_window | content_policy | 429 | timeout
<strategy> simple_shuffle | usage_based | latency_based | cost_based | least_busy
<dimension> latency | throughput | session_anomaly | memory | idle_memory (perf only; SLO/threshold assertion, not binary)
<dimension> latency | throughput | session_anomaly | memory | idle_memory (perf only, SLO/threshold assertion, not binary)
assertion : routes_to_fallback | succeeds_within_retries | picks_under_tpm | returns_cached
| trips_then_recovers | under_slo
e.g. reliability.fallback.context_window.routes_to_fallback exercised_on=[chat_completions]
reliability.cooldown.429.trips_then_recovers exercised_on=[chat_completions, messages]
```
Quota Management - behavior features (entity- or config-driven caps and their accounting; endpoint is exercised_on)
Quota Management - behavior features (entity- or config-driven caps and their accounting, endpoint is exercised_on)
```
quota_management.<behavior>.<variant>.<assertion>
@ -286,7 +286,7 @@ quota_management.<behavior>.<variant>.<assertion>
quota_management.budget.key.blocks_over_limit exercised_on=[chat_completions]
```
Logging & Guardrails - behavior features (config-driven; endpoint is exercised_on)
Logging & Guardrails - behavior features (config-driven, endpoint is exercised_on)
```
logging.<integration>.<event>.<assertion>
@ -307,13 +307,13 @@ Other - holding pen (endpoint or behavior)
```
other.<area>.<case>.<assertion>
area : auth | lifecycle | config | ...
rule : audited periodically; a cluster here promotes to a new component
rule : audited periodically, and a cluster here promotes to a new component
e.g. other.auth.jwt.valid_token_allows
other.lifecycle.readiness.reports_db
```
## Hard Rules
- no unit tests of a product feature under `tests/e2e`, and no mock tests or monkeypatching of code anywhere in it: a product feature is proven end to end against a live proxy, never with a unit test. if a contributor asks you to write an end to end test, do NOT stage a unit test with it; if you find a product gap, call it out in the PR description. the harness's own plumbing is tested outside the suite, in `tests/e2e_harness/` (mirroring this folder's layout), because the Buildkite e2e run copies `tests/e2e/` into the runner image and runs every test in it, so a harness test in here would count as a product test in the nightly numbers. those tests run without a proxy and take their inputs as arguments or env vars (setting an env var through pytest's `monkeypatch` fixture is fine, patching a function, class, or module is not), and no coverage-registry or compat-matrix cell rests on them. judge a change inside one of them by that standard, not as a misplaced product test
- no unit tests of a product feature under `tests/e2e`, and no mock tests or monkeypatching of code anywhere in it: a product feature is proven end to end against a live proxy, never with a unit test. if a contributor asks you to write an end to end test, do NOT stage a unit test with it. If you find a product gap, call it out in the PR description. the harness's own plumbing is tested outside the suite, in `tests/e2e_harness/` (mirroring this folder's layout), because the Buildkite e2e run copies `tests/e2e/` into the runner image and runs every test in it, so a harness test in here would count as a product test in the nightly numbers. those tests run without a proxy and take their inputs as arguments or env vars (setting an env var through pytest's `monkeypatch` fixture is fine, patching a function, class, or module is not), and no coverage-registry or compat-matrix cell rests on them. judge a change inside one of them by that standard, not as a misplaced product test
- use model management endpoints to create new models for a test. this could be in a conftest / inline for each test. ask the user what they want.
@ -321,7 +321,7 @@ other.<area>.<case>.<assertion>
- when it comes to typing an input schema for an api endpoint, have it type X = A | B | C ... where X = exhaustive union of all supported input schemas and A, B, C typically are composed by a base type. types are only pretty for a api request / response body. make sure to compose types instead of repeating the same base attributes over and over again.
- spin up a local proxy by running the litellm proxy locally (`litellm --config <your-e2e-config>.yml --port 4000`; see CONTRIBUTING.md), make sure all tests pass. if a test fails due to an internally found issue, let users know to create a linear ticket for it.
- spin up a local proxy by running the litellm proxy locally (`litellm --config <your-e2e-config>.yml --port 4000`, see CONTRIBUTING.md), make sure all tests pass. if a test fails due to an internally found issue, let users know to create a linear ticket for it.
- do not use xfail markers, tests should be written in a form that the end user expects it to pass

View file

@ -195,7 +195,7 @@ If a step is missing, the test is not done. That is the whole pattern
## Style: lay the pattern down in a class
Keep the cases for one feature inside a class so the file reads as a spec for how that feature behaves in production. The class name says what is under test; each method is one behavior. Think of it as documenting the contract, with the rough intent being
Keep the cases for one feature inside a class so the file reads as a spec for how that feature behaves in production. The class name says what is under test. Each method is one behavior. Think of it as documenting the contract, with the rough intent being
```python
# pseudo-code to convey intent
@ -211,9 +211,9 @@ class TestPromptCompression:
assert response.cost == compressed_value # the cost was actually reduced
```
That snippet only conveys intent. What you actually write uses the real harness: the `client` fixture for your suite, the `scoped_key` fixture for an auto-deleted key, typed pydantic bodies from `models.py`, and `unwrap(...)` on the tagged-union result. `tests/e2e/llm_translation/test_custom_pricing_e2e.py` is the reference to copy from; it creates a scoped key, drives a real gemini call, polls `/spend/logs` to a deadline for the cost-breakdown row, then asserts the input and output costs match the configured custom rates and that a sibling deployment kept its own price. Read it before writing yours
That snippet only conveys intent. What you actually write uses the real harness: the `client` fixture for your suite, the `scoped_key` fixture for an auto-deleted key, typed pydantic bodies from `models.py`, and `unwrap(...)` on the tagged-union result. `tests/e2e/llm_translation/test_custom_pricing_e2e.py` is the reference to copy from. It creates a scoped key, drives a real gemini call, polls `/spend/logs` to a deadline for the cost-breakdown row, then asserts the input and output costs match the configured custom rates and that a sibling deployment kept its own price. Read it before writing yours
## Use the shared transport; never touch requests directly
## Use the shared transport, never touch requests directly
Every HTTP call goes through the shared transport, never through `requests.*` in a test. `e2e_http.py` is the only module permitted to call `requests.*`, and that is enforced in CI by `tests/code_coverage_tests/check_e2e_no_raw_requests.py`. A test that imports requests will fail the check
@ -227,7 +227,7 @@ The shape is layered so tests stay declarative
Each suite provides its own `client` fixture (see `llm_translation/passthrough_client.py`), a frozen dataclass that holds the shared `ProxyClient` (as `.proxy`) and adds suite-specific routes. Cleanup runs through that same `ProxyClient`, so whatever keys or customers your test creates get torn down by the `resources` fixture
Request and response bodies are typed pydantic models in `models.py`; only the fields a test reads are modelled, and nothing passes raw dicts. Outcomes come back as a `Result[R]` tagged union (`Success`, `NetworkError`, `UnauthorizedError`, `RateLimitedError`, `ValidationError`, `UnknownApiError`). Handle them with `match`, or call `unwrap(...)` when a non-success should fail the test. The harness hard-fails and never skips: a test marked `e2e` fails when no proxy answers its liveness probe, and once a request reaches the proxy any wrong behavior is likewise a hard failure, so a missing proxy turns the run red instead of being mistaken for a pass
Request and response bodies are typed pydantic models in `models.py`. Only the fields a test reads are modelled, and nothing passes raw dicts. Outcomes come back as a `Result[R]` tagged union (`Success`, `NetworkError`, `UnauthorizedError`, `RateLimitedError`, `ValidationError`, `UnknownApiError`). Handle them with `match`, or call `unwrap(...)` when a non-success should fail the test. The harness hard-fails and never skips: a test marked `e2e` fails when no proxy answers its liveness probe, and once a request reaches the proxy any wrong behavior is likewise a hard failure, so a missing proxy turns the run red instead of being mistaken for a pass
Mark live tests with `@pytest.mark.e2e` (on the class or the module). Pure coverage of the harness itself lives in `tests/e2e_harness/` and runs without a proxy (`LITELLM_MASTER_KEY=sk-harness uv run pytest tests/e2e_harness`). A test that needs proxy configuration the default stack does not carry goes behind an opt-in marker (`managed_files`, `prompt_caching_stack`, `weekly`), each deselected unless its env var is set; `OPT_IN_MARKERS` in `conftest.py` maps marker to env var, and the coverage collector counts such a cell only where the env var is set. Use `scoped_key` for a fresh all-models key that auto-deletes, `resources` when you need to create and tear down more than a key, and `unique_marker()` from `e2e_config` to keep prompts, tags, and customer ids from colliding across concurrent runs and the shared response cache

View file

@ -10,8 +10,8 @@ Run them from the repo root. `e2e_config` reads `LITELLM_MASTER_KEY` at import a
LITELLM_MASTER_KEY=sk-harness uv run pytest tests/e2e_harness
```
`pytest.ini` here puts `tests/e2e` and the suite folders whose modules are under test on the path, so imports look exactly as they do inside the suite (`from e2e_http import ...`, `from batch_cleanup import ...`). `claude_code/test_request_determinism.py` drives the real `claude` CLI; deselect it with `-m "not cli_determinism"` when the CLI is not installed
`pytest.ini` here puts `tests/e2e` and the suite folders whose modules are under test on the path, so imports look exactly as they do inside the suite (`from e2e_http import ...`, `from batch_cleanup import ...`). `claude_code/test_request_determinism.py` drives the real `claude` CLI. Deselect it with `-m "not cli_determinism"` when the CLI is not installed
Rules: no `e2e` marker and no `@meta`, since nothing here drives the proxy; `@pytest.mark.covers` only where the test proves the collector or the JUnit properties read it; inputs via arguments or env vars (setting an env var through pytest's `monkeypatch` fixture is fine, patching a function, class or module is not); and the same typing bar as the suite, `make lint-e2e-basedpyright` covers this folder and allows zero errors. The raw HTTP client ban (`tests/code_coverage_tests/check_e2e_no_raw_requests.py`) applies here too
Rules: no `e2e` marker and no `@meta`, since nothing here drives the proxy. `@pytest.mark.covers` only where the test proves the collector or the JUnit properties read it. Inputs via arguments or env vars (setting an env var through pytest's `monkeypatch` fixture is fine, patching a function, class or module is not). And the same typing bar as the suite, `make lint-e2e-basedpyright` covers this folder and allows zero errors. The raw HTTP client ban (`tests/code_coverage_tests/check_e2e_no_raw_requests.py`) applies here too
CI: the `python` job in `.github/workflows/test-linting.yml` runs this folder whenever anything under `tests/e2e/` (except `ui/`) or `tests/e2e_harness/` changes, with the `claude` CLI installed. The CircleCI `provider_replay_harness` job also runs the provider-edge and fixture tests at the root of this folder next to `tests/code_coverage_tests/test_provider_replay_harness.py`, which imports helpers from `test_provider_edge.py`

View file

@ -16,12 +16,12 @@ assert float(rows[0]["spend"]) == pytest.approx(10 * 0.001 + 5 * 0.0001 + 7 * 0.
```
`sleep(3)` fails on a slow runner and taxes every fast one. Assert the outbound body in the upstream
handler; a leaked field is invisible from the response. `monkeypatch.setenv` is fine; patching our own
handler. A leaked field is invisible from the response. `monkeypatch.setenv` is fine. Patching our own
function in a full stack is not
## Where it goes
By the domain a user would name: `pricing`, `spend`, `routing`, `mcp`. A file only needs to live in a
directory that a `GROUPS` entry in `run.py` selects; there is no manifest and no `covers` marker on new
directory that a `GROUPS` entry in `run.py` selects. There is no manifest and no `covers` marker on new
tests. A product bug the test exposes is `pytest.skip("BUG: <symptom>")` at the top of the body, not a
fix in the test and not a deletion. Needs no proxy, DB or Redis: `tests/unit`

View file

@ -45,21 +45,21 @@ tests/rust-python-harness/
└── suite_runner.py
```
- A strategy is a folder under `strategies/` with a one-line `AGENTS.md` and an `__init__.py` exporting exactly one `STRATEGY: StrategyDefinition`; its id must equal the folder name
- A strategy is a folder under `strategies/` with a one-line `AGENTS.md` and an `__init__.py` exporting exactly one `STRATEGY: StrategyDefinition`. Its id must equal the folder name
- `shared/reporting/strategy.py` is the contract: runnable module/suite specs, not-implemented/skipped specs, the runner protocol, and `StrategyDefinition`
- Every `STRATEGY` explicitly classifies every SDK function; surface-aware strategies declare their surfaces and classify the complete surface-by-function matrix
- Run locally only; no CI integration
- `python -m tests.rust-python-harness run <strategy>|all` runs the selected strategy; `--function` is common, while each strategy exposes only its supported options
- Every `STRATEGY` explicitly classifies every SDK function. Surface-aware strategies declare their surfaces and classify the complete surface-by-function matrix
- Run locally only. No CI integration
- `python -m tests.rust-python-harness run <strategy>|all` runs the selected strategy. `--function` is common, while each strategy exposes only its supported options
- Examples: `run e2e_parity --surface sdk --function ocr`, `run unit_tests_parity --function ocr --pytest-arg=-x`, or `run all --function ocr`
- `cli/catalog.py` discovers strategies, validates their Python definitions, and orders them; `cli/__init__.py` builds the Click command tree; `cli/commands.py` runs selected cases
- `cli/catalog.py` discovers strategies, validates their Python definitions, and orders them. `cli/__init__.py` builds the Click command tree. `cli/commands.py` runs selected cases
- `e2e_parity/` compares SDK objects, exceptions, callbacks, and streams, or gateway HTTP responses
- `trace_parity/` profiles the Python call stack and prints every collected Python call under `litellm/`; it never collects Rust spans and never rebuilds the native extension
- `trace_parity/` profiles the Python call stack and prints every collected Python call under `litellm/`. It never collects Rust spans and never rebuilds the native extension
- E2E and trace strategies load their registered module cases and run surface-specific execution from their folders
- `shared/unit_runners/contracts.py` owns the typed per-function unit contracts consumed by `unit_tests_parity` and `unit_tests_rust`
- `unit_tests_parity/runner.py` runs each contract's `unit_parity_scope` with `LITELLM_RUST=0` and `LITELLM_RUST=1` in separate processes and requires matching outcomes, including failures; exclusions require a reason in the contract
- `unit_tests_rust/runner.py` runs each contract's focused Cargo test suite; native Rust unit tests stay beside their implementation
- `unit_tests_parity/runner.py` runs each contract's `unit_parity_scope` with `LITELLM_RUST=0` and `LITELLM_RUST=1` in separate processes and requires matching outcomes, including failures. Exclusions require a reason in the contract
- `unit_tests_rust/runner.py` runs each contract's focused Cargo test suite. Native Rust unit tests stay beside their implementation
- `shared/unit_runners/suite_runner.py` runs typed suites registered in code with nodeids of the form `suite:<strategy_id>:<function>:<suite>`
- Every strategy declares its report sections and presentation in its own `reporting.py`; shared reporting code only provides reusable models and cell-formatting primitives
- Every strategy declares its report sections and presentation in its own `reporting.py`. Shared reporting code only provides reusable models and cell-formatting primitives
- `shared/` contains reusable parity, tracing, reporting primitives, and unit-runner machinery
- Keep fixtures with their owning API and existing Python tests in their current locations
- Each strategy folder carries an `AGENTS.md` one-liner stating what it should be doing

View file

@ -1 +1 @@
Prints every collected Python call under litellm/ from live traces against replayed HTTP responses. API-key and Vertex credentials scenarios exercise separate authentication paths; credentials scenarios replay the token exchange locally.
Prints every collected Python call under litellm/ from live traces against replayed HTTP responses. API-key and Vertex credentials scenarios exercise separate authentication paths. Credentials scenarios replay the token exchange locally.

View file

@ -23,7 +23,7 @@ with patch.object(streamer, "_group_by_date") as mock_group, patch.object(stream
assert mock_send.call_count == 2
```
Green if `send_batched` drops every row. pydantic doubles in 12 of 203 files, fastapi 11 of 594;
Green if `send_batched` drops every row. pydantic doubles in 12 of 203 files, fastapi 11 of 594.
`tests/test_litellm` 59 percent. Exception: the count is the behaviour
(`test_dual_cache_async_batch_get_cache_coalesces_concurrent_redis_reads`, fifty readers, `call_count == 1`)

View file

@ -4,9 +4,9 @@ Never put LiteLLM tokens or API keys in `localStorage`. `localStorage` survives
When you fix lint violations that are grandfathered in `eslint-suppressions.json`, run `eslint . --prune-suppressions` and commit the updated baseline so the gate ratchets down instead of leaving a stale suppression
`src/lib/http/schema.d.ts` is generated from the proxy's OpenAPI spec; never hand-edit it. After changing a backend route or response model that the dashboard consumes, run `npm run gen:api` and commit the result (CI `lint / ui-api-types` enforces this)
`src/lib/http/schema.d.ts` is generated from the proxy's OpenAPI spec. Never hand-edit it. After changing a backend route or response model that the dashboard consumes, run `npm run gen:api` and commit the result (CI `lint / ui-api-types` enforces this)
Tests come in three tiers, named by the standard definitions. `Foo.test.tsx` is a unit test: one module, collaborators replaced by doubles, no multi-component tree, and it should run in milliseconds. `Foo.integration.test.tsx` renders a real component tree with real children and only stubs the network boundary; it costs seconds per case, so it earns its place by proving wiring that a unit test cannot reach. Browser-level tests live in `tests/e2e/ui/` as Playwright specs against a live proxy
Tests come in three tiers, named by the standard definitions. `Foo.test.tsx` is a unit test: one module, collaborators replaced by doubles, no multi-component tree, and it should run in milliseconds. `Foo.integration.test.tsx` renders a real component tree with real children and only stubs the network boundary. It costs seconds per case, so it earns its place by proving wiring that a unit test cannot reach. Browser-level tests live in `tests/e2e/ui/` as Playwright specs against a live proxy
When a component holds logic worth asserting, extract the logic and unit-test it there rather than driving it through a render. `CreateMCPServer` is the worked example: its payload building lives in `createServerPayload.ts` with 46 unit tests that run in single-digit milliseconds, while `CreateMCPServer.integration.test.tsx` keeps only the cases that prove a form field reaches the right payload key. A test that renders a whole modal to assert the shape of one object belongs in the first category, not the second

View file

@ -13,7 +13,7 @@ This directory has drifted from its own spec before — a hand-rolled `<button>`
3. **Does it need a color for "selected" / "active" / "hover"?** Before picking a class:
- Read the actual values in `src/app/globals.css` for the classes you're about to use. Don't assume `accent` ≠ `secondary` ≠ `muted` — in this theme they're identical. Confirm contrast against the _actual container_ background, not the general Tailwind palette in your head.
- If the element lives inside the sidebar, use the `sidebar-*` token family (`bg-sidebar`, `bg-sidebar-accent`, `text-sidebar-accent-foreground`), not the generic tokens.
4. **Is the shadcn primitive you want to use not in `src/components/ui/`?** Check the "Known gaps" list in `design.md` first — if it's `Card`, `Textarea`, `sonner`, or a `Sidebar` block, use the documented substitute. If it's something else entirely missing, stop and flag it; don't hand-write a parallel implementation inside a chat component.
4. **Is the shadcn primitive you want to use not in `src/components/ui/`?** Check the "Known gaps" list in `design.md` first — if it's `Card`, `Textarea`, `sonner`, or a `Sidebar` block, use the documented substitute. If it's something else entirely missing, stop and flag it. Don't hand-write a parallel implementation inside a chat component.
5. **Still unsure?** Grep this directory and `src/app/(dashboard)/` for an existing instance of the same UI idea (e.g. "how does the rest of the dashboard render a data table") before inventing a new pattern for chat specifically. The chat UI should look like the rest of the app, not like its own product.
## Decision tree: "I'm fixing a bug in an existing component"
@ -34,7 +34,7 @@ Never deviate silently. If you improvise, the next agent (or you, in six months)
Run through this before considering a chat UI change done:
- [ ] Every clickable element is a `Button` variant, not a raw `<button>` — search your diff for `<button` and justify any that remain (the chat composer's textarea-adjacent icon buttons and similar truly-custom controls are the only accepted exceptions; nav items and dialog actions are not).
- [ ] Every clickable element is a `Button` variant, not a raw `<button>` — search your diff for `<button` and justify any that remain (the chat composer's textarea-adjacent icon buttons and similar truly-custom controls are the only accepted exceptions, nav items and dialog actions are not).
- [ ] Every color class you added actually contrasts against its container — checked in `globals.css`, not assumed.
- [ ] Sidebar-scoped elements use `sidebar-*` tokens, not generic `accent`/`secondary`/`muted`.
- [ ] No new `text-foreground/NN` opacity hacks — use `text-muted-foreground`.