chore(rust-gateway): drop out-of-scope skill and CLAUDE.md policy changes

Remove the .agents skill files and the litellm-rust/CLAUDE.md port-parity additions from this PR; they are repository-policy changes unrelated to the OCR transport. The custom_llm_provider provider-resolution code and tests stay.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
This commit is contained in:
Devin AI 2026-07-16 23:39:29 +00:00
parent 304501153d
commit c19e6382d6
3 changed files with 0 additions and 130 deletions

View file

@ -1,48 +0,0 @@
---
name: porting-python-to-rust
description: How to port LiteLLM Python behavior into the litellm-rust crates (litellm-core, litellm-ai-gateway, litellm-python-bridge) so the Rust path matches Python exactly. Use this whenever adding or changing any Rust route, provider transform, router, or config resolution, especially anything that resolves a provider from a model.
---
# Porting Python behavior into litellm-rust
The Rust workspace ports LiteLLM's Python behavior; it does not reinvent it. When Python already has an abstraction for something, mirror it; do not invent a local shortcut that diverges from Python. Read the Python source first, then port its contract and field names.
## Mirror Python abstractions; never hand-roll what Python centralizes
The most common mistake is reimplementing logic Python already owns in one place. Provider resolution is the canonical example:
- WRONG: hand-rolling `model.split('/')` inline in a route (a bespoke `split_provider`). This ignores the deployment's explicit `custom_llm_provider` and drifts from Python.
- RIGHT: deployments carry `custom_llm_provider` in `litellm_params` (mirroring Python's `litellm_params`), and provider resolution goes through `litellm_core::get_llm_provider::get_llm_provider`, the pure port of `litellm.get_llm_provider` (`litellm/litellm_core_utils/get_llm_provider_logic.py`).
`get_llm_provider(model, custom_llm_provider)` precedence, matching Python:
- an explicit `custom_llm_provider` wins; a matching `provider/` prefix is stripped from the model, otherwise the model is returned unchanged
- with no explicit provider, the `provider/model` prefix is split off
- empty or whitespace-only values are treated as absent (host-layer credential/config rule)
Before writing any new "figure out the provider / api_base / api_key from a model" logic, use `get_llm_provider`. If it is missing a Python behavior you need (api_base based inference, azure ai studio / cohere_chat aliasing), extend that one helper to match Python rather than branching in the caller.
## Deployment fields mirror Python litellm_params
`litellm_core::router::LiteLLMParams` mirrors Python's `litellm_params` dict. When Python reads a field from `litellm_params` (e.g. `custom_llm_provider`, `api_key`, `api_base`), add it here with `#[serde(default)]` so it deserializes straight from the proxy `model_list` (the gateway loads deployments as JSON via `litellm.proxy.read_model_list`). Do not resolve provider identity from the model string when the deployment already carries it explicitly.
## Crate boundaries (see litellm-rust/AGENTS.md + CLAUDE.md)
- `litellm-core`: pure translation only (types, route contracts, provider transforms, router, and resolution helpers like `get_llm_provider`). No network, env reads, filesystem, or server behavior. New pure port logic belongs here so it is unit-testable and reusable.
- `litellm-ai-gateway`: the only crate that does I/O. Axum routes, auth, config resolution, network. Routes stay thin (handler validates + delegates to a no-axum `service`); Axum types never leak into the service layer.
- `litellm-python-bridge`: thin PyO3 adapter; no business logic.
Do not add a crate for a new route or provider; add a module.
## Prove parity with tests
Port behavior comes with tests that would fail if the port drifts from Python. For provider/model resolution that means covering: explicit `custom_llm_provider` with and without a matching prefix, the `provider/model` split fallback, blank/whitespace provider, and the no-provider error. Add a gateway-level test that a deployment configured with an explicit `custom_llm_provider` and a bare model (no prefix) still routes correctly.
## Required checks before pushing
From `litellm-rust`:
```
cargo fmt --check
cargo clippy -p litellm-ai-gateway --all-targets --features server -- -D warnings
cargo clippy -p litellm-core -p litellm-python-bridge --all-targets -- -D warnings
cargo test --workspace
```

View file

@ -1,65 +0,0 @@
---
name: testing-rust-axum-gateway
description: How to build, run, and E2E-test the standalone Rust Axum gateway (litellm-ai-gateway), including the OCR routes (/v1/ocr, /ocr) against a live provider, plus the Python baseline proxy for parity.
---
# Testing the LiteLLM Rust Axum gateway (litellm-ai-gateway)
## Build the gateway binary (with config loading)
Loading `model_list` from a proxy YAML requires the `python-config` feature, which links libpython. From `litellm-rust`:
```
PYO3_PYTHON=<repo>/.venv/bin/python \
RUSTFLAGS="-L native=/usr/lib/python3.10/config-3.10-x86_64-linux-gnu" \
cargo build --release -p litellm-ai-gateway --bin litellm-ai-gateway --features "server python-config"
```
- `PYO3_PYTHON` must point at a python with a shared libpython.
- The `RUSTFLAGS` link-search is needed because the dev `libpython3.10.so` symlink lives under that config dir.
- If you hit `undefined symbol: Py...` from a stale extension-module build, run `cargo clean -p pyo3 -p pyo3-ffi` first.
## Python deps the embedded interpreter needs
The gateway's config reader (`litellm.proxy.read_model_list`) and the Python baseline proxy both need `litellm[proxy]`
extras. The bundled `.venv` is often missing them. The `.venv` has no `pip`; use uv WITHOUT touching uv.lock:
```
VIRTUAL_ENV=<repo>/.venv uv pip install orjson apscheduler uvloop python-dotenv \
gunicorn uvicorn fastapi starlette backoff pyyaml rq fastapi-sso PyJWT python-multipart \
cryptography pynacl websockets boto3 mcp RestrictedPython rich pydantic-settings expression litellm-proxy-extras
```
Do NOT let `uv run`/`uv sync` rewrite/commit `uv.lock`. Blueprint `uv sync --inexact --frozen` does NOT pull the
proxy extra, so these must be added (ideally `uv sync ... --extra proxy`).
## Minimal config (avoids redis/DB/callbacks)
```yaml
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
model_list:
- model_name: rust-ocr-mistral
litellm_params:
model: mistral/mistral-ocr-latest
api_key: os.environ/MISTRAL_API_KEY
```
## Run both servers
Gateway (expect boot log `loaded model_list from <path> via python config reader`):
```
LITELLM_MASTER_KEY=sk-e2e-local HOST=127.0.0.1 PORT=4001 LITELLM_CONFIG_PATH=<minimal.yml> \
PYTHONPATH=<repo>/.venv/lib/python3.10/site-packages:<repo> \
./litellm-rust/target/release/litellm-ai-gateway
```
Python baseline:
```
LITELLM_MASTER_KEY=sk-e2e-local PYTHONPATH=<repo>:<repo>/.venv/lib/python3.10/site-packages \
.venv/bin/python litellm/proxy/proxy_cli.py --config <minimal.yml> --port 4000 --host 127.0.0.1
```
## Sanity checks / assertions for OCR
- `GET /health/readiness` -> 200; `POST /v1/ocr` with no bearer -> 401 (fails closed).
- OCR success: HTTP 200, `object=="ocr"`, non-empty `model`, `pages` list len>0 with `pages[0].markdown`, and
non-empty `usage_info` (Mistral serializes usage as `usage_info`, not `usage`).
- Gateway and Python baseline both echo `model` as the requested alias (e.g. `rust-ocr-mistral`); the gateway
normalizes the public response `model` back to the alias after the provider call, matching the baseline.
- Never print the provider API key, base64 payloads, raw provider bodies, or full OCR text. Pipe curl through a jq
projection like `{object, model, pages_len:(.pages|length), md_head:(.pages[0].markdown[0:40]), usage_info}`.
## Devin Secrets Needed
- `MISTRAL_API_KEY` (live Mistral OCR calls)

View file

@ -39,23 +39,6 @@ Not allowed in `core`:
Python owns rollout state and fallback while Rust is being introduced. Rust
paths must be off by default until parity tests prove equivalence with Python.
## Port parity: mirror Python, do not hand-roll
This workspace ports LiteLLM's Python behavior; it does not reinvent it. When
Python already centralizes a piece of logic, mirror that abstraction and its field names;
do not reimplement a local shortcut that diverges from Python. See the
`porting-python-to-rust` skill and always apply it when adding or changing a Rust
route, provider transform, router, or config resolution.
Provider resolution is the canonical rule: a deployment's provider comes from its
`custom_llm_provider` (in `litellm_params`), resolved through
`litellm_core::get_llm_provider` (the pure port of `litellm.get_llm_provider`).
Never hand-roll `model.split('/')` in a route (no bespoke `split_provider`): that
ignores the explicit `custom_llm_provider` and drifts from Python. `LiteLLMParams`
carries `custom_llm_provider` so it deserializes straight from the proxy
`model_list`. If `get_llm_provider` lacks a Python behavior you need, extend that
one helper to match Python rather than branching in the caller.
## Production Bar
Rust code in this workspace is held to a strict parity and robustness bar from