From c19e6382d6539916fffcdc570e5499c05bf430f7 Mon Sep 17 00:00:00 2001 From: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Date: Thu, 16 Jul 2026 23:39:29 +0000 Subject: [PATCH] chore(rust-gateway): drop out-of-scope skill and CLAUDE.md policy changes Remove the .agents skill files and the litellm-rust/CLAUDE.md port-parity additions from this PR; they are repository-policy changes unrelated to the OCR transport. The custom_llm_provider provider-resolution code and tests stay. Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> --- .../skills/porting-python-to-rust/SKILL.md | 48 -------------- .../skills/testing-rust-axum-gateway/SKILL.md | 65 ------------------- litellm-rust/CLAUDE.md | 17 ----- 3 files changed, 130 deletions(-) delete mode 100644 .agents/skills/porting-python-to-rust/SKILL.md delete mode 100644 .agents/skills/testing-rust-axum-gateway/SKILL.md diff --git a/.agents/skills/porting-python-to-rust/SKILL.md b/.agents/skills/porting-python-to-rust/SKILL.md deleted file mode 100644 index 55360133ddd..00000000000 --- a/.agents/skills/porting-python-to-rust/SKILL.md +++ /dev/null @@ -1,48 +0,0 @@ ---- -name: porting-python-to-rust -description: How to port LiteLLM Python behavior into the litellm-rust crates (litellm-core, litellm-ai-gateway, litellm-python-bridge) so the Rust path matches Python exactly. Use this whenever adding or changing any Rust route, provider transform, router, or config resolution, especially anything that resolves a provider from a model. ---- - -# Porting Python behavior into litellm-rust - -The Rust workspace ports LiteLLM's Python behavior; it does not reinvent it. When Python already has an abstraction for something, mirror it; do not invent a local shortcut that diverges from Python. Read the Python source first, then port its contract and field names. - -## Mirror Python abstractions; never hand-roll what Python centralizes - -The most common mistake is reimplementing logic Python already owns in one place. Provider resolution is the canonical example: - -- WRONG: hand-rolling `model.split('/')` inline in a route (a bespoke `split_provider`). This ignores the deployment's explicit `custom_llm_provider` and drifts from Python. -- RIGHT: deployments carry `custom_llm_provider` in `litellm_params` (mirroring Python's `litellm_params`), and provider resolution goes through `litellm_core::get_llm_provider::get_llm_provider`, the pure port of `litellm.get_llm_provider` (`litellm/litellm_core_utils/get_llm_provider_logic.py`). - -`get_llm_provider(model, custom_llm_provider)` precedence, matching Python: -- an explicit `custom_llm_provider` wins; a matching `provider/` prefix is stripped from the model, otherwise the model is returned unchanged -- with no explicit provider, the `provider/model` prefix is split off -- empty or whitespace-only values are treated as absent (host-layer credential/config rule) - -Before writing any new "figure out the provider / api_base / api_key from a model" logic, use `get_llm_provider`. If it is missing a Python behavior you need (api_base based inference, azure ai studio / cohere_chat aliasing), extend that one helper to match Python rather than branching in the caller. - -## Deployment fields mirror Python litellm_params - -`litellm_core::router::LiteLLMParams` mirrors Python's `litellm_params` dict. When Python reads a field from `litellm_params` (e.g. `custom_llm_provider`, `api_key`, `api_base`), add it here with `#[serde(default)]` so it deserializes straight from the proxy `model_list` (the gateway loads deployments as JSON via `litellm.proxy.read_model_list`). Do not resolve provider identity from the model string when the deployment already carries it explicitly. - -## Crate boundaries (see litellm-rust/AGENTS.md + CLAUDE.md) - -- `litellm-core`: pure translation only (types, route contracts, provider transforms, router, and resolution helpers like `get_llm_provider`). No network, env reads, filesystem, or server behavior. New pure port logic belongs here so it is unit-testable and reusable. -- `litellm-ai-gateway`: the only crate that does I/O. Axum routes, auth, config resolution, network. Routes stay thin (handler validates + delegates to a no-axum `service`); Axum types never leak into the service layer. -- `litellm-python-bridge`: thin PyO3 adapter; no business logic. - -Do not add a crate for a new route or provider; add a module. - -## Prove parity with tests - -Port behavior comes with tests that would fail if the port drifts from Python. For provider/model resolution that means covering: explicit `custom_llm_provider` with and without a matching prefix, the `provider/model` split fallback, blank/whitespace provider, and the no-provider error. Add a gateway-level test that a deployment configured with an explicit `custom_llm_provider` and a bare model (no prefix) still routes correctly. - -## Required checks before pushing - -From `litellm-rust`: -``` -cargo fmt --check -cargo clippy -p litellm-ai-gateway --all-targets --features server -- -D warnings -cargo clippy -p litellm-core -p litellm-python-bridge --all-targets -- -D warnings -cargo test --workspace -``` diff --git a/.agents/skills/testing-rust-axum-gateway/SKILL.md b/.agents/skills/testing-rust-axum-gateway/SKILL.md deleted file mode 100644 index c597a633a59..00000000000 --- a/.agents/skills/testing-rust-axum-gateway/SKILL.md +++ /dev/null @@ -1,65 +0,0 @@ ---- -name: testing-rust-axum-gateway -description: How to build, run, and E2E-test the standalone Rust Axum gateway (litellm-ai-gateway), including the OCR routes (/v1/ocr, /ocr) against a live provider, plus the Python baseline proxy for parity. ---- - -# Testing the LiteLLM Rust Axum gateway (litellm-ai-gateway) - -## Build the gateway binary (with config loading) -Loading `model_list` from a proxy YAML requires the `python-config` feature, which links libpython. From `litellm-rust`: - -``` -PYO3_PYTHON=/.venv/bin/python \ -RUSTFLAGS="-L native=/usr/lib/python3.10/config-3.10-x86_64-linux-gnu" \ -cargo build --release -p litellm-ai-gateway --bin litellm-ai-gateway --features "server python-config" -``` -- `PYO3_PYTHON` must point at a python with a shared libpython. -- The `RUSTFLAGS` link-search is needed because the dev `libpython3.10.so` symlink lives under that config dir. -- If you hit `undefined symbol: Py...` from a stale extension-module build, run `cargo clean -p pyo3 -p pyo3-ffi` first. - -## Python deps the embedded interpreter needs -The gateway's config reader (`litellm.proxy.read_model_list`) and the Python baseline proxy both need `litellm[proxy]` -extras. The bundled `.venv` is often missing them. The `.venv` has no `pip`; use uv WITHOUT touching uv.lock: -``` -VIRTUAL_ENV=/.venv uv pip install orjson apscheduler uvloop python-dotenv \ - gunicorn uvicorn fastapi starlette backoff pyyaml rq fastapi-sso PyJWT python-multipart \ - cryptography pynacl websockets boto3 mcp RestrictedPython rich pydantic-settings expression litellm-proxy-extras -``` -Do NOT let `uv run`/`uv sync` rewrite/commit `uv.lock`. Blueprint `uv sync --inexact --frozen` does NOT pull the -proxy extra, so these must be added (ideally `uv sync ... --extra proxy`). - -## Minimal config (avoids redis/DB/callbacks) -```yaml -general_settings: - master_key: os.environ/LITELLM_MASTER_KEY -model_list: - - model_name: rust-ocr-mistral - litellm_params: - model: mistral/mistral-ocr-latest - api_key: os.environ/MISTRAL_API_KEY -``` - -## Run both servers -Gateway (expect boot log `loaded model_list from via python config reader`): -``` -LITELLM_MASTER_KEY=sk-e2e-local HOST=127.0.0.1 PORT=4001 LITELLM_CONFIG_PATH= \ -PYTHONPATH=/.venv/lib/python3.10/site-packages: \ -./litellm-rust/target/release/litellm-ai-gateway -``` -Python baseline: -``` -LITELLM_MASTER_KEY=sk-e2e-local PYTHONPATH=:/.venv/lib/python3.10/site-packages \ -.venv/bin/python litellm/proxy/proxy_cli.py --config --port 4000 --host 127.0.0.1 -``` - -## Sanity checks / assertions for OCR -- `GET /health/readiness` -> 200; `POST /v1/ocr` with no bearer -> 401 (fails closed). -- OCR success: HTTP 200, `object=="ocr"`, non-empty `model`, `pages` list len>0 with `pages[0].markdown`, and - non-empty `usage_info` (Mistral serializes usage as `usage_info`, not `usage`). -- Gateway and Python baseline both echo `model` as the requested alias (e.g. `rust-ocr-mistral`); the gateway - normalizes the public response `model` back to the alias after the provider call, matching the baseline. -- Never print the provider API key, base64 payloads, raw provider bodies, or full OCR text. Pipe curl through a jq - projection like `{object, model, pages_len:(.pages|length), md_head:(.pages[0].markdown[0:40]), usage_info}`. - -## Devin Secrets Needed -- `MISTRAL_API_KEY` (live Mistral OCR calls) diff --git a/litellm-rust/CLAUDE.md b/litellm-rust/CLAUDE.md index 971dbf6197d..7c723e570ef 100644 --- a/litellm-rust/CLAUDE.md +++ b/litellm-rust/CLAUDE.md @@ -39,23 +39,6 @@ Not allowed in `core`: Python owns rollout state and fallback while Rust is being introduced. Rust paths must be off by default until parity tests prove equivalence with Python. -## Port parity: mirror Python, do not hand-roll - -This workspace ports LiteLLM's Python behavior; it does not reinvent it. When -Python already centralizes a piece of logic, mirror that abstraction and its field names; -do not reimplement a local shortcut that diverges from Python. See the -`porting-python-to-rust` skill and always apply it when adding or changing a Rust -route, provider transform, router, or config resolution. - -Provider resolution is the canonical rule: a deployment's provider comes from its -`custom_llm_provider` (in `litellm_params`), resolved through -`litellm_core::get_llm_provider` (the pure port of `litellm.get_llm_provider`). -Never hand-roll `model.split('/')` in a route (no bespoke `split_provider`): that -ignores the explicit `custom_llm_provider` and drifts from Python. `LiteLLMParams` -carries `custom_llm_provider` so it deserializes straight from the proxy -`model_list`. If `get_llm_provider` lacks a Python behavior you need, extend that -one helper to match Python rather than branching in the caller. - ## Production Bar Rust code in this workspace is held to a strict parity and robustness bar from