mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
chore(rust-gateway): drop out-of-scope skill and CLAUDE.md policy changes
Remove the .agents skill files and the litellm-rust/CLAUDE.md port-parity additions from this PR; they are repository-policy changes unrelated to the OCR transport. The custom_llm_provider provider-resolution code and tests stay. Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
This commit is contained in:
parent
304501153d
commit
c19e6382d6
3 changed files with 0 additions and 130 deletions
|
|
@ -1,48 +0,0 @@
|
|||
---
|
||||
name: porting-python-to-rust
|
||||
description: How to port LiteLLM Python behavior into the litellm-rust crates (litellm-core, litellm-ai-gateway, litellm-python-bridge) so the Rust path matches Python exactly. Use this whenever adding or changing any Rust route, provider transform, router, or config resolution, especially anything that resolves a provider from a model.
|
||||
---
|
||||
|
||||
# Porting Python behavior into litellm-rust
|
||||
|
||||
The Rust workspace ports LiteLLM's Python behavior; it does not reinvent it. When Python already has an abstraction for something, mirror it; do not invent a local shortcut that diverges from Python. Read the Python source first, then port its contract and field names.
|
||||
|
||||
## Mirror Python abstractions; never hand-roll what Python centralizes
|
||||
|
||||
The most common mistake is reimplementing logic Python already owns in one place. Provider resolution is the canonical example:
|
||||
|
||||
- WRONG: hand-rolling `model.split('/')` inline in a route (a bespoke `split_provider`). This ignores the deployment's explicit `custom_llm_provider` and drifts from Python.
|
||||
- RIGHT: deployments carry `custom_llm_provider` in `litellm_params` (mirroring Python's `litellm_params`), and provider resolution goes through `litellm_core::get_llm_provider::get_llm_provider`, the pure port of `litellm.get_llm_provider` (`litellm/litellm_core_utils/get_llm_provider_logic.py`).
|
||||
|
||||
`get_llm_provider(model, custom_llm_provider)` precedence, matching Python:
|
||||
- an explicit `custom_llm_provider` wins; a matching `provider/` prefix is stripped from the model, otherwise the model is returned unchanged
|
||||
- with no explicit provider, the `provider/model` prefix is split off
|
||||
- empty or whitespace-only values are treated as absent (host-layer credential/config rule)
|
||||
|
||||
Before writing any new "figure out the provider / api_base / api_key from a model" logic, use `get_llm_provider`. If it is missing a Python behavior you need (api_base based inference, azure ai studio / cohere_chat aliasing), extend that one helper to match Python rather than branching in the caller.
|
||||
|
||||
## Deployment fields mirror Python litellm_params
|
||||
|
||||
`litellm_core::router::LiteLLMParams` mirrors Python's `litellm_params` dict. When Python reads a field from `litellm_params` (e.g. `custom_llm_provider`, `api_key`, `api_base`), add it here with `#[serde(default)]` so it deserializes straight from the proxy `model_list` (the gateway loads deployments as JSON via `litellm.proxy.read_model_list`). Do not resolve provider identity from the model string when the deployment already carries it explicitly.
|
||||
|
||||
## Crate boundaries (see litellm-rust/AGENTS.md + CLAUDE.md)
|
||||
|
||||
- `litellm-core`: pure translation only (types, route contracts, provider transforms, router, and resolution helpers like `get_llm_provider`). No network, env reads, filesystem, or server behavior. New pure port logic belongs here so it is unit-testable and reusable.
|
||||
- `litellm-ai-gateway`: the only crate that does I/O. Axum routes, auth, config resolution, network. Routes stay thin (handler validates + delegates to a no-axum `service`); Axum types never leak into the service layer.
|
||||
- `litellm-python-bridge`: thin PyO3 adapter; no business logic.
|
||||
|
||||
Do not add a crate for a new route or provider; add a module.
|
||||
|
||||
## Prove parity with tests
|
||||
|
||||
Port behavior comes with tests that would fail if the port drifts from Python. For provider/model resolution that means covering: explicit `custom_llm_provider` with and without a matching prefix, the `provider/model` split fallback, blank/whitespace provider, and the no-provider error. Add a gateway-level test that a deployment configured with an explicit `custom_llm_provider` and a bare model (no prefix) still routes correctly.
|
||||
|
||||
## Required checks before pushing
|
||||
|
||||
From `litellm-rust`:
|
||||
```
|
||||
cargo fmt --check
|
||||
cargo clippy -p litellm-ai-gateway --all-targets --features server -- -D warnings
|
||||
cargo clippy -p litellm-core -p litellm-python-bridge --all-targets -- -D warnings
|
||||
cargo test --workspace
|
||||
```
|
||||
|
|
@ -1,65 +0,0 @@
|
|||
---
|
||||
name: testing-rust-axum-gateway
|
||||
description: How to build, run, and E2E-test the standalone Rust Axum gateway (litellm-ai-gateway), including the OCR routes (/v1/ocr, /ocr) against a live provider, plus the Python baseline proxy for parity.
|
||||
---
|
||||
|
||||
# Testing the LiteLLM Rust Axum gateway (litellm-ai-gateway)
|
||||
|
||||
## Build the gateway binary (with config loading)
|
||||
Loading `model_list` from a proxy YAML requires the `python-config` feature, which links libpython. From `litellm-rust`:
|
||||
|
||||
```
|
||||
PYO3_PYTHON=<repo>/.venv/bin/python \
|
||||
RUSTFLAGS="-L native=/usr/lib/python3.10/config-3.10-x86_64-linux-gnu" \
|
||||
cargo build --release -p litellm-ai-gateway --bin litellm-ai-gateway --features "server python-config"
|
||||
```
|
||||
- `PYO3_PYTHON` must point at a python with a shared libpython.
|
||||
- The `RUSTFLAGS` link-search is needed because the dev `libpython3.10.so` symlink lives under that config dir.
|
||||
- If you hit `undefined symbol: Py...` from a stale extension-module build, run `cargo clean -p pyo3 -p pyo3-ffi` first.
|
||||
|
||||
## Python deps the embedded interpreter needs
|
||||
The gateway's config reader (`litellm.proxy.read_model_list`) and the Python baseline proxy both need `litellm[proxy]`
|
||||
extras. The bundled `.venv` is often missing them. The `.venv` has no `pip`; use uv WITHOUT touching uv.lock:
|
||||
```
|
||||
VIRTUAL_ENV=<repo>/.venv uv pip install orjson apscheduler uvloop python-dotenv \
|
||||
gunicorn uvicorn fastapi starlette backoff pyyaml rq fastapi-sso PyJWT python-multipart \
|
||||
cryptography pynacl websockets boto3 mcp RestrictedPython rich pydantic-settings expression litellm-proxy-extras
|
||||
```
|
||||
Do NOT let `uv run`/`uv sync` rewrite/commit `uv.lock`. Blueprint `uv sync --inexact --frozen` does NOT pull the
|
||||
proxy extra, so these must be added (ideally `uv sync ... --extra proxy`).
|
||||
|
||||
## Minimal config (avoids redis/DB/callbacks)
|
||||
```yaml
|
||||
general_settings:
|
||||
master_key: os.environ/LITELLM_MASTER_KEY
|
||||
model_list:
|
||||
- model_name: rust-ocr-mistral
|
||||
litellm_params:
|
||||
model: mistral/mistral-ocr-latest
|
||||
api_key: os.environ/MISTRAL_API_KEY
|
||||
```
|
||||
|
||||
## Run both servers
|
||||
Gateway (expect boot log `loaded model_list from <path> via python config reader`):
|
||||
```
|
||||
LITELLM_MASTER_KEY=sk-e2e-local HOST=127.0.0.1 PORT=4001 LITELLM_CONFIG_PATH=<minimal.yml> \
|
||||
PYTHONPATH=<repo>/.venv/lib/python3.10/site-packages:<repo> \
|
||||
./litellm-rust/target/release/litellm-ai-gateway
|
||||
```
|
||||
Python baseline:
|
||||
```
|
||||
LITELLM_MASTER_KEY=sk-e2e-local PYTHONPATH=<repo>:<repo>/.venv/lib/python3.10/site-packages \
|
||||
.venv/bin/python litellm/proxy/proxy_cli.py --config <minimal.yml> --port 4000 --host 127.0.0.1
|
||||
```
|
||||
|
||||
## Sanity checks / assertions for OCR
|
||||
- `GET /health/readiness` -> 200; `POST /v1/ocr` with no bearer -> 401 (fails closed).
|
||||
- OCR success: HTTP 200, `object=="ocr"`, non-empty `model`, `pages` list len>0 with `pages[0].markdown`, and
|
||||
non-empty `usage_info` (Mistral serializes usage as `usage_info`, not `usage`).
|
||||
- Gateway and Python baseline both echo `model` as the requested alias (e.g. `rust-ocr-mistral`); the gateway
|
||||
normalizes the public response `model` back to the alias after the provider call, matching the baseline.
|
||||
- Never print the provider API key, base64 payloads, raw provider bodies, or full OCR text. Pipe curl through a jq
|
||||
projection like `{object, model, pages_len:(.pages|length), md_head:(.pages[0].markdown[0:40]), usage_info}`.
|
||||
|
||||
## Devin Secrets Needed
|
||||
- `MISTRAL_API_KEY` (live Mistral OCR calls)
|
||||
|
|
@ -39,23 +39,6 @@ Not allowed in `core`:
|
|||
Python owns rollout state and fallback while Rust is being introduced. Rust
|
||||
paths must be off by default until parity tests prove equivalence with Python.
|
||||
|
||||
## Port parity: mirror Python, do not hand-roll
|
||||
|
||||
This workspace ports LiteLLM's Python behavior; it does not reinvent it. When
|
||||
Python already centralizes a piece of logic, mirror that abstraction and its field names;
|
||||
do not reimplement a local shortcut that diverges from Python. See the
|
||||
`porting-python-to-rust` skill and always apply it when adding or changing a Rust
|
||||
route, provider transform, router, or config resolution.
|
||||
|
||||
Provider resolution is the canonical rule: a deployment's provider comes from its
|
||||
`custom_llm_provider` (in `litellm_params`), resolved through
|
||||
`litellm_core::get_llm_provider` (the pure port of `litellm.get_llm_provider`).
|
||||
Never hand-roll `model.split('/')` in a route (no bespoke `split_provider`): that
|
||||
ignores the explicit `custom_llm_provider` and drifts from Python. `LiteLLMParams`
|
||||
carries `custom_llm_provider` so it deserializes straight from the proxy
|
||||
`model_list`. If `get_llm_provider` lacks a Python behavior you need, extend that
|
||||
one helper to match Python rather than branching in the caller.
|
||||
|
||||
## Production Bar
|
||||
|
||||
Rust code in this workspace is held to a strict parity and robustness bar from
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue