diff --git a/litellm-rust/Cargo.lock b/litellm-rust/Cargo.lock index 6ae8fe727dc..a8f69829cd5 100644 --- a/litellm-rust/Cargo.lock +++ b/litellm-rust/Cargo.lock @@ -4003,45 +4003,30 @@ dependencies = [ name = "litellm-inference" version = "0.1.0" dependencies = [ - "base64 0.22.1", "bytes", "futures-util", "litellm-auth", - "litellm-auth-aws", - "litellm-auth-gcp", "litellm-cache", "litellm-cache-memory", "litellm-cache-response", "litellm-core-utils", "litellm-framer", "litellm-host", - "litellm-host-native", "litellm-http", "litellm-inference", "litellm-llms", - "litellm-llms-types", "litellm-secrets", "litellm-tracing", - "mime_guess", - "moka", - "rand 0.8.7", "reqwest 0.12.28", "rstest", - "rstest_reuse", "serde", "serde_json", - "sha2 0.10.9", - "strum", - "subtle", "thiserror 2.0.19", "time", "tokio", - "tokio-tungstenite", "tokio-util", "tracing", "url", - "veil", - "wiremock", ] [[package]] @@ -6186,17 +6171,6 @@ dependencies = [ "unicode-ident", ] -[[package]] -name = "rstest_reuse" -version = "0.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b3a8fb4672e840a587a66fc577a5491375df51ddb88f2a2c2a792598c326fe14" -dependencies = [ - "quote", - "rand 0.8.7", - "syn 2.0.119", -] - [[package]] name = "rusqlite" version = "0.39.0" diff --git a/litellm-rust/crates/gateway-inference/AGENTS.md b/litellm-rust/crates/gateway-inference/AGENTS.md index a03cd9cbbb6..73340982811 100644 --- a/litellm-rust/crates/gateway-inference/AGENTS.md +++ b/litellm-rust/crates/gateway-inference/AGENTS.md @@ -1,6 +1,6 @@ - Expose a mountable Axum router; listener binding, server lifecycle, and shared inbound middleware belong to `gateway` - Own endpoint paths, request parsing, model alias resolution, and API-specific response and SSE error formats; delegate hosted call execution and HTTP body delivery to host-http -- Delegate inference execution to `core` and provider transformations and authentication to `llms` and the auth crates; do not duplicate them in handlers -- Let core validate inference fields and supported features, then map its errors to HTTP responses; do not add gateway checks for temporary core limitations +- Delegate inference execution to the `inference-` crates and provider transformations and authentication to `llms` and the auth crates; do not duplicate them in handlers +- Let the `inference-` crates validate inference fields and supported features, then map their errors to HTTP responses; do not add gateway checks for temporary inference limitations - Use injected deployments, HTTP pools, settings, and secret sources; do not load process configuration or construct independent clients in handlers -- Test HTTP contracts here, including status codes, forwarded headers, error envelopes, and streaming behavior; keep core and provider tests in their owning crates +- Test HTTP contracts here, including status codes, forwarded headers, error envelopes, and streaming behavior; keep inference and provider tests in their owning crates diff --git a/litellm-rust/crates/gateway/AGENTS.md b/litellm-rust/crates/gateway/AGENTS.md index 10a42b9ec3b..3c417111233 100644 --- a/litellm-rust/crates/gateway/AGENTS.md +++ b/litellm-rust/crates/gateway/AGENTS.md @@ -1,5 +1,5 @@ - Keep this crate a thin composition layer: mount endpoint routers and serve the supplied listener - Server lifecycle and shared inbound middleware belong here, including client authentication, rate limiting, and request logging -- Endpoint paths, request handling, model resolution, and response encoding belong to the mounted crates; provider execution belongs to `core` and `llms` +- Endpoint paths, request handling, model resolution, and response encoding belong to the mounted crates; provider execution belongs to the `inference-` crates and `llms` - Inject shared state and infrastructure; avoid global runtimes, duplicate client pools, and abstractions for hypothetical endpoint groups - Test mounting and server lifecycle through public HTTP behavior; test endpoint semantics in the owning crate diff --git a/litellm-rust/crates/inference-chat/AGENTS.md b/litellm-rust/crates/inference-chat/AGENTS.md new file mode 100644 index 00000000000..36dd72fc011 --- /dev/null +++ b/litellm-rust/crates/inference-chat/AGENTS.md @@ -0,0 +1,3 @@ +- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md) +- This crate owns Chat Completions call orchestration; payloads belong in `litellm-llms-types`, transformations in `llms/src/base_llm/chat` and `llms/src//chat` +- Chat Completions is currently non-streaming: the route returns its completed response, cached through `litellm_inference::caching::execute_unary` diff --git a/litellm-rust/crates/inference-messages/AGENTS.md b/litellm-rust/crates/inference-messages/AGENTS.md index 0feff9c30c2..c992fe60fe9 100644 --- a/litellm-rust/crates/inference-messages/AGENTS.md +++ b/litellm-rust/crates/inference-messages/AGENTS.md @@ -1,7 +1,14 @@ -This directory owns provider-independent Messages call orchestration: the entrypoint, call envelopes, provider selection, credential resolution, transport coordination, hooks, and stream lifecycle. Shared API data contracts belong in `litellm-llms-types::formats::messages`, adapter contracts and execution inputs in `llms/src/base_llm/messages`, and provider implementations in `llms/src//messages` - -Select concrete provider adapters and invoke their contracts. Delegate authentication policy, beta selection, payload rewriting, and response interpretation to those adapters. Keep provider policy out of request preparation and transport handlers. Calling a concrete provider helper for every provider is still a policy dependency - -Route types such as `MessagesCall`, prepared requests, and response wrappers containing live streams describe execution. Reuse the shared Messages payload types inside them instead of defining another request or response schema here - -Preserve the order of validation, normalization, caller-requested parameter removal, and provider transformation when that order affects observable behavior. Test provider dispatch, auth precedence, header handling, transformations, and responses through behavior, not source structure +- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md) +- This crate owns provider-independent Messages call orchestration: the entrypoint, call envelopes, provider selection, credential resolution, transport coordination, hooks, and stream lifecycle + - shared API data contracts belong in `litellm-llms-types::formats::messages` + - adapter contracts and execution inputs in `llms/src/base_llm/messages` + - provider implementations in `llms/src//messages` +- Select concrete provider adapters and invoke their contracts + - delegate authentication policy, beta selection, payload rewriting, and response interpretation to those adapters + - keep provider policy out of request preparation and transport handlers + - calling a concrete provider helper for every provider is still a policy dependency +- Route types such as `MessagesCall`, prepared requests, and response wrappers containing live streams describe execution; reuse the shared Messages payload types inside them instead of defining another request or response schema here +- The route returns `litellm_host::call::CallOutput`: a completed response, or a stream head and chunks +- Per-call dependencies are grouped in `litellm_inference::context::CallContext`; `src/lib.rs` explicitly sequences cache lookup, provider execution, result acceptance (`CallContext::result_ready`), and cache storage +- Preserve the order of validation, normalization, caller-requested parameter removal, and provider transformation when that order affects observable behavior +- Test provider dispatch, auth precedence, header handling, transformations, and responses through behavior, not source structure diff --git a/litellm-rust/crates/inference-ocr/AGENTS.md b/litellm-rust/crates/inference-ocr/AGENTS.md new file mode 100644 index 00000000000..c3644e850e0 --- /dev/null +++ b/litellm-rust/crates/inference-ocr/AGENTS.md @@ -0,0 +1,6 @@ +- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md) +- This crate owns OCR call orchestration; payloads belong in `litellm-llms-types`, transformations and the request handler in `llms/src/base_llm/ocr` and `llms/src//ocr` +- The route returns its completed response directly +- Route closures supply OCR-specific host capabilities, such as the caller's Azure AD token provider (`src/route.rs`) +- Provider code reaches the caller's hooks mid-call only through `litellm_llms::base_llm::ocr::handler::CallHooks`, which this crate implements over its host until it folds into `litellm_host::interceptors::Interceptors` +- OCR failures use `base_llm/ocr/error.rs`, the recorded exception to `litellm_inference::RouteError` until OCR folds into `litellm_llms::Error` diff --git a/litellm-rust/crates/inference-responses/AGENTS.md b/litellm-rust/crates/inference-responses/AGENTS.md new file mode 100644 index 00000000000..d3e49a1579a --- /dev/null +++ b/litellm-rust/crates/inference-responses/AGENTS.md @@ -0,0 +1,4 @@ +- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md) +- This crate owns Responses API call orchestration; payloads belong in `litellm-llms-types`, transformations in `llms/src/base_llm/responses` and `llms/src//responses` +- The HTTP route returns `litellm_host::call::CallOutput`: a completed response, or a stream head and chunks, cached through `litellm_inference::caching::execute_streaming` +- WebSocket sessions (`src/websocket.rs`) stay separate from the HTTP call driver because a connection can accept multiple requests while receiving events diff --git a/litellm-rust/crates/inference-responses/Cargo.toml b/litellm-rust/crates/inference-responses/Cargo.toml index 5aaa4a37b4c..f4622ba5db9 100644 --- a/litellm-rust/crates/inference-responses/Cargo.toml +++ b/litellm-rust/crates/inference-responses/Cargo.toml @@ -23,14 +23,10 @@ tokio-tungstenite.workspace = true tracing.workspace = true [dev-dependencies] -futures-util.workspace = true litellm-cache.workspace = true litellm-cache-memory.workspace = true -litellm-cache-response.workspace = true litellm-host-native.workspace = true -litellm-http.workspace = true litellm-inference = { workspace = true, features = ["test-support"] } litellm-tracing.workspace = true rstest.workspace = true -tokio.workspace = true wiremock.workspace = true diff --git a/litellm-rust/crates/inference-transcription/AGENTS.md b/litellm-rust/crates/inference-transcription/AGENTS.md new file mode 100644 index 00000000000..632e5cea92c --- /dev/null +++ b/litellm-rust/crates/inference-transcription/AGENTS.md @@ -0,0 +1,3 @@ +- Shared base and layering rules: [`../inference/AGENTS.md`](../inference/AGENTS.md) +- This crate owns audio transcription call orchestration; transformations belong in `llms/src/base_llm/audio_transcription` and `llms/src//audio_transcription` +- `AudioTranscriptionRoute::execute` returns the completed response directly; there is no hosted route machine or response cache for transcription yet diff --git a/litellm-rust/crates/inference-transcription/Cargo.toml b/litellm-rust/crates/inference-transcription/Cargo.toml index a69bdb127ba..62522161737 100644 --- a/litellm-rust/crates/inference-transcription/Cargo.toml +++ b/litellm-rust/crates/inference-transcription/Cargo.toml @@ -17,7 +17,6 @@ serde_json.workspace = true tracing.workspace = true [dev-dependencies] -litellm-http.workspace = true litellm-inference = { workspace = true, features = ["test-support"] } litellm-tracing.workspace = true rstest.workspace = true diff --git a/litellm-rust/crates/inference/AGENTS.md b/litellm-rust/crates/inference/AGENTS.md index 6baae6d5a48..ad8e24e35c0 100644 --- a/litellm-rust/crates/inference/AGENTS.md +++ b/litellm-rust/crates/inference/AGENTS.md @@ -1,49 +1,83 @@ -litellm-inference owns route orchestration. Messages and HTTP Responses return `litellm_host::call::CallOutput`, containing either a completed response or a stream head and chunks. OCR and currently non-streaming Chat Completions return their completed response directly +## Layering -Hosts assemble route objects from shared `CoreResources`, HTTP settings, and secret sources. Each route owns its provider client and authentication dependencies. Gateway routes live for the gateway lifetime; Python assembles routes per call from its settings snapshot +- `gateway-inference` (serving) -> `inference-` (one API's call) -> `inference` (shared base) +- This crate is the shared base every format crate builds on + - `RouteError` (`src/error.rs`), `CallOptions` (`src/lib.rs`), `CallContext` (`src/context.rs`) + - diagnostic spans (`src/diagnostic.rs`), outbound send and signing (`src/outbound.rs`), provider resolution (`src/provider.rs`), `CoreResources` (`src/resources.rs`) + - response caching (`src/caching.rs`): `Cachable`, `StreamCachable`, `CacheRequest`, `CallCache`, `execute_unary`, `execute_streaming`, stream capture + - shared test helpers behind the `test-support` feature (`src/test_support.rs`) +- Nothing here names an API format; format-specific code, constants, tests and test builders live in their `inference-` crate +- Never depend on an `inference-*` crate from here -Chat Completions, Messages, Responses, and OCR execute through their route objects. Calls pass `Interceptors` and an optional `ObservationSender` separately; use `&()` for no hooks and `None` for no observer. Construction does no work; preparation and lifecycle observation begin when the future is polled. Handlers accept `Interceptors`, never a concrete `ChannelInterceptors`. Native observers receive start and terminal events through the shared call runner; a stream retains its lifecycle until exhaustion, error, or drop. Hosted routes leave terminal observation to their driver +## Format crates -`route.rs` declares the concrete `Protocol` and implements a route method that accepts a typed request and constructs a `litellm_host::call::HostedMachine` with `hosted_call`. The shared call plumbing owns stream opening, delivery, backpressure, and detachment. Request decoding belongs to the boundary before the machine starts. Route closures only supply execution dependencies and route-specific host capabilities such as an OCR token provider. Use `run_hosted` for a native host so detachment is reported as cancellation. Python uses its own shared driver and preserves caller-task callback execution +- `inference-chat`, `inference-messages`, `inference-responses`, `inference-ocr`, `inference-transcription`, one per API +- Each owns its entrypoint, route request types (`*Request<'a>`), credential fallback, provider dispatch, and the handler glue that runs a provider config +- Each owns its `Protocol`, its `Cachable` / `StreamCachable` impl, its constants, and its tests +- `inference-*` is orchestration only + - shared API data contracts belong in `litellm-llms-types` + - adapter contracts and shared transformation machinery in `llms/src/base_llm//` + - provider policy in `llms/src///` +- Select concrete adapters, then invoke their contracts instead of applying one provider's policy to every call +- Route types describe call envelopes and execution state, not duplicate public payload schemas +- Format-specific rules go in that crate's `AGENTS.md` -Responses WebSocket sessions remain separate from the HTTP call driver because a connection can accept multiple requests while receiving events +## Routes -## Crate layering +- Hosts assemble route objects from shared `CoreResources`, HTTP settings, and secret sources + - each route owns its provider client and authentication dependencies + - gateway routes live for the gateway lifetime; Python assembles routes per call from its settings snapshot +- Calls pass `Interceptors` and an optional `ObservationSender` separately + - use `&()` for no hooks and `None` for no observer + - handlers accept `Interceptors`, never a concrete `ChannelInterceptors` +- Construction does no work; preparation and lifecycle observation begin when the future is polled +- Native observers receive start and terminal events through the shared call runner; a stream retains its lifecycle until exhaustion, error, or drop. Hosted routes leave terminal observation to their driver +- A hosted format's `route.rs` declares the concrete `Protocol` and a route method that takes a typed request and builds a `litellm_host::call::HostedMachine` with `hosted_call`; audio transcription has no hosted machine and returns its response directly + - the shared call plumbing owns stream opening, delivery, backpressure, and detachment + - request decoding belongs to the boundary before the machine starts + - route closures only supply execution dependencies and route-specific host capabilities + - Python uses its own shared driver and preserves caller-task callback execution +- Not here: serving HTTP (axum routes, extractors), config file reading, rollout state, databases, or callback execution of any kind. Routes run as machines that yield host operations and call events; which integrations consume those events is the host's business -For Messages, Responses, Chat Completions, OCR, and other API formats, `inference/src//` owns orchestration. Shared API data contracts belong in `litellm-llms-types`, adapter contracts and shared transformation machinery in `llms/src/base_llm//`, and provider policy in `llms/src///`. A repeated format directory name does not imply interchangeable responsibilities. Select concrete adapters here, then invoke their contracts instead of applying one provider's policy to every call. Route types describe call envelopes and execution state, not duplicate public payload schemas - -Crates separate API data, transformations, transport, and orchestration. Python package names identify counterparts, not ownership. Dependencies only point down: +## Crate boundaries +- Dependencies only point down; Python package names identify counterparts, not ownership - `litellm-llms-types` owns shared inference API contracts, grouped by format: pure serde data and shape validation, no I/O - `litellm-core-utils` mirrors `litellm/litellm_core_utils/`: pure helpers (provider resolution, prompt factory, call arguments, settings lookup and layer merge), no network I/O -- `litellm-http` is Rust-only and route-neutral: settings resolution, the pooled `reqwest` clients, TLS, proxies, the SSRF-safe media fetcher, request and header helpers, and transport errors. Python's `litellm/llms/custom_httpx/` is split by responsibility instead of mirrored: its transport half lives here, its OCR handler in `litellm-llms` -- `litellm-llms` mirrors `litellm/llms/`: `base_llm//transformation.rs`, `//transformation.rs`, and `base_llm/ocr/handler.rs` (the OCR request handler) -- `litellm-inference` mirrors the route packages (`litellm/ocr/`, `litellm/messages/`, ...): entrypoints, route request types, provider dispatch, the route machine, and hooks - -A route module owns the call entrypoint, route request types (`*Request<'a>`), credential fallback, provider dispatch, and the handler glue that runs a provider config. Provider code never imports from inference; when it needs the caller's hooks mid-call it goes through `litellm_llms::base_llm::ocr::handler::CallHooks`, the provider-level hooks OCR implements over its host until it folds into `litellm_host::interceptors::Interceptors`. Import every item from its canonical path. Never re-export another crate's items or give an item a second public path; the only re-export allowed is a private submodule surfacing its item at its module root (`mod error; pub use error::Error;`). Handlers belong in core or llms, never in a host crate +- `litellm-http` is Rust-only and route-neutral: settings resolution, pooled `reqwest` clients, TLS, proxies, the SSRF-safe media fetcher, request and header helpers, and transport errors +- `litellm-llms` mirrors `litellm/llms/`: `base_llm//transformation.rs`, `//transformation.rs`, and `base_llm/ocr/handler.rs` +- `litellm-inference-` mirrors the route packages (`litellm/ocr/`, `litellm/messages/`, ...) +- Provider code never imports from `inference` or `inference-*` +- Handlers belong in `inference-*` or `llms`, never in a host crate +- Import every item from its canonical path. Never re-export another crate's items or give an item a second public path; the only allowed re-exports are a private submodule surfacing its item at its module root (`mod error; pub use error::Error;`) and each format crate's `pub use litellm_inference::RouteError as Error;`, which keeps the pre-split `::Error` name ## Error placement -The workspace `Error definitions` rules shape each crate's error; this section decides which crate and module a failure belongs to +- The workspace `Error definitions` rules shape each crate's error; this section decides which crate and module a failure belongs to +- A failure is declared once, by the lowest crate that raises it + - every crate above nests it unchanged (`#[error(transparent)] Auth(#[from] litellm_auth::Error)`) or maps it once at its boundary, as `src/error.rs` does for `litellm_llms::Error` + - `RouteError` collects route failures for every format crate and never re-declares a variant a lower crate raises +- Scope follows the concept, not the first caller + - an error type under `litellm-llms`'s `/` is private to that provider: no other provider and nothing in `base_llm` may import it + - a failure two providers or two routes can hit (wire framing, stream event decoding, a malformed provider response) belongs to the crate that owns the concept: `litellm-framer` for framing, `litellm_llms::Error` for the transformation layer +- `litellm_llms::Error` (`crates/llms/src/error.rs`) is the one transformation error for every provider and API; `base_llm/ocr/error.rs` is the recorded exception until OCR folds into it -A failure is declared once, by the lowest crate that raises it. Every crate above nests that error unchanged (`#[error(transparent)] Auth(#[from] litellm_auth::Error)`) or maps it once at its boundary, as `src/error.rs` does for `litellm_llms::Error`. `RouteError` collects route failures and never re-declares a variant a lower crate raises +## Response caching and accounting -Scope follows the concept, not the first caller. An error type under `litellm-llms`'s `/` is private to that provider: no other provider and nothing in `base_llm` may import it. A failure two providers or two routes can hit, such as wire framing, stream event decoding, or a malformed provider response, belongs to the crate that owns the concept: `litellm-framer` for framing, `litellm_llms::Error` for the transformation layer - -`litellm_llms::Error` (`crates/llms/src/error.rs`) is the one transformation error for every provider and API. `base_llm/ocr/error.rs` is the recorded exception until OCR folds into it - -Not here: serving HTTP (axum routes, extractors), config file reading, rollout state, databases, or callback execution of any kind. Core runs each route as a machine that yields host operations and call events; which integrations consume those events is the host's business. - -## Response caching and accounting boundary - -Attach a `litellm_cache_response::ScopedCache` with `route.with_cache(cache)`. Cached and uncached routes use the same `execute` and `machine` methods. `CallOptions` carries a scope-free `CachePolicy` and observation; per-call policy never replaces the attached scope or service - -Messages groups per-call dependencies in `CallContext` and explicitly sequences cache lookup, provider execution, result acceptance, and cache storage. Provider transport does not own cache orchestration. Stream capture remains in the shared cache implementation - -Core owns request identity, typed response reconstruction and stream capture/replay. `cache-response` owns cache policy, namespacing, scope encoding, versioned envelopes and freshness. The SDK explicitly chooses shared scope. The gateway derives isolated scope from authenticated identity before attaching its service - -Core delivers `ExecutionFacts` through the awaited `ResultReady` host operation for both provider and cached results, before public response processing or stream opening. Facts carry resolved model/provider and result source, including the hit key. Usage remains in the typed response or delivered stream, where completion and cancellation determine what was actually reported. Passive observation is not an accounting delivery mechanism - -Core does not calculate prices, charge budgets, or update rate-limit counters. The legacy Python callback adapter translates execution facts into the existing Python logging contract; Python remains the accounting owner on that path. Native gateway accounting belongs to gateway dependencies, independently of `host-python`. Response-cache services expose no coordination counters or reservation APIs. A shared Redis deployment does not make response storage and accounting coordination the same dependency - -Cache lookup follows provider preparation, credential resolution and the request interceptor. Keys describe the effective provider URL, authenticated headers and rewritten body. Signed requests bypass caching until the signing identity has a stable cache representation +- Attach a `litellm_cache_response::ScopedCache` with `route.with_cache(cache)`; cached and uncached routes use the same `execute` and `machine` methods +- `CallOptions` carries a scope-free `CachePolicy` and observation; per-call policy never replaces the attached scope or service +- Provider transport does not own cache orchestration; stream capture stays in `src/caching.rs` +- This crate and the format crates own request identity, typed response reconstruction, and stream capture and replay +- `cache-response` owns cache policy, namespacing, scope encoding, versioned envelopes, and freshness +- The SDK explicitly chooses shared scope; the gateway derives isolated scope from authenticated identity before attaching its service +- `ExecutionFacts` are delivered through the awaited `ResultReady` host operation for both provider and cached results, before public response processing or stream opening + - facts carry resolved model/provider and result source, including the hit key + - usage remains in the typed response or delivered stream, where completion and cancellation determine what was actually reported + - passive observation is not an accounting delivery mechanism +- Inference does not calculate prices, charge budgets, or update rate-limit counters + - the legacy Python callback adapter translates execution facts into the existing Python logging contract; Python remains the accounting owner on that path + - native gateway accounting belongs to gateway dependencies, independently of `host-python` + - response-cache services expose no coordination counters or reservation APIs; a shared Redis deployment does not make response storage and accounting coordination the same dependency +- Cache lookup follows provider preparation, credential resolution, and the request interceptor + - keys describe the effective provider URL, authenticated headers, and rewritten body + - signed requests bypass caching until the signing identity has a stable cache representation diff --git a/litellm-rust/crates/inference/Cargo.toml b/litellm-rust/crates/inference/Cargo.toml index e89ed8d94e7..0484749a349 100644 --- a/litellm-rust/crates/inference/Cargo.toml +++ b/litellm-rust/crates/inference/Cargo.toml @@ -14,41 +14,26 @@ litellm-cache-response.workspace = true litellm-framer.workspace = true tokio-util = { version = "0.7", features = ["codec"] } litellm-secrets.workspace = true -litellm-llms-types.workspace = true litellm-core-utils.workspace = true litellm-host.workspace = true bytes.workspace = true futures-util.workspace = true -base64.workspace = true litellm-auth = { workspace = true, features = ["aws", "azure", "gcp"] } -litellm-auth-aws.workspace = true litellm-http.workspace = true litellm-llms.workspace = true litellm-tracing.workspace = true tracing.workspace = true -moka.workspace = true -mime_guess = "2.0.5" -rand.workspace = true reqwest.workspace = true serde.workspace = true serde_json = { workspace = true, features = ["preserve_order"] } -strum.workspace = true -subtle.workspace = true tokio = { workspace = true, features = ["sync"] } -tokio-tungstenite.workspace = true thiserror.workspace = true time.workspace = true -sha2.workspace = true url.workspace = true -veil.workspace = true [dev-dependencies] litellm-inference = { workspace = true, features = ["test-support"] } litellm-cache-memory.workspace = true litellm-http = { workspace = true, features = ["test-support"] } -litellm-auth-gcp.workspace = true -litellm-host-native.workspace = true litellm-llms = { workspace = true, features = ["test-support"] } rstest.workspace = true -rstest_reuse.workspace = true -wiremock.workspace = true diff --git a/litellm-rust/crates/llms-types/AGENTS.md b/litellm-rust/crates/llms-types/AGENTS.md index 826aa009bec..0d418822d06 100644 --- a/litellm-rust/crates/llms-types/AGENTS.md +++ b/litellm-rust/crates/llms-types/AGENTS.md @@ -35,14 +35,14 @@ The same ownership rule applies to Messages, Responses, Chat Completions, OCR, a - Clamping effort, choosing a thinking budget, rewriting content, mapping finish reasons, computing normalized usage, and translating between API formats are policy or transformations and belong outside this crate, even when they are pure functions - Keep call envelopes and execution state in their owning crates - - `MessagesCall`, `MessagesShaping`, prepared provider requests, and the response wrapper containing a live stream belong in `core` + - `MessagesCall`, `MessagesShaping`, prepared provider requests, and the response wrapper containing a live stream belong in `inference-messages` - Provider config traits, `MessagesTransformContext`, `MessagesModelCapabilities`, `ThinkingBudgets`, `StreamShape`, and transformer state belong in `llms` - Catalog records and pricing belong in `model-catalog`, which may reuse wire enums such as `ReasoningEffort` - Host hooks, Python objects, credentials, clients, timeouts, and routing decisions do not become API payload types merely because they cross a crate boundary - Legacy logging operation selection belongs in `callbacks-legacy-python`, not this crate - Stream-event data belongs here, but live streams, decoders, framing, buffering, and stream lifecycle decisions do not - - Keep SSE and AWS framing in `framer`, provider decoding and conversion in `llms`, and call orchestration in `core` + - Keep SSE and AWS framing in `framer`, provider decoding and conversion in `llms`, and call orchestration in the `inference-` crates - `ResponsesWsEvent` belongs here - `ResponsesWsTransformResult` wraps the output of a provider transformation rather than a wire event and lives in `llms::base_llm::responses::transformation` - Protocol error payloads may live here, while operational errors remain in the crate that raises them diff --git a/litellm-rust/crates/llms/AGENTS.md b/litellm-rust/crates/llms/AGENTS.md index e833652797a..a34a962101d 100644 --- a/litellm-rust/crates/llms/AGENTS.md +++ b/litellm-rust/crates/llms/AGENTS.md @@ -16,7 +16,7 @@ Shared OCR document and response contracts live in `litellm-llms-types::formats: For Mistral, `async_transform_ocr_request` uses the base default in both languages. `resolve_headers` and `build_ocr_url` implement the respective environment and URL operations, and `normalize_response` implements the typed part of response transformation. Existing auth key/header handling and top-level response-extra preservation differ between languages; layout refactors must preserve those behaviors and verify them with the existing tests -For non-OCR pairs, order corresponding methods as parameter support/mapping, environment validation, URL construction, request transformation, and response transformation, followed by Rust-only runtime hooks. Auth resolution remains split between configs and route preparation in litellm-inference. Chat `supported_openai_param_mappings` describes accepted OpenAI/provider name pairs, unlike Python's `get_supported_openai_params` name list. Audio `map_transcription_params` remains a Rust filtering helper +For non-OCR pairs, order corresponding methods as parameter support/mapping, environment validation, URL construction, request transformation, and response transformation, followed by Rust-only runtime hooks. Auth resolution remains split between configs and route preparation in the `litellm-inference-` crates. Chat `supported_openai_param_mappings` describes accepted OpenAI/provider name pairs, unlike Python's `get_supported_openai_params` name list. Audio `map_transcription_params` remains a Rust filtering helper Azure Messages maps to `llms/azure_ai/anthropic/messages_transformation.py`; Bedrock Converse maps to `llms/bedrock/chat/converse_transformation.py`. `AnthropicConfig`, `AmazonConverseConfig`, and the non-OCR base traits are partial ports. `OpenAiResponsesApiConfig` implements WebSocket transformations and a direct HTTP Responses path. Its HTTP path does not implement Python model-specific parameter rewriting or Responses-to-Chat emulation. Preserve their acceptance gates, passthrough behavior, and host fallback contracts when aligning layout diff --git a/litellm-rust/crates/python-bridge/AGENTS.md b/litellm-rust/crates/python-bridge/AGENTS.md index 8ed8bc59a80..55837c4f32f 100644 --- a/litellm-rust/crates/python-bridge/AGENTS.md +++ b/litellm-rust/crates/python-bridge/AGENTS.md @@ -82,7 +82,7 @@ GIL handling to `litellm-host-python`. ## Bridge Shape - Prefer one stable method per top-level LiteLLM route, for example - `messages(...)`, calling the matching `litellm-inference` entrypoint. + `messages(...)`, calling the matching `litellm-inference-` entrypoint. - Do not add one exported PyO3 function per provider helper unless there is a measured reason. - Provider dispatch belongs in the `litellm-inference-*` route crate (e.g. diff --git a/litellm-rust/crates/router/README.md b/litellm-rust/crates/router/README.md index ed5b2966fe8..972d37d14be 100644 --- a/litellm-rust/crates/router/README.md +++ b/litellm-rust/crates/router/README.md @@ -2,4 +2,4 @@ Lookup is exact and returns `None` for an unknown name. This extraction preserves the gateway's existing behavior: the last entry wins when public names repeat. Multiple deployments per model group, routing strategies, retries, cooldowns, and fallbacks are not implemented yet -The router owns deployment configuration and selection. The gateway handles HTTP errors and responses, while `core` executes provider calls and resolves credentials +The router owns deployment configuration and selection. The gateway handles HTTP errors and responses, while the `inference-` crates execute provider calls and resolve credentials