Commit graph

23 commits

Author SHA1 Message Date
Devin AI
6734e64175 chore: merge OCR cutover updates into reqwest mapper
Some checks failed
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-18 18:12:58 +00:00
Devin AI
cd63b40255 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ocr_rust_default 2026-07-18 03:39:29 +00:00
Devin AI
c79f4891fb fix(ocr): bound document fetches and polling
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-18 03:39:24 +00:00
Devin AI
0cecdd3253 feat(ocr): port Rust cutover onto staging architecture
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-18 03:32:54 +00:00
yuneng-jiang
f3d20153b3
build(rust): raise pyo3 to 0.29 so the native bridge compiles on Python 3.14 (#33798)
pyo3 0.23.5 hard-caps the interpreter at Python 3.13, so building the
native bridge against a 3.14 interpreter aborts inside pyo3-ffi's build
script before anything links. This raises pyo3 and pyo3-async-runtimes
to 0.29 (currently the newest line, and the range starting at 0.26 that
supports 3.14) and migrates the three call sites whose APIs were renamed
across that range: Python::with_gil is now Python::attach and
Python::allow_threads is now Python::detach. On a GIL-enabled interpreter
those are pure renames with identical semantics, so behavior on 3.10
through 3.13 is unchanged

Verified by compiling the native module for cp313 and cp314 and driving
it directly on both interpreters: gil_stats reports exactly one GIL
release per sync OCR call and the async path completes, matching the
0.23.5 baseline. cargo fmt, clippy, and the workspace tests pass on both
3.13 and 3.14 with the lockfile locked, and the lock churn is confined to
the pyo3 crates

Part of #26343; addresses the pyo3 build failure reported in #33116
2026-07-18 00:40:15 +00:00
Devin AI
0d93689c6d merge litellm_ocr_rust_default into litellm_shared_reqwest_error_mapper
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-17 04:40:10 +00:00
devin-ai-integration[bot]
dc585235be
feat(vertex): add shared VertexAiBase Rust host authentication (#33602)
Some checks are pending
LiteLLM Rust / rustfmt, clippy, test (push) Waiting to run
* feat(vertex): add shared Rust Google authentication

Mint and refresh Vertex OAuth access tokens in the ai-gateway host layer via the official google-cloud-auth crate, preserving the inline service-account JSON, ADC and GOOGLE_APPLICATION_CREDENTIALS contract with library-managed caching/refresh and no hand-rolled signing. Core stays auth/IO-free: it only classifies the bearer source and rejects Google AIza API keys rather than sending them as OAuth. Also fix Vertex Mistral rawPredict endpoint construction so the global location targets aiplatform.googleapis.com.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(vertex): address review; SHA-256 cache key and env credentials

Move newly introduced Vertex constants into crate-level constants.rs, drop the doc/inline comments added by the auth PR, and replace the DefaultHasher u64 credential-cache key with a collision-resistant SHA-256 digest (sha2 is now a non-optional ai-gateway dependency so the non-server host path can use it).

When no explicit vertex_credentials optional param is supplied, the host now reads VERTEXAI_CREDENTIALS from the environment before falling back to standard ADC/GOOGLE_APPLICATION_CREDENTIALS, without exposing credential content. Adds tests covering env-based inline credential selection and that distinct credential sources cannot collide on one cache entry.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* build(vertex): pin google-cloud-auth to =1.13.0 for dependency-age policy

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(vertex): bound credential cache, reject AIza in auth header, data-minimize auth errors

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* build(vertex): pin google-cloud-auth to =1.9.0 and adopt MSRV-aware resolver for Rust 1.86

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(vertex): require well-formed Bearer scheme on caller Authorization header

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(vertex): require exactly one well-formed Bearer authorization header

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(vertex): key credential cache by content and pin rustls exactly

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* chore(bridge): remove accidentally committed native extension binary

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(vertex): split base auth and shared cache

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-16 21:37:28 -07:00
Devin AI
ed5cb16589 merge litellm_ocr_rust_default into litellm_shared_reqwest_error_mapper
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-17 04:08:17 +00:00
devin-ai-integration[bot]
4da89566eb
feat(ocr): resolve environment references in Rust (#33598)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-16 21:05:52 -07:00
devin-ai-integration[bot]
2cf209f265
fix(mistral-ocr): forward the complete public parameter contract (#33600)
* fix(mistral-ocr): forward the complete public parameter contract

The Rust OCR path already lists id in the supported Mistral parameters, but litellm/ocr/main.py ran filter_out_litellm_params before handing optional_params to the bridge. That generic filter drops every key in all_litellm_params, which includes id, so the supported public Mistral OCR field id was silently stripped and never reached the provider while the rest of the contract went through.

Restore id after the generic filter via an OCR-specific reserved set, keeping genuine internal params filtered. Add Rust core and ai-gateway contract tests that pin the exact serialized request body for the full supported contract and prove internal canaries are dropped, a Python regression test that id survives while internal params do not, and a deterministic capture E2E plus a live Mistral param test.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(mistral-ocr): drop unnecessary comments from parity tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(mistral-ocr): only forward reserved id when set; harden capture e2e

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(mistral-ocr): typed queue-injected capture harness for e2e

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(mistral-ocr): fully type E2E capture models and move gateway ocr tests

Replace the remaining Any/coarse object wire types in the OCR E2E module
with concrete frozen Pydantic models for the supported-param contract,
internal canaries, upstream capture request, and normalized response, and
build every request payload immutably per call. The capture fixture now
cleans up on every failure path (terminate then always wait, kill then
always wait, shutdown plus server_close, join the server thread) and
requires an HTTP 200 liveness probe. The live Mistral test now exercises
the complete public parameter contract with valid json_schema annotation
formats. Move the ai-gateway OCR test module out of io/ocr.rs into a
focused io/ocr/tests.rs.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(mistral-ocr): drop reserved id upstream when null and type the E2E harness

Extend map_ocr_params so a null-valued reserved public id is omitted from the
serialized request, matching the Python OCR boundary for the standalone Axum
path where a caller can submit JSON id:null; add core map and serialized-body
tests for both the absent and null id cases. Extract the deterministic capture
harness into a typed ocr_capture_proxy helper module with an exception-safe
fixture that cleans up on every failure path from capture-server creation
onward (terminate then always wait, kill then always wait, shutdown plus
server_close, join the server thread, close logs after the process exits) and
requires an HTTP 200 liveness probe. Serialize wire payloads with by_alias so
the annotation schema/additionalProperties keys match the provider contract.
Replace the lambda internal canary in the Python bridge regression with a
serializable value and cover the litellm_session_id and tags internal fields.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-16 20:25:22 -07:00
Devin AI
95a8312a1e refactor(ai-gateway): share one reqwest exception mapper across endpoints
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-17 02:57:08 +00:00
devin-ai-integration[bot]
6661462d5a
fix(ocr-errors): preserve public error and timeout contracts (#33605)
* fix(ocr-errors): preserve public error and timeout contracts

Classify reqwest timeouts as a typed CoreError::Timeout and add
CoreError::public_status_code() as the single exhaustive mapping from a
typed core error to its public HTTP status. The Python bridge raises a
typed RustOcrError carrying that status instead of a generic
RuntimeError, and litellm.ocr()/aocr() translate it into the matching
public exception so AuthenticationError/401, NotFoundError/404,
BadRequestError/4xx, InternalServerError/5xx and Timeout are preserved
end to end instead of collapsing to APIConnectionError/500.

Reject empty or whitespace-only 200 bodies so they fail loudly rather
than becoming an empty OCR success; invalid JSON already fails. Upstream
error bodies stay bounded and sanitized.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ocr-errors): map invalid OCR input to BadRequestError

Invalid caller input (bad document type, non-dict document, unusable
file input) was raised as a plain ValueError inside ocr()/aocr() and
collapsed to APIConnectionError/500 through the generic handler. Route
every OCR failure through one _map_ocr_exception host mapping: typed
RustOcrError keeps its status-based public exception, a plain ValueError
becomes BadRequestError/400, and a pydantic ValidationError (malformed
response, not client input) stays on the generic path.

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(ocr-errors): data-minimize public errors and preserve status-specific exceptions

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(ocr-errors): preserve exact unknown status and privatize input error

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(ocr-errors): exhaustive match mapper and typed error tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ocr-errors): sanitize InvalidRequest public message

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(ocr-errors): raise typed public union and drop NotFound provider miss

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ocr-errors): map malformed provider responses to sanitized 500 and hide input-error detail

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-16 19:49:26 -07:00
Devin AI
c264758eb8 fix(ocr): close cutover review gaps
Some checks are pending
LiteLLM Rust / rustfmt, clippy, test (push) Waiting to run
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-07-16 20:05:43 +00:00
ishaan-berri
bdafc9a008
feat(ocr): thin Rust OCR Python bridge (#31368)
* feat(ocr): thin Rust OCR Python bridge

* refactor(rust): group provider routing helpers
2026-06-25 18:42:59 -07:00
ishaan-berri
62f93a3343
feat: add Rust OCR providers (#31272)
* feat: port OCR providers to Rust gateway

* chore(deps): update langgraph checkpoint lock

* ci: scope ruff format check to changed files

* ci: fix OCR lint and patch coverage

* fix(ocr): block mapped IPv6 fetch targets

* test(ocr): include rust bridge coverage in OCR shard

* ci: rerun responses shard
2026-06-25 15:12:30 -07:00
Ishaan Jaff
1528c9b9d1
feat(ocr): route all OCR through Rust 2026-06-25 14:16:42 -07:00
Ishaan Jaff
e5eaf3487c
fix(ocr): block mapped IPv6 fetch targets 2026-06-25 13:02:20 -07:00
Ishaan Jaff
48b5c73f9a
feat: port OCR providers to Rust gateway 2026-06-25 13:01:33 -07:00
ishaan-berri
4efce809d0
feat(proxy): add POST /v1/callbacks/logs to replay logging payloads through callbacks (#31134)
* feat(proxy): add logging_endpoints package init

* feat(proxy): add POST /v1/callbacks/logs to replay logging payloads through the success/failure callback fan-out

* feat(proxy): register callback_logs_router

* test(proxy): add logging_endpoints test package init

* test(proxy): cover /v1/callbacks/logs replay, admin guard, and partial-failure handling

* refactor(proxy): move callback-logs request/response models to litellm/types/proxy

* refactor(proxy): wrap callback-logs replay in CallbackLogsReplayer class with payload logging

* test(proxy): update callback-logs tests for class-based replayer and separated types

* fix(proxy): cover /v1/callbacks/ in backend component allowlist

The new /v1/callbacks/logs route was dropped by both component
allowlists, failing test_gateway_plus_backend_covers_full_app. It's an
admin-only spend-logging route, so it belongs on the backend (control
plane) alongside the existing /callbacks family.

* refactor(proxy): use builtin dict/list generics in callback-logs endpoint

Switch Dict/List from typing to builtin dict/list to satisfy the ruff
strict-rule budget (UP006).

* refactor(proxy): use builtin dict/list generics in callback-logs types

UP006: builtin generics over typing.Dict/List.

* chore(ui): regenerate schema.d.ts for /v1/callbacks/logs

Run npm run gen:api to add the CallbackLogRecord/CallbackLogsRequest/
CallbackLogsResponse types and the /v1/callbacks/logs path, keeping the
dashboard types in sync with the proxy OpenAPI spec.

* fix(proxy): force stream=False when replaying callback logs

A replayed StandardLoggingPayload is a terminal, fully-aggregated event —
the producer (e.g. the rust realtime gateway) already collected the whole
session before POSTing. Marking the rebuilt Logging object as streaming made
async_success_handler wait for a complete_streaming_response that never
arrives, so the spend log was never written. Realtime sessions now land in
LiteLLM_SpendLogs.

* feat(litellm-rust): CustomLogger callback layer posting to /v1/callbacks/logs

integrations/ mirrors litellm/integrations/: a sync, typed CustomLogger trait
(base contract), a typed StandardLoggingPayload, and LiteLLMPythonProxyAPILogger
— the first concrete logger, owning a bounded channel + background worker that
batches and POSTs to the Python proxy's /v1/callbacks/logs.

* feat(litellm-rust): RealTimeStreaming per-session log collector

1:1 with Python's RealTimeStreaming: observe() accumulates O(1) usage/model/id
per event (never buffers frames); log_messages() builds one StandardLoggingPayload
on session close and fans out to the CustomLogger callbacks. request_id == the
OpenAI realtime session id (sess_…), with the gateway id as fallback.

* feat(litellm-rust): wire realtime logging into the splice (lock-free observe)

The collector is owned on the splice task and observed via a synchronous &mut
callback threaded through providers::realtime::realtime() — no Arc/Mutex/atomic
on the per-frame hot path. On session close the bridge flushes one payload.
AppState carries the registered loggers; main spawns the proxy logger.

* docs(litellm-rust): ai-gateway realtime logging architecture

* docs(litellm-rust): document request-log egress to the LiteLLM control plane

Add a 'Request logging' guide to the ai-gateway README: how to point the gateway
at a LiteLLM proxy via LITELLM_PROXY_BASE_URL (+ LITELLM_MASTER_KEY for the
admin-only /v1/callbacks/logs POST), and the non-blocking / one-payload-per-session
behavior.

* feat(litellm-rust): make log-egress tunables env-overridable

Channel capacity, batch size, and flush interval now read from
LITELLM_LOG_CHANNEL_CAPACITY / LITELLM_LOG_BATCH_SIZE / LITELLM_LOG_FLUSH_INTERVAL_MS,
falling back to the DEFAULT_* consts on missing/invalid/non-positive values.
Grouped behind an EgressTunables::from_env() read once at logger construction.

* docs(litellm-rust): document log-egress tuning env vars

* docs(litellm-rust): require constants in a crate-level constants.rs

Mirror of Python's litellm/constants.py rule — magic numbers and fixed strings
go in src/constants.rs, not inline in feature modules; env-overridable tunables
keep their DEFAULT_* value there.

* refactor(litellm-rust): move ai-gateway constants into constants.rs

Per the new rule: the log-egress defaults (proxy base, ingest path, channel
capacity, batch size, flush interval) and the realtime provider default move to
crates/ai-gateway/src/constants.rs; modules import from it.

* ci: run logging_endpoints tests in the proxy-infra coverage shard

tests/test_litellm/proxy/logging_endpoints wasn't in any coverage-uploading
job, so callback_logs_endpoints.py showed only import-level coverage (~35%) on
codecov/patch despite being ~98% covered locally. Add it to proxy-infra's
test-path so the test is exercised under --cov.

* fix(litellm-rust): hash the master key before logging — never send the raw credential

Greptile/Veria P1: user_api_key_hash was the plaintext LITELLM_MASTER_KEY, which
fans out to spend logs and every callback (Langfuse/Datadog) and could be
recovered from logs. SHA-256 it (auth::hash_token, matching the proxy's
hash_token); the field is named *_hash and the proxy stores it verbatim when it
isn't sk-prefixed, so the DB value is identical with zero plaintext exposure.

* fix(litellm-rust): observe realtime logging on upstream events only

Greptile P1: observe ran on the client->upstream arm too, so an authenticated
client could send a fabricated response.done and inflate its own spend log.
session.created/response.done are server->client events; observe the upstream
arm only.

* feat(proxy): bound callback-logs batch + return per-record failures

Greptile P2: cap /v1/callbacks/logs at MAX_CALLBACK_LOG_RECORDS (default 1000,
env-overridable) so one POST can't trigger an unbounded callback/DB fan-out; and
return per-record {index, error} failures so a caller (the rust gateway) can
distinguish a transient callback error from a structurally bad payload.

* chore(ui): regenerate schema.d.ts for CallbackLogFailure / failures field

* fix(constants): make MAX_CALLBACK_LOG_RECORDS a plain constant

It doesn't need to be env-configurable (only the rust egress tunables are). As an
os.getenv var it tripped tests/documentation_tests/test_env_keys.py, which requires
every env key to be documented in the (separate-repo) config_settings.md. Plain
constant → not scanned → code-quality + documentation checks pass.

* docs(litellm-rust): trim ai-gateway ARCHITECTURE.md to one diagram + notes

* docs(litellm-rust): tighten the README request-logging section

* docs(litellm-rust): ARCHITECTURE.md is just the diagram (gateway = inference, spend = callback)

* docs(litellm-rust): drop em-dashes from the request-logging section

---------

Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
2026-06-24 15:25:10 -07:00
ishaan-berri
bd759182ca
refactor(litellm-rust): dissolve providers into core + ai-gateway (strict 3-crate layers) (#31218)
* refactor(litellm-rust): move provider transforms into litellm-core + crate allowlist test

* feat(litellm-rust): ai-gateway absorbs route I/O (io/) with lib+server feature split

* refactor(litellm-rust): point python-bridge at litellm-ai-gateway

* build(litellm-rust): macOS pyo3 dynamic_lookup linker flag for cdylib builds

* docs(litellm-rust): 3-crate map in README/AGENTS + refresh CLAUDE boundary

* refactor(litellm-rust): update workspace members to the three crates
2026-06-24 12:22:43 -07:00
ishaan-berri
6072019b0d
perf: pre-warm upstream realtime connection pool to cut session-establishment latency (#31163)
* docs: realtime pre-warmed connection pool design + raw-passthrough follow-up

* docs: realtime pool benchmark, repro steps, and deploy guidance

* feat: expose Router::deployments() for host-side upstream enumeration

* refactor: split realtime dial/splice and add warm-handoff entry point

* feat: pre-warmed upstream realtime connection pool with fresh-dial fallback

* feat: add realtime pool handle to gateway AppState

* feat: try warm pooled upstream before fresh-dial in realtime service

* feat: thread realtime pool through the realtime route bridge

* feat: build and pre-warm the realtime pool at gateway startup

* build: lean Dockerfile for the realtime gateway (default features, env stand-in)

Minimal multi-stage image for load-testing the realtime pool: builds the
gateway with default features (no python-config, no libpython), runs on a
debian-slim base (~157MB), and reads model_list from the OPENAI_REALTIME_MODEL
env stand-in. No config.yaml or pip install needed. Build context is the repo
root; only litellm-rust/ is included via the sidecar .dockerignore.

* docs: add generic ai-gateway benchmarking skill

Teaches an agent how to benchmark any ai-gateway endpoint: deploy the gateway,
run one load generator against both provider-direct and the gateway with the
same protocol, phase-decompose latency (dial/session/first-token/total),
compare at scale, and report success% + p50/p95. Documents the
benchmarks/<endpoint>/ layout and the hard no-committed-keys rule.

* docs: drop per-endpoint realtime benchmark README

The measured results table lives in the PR description (numbers go stale in a
committed README). The benchmarks/realtime/ dir now holds only the sanitized
load-gen harness; the generic method is in benchmarks/SKILL.md.

* test: add sanitized realtime WS load-gen harness for gateway benchmarks

Copies the ws-bench Go load generator (main.go, go.mod, go.sum, Dockerfile,
run.sh) into benchmarks/realtime/. Measures dial/session/first-audio/total per
WebSocket connection against both OpenAI-direct and the gateway. No keys are
hardcoded — the bearer token comes from -key / $OPENAI_API_KEY; default host
is api.openai.com.

* test: add hosted-runner serve.sh wrapper for the realtime harness

Render one-off jobs don't surface stdout via the Logs API, so on a hosted
runner the long-lived web service runs the leg and PUBLISHES the result: serve.sh
decodes the base64 flag list, runs wsbench teeing output to /tmp/web/result.txt,
then serves it over HTTP so the result is fetchable at /result.txt. No secrets
are written to the served file (the -key is only in wsbench's argv). The
Dockerfile now copies both run.sh and serve.sh.

* test: make serve.sh publish result atomically and clear stale output

Remove any prior result.txt at startup and write the new run to a .partial file
that's atomically moved into place only once complete. Prevents a fetcher from
reading a previous run's numbers while the current run is still in flight.

* perf: refill the realtime pool concurrently so warm supply keeps up at scale

The replenisher dialed missing warm sockets sequentially, so a full refill cost
needed x handshake (~needed x 350ms). Under high connect rates the pool drained
faster than it refilled and ~85% of connects missed (measured: only ~15% pool
hits at 5000/500). Firing the dials together with join_all refills in ~one
handshake window, keeping warm supply close to peak concurrent connects so the
sub-millisecond warm handoff becomes the median rather than the lucky-hit tail.
Each warm_one dial is independent (no shared state until the final push under the
lock), so concurrent refill is safe. Pool unit tests unchanged and passing.

* build: drop Dockerfile.lean

The lean load-test image isn't worth carrying in the repo; deploy the gateway
however you normally do and set the pool env vars.

* docs: slim benchmarks/realtime to a README (harness moved to its own repo)

Drop the Go load-gen, Dockerfile, run.sh, serve.sh, and the generic SKILL.md from
the repo. The harness now lives at github.com/ishaan-berri/litellm-realtime-bench;
benchmarks/realtime/README.md carries the results table and links there for repro.

* docs: add realtime route README with pooling design + diagram

Pooling is now documented as a section in src/routes/realtime/README.md next to
the code (handoff diagram, sizing rule, config, notes) instead of the standalone
REALTIME_POOL_DESIGN.md RFC. Update the realtime_pool.rs doc pointer to it.

* fix: satisfy clippy manual_flatten on concurrent pool refill

Use .into_iter().flatten() instead of an if-let-Ok in the for loop over the
join_all results, and let rustfmt wrap it. Clears the CI clippy -D warnings
failure; fmt + clippy + cargo test all green locally.

---------

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
2026-06-23 21:49:56 -07:00
ishaan-berri
a6b7dcc7d6
build: add Dockerfile + render blueprint for rust ai-gateway (#31154)
* build: make rust ai-gateway Dockerfile config.yaml-based (repo-root context)

* build: add dockerignore to shrink repo-root context for ai-gateway

* build: add sample realtime config.yaml for rust ai-gateway

* build: point ai-gateway render blueprint at config.yaml + repo-root context

* docs: document config.yaml as primary path for rust ai-gateway

* build: run rust ai-gateway container as non-root user

* ci: re-trigger flaky otel fake-openai-endpoint cooldown
2026-06-23 20:41:50 -07:00
ishaan-berri
1d5ab42e14
feat: add minimal rust router + axum ai-gateway calling router.realtime (2/2) (#31135)
* add CoreError::Routing variant for deployment selection failures

* add minimal Rust Router (simple-shuffle) mirroring router.py spec

* add litellm-router crate manifest

* add ai-gateway POST /v1/realtime handler calling router.realtime

* add ai-gateway health routes

* wire ai-gateway routes into the axum app

* add ai-gateway AppState holding the shared router

* add ai-gateway axum server entrypoint

* add litellm-ai-gateway binary crate manifest

* docs: add ai-gateway folder-architecture AGENTS.md

* register router + ai-gateway crates and axum/rand deps in workspace

* update Cargo.lock for router + ai-gateway crates

* split router: extract model_list types into deployment module

* split router: extract routing policy into strategy module

* split router: move Router orchestration into router module

* router lib: wire submodules and re-export public API

* add read_model_list helper reusing ProxyConfig env/secret resolution

* add GIL-activity tracker (records acquisitions, 30s window)

* add GET /health/gil endpoint for polling GIL activity

* add pyo3 load_router_from_config bridge (feature-gated, load-time only)

* register /health/gil route in ai-gateway

* wire build_router: load from python config when feature enabled

* add optional pyo3 dep + python-config feature to ai-gateway

* update Cargo.lock for optional pyo3 dependency

* fix: satisfy strict ruff budget (FA100) in read_model_list

* test: cover read_model_list env resolution + empty config

* ai-gateway: bind localhost by default, warn on bad PORT/missing keys, wire gateway key

* ai-gateway: add gateway_key to AppState for realtime auth

* ai-gateway: require bearer auth + map unknown model to 404 on /v1/realtime

* ai-gateway: move python interop into python/ with load-time-only AGENTS.md

* ai-gateway: document auth, gil, and python folder in AGENTS.md

* core: add router module (model_list types + simple-shuffle selection)

* ai-gateway: dispatch realtime via core router + providers (drop router crate dep)

* update Cargo.lock: fold router into core

* workspace: drop crates/router member and litellm-router dep

* read_model_list: reuse ProxyConfig.get_config (includes + os.environ + DB) instead of thin yaml read

* ai-gateway: constant-time bearer compare + 500 (not 503) for unconfigured key

* ai-gateway: trim stored gateway key to match trimmed bearer token

* ai-gateway: add subtle dep for constant-time comparison

* workspace: add subtle dependency

* update Cargo.lock for subtle

* core router: make strategy a folder (one module per strategy, simple_shuffle)

* providers: make realtime() a streaming splice (client stream <-> OpenAI) instead of collect

* providers: add futures-channel dev-dep for the streaming live test

* ai-gateway: make /v1/realtime a WebSocket (auth before upgrade, splice typed events)

* ai-gateway: dispatch realtime as a stream splice

* ai-gateway: route /v1/realtime via GET (WebSocket), drop POST

* ai-gateway: enable axum ws feature + futures-util

* update Cargo.lock for ws feature + futures-channel

* core router: add has_deployment() for pre-flight model checks

* ai-gateway: extract auth into auth/ module (single master key, LITELLM_MASTER_KEY)

* ai-gateway routes: adopt router()-per-module template + merge in app()

* ai-gateway: document auth/ + routes template in AGENTS.md

* ai-gateway: realtime route as thin handler + service + transport

* ai-gateway: auth as a RequireMasterKey extractor (idiomatic axum FromRequestParts)

* ai-gateway: docs for auth extractor + simplified route template

* ai-gateway: collapse realtime route to mod.rs + service.rs; docs for extractor/template

* providers realtime: enforce idle timeout around the splice (reap stalled sessions)

* ai-gateway: rename realtime service timeout param to idle_timeout

---------

Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
2026-06-23 19:16:34 -07:00