mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-13 23:11:40 +00:00
7 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bdafc9a008
|
feat(ocr): thin Rust OCR Python bridge (#31368)
* feat(ocr): thin Rust OCR Python bridge * refactor(rust): group provider routing helpers |
||
|
|
62f93a3343
|
feat: add Rust OCR providers (#31272)
* feat: port OCR providers to Rust gateway * chore(deps): update langgraph checkpoint lock * ci: scope ruff format check to changed files * ci: fix OCR lint and patch coverage * fix(ocr): block mapped IPv6 fetch targets * test(ocr): include rust bridge coverage in OCR shard * ci: rerun responses shard |
||
|
|
4efce809d0
|
feat(proxy): add POST /v1/callbacks/logs to replay logging payloads through callbacks (#31134)
* feat(proxy): add logging_endpoints package init
* feat(proxy): add POST /v1/callbacks/logs to replay logging payloads through the success/failure callback fan-out
* feat(proxy): register callback_logs_router
* test(proxy): add logging_endpoints test package init
* test(proxy): cover /v1/callbacks/logs replay, admin guard, and partial-failure handling
* refactor(proxy): move callback-logs request/response models to litellm/types/proxy
* refactor(proxy): wrap callback-logs replay in CallbackLogsReplayer class with payload logging
* test(proxy): update callback-logs tests for class-based replayer and separated types
* fix(proxy): cover /v1/callbacks/ in backend component allowlist
The new /v1/callbacks/logs route was dropped by both component
allowlists, failing test_gateway_plus_backend_covers_full_app. It's an
admin-only spend-logging route, so it belongs on the backend (control
plane) alongside the existing /callbacks family.
* refactor(proxy): use builtin dict/list generics in callback-logs endpoint
Switch Dict/List from typing to builtin dict/list to satisfy the ruff
strict-rule budget (UP006).
* refactor(proxy): use builtin dict/list generics in callback-logs types
UP006: builtin generics over typing.Dict/List.
* chore(ui): regenerate schema.d.ts for /v1/callbacks/logs
Run npm run gen:api to add the CallbackLogRecord/CallbackLogsRequest/
CallbackLogsResponse types and the /v1/callbacks/logs path, keeping the
dashboard types in sync with the proxy OpenAPI spec.
* fix(proxy): force stream=False when replaying callback logs
A replayed StandardLoggingPayload is a terminal, fully-aggregated event —
the producer (e.g. the rust realtime gateway) already collected the whole
session before POSTing. Marking the rebuilt Logging object as streaming made
async_success_handler wait for a complete_streaming_response that never
arrives, so the spend log was never written. Realtime sessions now land in
LiteLLM_SpendLogs.
* feat(litellm-rust): CustomLogger callback layer posting to /v1/callbacks/logs
integrations/ mirrors litellm/integrations/: a sync, typed CustomLogger trait
(base contract), a typed StandardLoggingPayload, and LiteLLMPythonProxyAPILogger
— the first concrete logger, owning a bounded channel + background worker that
batches and POSTs to the Python proxy's /v1/callbacks/logs.
* feat(litellm-rust): RealTimeStreaming per-session log collector
1:1 with Python's RealTimeStreaming: observe() accumulates O(1) usage/model/id
per event (never buffers frames); log_messages() builds one StandardLoggingPayload
on session close and fans out to the CustomLogger callbacks. request_id == the
OpenAI realtime session id (sess_…), with the gateway id as fallback.
* feat(litellm-rust): wire realtime logging into the splice (lock-free observe)
The collector is owned on the splice task and observed via a synchronous &mut
callback threaded through providers::realtime::realtime() — no Arc/Mutex/atomic
on the per-frame hot path. On session close the bridge flushes one payload.
AppState carries the registered loggers; main spawns the proxy logger.
* docs(litellm-rust): ai-gateway realtime logging architecture
* docs(litellm-rust): document request-log egress to the LiteLLM control plane
Add a 'Request logging' guide to the ai-gateway README: how to point the gateway
at a LiteLLM proxy via LITELLM_PROXY_BASE_URL (+ LITELLM_MASTER_KEY for the
admin-only /v1/callbacks/logs POST), and the non-blocking / one-payload-per-session
behavior.
* feat(litellm-rust): make log-egress tunables env-overridable
Channel capacity, batch size, and flush interval now read from
LITELLM_LOG_CHANNEL_CAPACITY / LITELLM_LOG_BATCH_SIZE / LITELLM_LOG_FLUSH_INTERVAL_MS,
falling back to the DEFAULT_* consts on missing/invalid/non-positive values.
Grouped behind an EgressTunables::from_env() read once at logger construction.
* docs(litellm-rust): document log-egress tuning env vars
* docs(litellm-rust): require constants in a crate-level constants.rs
Mirror of Python's litellm/constants.py rule — magic numbers and fixed strings
go in src/constants.rs, not inline in feature modules; env-overridable tunables
keep their DEFAULT_* value there.
* refactor(litellm-rust): move ai-gateway constants into constants.rs
Per the new rule: the log-egress defaults (proxy base, ingest path, channel
capacity, batch size, flush interval) and the realtime provider default move to
crates/ai-gateway/src/constants.rs; modules import from it.
* ci: run logging_endpoints tests in the proxy-infra coverage shard
tests/test_litellm/proxy/logging_endpoints wasn't in any coverage-uploading
job, so callback_logs_endpoints.py showed only import-level coverage (~35%) on
codecov/patch despite being ~98% covered locally. Add it to proxy-infra's
test-path so the test is exercised under --cov.
* fix(litellm-rust): hash the master key before logging — never send the raw credential
Greptile/Veria P1: user_api_key_hash was the plaintext LITELLM_MASTER_KEY, which
fans out to spend logs and every callback (Langfuse/Datadog) and could be
recovered from logs. SHA-256 it (auth::hash_token, matching the proxy's
hash_token); the field is named *_hash and the proxy stores it verbatim when it
isn't sk-prefixed, so the DB value is identical with zero plaintext exposure.
* fix(litellm-rust): observe realtime logging on upstream events only
Greptile P1: observe ran on the client->upstream arm too, so an authenticated
client could send a fabricated response.done and inflate its own spend log.
session.created/response.done are server->client events; observe the upstream
arm only.
* feat(proxy): bound callback-logs batch + return per-record failures
Greptile P2: cap /v1/callbacks/logs at MAX_CALLBACK_LOG_RECORDS (default 1000,
env-overridable) so one POST can't trigger an unbounded callback/DB fan-out; and
return per-record {index, error} failures so a caller (the rust gateway) can
distinguish a transient callback error from a structurally bad payload.
* chore(ui): regenerate schema.d.ts for CallbackLogFailure / failures field
* fix(constants): make MAX_CALLBACK_LOG_RECORDS a plain constant
It doesn't need to be env-configurable (only the rust egress tunables are). As an
os.getenv var it tripped tests/documentation_tests/test_env_keys.py, which requires
every env key to be documented in the (separate-repo) config_settings.md. Plain
constant → not scanned → code-quality + documentation checks pass.
* docs(litellm-rust): trim ai-gateway ARCHITECTURE.md to one diagram + notes
* docs(litellm-rust): tighten the README request-logging section
* docs(litellm-rust): ARCHITECTURE.md is just the diagram (gateway = inference, spend = callback)
* docs(litellm-rust): drop em-dashes from the request-logging section
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
|
||
|
|
bd759182ca
|
refactor(litellm-rust): dissolve providers into core + ai-gateway (strict 3-crate layers) (#31218)
* refactor(litellm-rust): move provider transforms into litellm-core + crate allowlist test * feat(litellm-rust): ai-gateway absorbs route I/O (io/) with lib+server feature split * refactor(litellm-rust): point python-bridge at litellm-ai-gateway * build(litellm-rust): macOS pyo3 dynamic_lookup linker flag for cdylib builds * docs(litellm-rust): 3-crate map in README/AGENTS + refresh CLAUDE boundary * refactor(litellm-rust): update workspace members to the three crates |
||
|
|
6072019b0d
|
perf: pre-warm upstream realtime connection pool to cut session-establishment latency (#31163)
* docs: realtime pre-warmed connection pool design + raw-passthrough follow-up * docs: realtime pool benchmark, repro steps, and deploy guidance * feat: expose Router::deployments() for host-side upstream enumeration * refactor: split realtime dial/splice and add warm-handoff entry point * feat: pre-warmed upstream realtime connection pool with fresh-dial fallback * feat: add realtime pool handle to gateway AppState * feat: try warm pooled upstream before fresh-dial in realtime service * feat: thread realtime pool through the realtime route bridge * feat: build and pre-warm the realtime pool at gateway startup * build: lean Dockerfile for the realtime gateway (default features, env stand-in) Minimal multi-stage image for load-testing the realtime pool: builds the gateway with default features (no python-config, no libpython), runs on a debian-slim base (~157MB), and reads model_list from the OPENAI_REALTIME_MODEL env stand-in. No config.yaml or pip install needed. Build context is the repo root; only litellm-rust/ is included via the sidecar .dockerignore. * docs: add generic ai-gateway benchmarking skill Teaches an agent how to benchmark any ai-gateway endpoint: deploy the gateway, run one load generator against both provider-direct and the gateway with the same protocol, phase-decompose latency (dial/session/first-token/total), compare at scale, and report success% + p50/p95. Documents the benchmarks/<endpoint>/ layout and the hard no-committed-keys rule. * docs: drop per-endpoint realtime benchmark README The measured results table lives in the PR description (numbers go stale in a committed README). The benchmarks/realtime/ dir now holds only the sanitized load-gen harness; the generic method is in benchmarks/SKILL.md. * test: add sanitized realtime WS load-gen harness for gateway benchmarks Copies the ws-bench Go load generator (main.go, go.mod, go.sum, Dockerfile, run.sh) into benchmarks/realtime/. Measures dial/session/first-audio/total per WebSocket connection against both OpenAI-direct and the gateway. No keys are hardcoded — the bearer token comes from -key / $OPENAI_API_KEY; default host is api.openai.com. * test: add hosted-runner serve.sh wrapper for the realtime harness Render one-off jobs don't surface stdout via the Logs API, so on a hosted runner the long-lived web service runs the leg and PUBLISHES the result: serve.sh decodes the base64 flag list, runs wsbench teeing output to /tmp/web/result.txt, then serves it over HTTP so the result is fetchable at /result.txt. No secrets are written to the served file (the -key is only in wsbench's argv). The Dockerfile now copies both run.sh and serve.sh. * test: make serve.sh publish result atomically and clear stale output Remove any prior result.txt at startup and write the new run to a .partial file that's atomically moved into place only once complete. Prevents a fetcher from reading a previous run's numbers while the current run is still in flight. * perf: refill the realtime pool concurrently so warm supply keeps up at scale The replenisher dialed missing warm sockets sequentially, so a full refill cost needed x handshake (~needed x 350ms). Under high connect rates the pool drained faster than it refilled and ~85% of connects missed (measured: only ~15% pool hits at 5000/500). Firing the dials together with join_all refills in ~one handshake window, keeping warm supply close to peak concurrent connects so the sub-millisecond warm handoff becomes the median rather than the lucky-hit tail. Each warm_one dial is independent (no shared state until the final push under the lock), so concurrent refill is safe. Pool unit tests unchanged and passing. * build: drop Dockerfile.lean The lean load-test image isn't worth carrying in the repo; deploy the gateway however you normally do and set the pool env vars. * docs: slim benchmarks/realtime to a README (harness moved to its own repo) Drop the Go load-gen, Dockerfile, run.sh, serve.sh, and the generic SKILL.md from the repo. The harness now lives at github.com/ishaan-berri/litellm-realtime-bench; benchmarks/realtime/README.md carries the results table and links there for repro. * docs: add realtime route README with pooling design + diagram Pooling is now documented as a section in src/routes/realtime/README.md next to the code (handoff diagram, sizing rule, config, notes) instead of the standalone REALTIME_POOL_DESIGN.md RFC. Update the realtime_pool.rs doc pointer to it. * fix: satisfy clippy manual_flatten on concurrent pool refill Use .into_iter().flatten() instead of an if-let-Ok in the for loop over the join_all results, and let rustfmt wrap it. Clears the CI clippy -D warnings failure; fmt + clippy + cargo test all green locally. --------- Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com> |
||
|
|
a6b7dcc7d6
|
build: add Dockerfile + render blueprint for rust ai-gateway (#31154)
* build: make rust ai-gateway Dockerfile config.yaml-based (repo-root context) * build: add dockerignore to shrink repo-root context for ai-gateway * build: add sample realtime config.yaml for rust ai-gateway * build: point ai-gateway render blueprint at config.yaml + repo-root context * docs: document config.yaml as primary path for rust ai-gateway * build: run rust ai-gateway container as non-root user * ci: re-trigger flaky otel fake-openai-endpoint cooldown |
||
|
|
1d5ab42e14
|
feat: add minimal rust router + axum ai-gateway calling router.realtime (2/2) (#31135)
* add CoreError::Routing variant for deployment selection failures * add minimal Rust Router (simple-shuffle) mirroring router.py spec * add litellm-router crate manifest * add ai-gateway POST /v1/realtime handler calling router.realtime * add ai-gateway health routes * wire ai-gateway routes into the axum app * add ai-gateway AppState holding the shared router * add ai-gateway axum server entrypoint * add litellm-ai-gateway binary crate manifest * docs: add ai-gateway folder-architecture AGENTS.md * register router + ai-gateway crates and axum/rand deps in workspace * update Cargo.lock for router + ai-gateway crates * split router: extract model_list types into deployment module * split router: extract routing policy into strategy module * split router: move Router orchestration into router module * router lib: wire submodules and re-export public API * add read_model_list helper reusing ProxyConfig env/secret resolution * add GIL-activity tracker (records acquisitions, 30s window) * add GET /health/gil endpoint for polling GIL activity * add pyo3 load_router_from_config bridge (feature-gated, load-time only) * register /health/gil route in ai-gateway * wire build_router: load from python config when feature enabled * add optional pyo3 dep + python-config feature to ai-gateway * update Cargo.lock for optional pyo3 dependency * fix: satisfy strict ruff budget (FA100) in read_model_list * test: cover read_model_list env resolution + empty config * ai-gateway: bind localhost by default, warn on bad PORT/missing keys, wire gateway key * ai-gateway: add gateway_key to AppState for realtime auth * ai-gateway: require bearer auth + map unknown model to 404 on /v1/realtime * ai-gateway: move python interop into python/ with load-time-only AGENTS.md * ai-gateway: document auth, gil, and python folder in AGENTS.md * core: add router module (model_list types + simple-shuffle selection) * ai-gateway: dispatch realtime via core router + providers (drop router crate dep) * update Cargo.lock: fold router into core * workspace: drop crates/router member and litellm-router dep * read_model_list: reuse ProxyConfig.get_config (includes + os.environ + DB) instead of thin yaml read * ai-gateway: constant-time bearer compare + 500 (not 503) for unconfigured key * ai-gateway: trim stored gateway key to match trimmed bearer token * ai-gateway: add subtle dep for constant-time comparison * workspace: add subtle dependency * update Cargo.lock for subtle * core router: make strategy a folder (one module per strategy, simple_shuffle) * providers: make realtime() a streaming splice (client stream <-> OpenAI) instead of collect * providers: add futures-channel dev-dep for the streaming live test * ai-gateway: make /v1/realtime a WebSocket (auth before upgrade, splice typed events) * ai-gateway: dispatch realtime as a stream splice * ai-gateway: route /v1/realtime via GET (WebSocket), drop POST * ai-gateway: enable axum ws feature + futures-util * update Cargo.lock for ws feature + futures-channel * core router: add has_deployment() for pre-flight model checks * ai-gateway: extract auth into auth/ module (single master key, LITELLM_MASTER_KEY) * ai-gateway routes: adopt router()-per-module template + merge in app() * ai-gateway: document auth/ + routes template in AGENTS.md * ai-gateway: realtime route as thin handler + service + transport * ai-gateway: auth as a RequireMasterKey extractor (idiomatic axum FromRequestParts) * ai-gateway: docs for auth extractor + simplified route template * ai-gateway: collapse realtime route to mod.rs + service.rs; docs for extractor/template * providers realtime: enforce idle timeout around the splice (reap stalled sessions) * ai-gateway: rename realtime service timeout param to idle_timeout --------- Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com> |