Delete litellm/ocr/input.py and the native _ocr_file_document, _ocr_upload_document
and _ocr_mime_type helpers. File documents now project to a typed OcrDocumentInput
and the core lifecycle reads local paths, encodes bytes and asks the host to read
file-like objects through a ReadDocument operation before the provider request
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(python-bridge): split non-streaming bridge modules
* refactor(python-bridge): bring shared function tracing into route layer
* feat(dev): list Python route functions and call sites
* feat(dev): list Rust route functions and call sites
* docs(dev): record OCR parity gaps across Python and Rust
* feat(dev): list executed SDK calls with runtime tracing
* feat(dev): report Python vs Rust SDK pipeline steps in one CLI
* feat(dev): side-by-side pipeline step report in compare CLI
* fix(dev): drop invalid Final annotations in compare cell loop
* feat(dev): blue python-only and yellow rust-only steps in compare CLI
* feat(dev): vertical layout with section spacing in compare CLI
* fix(dev): validate SDK trace stages across sync and async routes
* refactor(rust): align SDK route call structure with Python
* refactor(python-bridge): share sync and async route call wrappers
* refactor(dev): split compare CLI into fixtures, runtime, and report modules
* fix(ci): run SDK trace tests and satisfy test lint
Adds a chat_completions route module to litellm-core, mirroring the messages
route, plus Anthropic Messages and Bedrock Converse provider configs. The
per-model `rust: true` opt-in now covers /chat/completions for both providers.
The core accepts an allowlisted subset (text conversations, non-streaming) and
returns CoreError::Unsupported for anything else, so tool calls, multimodal
content and streaming fall back to the Python path transparently.
Resolves LIT-5698
* feat(messages): route Azure Anthropic /messages through Rust behind rust:true
Adds an opt-in Rust path for non-streaming Azure Anthropic Messages. A
deployment sets rust: true in litellm_params to route litellm.messages()
and the proxy /v1/messages endpoint through the native Rust bridge; a
missing flag or rust: false keeps the existing Python path, and non-Azure
providers, streaming, an unavailable bridge, or a None result all fall
back to Python. Rust-backed responses carry an x-litellm-rust: true
response header so callers can see which path served the request.
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(docs): exclude LITELLM_USE_RUST_MESSAGES rollout flag from env-doc check
Mirrors the existing LITELLM_USE_RUST_OCR entry; the flag is an internal
rollout toggle that is intentionally not in the public environment settings
docs yet.
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(rust_bridge): isolate OCR enable flag and drop dead messages global toggle
use_litellm_rust only mutates the OCR enabled flag when configuring OCR (or
called with no bridge kwargs, preserving the legacy contract), so configuring
only the messages bridge no longer flips OCR state.
Remove the vestigial global enabled/env state from the messages bridge. Routing
is controlled per deployment by rust:true in the shared handler gate, so the
messages module never consulted the global toggle; drop it rather than leave a
no-op switch.
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* refactor(rust/messages): split Anthropic config into its own provider file and type the request/response contract
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* feat(messages): route eligible Azure Anthropic streaming through Rust via buffered fake-stream
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(messages): fold system-role messages for Azure Anthropic and fall back to Python on Rust bridge errors
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(rust_bridge): use Python::attach for amessages after pyo3 bump
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(proxy): mock get_configured_token_limits in model_info tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* ci: run rust_bridge unit tests in misc shard
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* Revert "ci: run rust_bridge unit tests in misc shard"
This reverts commit c86d861a03.
* test(anthropic): move rust messages bridge tests into misc-shard dir
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* feat(mistral): support Mistral OCR 4 (mistral-ocr-4-0)
Add the mistral/mistral-ocr-4-0 model to the cost map and reprice
mistral/mistral-ocr-latest, which now resolves to OCR 4 server-side,
at $4 / 1000 pages. Add the include_blocks param so callers can request
OCR 4's paragraph-level bounding boxes and typed content blocks.
OCR 4's new per-page response fields (blocks, confidence_scores, tables,
hyperlinks, header, footer) already pass through transform_ocr_response
via the extra="allow" config on OCRPage; add a regression test pinning
that behavior alongside cost and param coverage.
* fix(mistral): revert unverified OCR 4 annotation_cost_per_page bump
Mistral's published OCR 4 pricing lists $4/1000 pages for the API and no
separate annotation rate; the $5/1000 figure is the distinct Document AI
(Studio) tier. The earlier 0.003 -> 0.005 bump on annotation_cost_per_page
had no cited source, and ocr_cost() never reads that field (it bills off
ocr_cost_per_page), so the value is documentation-only.
Revert annotation_cost_per_page to the existing 0.003 convention for both
mistral-ocr-latest and mistral-ocr-4-0, keeping only the verified, tested
ocr_cost_per_page: 0.004 change.
* fix(mistral): set OCR 4 annotation_cost_per_page to verified $5/1000 rate
Verified against Mistral's authoritative sources: the pricing page, the
OCR 4 announcement, and the ocr-4-0 model card all list OCR 4 at $4/1000
pages for basic OCR and $5/1000 for annotated pages (Document AI). The
$5/1000 figure is the annotated-pages rate, which is exactly what
annotation_cost_per_page encodes, mirroring the original OCR entry's
0.001 basic / 0.003 annotated split.
Restore annotation_cost_per_page to 0.005 for mistral-ocr-latest and
mistral-ocr-4-0; the earlier revert to 0.003 was based on an incomplete
reading that treated Document AI as a separate product. ocr_cost_per_page
stays 0.004, which is the value billed by ocr_cost().
* fix(mistral-rust): include_blocks in Rust OCR supported params
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* refactor(litellm-rust): move provider transforms into litellm-core + crate allowlist test
* feat(litellm-rust): ai-gateway absorbs route I/O (io/) with lib+server feature split
* refactor(litellm-rust): point python-bridge at litellm-ai-gateway
* build(litellm-rust): macOS pyo3 dynamic_lookup linker flag for cdylib builds
* docs(litellm-rust): 3-crate map in README/AGENTS + refresh CLAUDE boundary
* refactor(litellm-rust): update workspace members to the three crates
* docs: realtime pre-warmed connection pool design + raw-passthrough follow-up
* docs: realtime pool benchmark, repro steps, and deploy guidance
* feat: expose Router::deployments() for host-side upstream enumeration
* refactor: split realtime dial/splice and add warm-handoff entry point
* feat: pre-warmed upstream realtime connection pool with fresh-dial fallback
* feat: add realtime pool handle to gateway AppState
* feat: try warm pooled upstream before fresh-dial in realtime service
* feat: thread realtime pool through the realtime route bridge
* feat: build and pre-warm the realtime pool at gateway startup
* build: lean Dockerfile for the realtime gateway (default features, env stand-in)
Minimal multi-stage image for load-testing the realtime pool: builds the
gateway with default features (no python-config, no libpython), runs on a
debian-slim base (~157MB), and reads model_list from the OPENAI_REALTIME_MODEL
env stand-in. No config.yaml or pip install needed. Build context is the repo
root; only litellm-rust/ is included via the sidecar .dockerignore.
* docs: add generic ai-gateway benchmarking skill
Teaches an agent how to benchmark any ai-gateway endpoint: deploy the gateway,
run one load generator against both provider-direct and the gateway with the
same protocol, phase-decompose latency (dial/session/first-token/total),
compare at scale, and report success% + p50/p95. Documents the
benchmarks/<endpoint>/ layout and the hard no-committed-keys rule.
* docs: drop per-endpoint realtime benchmark README
The measured results table lives in the PR description (numbers go stale in a
committed README). The benchmarks/realtime/ dir now holds only the sanitized
load-gen harness; the generic method is in benchmarks/SKILL.md.
* test: add sanitized realtime WS load-gen harness for gateway benchmarks
Copies the ws-bench Go load generator (main.go, go.mod, go.sum, Dockerfile,
run.sh) into benchmarks/realtime/. Measures dial/session/first-audio/total per
WebSocket connection against both OpenAI-direct and the gateway. No keys are
hardcoded — the bearer token comes from -key / $OPENAI_API_KEY; default host
is api.openai.com.
* test: add hosted-runner serve.sh wrapper for the realtime harness
Render one-off jobs don't surface stdout via the Logs API, so on a hosted
runner the long-lived web service runs the leg and PUBLISHES the result: serve.sh
decodes the base64 flag list, runs wsbench teeing output to /tmp/web/result.txt,
then serves it over HTTP so the result is fetchable at /result.txt. No secrets
are written to the served file (the -key is only in wsbench's argv). The
Dockerfile now copies both run.sh and serve.sh.
* test: make serve.sh publish result atomically and clear stale output
Remove any prior result.txt at startup and write the new run to a .partial file
that's atomically moved into place only once complete. Prevents a fetcher from
reading a previous run's numbers while the current run is still in flight.
* perf: refill the realtime pool concurrently so warm supply keeps up at scale
The replenisher dialed missing warm sockets sequentially, so a full refill cost
needed x handshake (~needed x 350ms). Under high connect rates the pool drained
faster than it refilled and ~85% of connects missed (measured: only ~15% pool
hits at 5000/500). Firing the dials together with join_all refills in ~one
handshake window, keeping warm supply close to peak concurrent connects so the
sub-millisecond warm handoff becomes the median rather than the lucky-hit tail.
Each warm_one dial is independent (no shared state until the final push under the
lock), so concurrent refill is safe. Pool unit tests unchanged and passing.
* build: drop Dockerfile.lean
The lean load-test image isn't worth carrying in the repo; deploy the gateway
however you normally do and set the pool env vars.
* docs: slim benchmarks/realtime to a README (harness moved to its own repo)
Drop the Go load-gen, Dockerfile, run.sh, serve.sh, and the generic SKILL.md from
the repo. The harness now lives at github.com/ishaan-berri/litellm-realtime-bench;
benchmarks/realtime/README.md carries the results table and links there for repro.
* docs: add realtime route README with pooling design + diagram
Pooling is now documented as a section in src/routes/realtime/README.md next to
the code (handoff diagram, sizing rule, config, notes) instead of the standalone
REALTIME_POOL_DESIGN.md RFC. Update the realtime_pool.rs doc pointer to it.
* fix: satisfy clippy manual_flatten on concurrent pool refill
Use .into_iter().flatten() instead of an if-let-Ok in the for loop over the
join_all results, and let rustfmt wrap it. Clears the CI clippy -D warnings
failure; fmt + clippy + cargo test all green locally.
---------
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>