mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-10 03:28:53 +00:00
596 B
596 B
ai-gateway architecture
The Rust ai-gateway does LLM inference (realtime WebSocket). Spend tracking is an API callback: it POSTs each finished session to the LiteLLM proxy, which records spend and runs the usual callbacks.
OCR provider execution lives in litellm-runtime. The gateway OCR module is a
compatibility host adapter for custom logger, guardrail, and request metadata
types; runtime has no dependency on the gateway.
flowchart LR
C[client] <--> G[Rust ai-gateway<br/>LLM inference]
G <--> O[OpenAI realtime]
G -. spend tracking callback .-> P[litellm proxy]