* refactor(rust): extract litellm-host-native as the shared Rust host driver Move service and hook dispatch out of host-http into a Driver that owns the machine and Rust handlers, returning at completion or a stream boundary and holding the demand reply until the consumer advances. Move the in-process runner onto the same driver. host-http now layers encoding, SSE, body polling and lifecycle observation over it. host-python keeps driving litellm-host directly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(rust): interrupt the machine when the in-process stream consumer fails Restores the pre-refactor interruption path for StreamConsumer errors via Driver::fail and ports the generic run lifecycle tests into host-native. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(rust): separate the machine contract from coroutine execution * auth update * refactor(rust): use standard flow control for host requests * style(rust): keep host driver imports formatted * chores * mostly relocation * refactor(rust): separate interceptors from queued observers * refactor(rust): centralize legacy callback mappings and lifecycle * docs: define Python host boundaries and migration plan * refactor: enforce Python host and bridge boundaries * refactor(rust): separate operations from callback composition * refactor(rust): compose SDK policy through call hooks --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
1 KiB
litellm-host-native is the Rust driver for hosted calls. Driver owns the machine, a HostCallHandler and a Interceptors; advance() answers services and hooks inline and returns at completion or at the next stream boundary, holding the Reply<ControlFlow<()>> until the consumer calls advance() or detach() again. Dropping the driver drops the machine and so cancels the call
The consumer decides demand, so the driver never spawns a producer task and never buffers chunks ahead of demand. litellm-host-http polls it from the response body; in_process::run_hosted polls it on behalf of a StreamConsumer. Both observe lifecycle terminals themselves, the driver reports none
Depend on litellm-host only. HTTP encoding stays in litellm-host-http; litellm-host-python drives the machine directly so Python callbacks stay in the caller's asyncio task
services.rs owns HostCallHandler and its borrowed and no-service implementations. This is the Rust driver's handler contract; the shared host crate owns the service request protocol