mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
3 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8596fe954d
|
feat(interactions): durable cross-pod settlement for background interaction billing (#41955)
* feat(interactions): durable cross-pod settlement for background interaction billing Background interaction billing lived only in the creating replica's memory, so a DELETE routed to another replica, or a restart of the creating one, never billed the completed provider work and the budget reservation was refunded at the poll timeout. The create now registers the billing context in a settlement store before returning, the proxy installs a Prisma-backed store at boot (LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica claims the row once through a conditional update before billing or releasing, startup resumes every unclaimed row with its remaining timeout, and a give-up records an unsettled outcome instead of silently reconciling to zero. The SDK keeps an in-memory store and behaves as before. * fix(interactions): survive a settlement install failure at boot and stop carrying request headers * fix(interactions): drop the stored request context once a settlement row is settled * fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline * fix(interactions): carry a missing model through the settlement context for agent-only background creates An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it. * fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials. * fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill * fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler. A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere. * fix(interactions): settle an unverified registration through the durable claim A create whose settlement-store write raised no longer bills through a private in-memory gate that a later boot's resume cannot see. The claim asks the durable store first and falls back to the local gate only when the store answers that no row exists, and a missing settlement table reads as no rows so a replica without the migration still settles in process. * test(proxy): keep the settlement test where the proxy-infra shard collects it The merge of main moved test_background_interaction_settlement.py under tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected whole by the proxy-infra shard, which is where the test ran before the merge. * fix(interactions): raise on a non-2xx Gemini interaction fetch AsyncHTTPHandler.get never raises for status and the Gemini GET transform only raised when the body was not JSON, so a 500 or 404 carrying Gemini's JSON error body parsed as an interaction with no status. A delete on a replica other than the creator then claimed the settlement as released and forwarded the delete instead of failing closed, and the bill was lost. The transform now raises GeminiError with the vendor's status, as the delete transform already does; the in-process poll already retries a fetch that raises * test(integration): audit durable background interaction settlement across replicas Twenty-six deterministic cells drive a one-worker creator and a two-worker settler against an owned scripted Gemini upstream: cross-replica deletes bill once, failed and cancelled interactions release, a later replica resumes unclaimed rows, custom deployment pricing bills at the deployment rate, a fetch the settler cannot make fails the delete closed, odd ids are refused, a missing settlement table keeps in-process billing, polling disabled registers nothing, the budget reservation is released by the settler, an upstream outage mid-burst fails closed and recovers, killed workers hand their polls to the respawned ones, and concurrent deletes on a slow upstream settle exactly once. The support upstream gains a scripted interaction store with per-id GET status and delay, and the process helper gains an owned upstream a test can stop and restart * test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate * chore(ui): regenerate dashboard API types after merging main * test(integration): accept the 422 budget refusal and a respawned worker's resume The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged * test(integration): delete the pinned key's interaction with a second key A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
c168199e33
|
test(ci): repair stale tests and move retired OpenAI text-completion fixtures (#43958)
* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups Request fakes now carry the scope a real Starlette request has, the GCS pub/sub spend-log golden gains the agent identity keys from #43722, the auto-router session tests follow the baseline_models contract from #43348, and the Interactions spec checks resolve the create body and resource paths from the live spec instead of hardcoded names * test(ci): move retired OpenAI text-completion fixtures to live vehicles OpenAI still serves native /v1/completions on the gpt-5.4 family, so the single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt batches and echo with logprobs now 500 on every OpenAI model, so those cases keep the same text-completion-openai transport pointed at Fireworks, which documents both. The optional-params test asserts the request body actually sent instead of a success callback whose assertions were swallowed * test(ci): use a serverless Fireworks model for the text-completion batch and echo cases gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not deployed; glm-5p3-flash is listed as serverless * test(ci): skip the ROI calculator repository listing in the security route sweep GET /roi-calculator/repositories (#43669) lists repositories from the configured GitHub API, api.github.com by default, so the S2 sweep's GET of every route made the owned proxy reach an external host and failed the egress check in 31 integration-security tests. It joins /get/latest_release_info in the deny list |
||
|
|
f6882246d4
|
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: rename fork-flag to unit-flag now that it applies on every event * test: move tests/test_litellm root and small trees into tests/unit Pure renames, no content changes. Follow-up commits in this PR fix references, merge the three files that already existed in tests/unit, keep live-provider tests in tests/test_litellm and wire CI. * test: carry tests/test_litellm conftest isolation into tests/unit Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS, proxy-URL and keychain env, and session-end client cleanup now reset for unit tests too. The environment isolation owns its MonkeyPatch so a test's own monkeypatch is undone before the model-cost teardown runs. * test: merge, split and prune the moved root and small-tree tests Merge batches/test_batch_utils.py and the chat_completions and messages dispatch tests into the files that already existed in tests/unit. Keep the live Gemini interactions tests, the async image-fetch format test and the OpenAI embedding scorer test in tests/test_litellm since they need real network or keys. Put test_router.py under tests/unit/test_router so the existing package no longer shadows it. Delete eight tests the audit found superseded by stronger ones kept in this move. * ci: run the moved root and small-tree tests under their legacy flags Add the misc and responses-caching-types flags to unit_selection.sh and CircleCI, extend enterprise-routing and mcp-integration, and point the legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest and change classifier at the new paths. * test: make the new tests/unit directories packages tests/unit/test_package_layout.py requires every directory to carry an __init__.py, and without one the moved and retained test_litellm_responses_bridge.py modules collide on import. * test: scope the unit socket block to tests/unit in shared sessions The GHA shards collect the legacy test-path and the unit selection in one pytest session. The unit conftest's loopback-only block leaked into legacy modules that reach the network at import. The legacy conftest now lifts the restriction at collect and setup time, and the unit conftest re-applies it when collecting its own modules. * test: give the shard-script tests their own GITHUB_OUTPUT They only passed where the runner set it. The CircleCI unit job's env allowlist drops it, so the script's redirect failed there. * test: point the router and module-deletion checks at tests/unit router_code_coverage and code_qa_check_tests only searched tests/test_litellm, so the moved router tests no longer counted. The two silent-experiment tests the audit deleted were the only direct callers of those methods; they are replaced with tests that assert the forwarded shadow request and the recursion guard. --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |