mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
* feat(interactions): durable cross-pod settlement for background interaction billing Background interaction billing lived only in the creating replica's memory, so a DELETE routed to another replica, or a restart of the creating one, never billed the completed provider work and the budget reservation was refunded at the poll timeout. The create now registers the billing context in a settlement store before returning, the proxy installs a Prisma-backed store at boot (LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica claims the row once through a conditional update before billing or releasing, startup resumes every unclaimed row with its remaining timeout, and a give-up records an unsettled outcome instead of silently reconciling to zero. The SDK keeps an in-memory store and behaves as before. * fix(interactions): survive a settlement install failure at boot and stop carrying request headers * fix(interactions): drop the stored request context once a settlement row is settled * fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline * fix(interactions): carry a missing model through the settlement context for agent-only background creates An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it. * fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials. * fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill * fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler. A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere. * fix(interactions): settle an unverified registration through the durable claim A create whose settlement-store write raised no longer bills through a private in-memory gate that a later boot's resume cannot see. The claim asks the durable store first and falls back to the local gate only when the store answers that no row exists, and a missing settlement table reads as no rows so a replica without the migration still settles in process. * test(proxy): keep the settlement test where the proxy-infra shard collects it The merge of main moved test_background_interaction_settlement.py under tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected whole by the proxy-infra shard, which is where the test ran before the merge. * fix(interactions): raise on a non-2xx Gemini interaction fetch AsyncHTTPHandler.get never raises for status and the Gemini GET transform only raised when the body was not JSON, so a 500 or 404 carrying Gemini's JSON error body parsed as an interaction with no status. A delete on a replica other than the creator then claimed the settlement as released and forwarded the delete instead of failing closed, and the bill was lost. The transform now raises GeminiError with the vendor's status, as the delete transform already does; the in-process poll already retries a fetch that raises * test(integration): audit durable background interaction settlement across replicas Twenty-six deterministic cells drive a one-worker creator and a two-worker settler against an owned scripted Gemini upstream: cross-replica deletes bill once, failed and cancelled interactions release, a later replica resumes unclaimed rows, custom deployment pricing bills at the deployment rate, a fetch the settler cannot make fails the delete closed, odd ids are refused, a missing settlement table keeps in-process billing, polling disabled registers nothing, the budget reservation is released by the settler, an upstream outage mid-burst fails closed and recovers, killed workers hand their polls to the respawned ones, and concurrent deletes on a slow upstream settle exactly once. The support upstream gains a scripted interaction store with per-id GET status and delay, and the process helper gains an owned upstream a test can stop and restart * test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate * chore(ui): regenerate dashboard API types after merging main * test(integration): accept the 422 budget refusal and a respawned worker's resume The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged * test(integration): delete the pinned key's interaction with a second key A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| _support | ||
| agent_tests | ||
| audio_tests | ||
| base_sdk_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| guardrails_tests | ||
| harness_e2e | ||
| image_gen_tests | ||
| integration | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| rust-python-harness | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| test_litellm_rust | ||
| unified_google_tests | ||
| unit | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _process_helpers.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _wait_helpers.py | ||
| _ws_vcr.py | ||
| AGENTS.md | ||
| capturing_transport.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_otel_thread_leak.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_rust_python_harness.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
| white_100x100.png | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/unit
This folder can only run mock tests.