mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-06 02:48:13 +00:00
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
Terraform Modules / fmt, validate, test (aws) (push) Waiting to run
Terraform Modules / fmt, validate, test (gcp) (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
* test(integration): saving echoed model_info never persists cost map pricing as a deployment override (Pylon #6870)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): budget_duration change on /budget/update recomputes budget_reset_at (Pylon #6913)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): count_tokens on a budgeted key reserves no budget and a later completion still succeeds (Pylon #6966)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): /cost/estimate reports configured prices for a deployment absent from the cost map (Pylon #7014)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cache the team member default budget in Redis as JSON (Pylon #7180)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): aggregated team daily activity reports whole-range team spend in one page (Pylon #7224)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): failed daily user rollup commits are retried so spend report and daily activity agree (Pylon #7268)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): CLI session token without org_id is charged to and capped by the team organization budget (Pylon #7291)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): gemini passthrough success releases its budget reservation from the spend counter (Pylon #7295)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): batch retrieval spend row sums reasoning tokens and counts output and error file failures (Pylon #7341)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): uncostable batches retire from the cost poll page so newer batches are costed (Pylon #7342)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): charge a team member added without any budget on its membership row (Pylon #7363)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): failed dispatched requests keep estimated input tokens in spend logs (Pylon #7519)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bedrock passthrough converse guardrail ignores tool definitions (Pylon #7524)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): explicit null budget_duration on /team/new is not replaced by default_team_params (Pylon #7536)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): PATCH /organization/update with a null limit clears it (Pylon #7577)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): ultrafast service_tier bills ultrafast rates without leaking pricing fields upstream (Pylon #7587)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): keep the selected model in the response and spend log for an Azure Model Router alias (Pylon #7636)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): disconnected Bedrock /v1/messages stream still bills terminal usage (Pylon #7685)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): databricks cached prompt tokens bill at cache rates (Pylon #7738)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): completed batch spend row records reasoning tokens and error file failures (Pylon #7928)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bill OCR annotation pages at annotation_cost_per_page (Pylon #7958)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): in-flight count tokens request reserves no key budget so a completion still reaches the provider (Pylon #7307)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): fail-closed key rejects known estimate over remaining budget before provider (Pylon #7691)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): streamed /v1/responses success callbacks keep provider response headers (Pylon #7775)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "test(integration): fail-closed key rejects known estimate over remaining budget before provider (Pylon #7691)"
This reverts commit 910348be7a.
* test(integration): reconcile contracts manifest for bundled regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): isolate proxy config writes in bundled regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): address review feedback on budget reset bounds and callback batch accumulation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): avoid rebinding the cache identity accumulator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): assert forwarded messages per cache identity call
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): make budget reset and team default tests deterministic
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
110 lines
4.8 KiB
Python
110 lines
4.8 KiB
Python
import uuid
|
|
from hashlib import sha256
|
|
from pathlib import Path
|
|
from typing import Final
|
|
|
|
import pytest
|
|
import yaml
|
|
from integration._support.client import Gateway, eventually, object_value, string_value
|
|
from integration._support.database import read_rows
|
|
from integration._support.process import owned_proxy
|
|
from integration._support.upstream import delete_scenario, register_scenario
|
|
from integration.cost_calculation.cost_tracking_case import JsonResponse
|
|
from pydantic import JsonValue
|
|
|
|
INPUT_COST_PER_TOKEN: Final = 0.000001
|
|
OUTPUT_COST_PER_TOKEN: Final = 0.001
|
|
PROMPT_TOKENS: Final = 10
|
|
CANDIDATE_TOKENS: Final = 5
|
|
COST_PER_CALL: Final = PROMPT_TOKENS * INPUT_COST_PER_TOKEN + CANDIDATE_TOKENS * OUTPUT_COST_PER_TOKEN
|
|
MAX_BUDGET: Final = 0.02
|
|
CALLS_WITHIN_BUDGET: Final = 4
|
|
|
|
|
|
def _key_spend(digest: str) -> float:
|
|
rows: Final = read_rows('SELECT spend FROM "LiteLLM_VerificationToken" WHERE token=%s', (digest,))
|
|
assert len(rows) == 1, rows
|
|
return float(rows[0]["spend"])
|
|
|
|
|
|
def _generate_content_request(model: str) -> dict[str, JsonValue]:
|
|
return {"contents": [{"role": "user", "parts": [{"text": f"budget {model}"}]}]}
|
|
|
|
|
|
def _generate_content_response(model: str) -> JsonResponse:
|
|
return JsonResponse(
|
|
content_type="application/json",
|
|
body={
|
|
"candidates": [
|
|
{
|
|
"content": {"parts": [{"text": f"scripted answer {model}"}], "role": "model"},
|
|
"finishReason": "STOP",
|
|
"index": 0,
|
|
}
|
|
],
|
|
"usageMetadata": {
|
|
"promptTokenCount": PROMPT_TOKENS,
|
|
"candidatesTokenCount": CANDIDATE_TOKENS,
|
|
"totalTokenCount": PROMPT_TOKENS + CANDIDATE_TOKENS,
|
|
},
|
|
"modelVersion": model,
|
|
},
|
|
)
|
|
|
|
|
|
def _served_call(gateway: Gateway, model: str, key: str, scenario_id: str, call: int) -> None:
|
|
digest: Final = sha256(key.encode()).hexdigest()
|
|
spend_before: Final = _key_spend(digest)
|
|
assert spend_before == pytest.approx((call - 1) * COST_PER_CALL) and spend_before < MAX_BUDGET
|
|
response: Final = gateway.request(
|
|
"POST",
|
|
f"/gemini/v1beta/models/{model}:generateContent",
|
|
_generate_content_request(model),
|
|
headers={"x-goog-api-key": key, "x-pass-x-scripted-scenario": scenario_id},
|
|
)
|
|
assert response.status_code == 200, f"call {call} with key spend {spend_before}: {response.text}"
|
|
assert response.json() == _generate_content_response(model).body, response.text
|
|
eventually(lambda: _key_spend(digest), lambda spend: spend >= call * COST_PER_CALL - 1e-9, seconds=70)
|
|
|
|
|
|
@pytest.mark.covers("spend.budget_reservation.gemini_passthrough_success_releases_reservation_from_spend_counter")
|
|
def test_repeated_gemini_passthrough_calls_stay_served_while_key_spend_is_below_max_budget(
|
|
gateway: Gateway, tmp_path: Path
|
|
) -> None:
|
|
config: Final = yaml.safe_load(Path("tests/integration/proxy_config.yaml").read_text())
|
|
config["environment_variables"] = {
|
|
"GEMINI_API_BASE": gateway.upstream_url,
|
|
"GEMINI_API_KEY": "scripted",
|
|
}
|
|
path: Final = tmp_path / "gemini-passthrough.yaml"
|
|
path.write_text(yaml.safe_dump(config))
|
|
with owned_proxy(gateway, tmp_path, {}, config=path) as candidate, candidate.scenario() as scenario:
|
|
model: Final = f"gemini-passthrough-{uuid.uuid4().hex}"
|
|
created: Final = candidate.post(
|
|
"/model/new",
|
|
{
|
|
"model_name": model,
|
|
"litellm_params": {
|
|
"model": "gemini/gemini-2.5-flash",
|
|
"api_key": "scripted",
|
|
"api_base": gateway.upstream_url,
|
|
"input_cost_per_token": INPUT_COST_PER_TOKEN,
|
|
"output_cost_per_token": OUTPUT_COST_PER_TOKEN,
|
|
},
|
|
"model_info": {"id": model, "max_output_tokens": 10},
|
|
},
|
|
)
|
|
scenario.cleanups.callback(scenario.delete_model, string_value(object_value(created["model_info"])["id"]))
|
|
handle: Final = register_scenario(f"sc-{model}", _generate_content_response(model))
|
|
scenario.cleanups.callback(delete_scenario, handle)
|
|
key: Final = scenario.key(models=[model], max_budget=MAX_BUDGET)
|
|
for call in range(1, CALLS_WITHIN_BUDGET + 1):
|
|
_served_call(candidate, model, key, handle.scenario_id, call)
|
|
assert _key_spend(sha256(key.encode()).hexdigest()) == pytest.approx(CALLS_WITHIN_BUDGET * COST_PER_CALL)
|
|
denied: Final = candidate.request(
|
|
"POST",
|
|
f"/gemini/v1beta/models/{model}:generateContent",
|
|
_generate_content_request(model),
|
|
headers={"x-goog-api-key": key, "x-pass-x-scripted-scenario": handle.scenario_id},
|
|
)
|
|
assert denied.status_code == 422 and denied.json()["error"]["type"] == "budget_exceeded", denied.text
|