Commit graph

19 commits

Author SHA1 Message Date
Yuneng Jiang
963cb4694d
fix(proxy): gate image-gen reservation strictly on model mode
The previous detection treated any model with input_cost_per_image
or output_cost_per_image as image generation. Several chat and
embedding models carry those fields to price multimodal vision input,
not generated images:

- gemini-3.1-pro-preview (mode=chat) has output_cost_per_image=0.00012
  alongside input/output token pricing.
- azure/gpt-realtime-* (mode=chat) has input_cost_per_image=5e-6.
- amazon.titan-embed-image-v1 (mode=embedding) has
  input_cost_per_image=6e-5.

For these models the image-gen branch fired first and reserved a
fraction of a cent per request, short-circuiting the token-priced
path entirely. Long Gemini chats reserved 1 × $0.00012 instead of
the true token cost.

Gate strictly on mode in {"image_generation", "image_edit"}. All 197
real image_generation entries and all 31 image_edit entries
(Flux Kontext, Stability inpaint/outpaint, etc.) carry the right mode,
so the field-presence fallback was unnecessary.

Adds regression tests for the chat-model-with-image-cost-field case
and for image_edit reservation.
2026-05-09 09:16:27 -07:00
Yuneng Jiang
0d551ac4f0
fix(proxy): reserve per-image cost for image-generation requests
Image-generation routes (dall-e-3, flux, etc.) have no per-token output
cost so they fell through to the no-reservation read-time-only path.
Concurrent image requests against a depleted budget could all pass
common_checks (counter exactly at max_budget passes the strict-`>`
gate) and reach the provider before reconciliation caught up.

Add per-image reservation in _estimate_request_max_cost_for_model:
when the model has a per-image cost field, reserve `n × cost_per_image`
upfront. The atomic counter increment serializes concurrent admissions,
so the second request sees the post-first-reservation counter and
raises BudgetExceededError instead of silently leaking through.

Both `output_cost_per_image` and `input_cost_per_image` are honored —
naming is inconsistent across providers (OpenAI dall-e-3 uses
input_cost_per_image, aiml/dall-e-3 uses output_cost_per_image for
the same per-generated-image price).

Per-pixel pricing (DALL-E 2 size variants) and TTS/STT routes still
fall through to read-time enforcement; those are follow-ups.
2026-05-08 21:08:55 -07:00
Yuneng Jiang
adc41ade8c
fix(proxy): bound budget reservation per request instead of pinning to remaining headroom
reserve_budget_for_request fell back to reserving the entire remaining
team/key/user headroom whenever a request omitted max_tokens, which
pinned the spend counter at max_budget for the duration of the
in-flight request and false-positive-blocked every concurrent or
back-to-back request until the success callback reconciled. Surfaced
as an integration-test team being budget-blocked at its $2000 cap
while DB spend was $0.144.

Switch the missing-max_tokens path to a fixed default of 16384 output
tokens (mirrors parallel_request_limiter_v3's DEFAULT_MAX_TOKENS_ESTIMATE
precedent), and clamp explicit max_tokens at the model's
max_output_tokens for reservation accounting only. The outbound request
body is unchanged, so providers see whatever the caller actually sent;
only the local integer used to compute reservation cost is bounded.
This also prevents a hostile max_tokens=999999999 from inflating one
request's reservation up to the entire team headroom.

For Opus 4.7 (output $25/M, max_output 128K) on a $2000 budget the
worst-case per-request reservation drops from "everything left" to
$3.20, raising admittable concurrency from 1 to ~625.
2026-05-08 20:18:31 -07:00
user
83ed317c50 track reservation entry before counter write 2026-05-01 00:09:51 -07:00
user
403bbc3b88 degrade budget reservation cache failures 2026-04-30 23:53:36 -07:00
user
0b1ea9eb8f harden budget reservation edge cases 2026-04-30 21:49:31 -07:00
user
dcfde1b899 fallback to plain org cache for spend counters 2026-04-30 20:04:16 -07:00
user
405de46329 cap budget reservations to remaining headroom 2026-04-30 19:24:48 -07:00
user
46068be6f6 skip invalid budget window reservations 2026-04-30 18:29:14 -07:00
user
38ebd4de3d harden partial budget reservation cleanup 2026-04-30 18:07:54 -07:00
user
fce86d1334 fix budget reservation greptile findings 2026-04-30 17:08:45 -07:00
user
e034935b53 fix budget reservation window fallback races 2026-04-30 16:26:43 -07:00
user
1373ae1021 fix budget tag spend counter reconciliation 2026-04-30 16:09:38 -07:00
user
719e891c3a coalesce malformed window reservation seeding 2026-04-30 14:38:37 -07:00
user
0b71282985 address budget reservation review findings 2026-04-30 14:06:42 -07:00
user
0794ae67be avoid direct budget reservation db lookups 2026-04-30 13:49:59 -07:00
user
ca50868b75 harden end-user and tag budget reservations 2026-04-30 13:37:10 -07:00
user
09503ebb8f harden budget reservation recovery 2026-04-29 21:06:30 -07:00
user
5a619cf879 tighten budget spend admission 2026-04-29 20:30:09 -07:00