Commit graph

9502 commits

Author SHA1 Message Date
mateo-berri
8ccbd82bd2 fix: authorize custom-logger spend payload against its own owner
The custom-logger detail branch reads the payload from cold storage, which is
written independently of the spend-log table and can outlive its row. The DB
owner pre-check then has nothing to verify for an id lookup that matches no row,
so a foreign tenant's stored payload could be returned. Authorize the returned
payload against the owner recorded inside it (metadata user/team id), failing
closed when none is recorded. Also fold the three identical 403 raises into one
helper.
2026-09-01 12:13:26 -07:00
mateo-berri
61cef45d8e fix: re-verify ownership on custom-logger payload branch for id lookups 2026-09-01 11:33:20 -07:00
mateo-berri
5382f9b720 fix: re-verify ownership on fetched spend log rows for id lookups 2026-09-01 11:20:36 -07:00
mateo-berri
c0adca7c94 fix(spend-logs): authorize request_id lookups over distinct owners, build call id index concurrently 2026-09-01 10:20:29 -07:00
mateo-berri
b35a1321eb fix(spend-logs): require every ambiguous request_id match to be owned
litellm_call_id is populated from the client-settable x-litellm-call-id
header, so a request_id lookup can match more than one row across
tenants. Authorizing on a single arbitrary match let an attacker reuse
a victim's request_id as their own call id and read the victim's spend
log row. Widen the ownership check to require every matching row to
belong to the caller, failing closed on any foreign match.
2026-08-31 22:14:49 -07:00
mateo-berri
f1dea17be1 fix(spend_logs): store litellm_call_id and match it in request_id lookups
Success spend rows are keyed by the upstream provider response id, so the
x-litellm-call-id response header value never found them. Add a nullable
indexed litellm_call_id column to LiteLLM_SpendLogs, populate it at write
time, and widen every request_id lookup surface (/spend/logs,
/spend/logs/ui, request details, ownership check) to match either id.
2026-08-31 21:37:27 -07:00
Kolade Fajimi
d1320404fe
fix(redis): coerce env var string types and fix param discovery through decorator wrappers (#30644)
* fix(redis): coerce env var string types and fix param discovery through decorator wrappers

inspect.getfullargspec doesn't work on redis.Redis/redis.RedisCluster because
their __init__ is wrapped by @deprecated_args, which replaces the explicit
signature with *args/**kwargs internally. getfullargspec returns an empty arg
list, so _get_redis_kwargs and _get_redis_cluster_kwargs silently dropped every
real constructor parameter not in their hand-picked include_args set --
cluster_error_retry_attempts and connection_error_retry_attempts among them, so
an operator's configured retry bound never reached the Redis Cluster client and
it fell back to redis-py's own default instead.

Rebased onto litellm_internal_staging, which had independently added
_init_arg_names (MRO-walking, inspect.unwrap-based) for the same class of bug
in _get_redis_url_kwargs. Reused that pattern (as _unwrapped_init_args, without
the MRO walk: redis.Redis/RedisCluster declare every real parameter directly on
their own __init__, and MRO-walking breaks the tests here that mock the class
with autospec=True, since inspect.getmro needs a real __mro__) rather than
introducing a second, differently-shaped fix for the same problem.
_get_redis_cluster_kwargs now also honors its own client argument instead of
ignoring it, so the async cluster client's own extra constructor kwargs
(cluster_error_retry_attempts, connection_error_retry_attempts,
decode_responses, ...) are no longer filtered out by introspecting the sync
class regardless of which client is actually built.

Also fixes environment variables and Helm --set values always arriving as
strings: redis-py 8.x changed health_check_interval's arithmetic to require a
real number, so a stringified value raised TypeError on every Redis operation
instead of connecting. _coerce_redis_kwargs_types coerces to each parameter's
declared type at the end of _get_redis_client_logic, with an explicit type
table for max_connections/socket_timeout/socket_connect_timeout since redis-py
8.x changed the timeout defaults from None to int 5, which would otherwise
make a fractional value fail int() and get dropped.

Co-authored-by: mangabits <1457532+mangabits@users.noreply.github.com>

* ci: verify redis-py client version compatibility across a version matrix

* test(redis): assert an async-only cluster kwarg every matrix version declares

connection_error_retry_attempts is on the async cluster constructor in redis-py
5.x only; 6.0 removed it in favor of retry. The 6.4.0, 7.4.1 and 8.0.1 legs were
failing on that missing parameter name rather than on the behavior under test,
while the allow-list itself was doing the right thing on all four versions.

decode_responses is async-cluster-only on every version the matrix covers, so it
stands in for the same property: the sync cluster class takes it through **kwargs
and never names it in its signature. Reverting _get_redis_cluster_kwargs to ignore
its client argument still fails both tests on 5.3.1 and 8.0.1.

test_async_cluster_passes_async_only_kwargs now builds the real async cluster
client and reads connection_kwargs off it, so it no longer needs a patched class
factory; the constructor does no I/O. The retry-attempts test keeps its patch,
since redis-py >= 6 stores no cluster_error_retry_attempts attribute on the built
client and the constructor call is the only place the forwarded value shows up.

The _get_redis_cluster_kwargs docstring cited the same two parameters as its
examples of async-only kwargs, which is what made the test look reasonable;
cluster_error_retry_attempts is on both classes and connection_error_retry_attempts
is gone from 6.0 on, so it now names decode_responses instead.

* test(redis): drop internal patches from the kwarg coercion tests

The test-quality gate flagged the new patch() calls on litellm internals these
tests added. Three of them faked litellm._redis.inspect.signature with a MagicMock
to hand _coerce_redis_kwargs_types a synthetic parameter; that function already
takes a client argument, so they pass stub functions instead, matching the
_redis_signature_8x idiom the file uses elsewhere. The fourth patched
_redis_kwargs_from_environment to {} to prove _get_redis_client_logic raises
without a host or url, which clearing the real env keys through
_get_redis_env_kwarg_mapping does without pinning the test to that call.

Both files now sit one TQ008 below the merge base rather than six above it.

* fix(redis): keep the sync client construction inside the basedpyright budget

_get_redis_client_logic now returns dict[str, object] rather than an untyped
dict, which is the honest type for operator-supplied config, but it turns the
33 reportUnknownArgumentType errors at redis.Redis(**redis_kwargs) into 33
reportArgumentType errors plus one reportCallIssue, both over their budget.
No static type fits: redis-py's constructor declares 40-odd differently typed
parameters and the values arrive from config and env, so the allow-list and
coercion above are derived from that same signature and redis-py validates each
value itself at runtime.

The two suppressions name their exact rule and carry that reason. The file ends
up 42 basedpyright errors below the merge base, with reportArgumentType and
reportCallIssue back at the base counts of 3 and 0.

* fix(redis): coerce cluster-only and None-default bool kwargs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mangabits <1457532+mangabits@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 20:51:31 -07:00
tin-berri
8d6d7f9ce9
feat(complexity_router): opt-in modality-based capability routing for image requests (#39032) 2026-08-31 19:51:41 -07:00
yucheng-berri
ccd76dac50
fix(proxy): wire team-level logging callbacks into passthrough endpoints (#38979)
* fix(proxy): wire team-level logging callbacks into passthrough endpoints

LIT-5152: passthrough routes now wire dynamic team-level callbacks
(success_callback, failure_callback, callback_vars) into Logging constructor,
mirroring the add_litellm_data_to_request behavior. Three hardening fixes:

1. Catch TypeError/AttributeError in _get_validated_callback_metadata when
   team logging metadata has wrong shape (e.g., logging list instead of dict),
   preventing HTTP 500 on passthrough routes with malformed config.

2. Wrap websocket passthrough logging initialization in try/except, since
   the socket is already accepted at that point; errors after accept() yield
   abrupt close (1006/1011) rather than clean HTTP error response.

3. Handle malformed deprecated callback_settings gracefully with try/except.

4. Wrap HTTP passthrough callback resolution in try/except to prevent 500 on
   malformed team metadata (backward-compatibility fix).

Changes:
- pass_through_endpoints.py: wire dynamic callbacks in HTTP+WS paths, handle
  malformed metadata gracefully with try/except fallbacks
- litellm_pre_call_utils.py: expand exception handling in validators
- test file: regression test for happy-path team callback wiring

* refactor(proxy): share passthrough team-callback resolution and cover its fail-open path

Collapse the duplicated callback wiring on the HTTP and websocket passthrough
paths into one helper that returns a frozen wiring value, log resolution
failures at error level so a broken logging config stays visible, and add
regression tests for malformed team metadata and an operational lookup failure.

Reverts the _get_validated_callback_metadata except widening: it changed
behavior for normal LLM routes, which is outside this ticket's scope.

* fix(proxy): keep passthrough alive when team callback vars hold env references

The deprecated team_metadata.callback_settings branch builds
TeamCallbackMetadata directly, skipping the AddTeamCallback validation
that strips os.environ/ references from the newer logging list. Stamping
those vars onto the Logging object made its constructor raise, so a team
on the legacy shape got HTTP 500 on every passthrough call. Validate the
resolved vars inside the fail-open boundary instead, so the request goes
through with dynamic callbacks skipped and the reason logged.

* fix(proxy): lint violations in team callback wiring helper

* style: format lint
2026-09-01 02:38:39 +00:00
yucheng-berri
3fadcd7155
fix(auth): quiet malformed virtual key rejections to stdout (#38838)
* fix(auth): quiet malformed virtual key rejections to stdout

Reduce noisy invalid-api-key error logs by classifying malformed virtual
keys and routing their rejections to stdout as WARNING instead of stderr
as ERROR. Suppressible via LITELLM_LOG=ERROR or log_client_error_tracebacks=true.

Changes:
- auth_utils: is_invalid_virtual_key_error() classifier and marker functions
- auth_exception_handler: log invalid keys as WARNING to child logger before
  identity seeding and callbacks, escalate non-401 transforms to ERROR
- user_api_key_auth: websocket early-raise WebSocketException(1008) to avoid
  double-logging at HTTP layer
- _logging: child logger verbose_proxy_stdout_logger with no handler/level;
  LevelRoutingStreamHandler routes its WARNING records to stdout; handler
  setLevel in _turn_on_json() closes JSON config handler level leak
- test_auth_exception_handler: new test case verifying malformed-key logs
  at WARNING with marker retention through transformations

Fixes LIT-5362

* fix(auth): classify malformed-key 401 by raise-site marker, not message text

Review round 1 (Greptile P2, veria Low):
- Move the marker attribute name to litellm/constants.py per the shared
  sentinel convention
- Stamp the marker on the malformed-key 401 where it is raised and classify
  only by it. Message text is caller-influenceable on other 401s (vector
  store ids, organization ids are interpolated into their messages), so a
  phrase match would let a request body demote an authorization failure to
  the quiet log path
- Regression test: a 401 carrying the phrase but not the marker stays at
  ERROR on stderr
2026-08-31 18:19:40 -07:00
ryan-crabbe-berri
50eed7efa2
Merge pull request #38851 from BerriAI/litellm_window_spend_seed_by_time
refactor(proxy): bound the budget window seed by time instead of request ids
2026-08-31 17:21:45 -07:00
Mateo Wang
99703a30f0
Merge pull request #36008 from nuernber/litellm_bedrock_messages_disconnect_billing
fix(anthropic_messages): drain upstream in a detached pump so client …
2026-08-31 16:54:41 -07:00
mateo-berri
a99f62d1bc fix(anthropic_messages): park deferred billing before the end-of-stream sentinel
At end of drain the pump enqueued the sentinel first and picked the
billing mode from client_detached afterward, so a client that consumed
the sentinel and tore the relay down before the pump resumed (possible
whenever the sentinel enqueue hit a full queue) had its fully delivered
response billed through the teardown path, skipping the proxy's
post-response hook. Bill or park before the sentinel goes out, and let
an unconsumed sentinel fall back to dispatching the parked billing.
2026-08-31 16:46:12 -07:00
tin-berri
3829418878
feat(shadow_eval): target teams and users so JWT-auth traffic can be evaluated (#39015)
Shadow eval jobs previously targeted only virtual keys, so deployments on
pure JWT auth (which present no key at all) could never sample their
traffic. Jobs now carry a typed (target_type, target_id) pair covering
keys, teams, and users; sampling matches the identity every request
resolves to at auth time, so team and user jobs cover JWT traffic with
no client changes.

Resolves LIT-6578
2026-08-31 16:37:38 -07:00
tin-berri
f93d9b6b67
feat(complexity_router): escalate oversized prompts to a tier that fits before dispatch (#38844)
* feat(complexity_router): escalate oversized prompts to a tier that fits before dispatch

The classifier scores complexity and never prompt size, so a long agentic
session whose newest ask is trivial classifies SIMPLE onto a small-window
tier and the provider rejects it with a context-window 400 that nothing
retries. The gate runs after classification on every decision path
(classify tail and session-affinity pin), estimates prompt tokens
including the out-of-band carriers (top-level system, tools,
instructions), and when the decided tier provably cannot hold the prompt
moves the request to the lowest configured tier with a model whose
declared window fits, restricting the pick to fitting models when the
decided tier can keep it. Models with no resolvable window are never
escalated away from or onto, escalated decisions are never written as
session pins, and the decision records context_escalated plus the
original tier in spend logs.

Resolves LIT-6503

* fix(complexity_router): judge groups by smallest window, bound skips by bytes, filter adaptive picks

Review-round rework, one mechanism per finding. A group is judged by its
smallest resolvable deployment window, since the core router picks within
a group with no fit check. The counting skip is gated on UTF-8 byte
length, which BPE token counts can never exceed, so token-dense scripts
cannot slip past it; only a real tokenizer count ever moves a request and
a failed count leaves the placement alone. The fit facts now filter every
adaptive phase including cold start and the tier fallbacks. Window
questions adopt the declared provider and never resolve authenticating
providers, and a router instance without get_model_list degrades the gate
to a no-op. Tests rebuilt on real Router instances resolving deployment
model_info end to end, plus a full-path test through
async_get_available_deployment
2026-08-31 16:12:59 -07:00
Mateo Wang
81c8c93bef
Merge pull request #38873 from BerriAI/litellm_fix_model_block_response_500
fix(proxy): return 200 from /model/block and /model/unblock instead of 500
2026-08-31 15:56:21 -07:00
Mateo Wang
d60e77ae8c
Merge pull request #38819 from BerriAI/litellm_fix_gemini_tts_response_format
fix(speech): stop forwarding response_format as a chat param for Gemini TTS
2026-08-31 15:56:18 -07:00
Mateo Wang
9b87413540
Merge pull request #38878 from BerriAI/litellm_fix_master_key_rotation_blocked
fix(proxy): preserve model table columns on master key rotation
2026-08-31 15:56:03 -07:00
tin-berri
296bde0d0d
feat(complexity-router): add classification_mode to skip classifier on continuation turns (#38861) 2026-08-31 15:50:04 -07:00
ryan-crabbe-berri
5eff708d0f fix(proxy): keep persisted spend another pod has not incremented in the window seed
The batch start told the seed which LiteLLM_SpendLogs rows were its own, but
using it as a hard cutoff also dropped rows another pod had already persisted.
Those rows are only repaid by that pod's own increment, so if it died first the
window row stayed permanently under the recorded spend.

The seed now reads both sums in one scan and takes off this batch's own spend,
flooring at the pre-batch total for the case where its log rows have not landed
yet. Redis payloads keep an empty request_ids so a leader from before the field
was dropped can still merge what it pops during a rolling deploy.

Claude-Session: https://claude.ai/code/session_01QvQzYztinxj8ZuD5YxbVdL
2026-08-31 15:49:03 -07:00
Mateo Wang
7a02e4163f
Merge pull request #38913 from BerriAI/litellm_gigachat_passthrough_25886
feat(gigachat): add native API passthrough routes with spend logging
2026-08-31 15:47:27 -07:00
mateo-berri
eb00986f18 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886
# Conflicts:
#	osv-scanner.toml
2026-08-31 15:25:10 -07:00
Mateo Wang
5a0ed05765
Merge pull request #38479 from yatishgoel/litellm_fix_model_rename_router_sync
fix(router): apply model renames to the in-memory deployment list
2026-08-31 15:22:30 -07:00
Mateo Wang
ac206518a0
Merge pull request #38881 from Lee-Si-Yoon/friendli/glm-5.3
feat(friendli): add zai-org/GLM-5.3 model pricing
2026-08-31 15:18:45 -07:00
mateo-berri
91595780ec fix(gigachat): fold cached tokens back into prompt and total token counts
GigaChat reports prompt_tokens and total_tokens after subtracting cached
tokens (the docs example is prompt_tokens=1, precached_prompt_tokens=37,
total_tokens=5, so the fields are disjoint, not a subset). Map to the
OpenAI convention by adding precached_prompt_tokens back onto prompt and
total while still surfacing it as prompt_tokens_details.cached_tokens.
2026-08-31 15:16:44 -07:00
Mateo Wang
0c127caedb
Merge pull request #38940 from samtsai15/fix/anthropic-guardrail-image-sources
fix(guardrails): carry Anthropic url image sources through to guardrails
2026-08-31 15:16:04 -07:00
Mateo Wang
336269cfec
Merge pull request #38597 from BerriAI/litellm_fix_nova_sonic_realtime_user_asr_usage
fix(bedrock): surface Nova Sonic user transcripts, speech events, and usage in realtime API
2026-08-31 15:11:31 -07:00
mateo-berri
abbccd3fd6 Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into friendli/glm-5.3
# Conflicts:
#	litellm/model_prices_and_context_window_backup.json
#	model_prices_and_context_window.json
2026-08-31 15:11:06 -07:00
Mateo Wang
b3a1dd1115
Merge pull request #38880 from Lee-Si-Yoon/friendli/glm-5.3-flash-v2
feat(friendli): add zai-org/GLM-5.3-Flash model pricing
2026-08-31 15:07:06 -07:00
Devin AI
cc078edd1f test(mcp): isolate global MCP server registry in discoverable endpoints tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 21:56:57 +00:00
Devin AI
6809d537f0 merge: litellm_internal_staging into litellm_fix_nova_sonic_realtime_user_asr_usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 21:42:06 +00:00
Mateo Wang
b518be45fb
Merge pull request #38997 from BerriAI/litellm_add_responses_input_tokens_endpoint
feat(proxy): add /v1/responses/input_tokens token counting endpoint
2026-08-31 14:30:44 -07:00
mateo-berri
98b1e2e7b4 fix(gigachat): correct cached-token accounting, stream usage on all finish reasons, stop mutating cached request body
precached_prompt_tokens is a subset of prompt_tokens (OpenAI cached_tokens
semantics), so map it to prompt_tokens_details.cached_tokens instead of
adding it on top of prompt/total. Emit stream usage from any final chunk
carrying it rather than only finish_reason stop, which dropped tokens for
function_call and length streams. Merge auth metadata into a new dict in
the gigachat router handler instead of mutating the shared parsed-body
cache in place.
2026-08-31 14:28:30 -07:00
devin-ai-integration[bot]
40edeaaecb
fix(otel): emit cache token counts on OTel v2 LLM spans (#38716)
* fix(otel): emit cache token counts on OTel v2 LLM spans

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): trim comment in LLMUsage adapter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): drop casts in LLMUsage cache token adapter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(deps): bump restrictedpython to 8.3 for GHSA-ffg3-p8fm-mjx2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 13:59:59 -07:00
mateo-berri
b7da471784 fix(count_tokens): price an inline file block instead of raising on it
`ChatCompletionFileObject` is in the union `_count_content_list` accepts, but
`file` was missing from its match, so every local count of a Responses
`input_file` raised `Invalid content item type: file`. On
/v1/responses/input_tokens that surfaced as an opaque 500 whenever the model's
provider counting API refused the block and the local tokenizer took over.

Count it the way the module already counts the same thing in Anthropic's
dialect: the filename like a document title, the inline bytes through the
image pricer.
2026-08-31 13:53:22 -07:00
Mateo Wang
be3386f3ce
Merge pull request #38995 from BerriAI/litellm_openai_wif
feat(openai): support workload identity federation (OIDC token exchange)
2026-08-31 13:37:30 -07:00
mateo-berri
bd794f9f18 fix(friendli): track GLM-5.3 discounted live pricing and declare effort levels 2026-08-31 13:26:37 -07:00
mateo-berri
a90fb538bf fix(friendli): declare GLM-5.3-Flash reasoning efforts as explicit levels 2026-08-31 13:25:51 -07:00
Mateo Wang
66295e7da7
Merge pull request #34696 from cat0825/fix/34379-unblock-customer
fix(proxy): allow unblocking customers via /customer/update
2026-08-31 13:21:00 -07:00
mateo-berri
c9908ffabb fix(responses): count input_file tokens instead of silently dropping the file
The Responses-to-chat transform dropped the filename OpenAI requires next to
file_data, so a request carrying an inline PDF counted 13 tokens instead of 36
and a real completion through the chat bridge got a 400.
2026-08-31 13:17:43 -07:00
mateo-berri
59732f068b Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gigachat_passthrough_25886 2026-08-31 13:16:41 -07:00
mateo-berri
ed5ee51dd2 fix(passthrough): map sync streaming errors, keep router streaming responses unwrapped, and resolve gigachat from api base
- sync llm_passthrough_route: read and close an error-status streaming
  response before mapping it, so upstream 4xx/5xx surface as the provider
  error instead of httpx.ResponseNotRead
- AsyncPassthroughStreamingResponse: expose aiter_bytes() and carry
  _hidden_params so the router attaches headers in place instead of
  wrapping the stream in HiddenParamsAsyncIteratorWrapper, which 500'd
  every streaming azure router-model passthrough request
- logging: swap the passthrough httpx result for the transformed
  ModelResponse/EmbeddingResponse when firing success callbacks
- get_llm_provider: resolve gigachat from its api base and drop the dead
  gigachat_models elif branch
- constants: register the gigachat api base in openai_compatible_endpoints
2026-08-31 13:16:40 -07:00
mateo-berri
fe90c6f6fc fix(count_tokens): keep assistant turns on the provider counting API
Assistant list content was forwarded to /v1/responses/input_tokens as chat
`text` blocks, which the Responses API rejects (it accepts only output_text
and refusal inside an assistant turn). The 400 sent the whole request to the
local tokenizer, so any conversation with an assistant turn silently lost
provider-exact counting, including the image counting added in 73ab647b1c.

Assistant content now collapses to the plain string the Responses API counts
identically, and image parts are kept to user turns where they are legal.
2026-08-31 12:59:40 -07:00
devin-ai-integration[bot]
1249f84b10
fix(vertex_ai): graft default vertex path when api_base has a version-only path (#38986)
* fix(vertex_ai): graft default vertex path when api_base has a version-only path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): keep query and fragment placement when grafting vertex path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): merge alt=sse into existing query when streaming

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-31 12:56:34 -07:00
Cursor Agent
ae83444a3e
fix(openai): treat empty api_key as unset for WIF resolution 2026-08-31 19:55:29 +00:00
mateo-berri
e7dc0213bd fix(openai): treat empty api key values as unset for workload identity 2026-08-31 12:52:55 -07:00
Mateo Wang
0c21b30cb7
feat(spend_tracking): persist router metadata in spend logs for internal router models (#39001)
* feat(spend_tracking): persist router metadata in spend logs for internal router models

* test(spend_tracking): expect router_metadata key in exact-payload tests, type the routed-kwargs helper
2026-08-31 12:52:34 -07:00
Ashton Sidhu
9f9236e8d5
fix(guardrails): exclude images from HiddenLayer v1 scans (#29210)
* Don't scan images

* Fix failing tests

* Fix lint: typed image-part filter, restore monkeypatch-based tests

---------

Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
2026-08-31 12:50:42 -07:00
mateo-berri
46e090d2f3 fix(anthropic_messages): bill partial spend when a queued pump error is never consumed
When the upstream errors while the client is still connected, the pump
forwards the exception through the relay queue so the proxy's failure
handling re-raises it. If the client disconnects before consuming that
queued exception, neither the failure hook nor billing ran and the spend
row was lost. The pump now waits for client detach and, if the exception
was never consumed, salvages partial spend like the post-disconnect
error path.

Also rewrites the bedrock disconnect logging test to the detached-pump
contract: billing fires after the upstream drain completes, not
synchronously at aclose().
2026-08-31 12:47:26 -07:00
mateo-berri
73ab647b1c fix(count_tokens): preserve image inputs when counting Responses API tokens
The chat-to-Responses reverse transform kept only text blocks, so an image
input was dropped before the count went to OpenAI. A 256x256 image request
counted 13 tokens instead of 268.
2026-08-31 12:43:17 -07:00