Bedrock invoke /v1/messages streaming reports cache_read_input_tokens and
cache_creation_input_tokens on message_stop.usage while attaching
amazon-bedrock-invocationMetrics to the same chunk. The stream decoder
rebuilt that chunk's usage block from inputTokenCount/outputTokenCount
alone, which exclude cache reads and writes, so the cache breakdown was
destroyed before _promote_message_stop_usage could surface it and cache
tokens were billed at $0. Merge instead of replace, and also map
cacheReadInputTokenCount/cacheWriteInputTokenCount when Bedrock reports
the cache itemization inside the invocation metrics.
Co-authored-by: Brian Cox <3924351+brian5021@users.noreply.github.com>
Azure rejects the legacy `max_tokens` key for the whole gpt-5 name family, but
`AzureOpenAIGPT5Config.is_model_gpt_5_model` deliberately excludes `gpt-5-chat*`
so those deployments fall through to `AzureOpenAIConfig`, which sends `max_tokens`
verbatim and gets a 400 back on every request that carries it, `/health` probes
included.
One predicate was answering two independent questions. Split it: the new
`AzureOpenAIConfig.requires_max_completion_tokens` covers the whole gpt-5 name
family and drives only the rename, while `is_model_gpt_5_model` keeps keying
reasoning_effort, the temperature clamp and the dropped penalties off the
reasoning question, so #13781 stays fixed.
Adds an opt-in operator allow-list, litellm_settings::bedrock_request_metadata_fields, that forwards LiteLLM key, team and end-user identity plus client spend_logs_metadata into Bedrock request metadata so Bedrock spend can be grouped in AWS Cost Explorer.
Covers all three Bedrock surfaces: the Converse body requestMetadata field, and a signed X-Amzn-Bedrock-Request-Metadata header on Invoke chat completions and on Invoke /v1/messages, where the header is the only viable leg.
The resolver reads both metadata variable names, reserves the whole user_api_key_ prefix against caller-supplied keys, caps the client slot budget explicitly at 16 minus the reserved count, and drops rather than rejects auto-injected values that violate Bedrock constraints. Caller-supplied requestMetadata keeps its existing 400 semantics.
The request-metadata field and header are proxy-owned whenever forwarding is enabled. A caller-supplied value, reachable through the generic extra_headers passthrough, is dropped unconditionally and compared case-insensitively, and is replaced only by the proxy's own value, so identity in the AWS billing record cannot be forged. Absence of a resolved value still means absence on the wire rather than a fallback to the caller's. The guardrail headers keep their existing no-displace behaviour.
* fix(guardrails): scan text on /guardrails/apply_guardrail for Azure Content Safety
The two Azure Content Safety guardrails never implemented apply_guardrail, so the
endpoint fell through to the base no-op and answered 200 with the caller's text
echoed back, having scanned nothing.
Implementing that method also flips the proxy's unified-vs-native dispatch, which
would move request traffic off these guardrails' own hooks. Add an opt-out that
keeps every lifecycle event on the native hooks, so only the endpoint changes.
* test(guardrails): cover the remaining native-hook opt-out dispatch sites
Adds regression tests for the parallel post-call path, the MCP post-call hook, and
the policy engine step, so every read of the opt-out flag fails when removed.
Callers can opt into the provider's raw operation response on /v1/ocr with the x-req-format: native header (or req_format in the body) while page-based cost tracking keeps reading usage_info off the normalized response.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): link key info header to its user, creator, team, and organization
The key info page showed the owning user and creator as plain text and never surfaced the team or organization at all, so walking from a key to its parent entities meant copying ids into other pages. The header now renders User and Created By as links to the user detail page, and gains a far-right column with Team and Organization links (alias when known, id otherwise). Client-side navigation logic shared by BadgeLink and IdentityCell moves into a reusable EntityLink so all entity links behave the same
* test(ui): mock next/navigation in VirtualKeysTable test
KeyInfoView now renders EntityLink, which calls useRouter, so the table test that opens the key detail needs the app router mocked
* Add default model pin to complexity router UI
A complexity router's default model was only ever derived from the tiers, so
operators had no way to point the fallback at a model that is not first in the
Simple or Medium tier. Adds a Default Model select that records an explicit pin.
The pin is stored in complexity_router_config.default_model, which the backend
already reads, and mirrored onto complexity_router_default_model on save. Both
paths resolve through one helper that mirrors init_complexity_router_deployment:
a pin wins, otherwise MEDIUM or SIMPLE. Recording the pin in the config keeps it
distinguishable from a derived value, so a pin that happens to match the tiers
survives a round trip instead of being read back as tier tracking.
The edit modal only requires one non-empty tier, so a router with models in
COMPLEX alone can reach save with nothing the backend would pick. That now
blocks with an inline message rather than saving a router that raises at init.
* fix(UI): probe the pinned default model in the auto router connection test
The connection test built its targets from the tiers and the embedding model
only, so a Default Model pin outside every tier was never reached and a green
result could hide an unreachable default. model_info_view had already hand
rolled the dedupe and append locally, so the rule moved into
buildAutoRouterTestTargets and both call sites now share it.
* fix(ui): mirror backend precedence when resolving a complexity router default
The edit modal only recognized a pin stored in complexity_router_config.default_model,
so a router whose default lived solely in litellm_params.complexity_router_default_model
lost it on the next save. That field cannot be trusted outright either: before this PR
every save wrote a tier-derived value into it, so treating any value as a pin would
freeze legacy routers away from their tiers. Hydration now takes the config marker as
authoritative and falls back to litellm_params only when it diverges from what the tiers
alone derive, which is only reachable through an external API or config write.
Test Connection had the mirror-image bug: it fell back to complexity_router_config.default_model,
a UI-only marker init_complexity_router_deployment never reads, so it could probe a model
the router would never call. It now follows router.py exactly: litellm_params, else pure
tier-derivation.
Also reword a tooltip that hardcoded the Default Model select's position on the page, and
document the dual write and the create-vs-edit validation asymmetry.
Suppress the mutation/private-access lints the new
batch_successful_requests/batch_failed_requests plumbing triggers,
matching the existing suppressed pattern already used for
response_cost/batch_models on the same lines.
Batch retrieval already computed cost/usage on completion, but silently
dropped reasoning tokens and never counted per-line success/failure.
Adds BatchCostUsageResult (replacing bare cost/usage/models tuples) with
successful_requests/failed_requests, and threads reasoning_tokens through
the aggregated Usage. Both surface on SpendLogs the same way batch_models
already does.
Replace _attach_budget_limits_usage, which rewrote the caller's key_info dict, with _budget_limits_with_usage returning a new list. Callers assign the result once. Keeps the response shape and spend-counter read path identical while following the repo's no-mutation rule.
Shape detection and block normalization sat in the generic batch layer, which
let batch and live parsing of the same wire format drift apart. Both now live on
AmazonConverseConfig as is_converse_usage_shape and usage_from_batch_output, so
batch_utils asks the provider adapter rather than knowing Bedrock's field names.
Adds direct coverage for the shape predicate, the completion of an incomplete
block, cache-count inflation, and the streaming usage event that shares the
public transform. Drops the narrative banner from the batch tests.
The Final annotations on the new regression vars make the PR's net LIT010
delta -1 (one fewer than base), so the earlier bump to 16743 was an
over-estimate. Reset the limit to the base value 16715 so the one-way
budget ratchet passes; the codebase-wide total (16714) stays under it.
Address the Greptile review on #36397. The earlier commit only ratcheted
the LIT010 budget; it never applied the annotations, so
duplicates_in_one_item and duplicates_across_items were still bound
without a Final declaration (LIT010) and the first fixture line was at the
120-char ceiling. Annotate both with `: Final` and wrap the long literal.
RED -> GREEN: check_type_discipline flagged both vars LIT010 before ->
LIT010 gone after (file total 551 -> 549, LIT002 unchanged at 953);
test_calculate_web_search_requests_counts_unique_queries still passes.
Address Greptile review on #36397: duplicates_in_one_item and
duplicates_across_items lacked Final declarations (LIT010). Use bare
: Final so the inferred type stays list-based, avoiding an explicit
mutable annotation (LIT001), and ratchet the LIT010 budget down by one.
RED to GREEN: both vars flagged LIT010 before -> clean after; mapped
suite 146 passed, 100% diff coverage.
Gemini 3 per_query grounding is billed per unique search query the model
executes, ignoring empty queries. _calculate_web_search_requests summed every
non-empty webSearchQueries string across grounding metadata items, so repeated
queries within a request inflated web_search_requests and overstated cost. Count
distinct non-empty queries across items instead.
Fixes#36377
- decode upstream first frame as utf-8 instead of ascii
- reject model-restricted keys at connect to match HTTP model enforcement
- log the actual request path for /openai_passthrough traffic
ResponseAPIUsage.parse_cost already flattens Perplexity's
usage.cost.total_cost dict down to a float before it reaches the
perplexity cost calculator, so the isinstance(cost_info, dict) check
was always False on that path. Every Responses-mode Perplexity model
was silently falling back to manual token-rate calculation and
recording $0 spend whenever static per-token rates were missing.
`/health` already stripped `api_key` from each deployment row via
`ILLEGAL_DISPLAY_PARAMS`, but `extra_headers`, `headers`, and `aws_session_token`
were never added to that list, so `GET /health` leaked provider credentials
(Azure `api-key`, Google `x-goog-api-key`, Bearer tokens, AWS session tokens) in
plaintext to any caller, even without a master key.
Add those three fields to `ILLEGAL_DISPLAY_PARAMS` so `_clean_endpoint_data()`
omits them for all callers, matching how `api_key` is already handled.
Fixes#36898
Embeddings rows were identified by body shape (has `input`, no
`messages`/`prompt`), which also matches a `/v1/responses` batch row
and reserved zero output tokens for it -- letting a project caller run
large Responses generations against a quota-limited model without
consuming OTPM. Classify embeddings by the row's own `url` instead,
and read `max_output_tokens` as a Responses output cap alongside
`max_tokens`/`max_completion_tokens`.
Co-authored-by: Cursor <cursoragent@cursor.com>
Resolves conflicts from the upstream merge and addresses the Veria-AI
review comment on this PR: batch rows could bypass a project's
per-model ITPM/OTPM quota when the batch's file-bound/routing model
had no quota configured. Charges each row's own model against its own
project quota instead of only the routing model's, and fixes rate
limit error messages to attribute the correct model via a new
descriptor_value field on RateLimitStatus/AtomicCounterMeta. Also
re-syncs the ruff-strict, type-discipline, and basedpyright budgets
against the correct (non-stale) merge base.
Co-authored-by: Cursor <cursoragent@cursor.com>
Every bedrock batch output line went through the Anthropic usage parser, which
reads snake_case input_tokens/output_tokens. Converse-family models (Nova and
friends) report camelCase inputTokens/outputTokens, so their usage came back
0/0/0 and the batch billed $0 despite real token consumption.
Usage is now selected by the shape of the payload: a Converse-shaped block goes
through the same transform the live Converse path uses, so a batch and an
equivalent non-batch call agree on tokens, including cache reads and writes.
Anthropic-shaped bedrock output is unchanged.
A shape neither parser understands (an InvokeModel-native payload from Titan,
Cohere, or Llama, which name their counts differently again) still reads zero,
but now warns with the keys it saw instead of silently billing $0.
Exposes the Converse usage transform as public, since batch parsing is a second
legitimate caller; that also removes the private-member access invoke_handler
was already making.
Replace Any-typed seams with real types in files carrying the highest
remaining reportAny/reportExplicitAny density after #34745: the proxy
server and its utils, the router, the streaming handler and chunk builder,
litellm_logging, the redis cache, the MCP db/tool-registry/spend-writer
layer, the anthropic pass-through adapters and guardrail translation, the
lasso and presidio guardrail hooks, the azure_ai agents handler, the
management endpoints (keys, users, ui_sso, model access groups, config
override, MCP, projects), the responses MCP handlers, response polling
background streaming, and the containers and vector stores mains
No casts, no type: ignore, no noqa, no new suppressions, and no Any
annotations that were not already at base. Whole-tree basedpyright:
reportAny 14,610 -> 14,009, reportExplicitAny 5,100 -> 4,780, total
144,743 -> 143,471, with no rule increasing repo-wide or in any file.
Budgets ratcheted: basedpyright -1,272 across 48 rules, ruff-strict -85,
type-discipline -37
The log details drawer moved off Ant Design in 03d2b16bc, so its section
header renders lucide ChevronUp/ChevronDown rather than antd's UpOutlined
and DownOutlined. The collapse test still waited on .anticon-up and
.anticon-down, which no longer exist anywhere under view_logs, so it
failed on every run and burned all three attempts identically.
Point the three assertions at .lucide-chevron-up and .lucide-chevron-down,
matching how the dashboard's other suites address lucide icons.
Resolves the transform_create_file_response conflict by keeping the
_uploaded_object_size handoff over the response Content-Length read,
and adds the rebind-ok justification LIT011 now requires for the
upload-size litellm_params handoff after the base budget ratcheted.