Callers can opt into the provider's raw operation response on /v1/ocr with the x-req-format: native header (or req_format in the body) while page-based cost tracking keeps reading usage_info off the normalized response.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): link key info header to its user, creator, team, and organization
The key info page showed the owning user and creator as plain text and never surfaced the team or organization at all, so walking from a key to its parent entities meant copying ids into other pages. The header now renders User and Created By as links to the user detail page, and gains a far-right column with Team and Organization links (alias when known, id otherwise). Client-side navigation logic shared by BadgeLink and IdentityCell moves into a reusable EntityLink so all entity links behave the same
* test(ui): mock next/navigation in VirtualKeysTable test
KeyInfoView now renders EntityLink, which calls useRouter, so the table test that opens the key detail needs the app router mocked
* Add default model pin to complexity router UI
A complexity router's default model was only ever derived from the tiers, so
operators had no way to point the fallback at a model that is not first in the
Simple or Medium tier. Adds a Default Model select that records an explicit pin.
The pin is stored in complexity_router_config.default_model, which the backend
already reads, and mirrored onto complexity_router_default_model on save. Both
paths resolve through one helper that mirrors init_complexity_router_deployment:
a pin wins, otherwise MEDIUM or SIMPLE. Recording the pin in the config keeps it
distinguishable from a derived value, so a pin that happens to match the tiers
survives a round trip instead of being read back as tier tracking.
The edit modal only requires one non-empty tier, so a router with models in
COMPLEX alone can reach save with nothing the backend would pick. That now
blocks with an inline message rather than saving a router that raises at init.
* fix(UI): probe the pinned default model in the auto router connection test
The connection test built its targets from the tiers and the embedding model
only, so a Default Model pin outside every tier was never reached and a green
result could hide an unreachable default. model_info_view had already hand
rolled the dedupe and append locally, so the rule moved into
buildAutoRouterTestTargets and both call sites now share it.
* fix(ui): mirror backend precedence when resolving a complexity router default
The edit modal only recognized a pin stored in complexity_router_config.default_model,
so a router whose default lived solely in litellm_params.complexity_router_default_model
lost it on the next save. That field cannot be trusted outright either: before this PR
every save wrote a tier-derived value into it, so treating any value as a pin would
freeze legacy routers away from their tiers. Hydration now takes the config marker as
authoritative and falls back to litellm_params only when it diverges from what the tiers
alone derive, which is only reachable through an external API or config write.
Test Connection had the mirror-image bug: it fell back to complexity_router_config.default_model,
a UI-only marker init_complexity_router_deployment never reads, so it could probe a model
the router would never call. It now follows router.py exactly: litellm_params, else pure
tier-derivation.
Also reword a tooltip that hardcoded the Default Model select's position on the page, and
document the dual write and the create-vs-edit validation asymmetry.
Shape detection and block normalization sat in the generic batch layer, which
let batch and live parsing of the same wire format drift apart. Both now live on
AmazonConverseConfig as is_converse_usage_shape and usage_from_batch_output, so
batch_utils asks the provider adapter rather than knowing Bedrock's field names.
Adds direct coverage for the shape predicate, the completion of an incomplete
block, cache-count inflation, and the streaming usage event that shares the
public transform. Drops the narrative banner from the batch tests.
- decode upstream first frame as utf-8 instead of ascii
- reject model-restricted keys at connect to match HTTP model enforcement
- log the actual request path for /openai_passthrough traffic
ResponseAPIUsage.parse_cost already flattens Perplexity's
usage.cost.total_cost dict down to a float before it reaches the
perplexity cost calculator, so the isinstance(cost_info, dict) check
was always False on that path. Every Responses-mode Perplexity model
was silently falling back to manual token-rate calculation and
recording $0 spend whenever static per-token rates were missing.
_get_or_start_block trusted the item_id -> block index map without checking
whether that block was still open, so a provider that reuses one item id for
a whole run and interleaves channels got a delta addressed to a stopped
block. Replaying reasoning, text, reasoning, text produced
content_block_delta index=0 after content_block_stop index=0, which is not a
valid Anthropic stream.
Treat the mapping as valid only while it points at the open block, and
rebind the item to a fresh block otherwise. Items registered through
response.output_item.added still reuse their block, since that block is the
open one while its deltas arrive.
Apodex Deep Research is what surfaced this: it labels every reasoning delta
of a run rs_<response_id> and every answer delta msg_<response_id>, so any
interleaving hits the stale mapping.
A Deep Research run streams two agents. The worker emits its chain of
thought on the `reasoning` channel and a draft answer on a channel-less
delta; the reporter emits its own reasoning plus the single `output_text`
delta that matches the final response.completed snapshot.
Only `output_text` was mapped, so 176 of 181 deltas in a sample run
surfaced as GenericEvent and the reasoning was effectively lost. Map the
`reasoning` channel to response.reasoning_summary_text.delta, which
LiteLLM already translates into an Anthropic thinking_delta on the
/v1/messages route Deep Research takes, and give it its own item id.
The channel-less deltas stay unclaimed on purpose: splicing the worker's
draft into the answer would corrupt the text. The remaining
response.swarm.* lifecycle events keep passing through, since
transform_streaming_response has no way to drop a chunk and run_finished
carries the final content.
Embeddings rows were identified by body shape (has `input`, no
`messages`/`prompt`), which also matches a `/v1/responses` batch row
and reserved zero output tokens for it -- letting a project caller run
large Responses generations against a quota-limited model without
consuming OTPM. Classify embeddings by the row's own `url` instead,
and read `max_output_tokens` as a Responses output cap alongside
`max_tokens`/`max_completion_tokens`.
Co-authored-by: Cursor <cursoragent@cursor.com>
Resolves conflicts from the upstream merge and addresses the Veria-AI
review comment on this PR: batch rows could bypass a project's
per-model ITPM/OTPM quota when the batch's file-bound/routing model
had no quota configured. Charges each row's own model against its own
project quota instead of only the routing model's, and fixes rate
limit error messages to attribute the correct model via a new
descriptor_value field on RateLimitStatus/AtomicCounterMeta. Also
re-syncs the ruff-strict, type-discipline, and basedpyright budgets
against the correct (non-stale) merge base.
Co-authored-by: Cursor <cursoragent@cursor.com>
A live stream against apodex-1-1-deep-research shows the event is real and
load-bearing: 181 of the 193 events are response.swarm.llm_delta and
response.output_text.delta never appears, so without the mapping the answer
text only arrives in the final response.completed snapshot.
Restores the transform with the provenance recorded in a docstring, and
covers it with the payload shape captured off the wire, including the
reasoning channel that carries 176 of those deltas and must not be mistaken
for the answer.
Cross-checked the provider against platform.apodex.ai/docs and a live
GET /v1/models call.
- apodex-1.1 and apodex-1.1-mini advertised 256K max output; /v1/models
reports 65536, and max_tokens is the legacy alias of max_output_tokens
- apodex-1-1-deep-discover is Responses-API-only; /v1/chat/completions
answers 400 unsupported_api for the Discover tiers
- core models do not support response_format, so state it explicitly
- transform_cancel_response_api_response carried Content-Encoding over to
a response whose body it had already replaced, so httpx tried to
decompress plain JSON on read. A non-JSON body (the gateway answers a
timed-out cancel with an HTML 504) also escaped as a pydantic
ValidationError instead of the provider error
- drop the undocumented response.swarm.llm_delta mapping
- only the Deep Research tiers default stream to true; the core models
follow OpenAI and default it to false
The JSON provider path applies one contract to a whole provider, which is wrong
for Apodex: its two model families take different parameters. Replaces the
providers.json entry with litellm/llms/apodex/, reverting the shared
openai_like machinery to its original state.
/v1/responses is now keyed off the model. Core models are a stateless subset,
so store is pinned false and previous_response_id / background are rejected
rather than passed upstream to fail with a 400. The deep research tiers keep
all three, so background survives a client disconnect.
/v1/messages resolves per model too. Apodex serves the protocol natively for
the core models only, so the deep research tiers get no native config and fall
back to translation instead of hitting a path that does not serve them.
Chat completions pin stream to false for both families, drop tool params on the
deep research tiers, and rename max_completion_tokens to max_tokens. The
responses config also stops inheriting OpenAI's OPENAI_API_KEY fallback, which
would otherwise forward an unrelated OpenAI key to Apodex.
Tests live under tests/test_litellm/llms/apodex/ and touch no existing test file.
Registers apodex via providers.json with /v1/chat/completions, /v1/responses
and native /v1/messages, plus price map entries for the two core models and
the six deep research tiers.
Apodex defaults `stream` to true on both /v1/chat/completions and /v1/responses,
so a non-streaming litellm call would get SSE back and fail to parse it. Adds a
`send_explicit_stream_false` special-handling flag that pins the field on the
wire, and rewrites the JSON provider param mapping to build its result instead
of mutating the caller's dict.
Every bedrock batch output line went through the Anthropic usage parser, which
reads snake_case input_tokens/output_tokens. Converse-family models (Nova and
friends) report camelCase inputTokens/outputTokens, so their usage came back
0/0/0 and the batch billed $0 despite real token consumption.
Usage is now selected by the shape of the payload: a Converse-shaped block goes
through the same transform the live Converse path uses, so a batch and an
equivalent non-batch call agree on tokens, including cache reads and writes.
Anthropic-shaped bedrock output is unchanged.
A shape neither parser understands (an InvokeModel-native payload from Titan,
Cohere, or Llama, which name their counts differently again) still reads zero,
but now warns with the keys it saw instead of silently billing $0.
Exposes the Converse usage transform as public, since batch parsing is a second
legitimate caller; that also removes the private-member access invoke_handler
was already making.
Replace Any-typed seams with real types in files carrying the highest
remaining reportAny/reportExplicitAny density after #34745: the proxy
server and its utils, the router, the streaming handler and chunk builder,
litellm_logging, the redis cache, the MCP db/tool-registry/spend-writer
layer, the anthropic pass-through adapters and guardrail translation, the
lasso and presidio guardrail hooks, the azure_ai agents handler, the
management endpoints (keys, users, ui_sso, model access groups, config
override, MCP, projects), the responses MCP handlers, response polling
background streaming, and the containers and vector stores mains
No casts, no type: ignore, no noqa, no new suppressions, and no Any
annotations that were not already at base. Whole-tree basedpyright:
reportAny 14,610 -> 14,009, reportExplicitAny 5,100 -> 4,780, total
144,743 -> 143,471, with no rule increasing repo-wide or in any file.
Budgets ratcheted: basedpyright -1,272 across 48 rules, ruff-strict -85,
type-discipline -37
The log details drawer moved off Ant Design in 03d2b16bc, so its section
header renders lucide ChevronUp/ChevronDown rather than antd's UpOutlined
and DownOutlined. The collapse test still waited on .anticon-up and
.anticon-down, which no longer exist anywhere under view_logs, so it
failed on every run and burned all three attempts identically.
Point the three assertions at .lucide-chevron-up and .lucide-chevron-down,
matching how the dashboard's other suites address lucide icons.
Resolves the transform_create_file_response conflict by keeping the
_uploaded_object_size handoff over the response Content-Length read,
and adds the rebind-ok justification LIT011 now requires for the
upload-size litellm_params handoff after the base budget ratcheted.