* test(integration): point the scratch upgraded proxy's read replica at the scratch database
* test(integration): check the scratch upgraded proxy's reader role is connected to the scratch database
* test(integration): assert every configured proxy role holds a scratch connection without a mode branch
* test(integration): skip backend workers without a role in the scratch connection scan
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(router): retry a /v1/messages stream the provider drops before the first content chunk
A /v1/messages stream that the upstream closed before any content reached the
client answered an error event after a single attempt, so the router's
num_retries never applied to that drop. The pre-content failure is now retried
within the model group before the fallback chain runs, with the budget resolved
the way a failure raised before the stream opened resolves it: a retry policy
that names the error class, then the request's num_retries, then the
deployment's, then the router's. A drop after content reached the client keeps
surfacing the provider's error after one attempt.
Fixes#44238
* fix(router): hand a retry's non-retriable error to the fallback chain and type the retry helpers
A retry that failed before its stream opened with an error no retry covers raised straight to the
client, skipping a fallback the first attempt would have used. assert_never now comes from
typing_extensions so the router imports on Python 3.10, and the retry helpers read their kwargs
through typed narrowing instead of Mapping[str, Any]
* fix(router): cast the untyped router fallback defaults the stream retry gate reads
The retry gate passed the router's fallback attributes, declared without element types, to the
typed request override helper, which basedpyright counted as new unknown-argument errors
* fix(router): consult context_window_fallbacks when a retried /v1/messages stream overflows
A retry attempt raising ContextWindowExceededError reached the fallback chain inside its
mid-stream envelope, so only the regular fallbacks list matched. The fallback attempt now
unwraps it the way it unwraps a content policy error. The new router helpers are covered for
the router code coverage check with two direct-call tests and named covering tests
* fix(router): retry a 408 raised by a /v1/messages retry and honor deployment num_retries before the stream opens
* fix(router): attribute a retried /v1/messages stream to the deployment that served it and bound the retry-policy hold
* fix(router): retry /v1/messages error frames under their retry-policy class and keep the first drop's committed budget
An `event: error` frame that arrives before the first content delta now raises the exception class the pre-stream mapping gives an HTTP answer with the same status (429 RateLimitError, 500 and 529 InternalServerError, 503 ServiceUnavailableError, 504 Timeout), so a retry policy's per-class budget governs it the way it governs the error before the stream opened. The status the client sees is unchanged
A retry that lands on a sibling deployment keeps the budget the first drop committed to, read back from the request's attempted_retries and max_retries, instead of recomputing it from the new deployment's num_retries, matching the pre-stream retry loop
* refactor(anthropic): keep the error-frame exception mapping under llms and type the retry test helper
The status-to-exception mapping an `event: error` frame gets before the retry policy is consulted now lives next to the Anthropic error status map in llms/anthropic/common_utils.py, with its own unit test, and the two-deployment retry test helper takes explicit typed parameters instead of a bare dict and untyped kwargs
* refactor(anthropic): map an error frame's status with explicit returns on every path
* fix(router): map stream error frames through the pre-stream exception mapping
An overloaded `event: error` frame on a /v1/messages stream now raises the InternalServerError a 529 answer maps to, built by exception_type from the frame's own body, so one retry policy class governs the error before and after the first byte; a failed fallback after such a frame answers 500 like every other litellm path instead of the frame map's 503
A model_group_retry_policy that does not parse (a non-integer budget, an entry that is not a mapping) no longer fails every healthy stream of that group before its first attempt: the stream runs with no policy and the plain num_retries budget, with a warning naming the group
* fix(router): forward an error frame nothing can take over for as the provider sent it
A pre-content error frame whose class the retry policy grants no retry, with no fallback configured, raised an HTTP error only on the first attempt while the same frame after exhausted retries reached the client verbatim. Both now pass through as sent, the way the merge base forwarded every frame.
* test(integration): audit /v1/messages pre-content retry across routes and budgets
Adds the /audit cells for the pre-content stream retry: the native Anthropic route
(drops and error frames before content, HTTP rejections before the stream opens, SDK
sync and async, after-content and non-retriable controls, budget exhaustion, cache
twin, spend row and headers), the chat and responses bridges, the generic routes
(responses, chat, vllm pass-through, Gemini generateContent, fine-tuning jobs list),
owned two-worker proxies for router-level budgets, retry policies and fallbacks, and
two chaos cells (a worker killed mid burst, an outage on every first attempt). Shared
helpers for scripted Anthropic SSE upstreams and OpenAI-compatible wire replies live
in tests/integration/_support
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(spend-tracking): stop caching failed spend-log metadata lookups as confirmed misses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(spend-tracking): share the short-lived miss cache write
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): cover key alias recovery after spend log lookup failures across usage routes
* test(spend): bound outage alias lookups per miss window instead of a fixed count
* fix(spend-tracking): treat any spend-log lookup failure as a short-lived miss
The Prisma client raises a plain AttributeError when the database drops
the connection mid-query, so the PrismaError catch let it through and
the whole usage call answered 500. Any failure now keeps the 30 second
backoff only, and the integration proxy patches its test entitlement at
import so uvicorn's spawned workers inherit it
* test(integration): audit spend-log metadata recovery under timeouts and dropped connections
Cover the daily activity routes, the usage AI chat, the Vantage and
CloudZero dry runs and exports under a locked spend-log table and under
a database connection dropped mid-lookup, on a two-worker proxy, with
the recovery after the outage asserted through the proxy's own miss TTL.
Add a dropped_connection_relay that closes only the connection whose
bytes carry a trigger, so a cell can drop the one connection the
recovery query runs on while the rest of the pool keeps serving. Rewrite
the sweep and JWT cells for the merged main: the export route reads
metadata by SQL join and never calls the recovery, the search routes
answer key rows and find deleted keys by alias, and the daily-spend
owner recovery names the user while the alias stays blank. The sweep
cell now times out a second lookup under the same lock, which pins the
keys blank on the merge base and recovers on this branch.
* test(integration): match a dropped-connection trigger split across two reads
The dropped-connection relay checked each TCP read on its own, so a SQL
marker that straddled two reads never tripped it and the outage cells
would run without the outage they meant to exercise. Carry the tail of
the previous read into the next check, as the held-statement relay
already does, and pin that with a unit test that splits the trigger
across two writes.
* test(integration): scan relay triggers through an in-process helper
The dropped-connection relay now matches its SQL trigger through a TriggerScanner that carries the previous read's tail, and the unit test exercises that scanner directly instead of opening loopback sockets, which tests/unit forbids. The relay's end to end behavior stays covered by the integration cells
---------
Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(integration): azure-route basic translation cases on messages, chat completions and responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): use the three-line form for the azure chat LIT-9235 skip
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(sso): let CLI and Claude Code gateway sign-in through on DISABLE_ADMIN_UI nodes
* test(proxy): cover CLI SSO sign-in on a UI-disabled node
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(proxy): add models column to the end user table
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): enforce the end user models allowlist in model access checks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): accept and return models on the customer endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover the customer models allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): resolve team aliases before the end user model check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): bill per-second transcription models outside chat modes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): restore request-time billing for per-second transcription models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): bill audio length for per-second transcription models when known
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(openai/realtime): drop model from upstream URL for intent=transcription
* fix(openai/realtime): keep model on the upstream URL for non-OpenAI hosts
A custom api_base on the openai provider can be a gateway that routes on
the model query param, so the transcription model drop now applies only
to OpenAI's own hosts. xAI's URL is unchanged again
* test(integration): audit realtime transcription upstream URL on a custom api_base
Twenty-two integration cells cover transcription and conversation sessions on every realtime route, the OpenAI SDK sync and async clients, malformed and duplicated query params, unauthenticated and refused upgrades, an unknown model, idempotent spend logging, a twenty-session burst, an upstream outage mid burst, a worker kill, and a proxy restart, each asserting the exact query pairs the upstream received. The scripted upstream now records every websocket upgrade as an observation
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(integration): add anthropic-route /v1/messages basic cases for haiku-4-5, sonnet-5, opus-4-8 and sonnet-5-5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): add anthropic-route /v1/chat/completions and /v1/responses basic cases with generated-field matchers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): skip anthropic-route /v1/responses basic cases on LIT-9231 and expect the echoed request params
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): write LiteLLM-generated ids and timestamps as mock.ANY and drop matchers.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: align integration fixtures with current behavior
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for the spend flush before asserting its trace placement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): key response-cache entries by model group from litellm_metadata
/v1/responses routes through the router with model_group in litellm_metadata, which the cache key ignored, so identical requests to different model groups sharing one underlying model hit each other's cached responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: move Postgres, MCP and Redis suites to CircleCI integration
* ci: throwaway, drop tests/proxy_behavior from its CircleCI job to show assert-ci-coverage fails
* ci: revert throwaway assert-ci-coverage check
* ci: keep the e2e helpers the gate tests still use
* ci: move the roi-database Postgres shard to CircleCI integration
* ci: run redis-compat without CircleCI's Azure and cassette env, cover postgres_suite test_path
* ci: match the GitHub env for the moved Postgres and Redis jobs
* ci: unset provider keys in the CircleCI MCP job and drop unused e2e-stack helpers
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(guardrails): run end-of-stream post_call scan when the client disconnects mid-stream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): close the guardrail stream chain in async_data_generator on client disconnect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): leave the raw upstream response to the shielded finalizer on client disconnect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep disconnect cleanup going when a streaming callback cleanup raises
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): assert the refund through a recorder instead of the mock
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): inspect tool calls released before a disconnect under incremental_diff and record a failed scan marker
The incremental_diff transform stream now scans tool calls it already released when the client disconnects, and a disconnect scan whose translation raises after the guardrail recorded success also records guardrail_failed_to_respond
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): pin that a text-only disconnect scan is not handed a tool_calls finish
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover disconnect scans on every streaming endpoint and client, plus outage, worker-kill and cache-hit cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): prove the cache-hit twin is served from the cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): type the disconnect-close streams so basedpyright stops reporting unknown arguments
* fix(guardrails): give the guardrail metadata cast a reason so the type discipline gate accepts it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan released Messages and Responses tool calls on disconnect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(guardrails): format the disconnect scan unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(guardrails): drop mutable-ok markers that no longer suppress a rule
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan released Responses output after a finished item and end only in-flight Chat choices
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): type the disconnect scan test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): type request_data in the disconnect scan helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): type request_data in the disconnect scan test doubles
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): pin that chat streams with no tool call in flight end as released
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan only released chunks on disconnect and skip it once a block owns the verdict
The disconnect scan now uses the chunks actually yielded to the client, copies them before scanning,
skips when a mid-stream block or HTTP error already settled the verdict, and the iterator wrapper
only closes hooks that are async generators so plain async iterator hooks keep working
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): pin that a delivered guardrail error or final chunk settles the disconnect verdict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): close any hook iterator that exposes aclose when the stream ends early
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): accept a synchronous aclose on custom streaming hook iterators
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): swallow callback aclose errors at end of stream
A custom callback whose async_post_call_streaming_iterator_hook returns a
non-generator async iterator with a raising aclose() failed the finished
stream: content plus usage reached the client and then the stream surfaced
an error SSE with no [DONE], or aborted a post_call pipeline's buffering
loop into a 500 with an empty body. Wrap the aclose invocation in
_wrap_streaming_iterator_with_enrichment in try/except and log a warning
naming the callback and the cleanup error, matching close_guarded_stream
and _close_guarded_layers. Iteration-time hook exceptions still propagate.
* fix(proxy): log only the error type when a callback aclose raises
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): LITELLM_FIPS_MODE startup gate with provider assertion, TLS verify guard and loud password migration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): match ssl_verify off detection to runtime str_to_bool semantics
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): drop tautological fips probe test and satisfy CodeQL return checks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject the fake Prisma client through the module boundary instead of patching _setup_prisma_client
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* update logic that marks a logging callback as complete
* test(logging): cover streaming failure dedupe in mark_logging_complete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): cover streaming failure dedupe in S3 and DataDog
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(logging): keep has_run_logging as a deprecated alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): audit streaming failure dedupe across surfaces, fallbacks and sink outage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): assert anthropic upstream path in streaming failure audit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): count only provider posts in streaming audit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): assert the sink outage rejects uploads in burst audit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(logging): configure datadog retries with router override
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ci): keep the security sweep off the observed ROI GitHub route and give Lens integration tests a release identity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ci): expect the credential JWKS export to 404 for non-federation credentials in the security sweep
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(mcp): hand listed-tool metadata to pre-call hooks with per-caller catalog identity
Track the tools each MCP server listed per caller identity so pre_mcp_call and during_mcp_call hooks receive the tool description and input schema the client saw. Servers with no caller-dependent inputs share one slot; user identity, forwarded headers, stdio env, relayed bearers, and server-specific auth get their own. Local registry and OpenAPI paths pass the registered metadata and admin description overrides. The Agent 365 guardrail reads the new fields into its evaluate payload.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): drop the listed-tools empty sentinel and routine test docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): mark the listed-tools cache digest as a non-security hash
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key the listed-tools cache by the OBO subject token
token_exchange servers list upstream with the caller's own Entra bearer, so two callers on one
LiteLLM key with different subjects were sharing a catalog slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): resolve the BYOK credential before keying the listed-tools slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): drop the OAuth discovery cache when a server definition changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): drop a diff-narrating comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): never validate a supplied header on the tools/list BYOK path
The pre-listing resolver ran the tool-call byok_auth_required check even
when the caller already supplied x-mcp-auth, and it ran outside the
per-server error boundary, so a single deprecated-header caller dropped
the server from the aggregate list. Listing now returns a supplied header
unchanged and falls back to the stored credential without raising
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): assert the BYOK listing lands in the caller's listed-tool slot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): cover the deprecated string x-mcp-auth header on a BYOK tools/list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key the per-caller listed-tool slot by the hashed token
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key discovery cache by the hashed token
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates
Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): derive the listed-tool caller identity from the discovery cache key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys
An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.
The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.
Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep OAuth metadata generations only while a fetch is in flight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): return one masked text per scanned string in the selected-guardrail REST test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): fold the signed caller into the discovery digest instead of a second key hash
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): satisfy type discipline gate on listed-tool identity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): hand tools/call hooks the exact catalog entry tools/list served
get_listed_tool re-applied the admin description override on top of the cached listing, so a
guardrail-masked description was restored to its original wording at call time, and the OpenAPI /
local-registry call path built its metadata from the registry instead of the guarded caller catalog.
Both paths now return the cached entry as served, falling back to the registry only when no listing
was recorded
Adds tests/integration/mcp/test_mcp_listed_tool_metadata.py (red on the prior head for the two
regressions, red on the merge base for the feature, green on this head)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): key OpenAPI listed-tool entries per caller so tools/call reads its own guarded listing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): align listed-tool slot tests with per-caller keying
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep oauth2 listing on the minted or signed credential, not the stored BYOK secret
The listing helper that keys the per-caller catalog by the stored BYOK credential also handed that
credential to the upstream client, which on an oauth2 server short-circuited the client_credentials
mint and the MCPJWTSigner gate. Split the two: the catalog identity keeps the stored credential so
tools/call finds the caller's slot, while an oauth2 server's tools/list sends only the per-request
header, letting the M2M mint or signed JWT proceed as on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): oauth2 BYOK listing sends the minted token, not the stored secret, through the real proxy
Integration cell for the listing fix: a client_credentials BYOK server with a stored user credential,
one tools/list as that user, the peer must see a live minted bearer and one /token mint. Red at the
pre-fix tip (zero mints, stored secret upstream), green at the fixed head and at the merge base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep the stored BYOK credential for catalog identity only on tools/list
Listing used the resolved stored credential both to key the caller's catalog slot and as the
upstream transport header, so REST api_key and bearer_token listings sent the user's secret
instead of the server's static token and the MCPJWTSigner gate went quiet. The upstream client
and the signer gate now read the caller-supplied mcp_auth_header for every auth type, exactly
as before the catalog existed, and the stored credential only names the slot tools/call reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): type the listed-tool metadata read from pre-call kwargs for the basedpyright gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): hand never-listed tools/call hooks name and arguments only
The local-registry call path fell back to the registry entry with the admin description override
when no tools/list had been recorded for the caller, so a pre_mcp_call guardrail scanned a
description the caller was never served and blocked OpenAPI calls that passed before, and base's
own selected-guardrail REST test failed on the two-text redaction. _registered_tool_metadata now
returns the listed entry or None, so a tools/call with no prior listing sends name and arguments
only as promised, and that REST test double goes back to its base shape
* fix(mcp): keep during_mcp_call hooks on name and arguments only
call_tool handed the caller's listed entry to the during-hook task as well, so during_mcp_call
guardrails scanned the description line and schema leaves of any listed tool after the upstream
call had already run, blocking calls that passed before whenever the policy matched the
description, returned a fixed-length texts list, or hit the depth guard on a deep schema. The
listed entry is only disclosed for pre_mcp_call, so the during task no longer receives it and its
request object carries no description or schema, as before
* fix(mcp): key the BYOK catalog slot by the client's header, not the stored credential
tools/list resolved the stored BYOK credential to pick the caller's catalog slot, which read the
credential store before the classified try block. With Postgres down and a cold per-worker cache
that made every REST tools/list on an is_byok server fail with tools=[] and no upstream call, and
the read seeded the per-worker cache (including a negative entry), so a tools/call on another
worker after a store, rotate or revoke on this one kept using the stale value.
The slot is now keyed by what the client supplied plus the caller's hashed key, on both sides.
_get_tools_from_server and call_tool take a keyword-only catalog_auth_header that defaults to
mcp_auth_header as received (the default is the builtin Ellipsis so it survives a module reload).
The /mcp fan-out and execute_mcp_tool, which swap the resolved credential into mcp_auth_header,
pass the client's value explicitly. What goes upstream is unchanged. _byok_catalog_auth_header is
gone.
* fix(mcp): drop a listed catalog recorded across a server save
_record_listed_tools ran after the awaited upstream fetch, so a PUT /v1/mcp/server that landed
mid-fetch had its invalidation undone when the fetch completed: hooks then saw the pre-save
description next to the post-save definition until the next listing, instead of name and
arguments only.
The manager now keeps a per-server listed-tools generation, bumped by
_invalidate_server_definition_caches. _get_tools_from_server reads it before the fetch and
_record_listed_tools skips the write when it moved; the next listing records normally.
* fix(mcp): drop the catalog again once a saved OpenAPI server's registry is rebuilt
add_server and update_server publish the saved definition before the OpenAPI registry entries are
rebuilt from the spec, so a listing recorded during that fetch held the pre-save entries under the
new generation. The generation is bumped a second time after the registry refresh.
The during-hook task no longer accepts a listed entry, the one-line wrapper over get_listed_tool is
inlined at its two call sites, and the per-server generation map is a plain dict.
* fix(mcp): keep discovery and OAuth metadata caches across an OpenAPI spec re-read
add_server and update_server ran the full server-definition invalidation a
second time after the awaited OpenAPI spec fetch, which also dropped the
prompts/resources/templates discovery entries and the OAuth protected-resource
metadata filled under the already-published definition, so the next request
went upstream again. Only the listed-tool catalog recorded during the fetch
holds pre-save entries, so the post-fetch pass now drops just that catalog
and bumps its generation via the new _drop_listed_tools helper, which the
full invalidation also calls.
* fix(mcp): look a called tool up in the listed catalog by its bare name only
get_listed_tool stripped the server prefix a second time when the exact
name was absent from the caller's listing, so a never-listed upstream tool
whose bare name starts with the server prefix resolved to the listed
sibling and that sibling's description and input schema reached the
pre-call hooks for a call to a different tool. Every caller already passes
the once-stripped bare name, so the lookup is now exact.
Tests that looked the catalog up by a prefixed name now use the bare name
the callers pass; two new tests pin the never-listed sibling case at the
manager and at the tools/call path.
* fix(mcp): record a listed-tool catalog only for a listing the caller is served
_get_tools_from_server now records the catalog into the caller's
listed-tools slot only when asked (record_listing=True), which the
served listings pass: the /mcp and Responses API tools/list handlers via
_get_tools_from_mcp_servers, MCPServerManager.list_tools, and the REST
listing via _list_server_tools. Four internal listings stop recording,
so a later tools/call hands pre_mcp_call hooks name and arguments only,
as on main:
- _list_tools_before_first_call, the implicit listing inside tools/call
when this worker does not yet expose the tool
- fetch_pinnable_tool_catalog, the admin pin snapshot listed without the
catalog guard and without description overrides
- _initialize_tool_name_to_mcp_server_name_mapping, the startup fill
- get_tools_for_server, used by the semantic tool filter
_create_prefixed_tools returns to its tool-name mapping job only; the
record follows it in _get_tools_from_server.
* fix(mcp): opt every listing out of catalog recording unless it is served
The aggregate listing and _list_mcp_tools now default to record_listing=False,
so a catalog fetched inside a tools/call no longer fills the caller's
listed-tools slot. The /mcp/proxy meta-tools (call_tool, search_tools,
get_tool_schema) and the tool-search virtual tool stop recording: /mcp/proxy
serves only the meta-tools and the search serves only its hits, so a later
pre_mcp_call hook was reading a description the caller never listed.
The tools/list handler, the Responses MCP handler and the /v1/mcp/tools
management listing opt in with record_listing=True, since each serves the
catalog to the caller.
* fix(mcp): key the listed-tool slot by the caller's admission identity and forwarded bearer
The slot a tools/list records for a later tools/call was keyed by (user_id, api_key)
only, so every team-only JWT caller shared one slot and one JWT user acting in two
teams shared a slot; a tools/call then handed pre_mcp_call hooks a description another
caller was served. The slot is now keyed by the hashed key, user, team and organization,
plus the admission credential of a caller admitted with neither a key nor a user.
The caller bearer split the slot only on client-forwarded-token and token-exchange
servers; a legacy delegated oauth2 server (delegate_auth_to_upstream without client
credentials) also forwards it upstream and served a different catalog per bearer into
one slot. The bearer now splits the slot on every server whose egress forwards it
(_consumes_caller_authorization) or exchanges it.
* fix(mcp): record only tools served by the bridge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): keep bridge tool metadata request-local
* refactor(mcp): centralize listed catalog recording guard
* fix(mcp): preserve base TPM reservations for listed tool calls
* fix(mcp): preserve project token reservations for listed calls
* fix(mcp): record served catalogs and preserve call message bytes
* test(mcp): align listing expectation with deferred recording
* test(mcp): audit listed metadata across callers and bridge lifecycles
* test(mcp): preserve guardrail fixture worker affinity
* refactor(mcp): expose listed catalog recording API
* feat(mcp): pass served_tools through the anthropic messages bridge
/v1/messages auto-execution now hands this request's resolved tool
definitions to _execute_tool_calls, matching the Responses and chat
completions bridges: the pre_mcp_call hook receives the description
and input schema the model was shown for that call. Request-local
only; the shared listed-tools catalog is untouched.
* test(mcp): pin served_tools handoff on the anthropic messages bridge
Mirrors the credentials-forwarding test: the request's resolved tool
definitions must reach _execute_tool_calls under served_tools so
pre_mcp_call hooks judge the call on the description and input schema
the model was shown. Fails without the previous commit's one-liner.
* style(mcp): sort the local import block ruff flagged
* fix(mcp): keep the admin include_disabled_tools view off the listed-tools catalog
GET /mcp-rest/tools/list?include_disabled_tools=true is the admin-only
configuration view: apply_tool_filters is False, so it serves the full
server catalog. Recording that response into the caller's listed-tools
slot warmed tools/call metadata no runtime listing ever served,
breaking the only-a-served-listing-records invariant (Bugbot).
The record is now gated on apply_tool_filters; disabled tools stay
unreachable (the call-time allowlist 403 fires before hooks), so the
observable fix is the slot no longer warming from a settings view.
Verified live: the new test fails on the unfixed head and passes here,
and the rest of the listed-tool-metadata suite is unchanged.
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): log v1 migration failures at ERROR so LITELLM_LOG=ERROR shows them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): reuse litellm secret redaction and mask configured DB passwords exactly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy-extras): name the password alternation in _redact_credentials
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(proxy-extras): wrap v1 migration ERROR lines at 120 columns
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy-extras): integration cells for v1 migration ERROR logging and password redaction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy-extras): integration cells for component DB env vars, JSON logs, migration Job and v2 resolver
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy-extras): give every v1 migration integration cell the 900s timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy-extras): bind the recovery forwarder before migrating so the retry cannot race it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy-extras): drop the slow P3005 integration cell and bound migration subprocesses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): keep command repr and tolerate non-sequence cmd in migration error logs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy-extras): double the subprocess boundary in the cmd=None retry test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* test(ci): fix four CircleCI regressions on main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): stop reloading auth_checks in unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(constants): cover CLI JWT expiry env parsing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(gateway): restore proxy lifespan after importing gateway.main in launch tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): count tool_call cache_control marks in the injection census
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): remove cache census casts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): only skip injection on message or content marks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip injection on messages whose tool calls carry marks
Reverts 9f08d8aef8. A default 5m mark injected on assistant text lands
before the client's 1h tool_use mark, which Anthropic rejects with a 400
because a 1h breakpoint must not follow a 5m one. Keeping the full census
in the skip check leaves the client's tool_call breakpoint as the only
one on that message.
* fix(caching): count every client tool_call cache mark in the breakpoint census
The census gated tool_call marks on type function and dict shape, so a client mark on a call without a type or with a string cache_control slipped past the count and injection overflowed the 4 breakpoint cap. Count any non-None tool_call mark, keep the server tool exclusion, and add integration cells for the capped surfaces, the yaml stand-down, Bedrock and Gemini, and router affinity
* test(integration): hold the upstream so the worker kill lands mid-burst
* refactor(caching): reuse the transform's server tool lookup in the breakpoint census
The census now calls the same helper the Anthropic transform uses to decide
whether a marked tool call becomes a server tool block, so the two cannot
drift apart. The owned-proxy burst test waits up to 90 seconds for the burst
to reach the wire before it kills a worker
* refactor(anthropic): move the server tool rebuild check under llms/anthropic
* test(integration): audit the tool call mark census across chat, messages, responses, and chaos
Adds the /audit cells for the breakpoint census on assistant tool_calls marks: the
Responses stream bridge, the OpenAI and Anthropic SDK clients, in-process Pydantic
messages, response cache twins, request-level points, a provider 401 on a capped request,
malformed provider_specific_fields and tool_call ids, null or empty points, a points
update mid-burst, a proxy restart mid-burst, and a worker SIGKILL that picks the worker
holding the burst's upstream connections
* test(integration): close the SDK clients the cache census cells open
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): compare every non-transport provider header and check for late provider requests at session end
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): name TranslationTestCase fields after litellm and provider sides and drop regressions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): prefix checked TranslationTestCase fields with expected_ and name the fake reply mock_provider_response
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(integration): name TranslationTestCase fields in the translation README
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): add a claude-opus-5-5 base case and deployment next to claude-sonnet-4-6
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): name translation cases <MODEL>_TEST_CASE and document the naming rule
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(integration): move translation test rules into tests/integration/translation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tests): match the OS bind error in the owned-proxy port-race retry
* test(integration): keep the port-race predicate pure so its unit tests stay in-process
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(vertex_ai): forward system and tools to partner model count_tokens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vertex_ai): avoid mutable token request construction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vertex_ai): return partner count_tokens provider errors as values so the proxy falls back locally
* test(integration): cover Vertex AI partner count_tokens forwarding and local fallbacks
Add wire-level cells for /v1/messages/count_tokens, /utils/token_counter,
/v1/responses/input_tokens and the Gemini countTokens route on a Vertex AI
Claude deployment: the system prompt and tools reach the partner
count-tokens endpoint verbatim, null fields stay out of the body, malformed
tools are rejected before any peer call, peer, token-endpoint and connection
failures fall back to the local tokenizer unless disable_token_counter is
set, generation on the same deployment keeps working, and concurrent bursts
survive a peer outage, a slow peer and a worker SIGKILL. The sdk cells cover
litellm.acount_tokens the same way.
The _support/process.py and _support/client.py harness files are brought to
main's content so the self-booting cells read INTEGRATION_PROXY_READY_SECONDS
instead of a fixed 70 s boot budget.
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): key the zero-cost cache by the resolved model group
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): keep the zero-cost verdict per requested name and include hidden aliases
* test(integration): audit the zero-cost bypass through hidden model_group_alias names
* test(router): cover the extracted routing strategy switch
* test(integration): record a pre-flip burst before the alias flip
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(ui): prototype observed engineering ROI dashboard
* feat(roi): replace effort estimates with measured repository metrics
* fix(roi): finish connection recovery and generated API contracts
* fix(roi): show merged changes before accounts are linked
* fix(roi): preserve selected report tab across refreshes
* fix(roi): recover app authorization and keep detail values readable
* fix(roi): reuse the shared OAuth HTTP client
* feat(roi): combine providers and compare equal reporting periods
* docs: explain ROI metrics for first-time readers
* fix(roi): preserve connections and scheduled reports during setup
* ci(roi): assign database contracts to the active Postgres shard
* fix(roi): preserve issue counts and normalized connections
* fix(ui): compact ROI dashboard header and metrics
* fix(ui): show ROI repository count with expandable list
* fix(ui): wrap ROI controls within narrow panels
* fix(roi): restore sample report preview and simplify setup
* fix(bedrock_mantle): route Claude chat completions to the native Messages endpoint
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock_mantle): price Claude chat on the Mantle row and route region-prefixed ids
A Mantle request always carries a region, so a Claude id with no bedrock_mantle/<region>/ row fell through the model-info lookup to the bare Bedrock row, which the bedrock provider family also matches, and billed about 10 percent under the Mantle price. The lookup now tries the provider's region-free row before the bare model. The Claude route test asserts the Mantle row, a region-prefixed Claude id is covered end to end, and the provider config map references the Mantle config directly.
* test(bedrock_mantle): cover supported params for Claude and open-weight Mantle ids
* test(bedrock_mantle): audit the Claude chat bridge on the integration rig
Adds the deterministic cells from the /audit of the Mantle Claude chat
bridge: wire-level translation on every chat, responses, and messages
route, SigV4 and bearer auth, region prefixes, api_base suffixes,
unsupported params with and without drop_params, malformed model ids,
upstream errors, the response-cache hit, the health check, pricing from
the Mantle row for Claude and non-Claude ids with a bare Bedrock twin,
and chaos cells for a mixed burst, an upstream outage, slow streams,
and a worker kill on an owned two-worker proxy
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models
GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.
* fix(proxy): offer a Codex service tier only when every deployment of the model lists it
* fix(codex-catalog): an invalid service_tiers value offers no tier for the model
* fix(codex-catalog): read service tiers off the deployments the key's team can route to
A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them
The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default
* test(codex-catalog): drop the redundant module docstring and sort the imports
* test(integration): add the Codex catalog audit cells and the multi-worker convergence note
* test(integration): clean up every catalog test model and answer the refresh GET
* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers
Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.
* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut
The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns
The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(bedrock): propagate timeout to streaming requests
* test(bedrock): prove streaming fails at the request timeout against a slow upstream
* test(bedrock): simulate the slow upstream in process instead of over a local socket
* test(bedrock): audit the Converse and Invoke stream timeout on the proxy
Two integration files drive the per-request timeout on Bedrock streams
through the real proxy against an owned wire peer: the wire file covers
every surface (chat, messages, responses, invoke, pass-through), the
sad, edge and precedence rows, and the chaos file covers bursts, a
dropping upstream, a killed worker and a proxy stopped mid-burst.
The wire peer gains Reply.drop_connection so a cell can close the
socket before any response, and the harness's graceful stop grace is
now INTEGRATION_PROXY_STOP_SECONDS (default unchanged at 30), since a
two-worker supervisor's interpreter finalization takes longer than that
on a loaded box.
* test(bedrock): pin the fallback audit cell to one proxy worker
The fallback cell created both deployments through /model/new on one
worker and sent the chat request to the other, whose registry
read-through loads only the requested model, so the fallback target
was unknown there until the periodic DB poll. The cell now warms the
fallback model and sends the request over one keep-alive client, so
one TCP connection stays with one uvicorn worker, and it expects the
fallback upstream to see both requests.
---------
Co-authored-by: Sainyam Kapoor <hello@sainyam.me>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3
Bedrock prices Kimi K3 cache reads through implicit caching but Converse
rejects the explicit cachePoint marker ("This model doesn't support the
cachePoint field"), so any cache_control on the request answered 400.
Mark the three K3 rows supports_prompt_cache_breakpoint: false and have
bedrock_model_accepts_cache_points honor that flag before falling back to
supports_prompt_caching, keeping cached-token pricing intact.
* test(bedrock): assert cache points per request section
* fix(bedrock): honor a deployment's cache breakpoint flag for unmapped models
* fix(bedrock): read a converse-routed deployment's cache breakpoint flag
* refactor(bedrock): look up cache breakpoint flags by key
* test(bedrock): add the Kimi K3 cache point wire audit
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register typesafe as a provider so Jev deployments load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): move provider endpoints under llms and validate proxy bodies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add Cloudflare Clef and Strands Decider backends
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register decisions routes for managed agents and gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(decisions): use raw regex for cloudflare missing account match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): avoid cast in Cloudflare response unwrapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): default model, evaluation health probe, short Cloudflare names
The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.
* fix(decisions): let health_check_params override the evaluation probe
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit the decisions endpoint across providers, limits, health and chaos
Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).
The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.
* fix(decisions): send env API keys to a configured api_base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register the routes through the lazy feature registry
The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.
The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.
* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell
* fix(proxy): let a config pass-through beat a lazily registered route in eager mode
With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295)
This reverts commit ca1994e403.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): merge bridged tool calls into the same choice as the text
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): streamed alias matching a capability rule bills the deployment price
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): assert every streamed chunk carries the client alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(logging): log the client alias on the priced streamed response, the same as non-streamed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(roi): support GitLab and tagged branch costs
* fix(roi): count tagged branches independently of estimation status
* test(roi): capture live GitHub and GitLab report validation
* fix(roi): open estimate details at the start
* fix(roi): clarify cost views and unify report layout
* feat(roi): showcase per-PR costs in the sample report
* fix(roi): separate report tabs and preserve branch cost attribution
* fix(roi): preserve demo previews and align progress spacing
* fix(roi): isolate demo loading and parallelize fork lookups
Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression
* fix(roi): separate demo and live loading states
Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers
* fix(roi): ignore refreshes from a previous source
* fix: trust gateway context for ROI estimator exclusion
* fix: preserve historical ROI estimator exclusion
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run search_endpoints tests in proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(llm_http_handler): keep provider error text when re-raising mapped errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): allow promptless image edits and default search models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): default missing image edit image to None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): build image edit defaults without mutating request data
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject a fake router for the search default model test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on vector store and file lookup handlers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): keep the lookup handler raise block to a single statement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover missing required body params and provider lookup status codes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bind spend-row request id with partial to satisfy B023
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): only reject non-positive page_size on vector store list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): remove unreachable fine-tuning body validation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover streaming anthropic messages reaching the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): count only provider calls when asserting missing params never reach the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve merge-base request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve interaction completion model defaults
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* fix(vector_stores): return managed file ids from vector store file list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vector_stores): cover managed file list route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): only map round-trippable managed ids and index flat file ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): build managed file gin index concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy-extras): move the managed file gin index migration after main's newest
* fix(vector_stores): satisfy lint gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy): drop the stale no-index note on the raw file-id guard
* test(vector_stores): cover managed file ids on the vector store file list end to end
Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(interactions): durable cross-pod settlement for background interaction billing
Background interaction billing lived only in the creating replica's memory, so a
DELETE routed to another replica, or a restart of the creating one, never billed
the completed provider work and the budget reservation was refunded at the poll
timeout. The create now registers the billing context in a settlement store
before returning, the proxy installs a Prisma-backed store at boot
(LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica
claims the row once through a conditional update before billing or releasing,
startup resumes every unclaimed row with its remaining timeout, and a give-up
records an unsettled outcome instead of silently reconciling to zero. The SDK
keeps an in-memory store and behaves as before.
* fix(interactions): survive a settlement install failure at boot and stop carrying request headers
* fix(interactions): drop the stored request context once a settlement row is settled
* fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline
* fix(interactions): carry a missing model through the settlement context for agent-only background creates
An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it.
* fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere
The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials.
* fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill
* fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed
A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler.
A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere.
* fix(interactions): settle an unverified registration through the durable claim
A create whose settlement-store write raised no longer bills through a
private in-memory gate that a later boot's resume cannot see. The claim
asks the durable store first and falls back to the local gate only when
the store answers that no row exists, and a missing settlement table reads
as no rows so a replica without the migration still settles in process.
* test(proxy): keep the settlement test where the proxy-infra shard collects it
The merge of main moved test_background_interaction_settlement.py under
tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy
files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it
and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected
whole by the proxy-infra shard, which is where the test ran before the merge.
* fix(interactions): raise on a non-2xx Gemini interaction fetch
AsyncHTTPHandler.get never raises for status and the Gemini GET transform
only raised when the body was not JSON, so a 500 or 404 carrying Gemini's
JSON error body parsed as an interaction with no status. A delete on a
replica other than the creator then claimed the settlement as released and
forwarded the delete instead of failing closed, and the bill was lost. The
transform now raises GeminiError with the vendor's status, as the delete
transform already does; the in-process poll already retries a fetch that
raises
* test(integration): audit durable background interaction settlement across replicas
Twenty-six deterministic cells drive a one-worker creator and a two-worker
settler against an owned scripted Gemini upstream: cross-replica deletes
bill once, failed and cancelled interactions release, a later replica
resumes unclaimed rows, custom deployment pricing bills at the deployment
rate, a fetch the settler cannot make fails the delete closed, odd ids are
refused, a missing settlement table keeps in-process billing, polling
disabled registers nothing, the budget reservation is released by the
settler, an upstream outage mid-burst fails closed and recovers, killed
workers hand their polls to the respawned ones, and concurrent deletes on a
slow upstream settle exactly once. The support upstream gains a scripted
interaction store with per-id GET status and delay, and the process helper
gains an owned upstream a test can stop and restart
* test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate
* chore(ui): regenerate dashboard API types after merging main
* test(integration): accept the 422 budget refusal and a respawned worker's resume
The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged
* test(integration): delete the pinned key's interaction with a second key
A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(cost): price streamed aliases that only match a capability rule from the deployment model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): satisfy basedpyright delta after the cost-candidate sort
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(cost): skip the capability-rule check for exact cost-map keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): price streamed alias rows from the deployment's own rates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): isolate streamed alias deployments with per-run model names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(streaming): keep the unpriceable-stamp case on a truly unmapped model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): treat capability-rule matches as unmapped in every cost lookup
Route every model-info lookup on cost paths through get_priced_model_info
and _cached_get_priced_model_info_helper, which raise ModelNotMappedError
when the only match is a pricing-free fallback-generalizations rule. An
alias that matches a capability rule now falls through to the deployment's
real model instead of billing 0, and the earlier candidate-sorting fix is
reverted since the priced helper is the single choke point.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): assert the rule alias bills the same as the plain alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): treat router-registered rule-only model_cost entries as unmapped
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): count string rates and check pricing before the capability-rule match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(cost): repoint cost-path patches at get_priced_model_info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(cost): give the model_cost row cast a cast-ok reason
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): type the lazy get_priced_model_info export
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): narrow the fix to ordering rule-only cost candidates last
Drops get_priced_model_info and its call-site swaps, the lazy-import entry,
and the cost-path code check. Only the candidate sort in completion_cost and
pricing_entry_for_cost_calc stays, with its tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): cover rule-only base_model billing on every endpoint, client and failure mode
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>