* test(pricing): lock supports_prompt_caching on cache-priced grok rows
Check the catalog field itself so a price sync cannot drop the Vertex
Grok flag while the helper still passes via the xai/ fallback. Also
cover get_model_info on the full vertex_ai/xai/grok-4.6 key and keep
the backup map in lockstep.
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* test(pricing): assert grok cache rows via get_model_info
The helper can still pass via the bare xai/ grok row, so lock the
catalog flag through get_model_info instead of reading the JSON files
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* test(pricing): assert grok cache flags via get_model_info
Address Greptile P2 by checking supports_prompt_caching through
get_model_info on full catalog keys instead of raw JSON fields, so
the check cannot pass via the bare xai/grok-* helper path. Keep
backup/primary parity for vertex_ai/xai/grok-* entries.
* test(pricing): assert grok cache flags via runtime APIs only
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
get_user_object wrapped every failed read, a refused connection included, in
ValueError("User doesn't exist in db ..."), so JWT callers got a 401 naming a
missing user while Postgres was down and virtual-key callers got 503
no_db_connection for the same outage. A connection or transport error now
propagates as-is and the auth exception mapper answers 503 no_db_connection;
a genuinely missing row and query-level errors still answer 401.
The MCP auth and token-exchange docstrings and the exception-chain helper's
docstring described the old wrap and are updated to the new contract.
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* docs(rust): plan Python interop foundation
* fix(rust): preserve Python settings coercion at the native boundary
* chore(rust): drop interop planning note
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): resolve OCR provider secrets through an async SecretSource before transformation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): project the Python secret manager into the bridge and resolve OCR secrets through it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): drop premium_user from the secret manager snapshot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust-bridge): read the private key management globals once in the settings snapshot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): bound the bridge secret manager state cache to the active snapshot
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): inline coercion unit tests
* fix(rust): preserve Python secret manager bindings
* refactor(rust-bridge): let settings projectors own their contract specs
Each settings group now declares its SettingSpec rows next to the projector
that reads them, and the manifest test derives python_settings.json from those
tables instead of a hand-copied duplicate. Field carries (group, name) instead
of a dotted path, and coercion gains the dict-item reader plus the Redis
Boolean, certificate-requirement, non-empty string, and numeric adapters that
the cache configuration projection adopts next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(rust-bridge): capture the secret manager binding in one settings read
The secret_manager accessor now carries the live client and settings objects,
so the bridge classifies the binding from a single snapshot instead of
re-reading litellm globals. The unreachable native arm and the service alias
go away, the binding-to-state mapping moves next to the snapshot, and the
Python callback precomputes its key_manager name.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(rust-bridge): execute typed settings field declarations
* refactor(rust-bridge): compare cache backends by identity behind one exact trait
cache-response gains an object-safe ExactResponseCache so every exact-match
backend sits behind one pointer; WriteBuffer flushes through it. The bridge's
NativeResponseCache shrinks from nine variants and fifteen per-backend
accessors to an exact service plus the three semantic backends, and facade
mismatch detection compares BackendIdentity values instead of matching on
each backend type. Request projections move next to NativeRequest.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(rust-bridge): drive both Python-embedded semantic caches through one execution
Redis-semantic and Valkey-semantic operations now share one SemanticExecution
body: await the Python embedder, seed the task-local vector, run the native
backend, repeat per batch entry. Valkey drops its with_embedder path in favor
of the same seeded embedder, and each backend keeps its own embedding-failure
policy. PythonEmbedder exposes one call shape. Redis-semantic thresholds are
compared at the backend's f32 width, which un-breaks the redis-stack parity
tests that a 0.8 facade threshold failed before this branch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* wip
* feat(rust-bridge): complete response cache runtime surface
* fix(rust-bridge): preserve secret manager callback exceptions
* refactor(rust-bridge): unify route cache and secret rollout catalog
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
The gpt-5.6-responses_previous_response_id case sent a literal id the proxy never issued, which the Responses id security hook refuses with a 403 at production defaults. The case now primes a response through the proxy and chains the id it hands back, so the harness drops allow_unmanaged_response_ids and the security hook stays exercised
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(e2e-ui): check the MCP Tools tab against the upstream's own tools/list
DeepWiki renamed ask_question to ask_wiki_question, and the spec hardcoded the old name, so
e2e_ui_testing went red on main for something that is not a litellm regression. The spec now asks
the upstream server for its tool list with the official MCP TypeScript SDK and expects the tab to
show exactly those cards, so a vendor rename cannot turn the job red again.
* test(e2e-ui): cite the pinned DeepWiki tool name and drop the helper docstring
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(responses): drop client_metadata before bridging to chat completions
Codex CLI sends client_metadata on every /v1/responses call. For a
provider with no native Responses config the chat-completions bridge
forwarded the raw kwargs, so client_metadata reached the provider as a
chat body field and Databricks rejected the request with an unknown
field 400. The bridge now drops the Responses-only request fields
before calling completion while still passing every other kwarg
through, so deployment-level params such as chat_template_kwargs keep
reaching providers without a native config.
* fix(databricks): merge consecutive system messages for chat-template models
Codex sends instructions plus a leading developer item, which the Responses
bridge and the developer-to-system translation turn into two consecutive
system messages that Databricks chat-template models reject with "System
message must be at the beginning". Each run of consecutive system messages
is now merged into one before the request is built for non-Claude models.
Also keep client_metadata out of the bridged chat request even when
allowed_openai_params names it, so both bridge branches drop the same set.
* fix(databricks): skip empty system messages when merging consecutive ones
Databricks drops empty content before the merge, so a system message in a
run could carry no content key and the merge iterated None. Those messages
are now skipped; a run with no content at all keeps its first message.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* refactor(agentic-loop): build follow-up kwargs in one place so no executor can repeat a request param
The Responses and both chat completions follow-up executors each rebuilt the follow-up kwargs by hand and then expanded them next to the request params, so a plan whose kwargs repeated a request param raised a duplicate keyword TypeError. They now share build_agentic_followup_kwargs, which drops any key already sent as a request param (and the explicitly passed model/input/messages) from both the request kwargs and the plan kwargs. Each executor keeps its own internal-key filter unchanged, and the /v1/messages executor is untouched because it merges into a single dict and cannot hit this.
* test(agentic-loop): move follow-up regressions into their mapped test files
Greptile review: the executor regressions belong in test_llm_http_handler.py and test_chat_completion_agentic_loop.py rather than a split-off file, and the builder test helper returned a read-only mapping while promising a dict. The Responses overlap test is dropped because #41560 already added the same one to the mapped file.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): release unclaimed budget reservations at request end
* fix(proxy): release unclaimed budget reservations of websocket sessions too
* test(proxy): drop the structural middleware inheritance check
* fix(proxy): claim the budget reservation on streaming pass-through before its cost callback
The SSE chunk processor hands its success handler to the logging worker
after the response, so the request-end release freed the reservation
first and left the key unguarded until the worker drained. Claim it at
both end-of-stream hand-offs, the immediate enqueue and the coroutine
parked for deferred dispatch.
Give the xai realtime test double the litellm_params attribute every
real Logging object carries, since the wrapper now reads it.
* test(pass-through): give the vertex streaming test doubles a litellm_params dict
The spec'd Logging mocks in test_vertex_ai_anthropic_streaming_cost_injection.py
lacked the instance attribute the chunk processor now reads to claim the budget
reservation. Also restores main's _lazy_openapi_snapshot.json: the branch's copy
had been regenerated under Python 3.14, which dedents one docstring description
that the CI regeneration on Python 3.12 keeps indented, and the PR adds no lazily
loaded route, so main's file is the correct one.
* fix(pass-through): claim the budget reservation only after its cost callback is enqueued
Every pass-through success hand-off stamped callback_bound before handing the
coroutine to the logging worker. When that enqueue raised, the reservation stayed
claimed with no callback left to reconcile it, so the request-end release skipped it
and the reserved cost stayed pinned on the key's counter. Enqueue first, then claim,
so a failed hand-off leaves the reservation for the request-end release.
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(cost): warn and count $0 cost on billable requests
A request that carries usage but prices to $0 on a model whose pricing
entry has a non-zero rate now logs one warning naming the model, the
pricing entry, and the missing rate, and increments
litellm_zero_cost_requests_total{requested_model, model, model_id,
api_provider, reason}. Free models (every used rate is 0), requests
without usage, and unmapped models stay silent. The diagnostic rides on
the standard logging payload as zero_cost_diagnostic
* fix(cost): keep the zero-cost diagnostic importable on 3.10 and recursion-free
* fix(cost): warn once per request when a $0 result is priced again
* fix(cost): judge a free deployment by its own pricing and keep it silent on calculator errors
* fix(cost): warn once per request when a usage-less evaluation sits between two zero-cost findings
* fix(cost): judge zero-cost findings by the priced entry, skip cache hits, count failure rows
* test(cost): type the zero-cost diagnostic test helpers
* test(logging): flag a $0 terminal Responses stream event by its inner response
* chore: restore the lazy OpenAPI snapshot as CI's Python 3.12 generates it
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): drop cost-map metadata echoed back on model save
Filter unchanged cost-map fields from model-info save echoes while preserving edited overrides.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): drop a stored override when an echoed save resets it to the cost-map value
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): compare model_info echo against the deployment's cost-map lookup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): decrypt the stored model before the cost-map lookup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): treat a reset to the bundled catalog value as an echo even after router registration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): compare the reset against the catalog as loaded, not only the bundled backup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(types): type the catalog snapshot and echo filter parameters as Mapping[str, object]
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject the loaded catalog into update_db_model instead of patching the class
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): use contextlib.suppress for cost-map lookup miss to stay under BLE001 budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): price non-canonical image sizes from the nearest row and honour dump options
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fal_ai): drop monkeypatched mixed pricing case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(fal_ai): use the default dimensions constant directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): forward nested include and exclude when dumping image data
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): honour pydantic item selectors in image data serializer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): match negative item selectors in image data serializer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): opt-in litellm_call_id in JSON error bodies
Add general_settings.include_call_id_in_error_body. When true, the value
already on the x-litellm-call-id response header is copied into JSON error
bodies: as error.litellm_call_id on the OpenAI-shaped routes, /v1/messages,
and streaming first-chunk errors, and as a top-level litellm_call_id on
pass-through routes. Off by default, so error bodies stay byte-identical
unless an admin opts in
* chore(proxy): drop helper docstring and restore lazy OpenAPI snapshot
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(fal_ai): add flux-lora-depth image edits and moondream3 chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(fal_ai): retrigger codecov processing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): reject multi-turn and system messages for moondream3 chat
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): return 400 for invalid moondream3 chat requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): reject moondream3 responses missing output or usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fal_ai): reject streaming moondream3 requests before dispatch
stream never reaches optional_params, so the transform_request check could not fire; reject in _complete_fal_ai on ctx.stream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): sign batch S3 requests with s3_access_key_id and s3_secret_access_key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): keep S3 signer test additions scoped to new cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): drop e2e suite changes from the S3 signing fix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): build S3 credentials directly from the s3_* pair so ambient AWS_* env never mixes in
Restores the split-identity e2e coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
- fix(bedrock): send every Mantle beta in the anthropic-beta header on the bedrock/mantle route
- refactor(bedrock): type the Mantle header helper and build the header fields in one comprehension
* fix(guardrails): scan video prompts for key-attached guardrails on /v1/videos
/v1/videos dispatches call_type avideo_generation, which CallTypes did not
know and no guardrail translation handler covered, so the unified guardrail
hook returned the request unscanned. Add the video call types and an OpenAI
video guardrail translation package that scans the prompt for create, remix,
edit and extension requests
Resolves LIT-6685
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate api types for video call types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: skip avideo_generation in azure sdk client exhaustive check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): retry a leaked video job until the guardrail sync deadline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): satisfy the type-discipline gate in the video handler
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): gate the video e2e on a chat probe so a miss starts at most one paid job
Addresses Greptile review: typed RewritingGuardrail override, dropped routine docstrings, and the e2e waits for the key guardrail to sync via /chat/completions before its single /v1/videos call
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>