* feat(lens): add scripts/lens_dev.sh for one-command lens local dev
* feat: add make lens-dev target
* chore: gitignore .lens-dev local state
* docs(lens): mention make lens-dev in the developer note
* fix(lens): resolve a relative LENS_DEV_CONFIG against the caller's cwd
* fix(lens): random lens-dev master key, verify reused services, drop inherited REDIS_*
* test(lens): cover lens-dev worker token, master key, env and cleanup
* fix(lens): skip compose postgres when LENS_DEV_DATABASE_URL is set
* test(lens): external LENS_DEV_DATABASE_URL never starts compose postgres
* fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3
Bedrock prices Kimi K3 cache reads through implicit caching but Converse
rejects the explicit cachePoint marker ("This model doesn't support the
cachePoint field"), so any cache_control on the request answered 400.
Mark the three K3 rows supports_prompt_cache_breakpoint: false and have
bedrock_model_accepts_cache_points honor that flag before falling back to
supports_prompt_caching, keeping cached-token pricing intact.
* test(bedrock): assert cache points per request section
* fix(bedrock): honor a deployment's cache breakpoint flag for unmapped models
* fix(bedrock): read a converse-routed deployment's cache breakpoint flag
* refactor(bedrock): look up cache breakpoint flags by key
* test(bedrock): add the Kimi K3 cache point wire audit
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(lens): add dot field layout and live status helpers for the traces timeline
* test(lens): cover dot field layout, agent colors and live status
* feat(lens): draw the traces timeline as a live dot field
* feat(lens): add the sweep animation for the live traces timeline
* feat(lens): put the lens tabs in a compact header and fill the screen with traces
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register typesafe as a provider so Jev deployments load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): move provider endpoints under llms and validate proxy bodies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add Cloudflare Clef and Strands Decider backends
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register decisions routes for managed agents and gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(decisions): use raw regex for cloudflare missing account match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): avoid cast in Cloudflare response unwrapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): default model, evaluation health probe, short Cloudflare names
The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.
* fix(decisions): let health_check_params override the evaluation probe
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit the decisions endpoint across providers, limits, health and chaos
Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).
The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.
* fix(decisions): send env API keys to a configured api_base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register the routes through the lazy feature registry
The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.
The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.
* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell
* fix(proxy): let a config pass-through beat a lazily registered route in eager mode
With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them
Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.
A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.
A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.
Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* wip
* wip
* wip
* chore(trace): checkpoint ongoing Rust migration
* refactor(trace): group Python bridge under trace package
* refactor(traces): read span conventions through a Convention trait
Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.
The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.
* feat(trace): export Rust-owned wire schemas and enforce contract bounds
* fix(trace): bound quoted counts in ClickHouse wire schemas
* feat(trace): generate Python wire contracts with datamodel-code-generator
* test(trace): validate migrated callers and generated contracts at the native boundary
* refactor(traces): rename normalization convention to format
* fix(traces): reconcile spend evidence and preserve unknown costs
* feat(traces): normalize additional telemetry formats
* test(traces): cover captured normalization fixtures
* refactor(traces): isolate SDK normalization rules
* feat(tracing): seed all trace exports for local dashboard
* fix(clickhouse): preserve custom LiteLLM request metadata
* docs(traces): define normalization module boundaries
* docs(traces): define resolution and OTLP boundaries
* fix(ui): normalize nullable trace message names
* refactor(traces): split resolver modules and cover resolution behavior
* test(traces): replace normalization snapshots with behavior assertions
* fix(ui): align dashboard API contracts with generated types
* refactor(traces): type normalization and storage boundaries
* fix(traces): seed captured SDK spend and preserve provider identities
* wip
* test(traces): verify guide discovery and content ordering
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The interactive Lens demo used to build a fetch shim that encoded in-memory
fixtures as HTTP responses so the shared ApiClient could decode them again,
and every trace view branched on demo vs live to pick a URL. Lens and the
trace views now depend on two small service interfaces, LensApi and
TracesApi, with named operations. The live layer wraps the existing HTTP
calls, the demo layer reads fixtures directly, and a React context provides
whichever one the session runs on. Without a provider the hooks fall back to
the live implementation, so the live app and existing tests are unchanged.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): reorganize Lens dashboard components
* refactor(ui): align Lens forms with react-hook-form
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): keep Lens submit errors out of form validity
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): reduce Lens lint budget usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): use a query key factory for Lens queries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): simplify trace inspection in the gateway drawer
* feat(lens): add a conversation view for full traces
* test(lens): keep normalized message fixtures type safe
* refactor(lens): make conversation view read like a chat
* refactor(lens): use a quiet trace view menu
* refactor(lens): make trace view tabs explicit
* refactor(lens): restore compact trace view switch
* fix(lens): preserve complete conversation history and tool types
* feat(lens): add trace full-screen and close controls
* fix(lens): show forwarded answers and agent errors once
* fix(lens): reset full screen when closing a trace
* fix(roi): estimate linked authors and clarify model selection
* fix: preserve trace errors and ROI results across partial failures
* fix(roi): correct pagination variable typing
* fix(ui): place loaded root failures in conversation order
* fix(roi): read estimator recommendations from model catalog
* revert: remove catalog-driven ROI recommendations
* revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295)
This reverts commit ca1994e403.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(responses): merge bridged tool calls into the same choice as the text
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The floating bottom-right LiteAdmin button covered page controls such as
the Logs pagination buttons, and Playground had to hide it entirely.
Render the trigger as a pill in the header tools ahead of Docs and open
LiteAdmin as a panel docked beside the content column, which narrows the
page instead of covering it. Add a Cmd/Ctrl+J toggle and drop the
Playground override.
The Logs and trace drawers treated Cmd+J as a plain J and advanced the
selection, so they now share RunDrawer's rule that letter shortcuts
yield to modified presses and typing.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(dashboard): migrate to zod 4 and openai 6
Bump the dashboard to real zod 4.6.5 and openai 6.49.0 so every module
imports from bare "zod" instead of "zod/v4". Ports the ten files that
still used the zod 3 API (error params, record, passthrough, strict,
email/date validators, union discriminator codes) and adapts the
LiteAdmin tool schemas and form plumbing where openai 6 and zod 4
changed behaviour.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(dashboard): prettier format zod 4 schema files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(dashboard): restore system one missing state message under zod 4
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): streamed alias matching a capability rule bills the deployment price
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): assert every streamed chunk carries the client alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(logging): log the client alias on the priced streamed response, the same as non-streamed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): mark azure us/eu responses-only models as mode responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model-prices): keep azure us/eu o3-deep-research on chat, which Azure lists as chat capable
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(roi): support GitLab and tagged branch costs
* fix(roi): count tagged branches independently of estimation status
* test(roi): capture live GitHub and GitLab report validation
* fix(roi): open estimate details at the start
* fix(roi): clarify cost views and unify report layout
* feat(roi): showcase per-PR costs in the sample report
* fix(roi): separate report tabs and preserve branch cost attribution
* fix(roi): preserve demo previews and align progress spacing
* fix(roi): isolate demo loading and parallelize fork lookups
Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression
* fix(roi): separate demo and live loading states
Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers
* fix(roi): ignore refreshes from a previous source
* fix: trust gateway context for ROI estimator exclusion
* fix: preserve historical ROI estimator exclusion
* fix(search_tools): encrypt search tool litellm_params at rest
Encrypt every string value of a search tool's litellm_params on create and
update and decrypt on every DB read, so legacy plaintext rows load unchanged.
Include the table in master key rotation, LITELLM_MIGRATE_FROM_MASTER_KEY and
the migrate-encryption scan.
* fix(search_tools): keep edits made while the master key rotates
Write each rotated search tool row only if it still holds the litellm_params
that were read, and re-read and rotate it again if it was edited in between,
so a PUT that lands during /key/regenerate is not overwritten.
* fix(search_tools): retry rotation writes until the row stops changing
Rotate a search tool row again for as long as it keeps being edited instead of
giving up after five attempts, and stop with a warning only when the conditional
write fails on an unchanged row. Build the decrypted read result without
mutating it in place.
* refactor(search_tools): rotate edited rows in a loop, drop the step comment
Retry the conditional rotation write in a loop instead of recursion so sustained
edits cannot deepen the call stack, drop the step comment on the rotation call,
and stop mutating local state in the rotation tests.
* test(search_tools): drop the rotation test docstring
* Store search tool params as written when no encryption key is configured
* Rotate search tools under the salt key, keep non-ciphertext values and loaded tools that do not decrypt
* Treat a search tool as undecryptable only when its provider is ciphertext-length
* Drop suppressions the type discipline gate on main now reports as unused
* Show the loaded search tool in the admin list and info views when its DB params do not decrypt
* Keep the DB row's other fields when the admin views substitute loaded params
* fix(guardrails): encrypt guardrail litellm_params secrets at rest
* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation
- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals
* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor
- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes
* test(guardrails): drive the real guardrail rotator from the master key rotation test
- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key
* Annotate guardrail param encryption collections for type-discipline gate
* Type guardrail param recursion through validated JSON containers
* Type guardrail registry test helpers and drop section comment
* Reject client-supplied encrypted values in guardrail litellm_params
* Allow depth-bounded contains_encrypted_marker in the recursion detector
* Keep a loaded guardrail when its DB params do not decrypt with the current key
* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models
* Keep the loaded guardrail when an undecryptable param has no loaded value
* Drop suppressions the type discipline gate on main now reports as unused
* Assert what the reinitialized guardrail holds after an edit to an undecryptable one
* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize
* Type the rotation test helpers and drop the new test docstrings
* fix(guardrails): refuse to approve a submission whose params do not decrypt
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run search_endpoints tests in proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(llm_http_handler): keep provider error text when re-raising mapped errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): allow promptless image edits and default search models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): default missing image edit image to None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): build image edit defaults without mutating request data
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject a fake router for the search default model test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on vector store and file lookup handlers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): keep the lookup handler raise block to a single statement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover missing required body params and provider lookup status codes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bind spend-row request id with partial to satisfy B023
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): only reject non-positive page_size on vector store list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): remove unreachable fine-tuning body validation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover streaming anthropic messages reaching the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): count only provider calls when asserting missing params never reach the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve merge-base request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve interaction completion model defaults
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* fix(vector_stores): return managed file ids from vector store file list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vector_stores): cover managed file list route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): only map round-trippable managed ids and index flat file ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): build managed file gin index concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy-extras): move the managed file gin index migration after main's newest
* fix(vector_stores): satisfy lint gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy): drop the stale no-index note on the raw file-id guard
* test(vector_stores): cover managed file ids on the vector store file list end to end
Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
aws_bedrock_project_id reached Bedrock Mantle as anthropic-workspace, a
header AWS ignores, so requests ran under account defaults and models
that require a project data-retention mode failed with a 400. Send the
header AWS documents, anthropic-workspace-id, on both Mantle routes.
Re-lands #31994 by @gunjanjaswal, merged into the since-deleted
litellm_oss_staging branch on 2026-08-22, which never reached main.
Fixes#31947
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(streaming): keep the finish_reason of the provider's last content chunk in the logged response
When a provider sends the last content or tool-call delta and the finish_reason in one chunk (vLLM does this when speculative decoding finishes a reply in one step), the wrapper strips the finish_reason from that chunk and re-emits it on a terminal chunk from finish_reason_handler(). The terminal chunk reached the caller but was never added to self.chunks, so the response built for callbacks and SpendLogs defaulted to finish_reason "stop": tool calls and replies truncated by max_tokens were logged as plain stops. Append the terminal chunk in both the sync and async end-of-stream paths.
* fix(streaming): only log the terminal chunk when the provider sent a finish_reason
Review feedback: when a stream ends before any finish_reason arrives (e.g. an Anthropic stream cut after message_start), finish_reason_handler() synthesizes "stop"; adding that chunk to self.chunks made the response builder take the placeholder usage instead of estimating it. Append the terminal chunk only when received_finish_reason or intermittent_finish_reason is set, and test that case.
---------
Co-authored-by: joeym82956 <244252723+joeym82956@users.noreply.github.com>
* feat(interactions): durable cross-pod settlement for background interaction billing
Background interaction billing lived only in the creating replica's memory, so a
DELETE routed to another replica, or a restart of the creating one, never billed
the completed provider work and the budget reservation was refunded at the poll
timeout. The create now registers the billing context in a settlement store
before returning, the proxy installs a Prisma-backed store at boot
(LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica
claims the row once through a conditional update before billing or releasing,
startup resumes every unclaimed row with its remaining timeout, and a give-up
records an unsettled outcome instead of silently reconciling to zero. The SDK
keeps an in-memory store and behaves as before.
* fix(interactions): survive a settlement install failure at boot and stop carrying request headers
* fix(interactions): drop the stored request context once a settlement row is settled
* fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline
* fix(interactions): carry a missing model through the settlement context for agent-only background creates
An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it.
* fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere
The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials.
* fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill
* fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed
A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler.
A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere.
* fix(interactions): settle an unverified registration through the durable claim
A create whose settlement-store write raised no longer bills through a
private in-memory gate that a later boot's resume cannot see. The claim
asks the durable store first and falls back to the local gate only when
the store answers that no row exists, and a missing settlement table reads
as no rows so a replica without the migration still settles in process.
* test(proxy): keep the settlement test where the proxy-infra shard collects it
The merge of main moved test_background_interaction_settlement.py under
tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy
files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it
and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected
whole by the proxy-infra shard, which is where the test ran before the merge.
* fix(interactions): raise on a non-2xx Gemini interaction fetch
AsyncHTTPHandler.get never raises for status and the Gemini GET transform
only raised when the body was not JSON, so a 500 or 404 carrying Gemini's
JSON error body parsed as an interaction with no status. A delete on a
replica other than the creator then claimed the settlement as released and
forwarded the delete instead of failing closed, and the bill was lost. The
transform now raises GeminiError with the vendor's status, as the delete
transform already does; the in-process poll already retries a fetch that
raises
* test(integration): audit durable background interaction settlement across replicas
Twenty-six deterministic cells drive a one-worker creator and a two-worker
settler against an owned scripted Gemini upstream: cross-replica deletes
bill once, failed and cancelled interactions release, a later replica
resumes unclaimed rows, custom deployment pricing bills at the deployment
rate, a fetch the settler cannot make fails the delete closed, odd ids are
refused, a missing settlement table keeps in-process billing, polling
disabled registers nothing, the budget reservation is released by the
settler, an upstream outage mid-burst fails closed and recovers, killed
workers hand their polls to the respawned ones, and concurrent deletes on a
slow upstream settle exactly once. The support upstream gains a scripted
interaction store with per-id GET status and delay, and the process helper
gains an owned upstream a test can stop and restart
* test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate
* chore(ui): regenerate dashboard API types after merging main
* test(integration): accept the 422 budget refusal and a respawned worker's resume
The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged
* test(integration): delete the pinned key's interaction with a second key
A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): use OpenAI workload identity federation tokens on /openai_passthrough
The OpenAI passthrough routes (HTTP and websocket) only looked up a static
OpenAI API key, so proxies authenticating to OpenAI through workload identity
federation failed with "Required 'OPENAI_API_KEY'". When no non-empty static
key is configured, exchange the workload identity subject token for a bearer
token, scoped to the passthrough's own OPENAI_API_BASE so the token is never
sent to a non-OpenAI host
* fix(proxy): close the OpenAI websocket passthrough cleanly when the workload identity exchange fails
A rejected, unreachable or unreadable workload identity exchange raised out of
the websocket route before the handshake was accepted, so clients saw a bare
handshake failure. Log the cause and close with 1011 and a fixed reason instead
* refactor(openai): resolve workload identity bearer tokens for an api base in the OpenAI provider module
The passthrough route only owns the static key lookup now and asks the OpenAI
provider module for a workload identity bearer token scoped to its api base
---------
Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
* feat(lens): add preset watch-for checks for common agent failures
* feat(lens): add keyboard-driven watch-for picker with lens dot animation
* feat(lens): use the watch-for picker in investigation setup
* feat(lens): show preset checks by name in the criteria tab
* test(lens): cover saving and editing watch-for presets
* feat(lens): shorten watch-for summaries and start with three presets on
* feat(lens): lay out watch-for presets as toggle tiles with a clear add-your-own button
* feat(lens): open a custom check from the watch-for picker
* test(lens): cover watch-for tiles and the add-your-own button
* fix(lens): draw the selected tile border inside the tile so the dialog edge cannot clip it
Bedrock GPT-5.6 rejects reasoning effort minimal with 400 unsupported_value. The model map already
sets supports_minimal_reasoning_effort=false for these models, but neither the Converse path nor the
native Responses path read it, so the value was forwarded. Drop it under drop_params and raise
UnsupportedParamsError otherwise, matching how the OpenAI GPT-5 path treats explicitly disabled
effort levels
Co-authored-by: Aasif-Multani <20943280+Aasif-Multani@users.noreply.github.com>
Per-router auto-router savings were missing on proxies with
disable_spend_logs and for requests without a session id under
missing_session_id: omit. The router-day row is written by the
auto-router turn, whose enqueue sat behind the spend-logs flag and whose
builder dropped sessionless requests, while LiteLLM_DailyUserSpend has
neither gate. That money then showed only as unattributed savings.
Enqueue the turn whether or not spend logs are kept, and keep a
sessionless turn with an empty session id that writes the router-day row
while the session upserts skip it. The day and session rows still commit
in one statement, so late baseline corrections keep their ordering.
Without spend logs, write only the router-day aggregate: the turn drops
its session id, so no per-session row is stored, and baseline capture is
skipped, since a baseline observation can only publish once its
request's spend log exists. The flag keeps its meaning of no per-request
or per-session data, while the daily per-router money matches
LiteLLM_DailyUserSpend.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(cost): price streamed aliases that only match a capability rule from the deployment model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): satisfy basedpyright delta after the cost-candidate sort
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(cost): skip the capability-rule check for exact cost-map keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): price streamed alias rows from the deployment's own rates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): isolate streamed alias deployments with per-run model names
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(streaming): keep the unpriceable-stamp case on a truly unmapped model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): treat capability-rule matches as unmapped in every cost lookup
Route every model-info lookup on cost paths through get_priced_model_info
and _cached_get_priced_model_info_helper, which raise ModelNotMappedError
when the only match is a pricing-free fallback-generalizations rule. An
alias that matches a capability rule now falls through to the deployment's
real model instead of billing 0, and the earlier candidate-sorting fix is
reverted since the priced helper is the single choke point.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): assert the rule alias bills the same as the plain alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): treat router-registered rule-only model_cost entries as unmapped
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): count string rates and check pricing before the capability-rule match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(cost): repoint cost-path patches at get_priced_model_info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(cost): give the model_cost row cast a cast-ok reason
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost): type the lazy get_priced_model_info export
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost): narrow the fix to ordering rule-only cost candidates last
Drops get_priced_model_info and its call-site swaps, the lazy-import entry,
and the cost-path code check. Only the candidate sort in completion_cost and
pricing_entry_for_cost_calc stays, with its tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): cover rule-only base_model billing on every endpoint, client and failure mode
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>