Commit graph

53570 commits

Author SHA1 Message Date
berriai-litellm-provider-info-sync[bot]
9fe6442172
fix(azure): add azure_ai/flux.2-pro input token limit from the models sold directly page (#44405)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 11:40:44 -07:00
devin-ai-integration[bot]
a57cdfddf8
fix(ci): give the pass-through auth regression test a real FastAPI app (#44403)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 18:31:48 +00:00
devin-ai-integration[bot]
a849093d89
test(cost): pin OpenAI reported web search count with mixed actions (#44414)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 11:31:45 -07:00
devin-ai-integration[bot]
6c32384d8c
fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3 (#44292)
* fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3

Bedrock prices Kimi K3 cache reads through implicit caching but Converse
rejects the explicit cachePoint marker ("This model doesn't support the
cachePoint field"), so any cache_control on the request answered 400.
Mark the three K3 rows supports_prompt_cache_breakpoint: false and have
bedrock_model_accepts_cache_points honor that flag before falling back to
supports_prompt_caching, keeping cached-token pricing intact.

* test(bedrock): assert cache points per request section

* fix(bedrock): honor a deployment's cache breakpoint flag for unmapped models

* fix(bedrock): read a converse-routed deployment's cache breakpoint flag

* refactor(bedrock): look up cache breakpoint flags by key

* test(bedrock): add the Kimi K3 cache point wire audit

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 18:11:46 +00:00
tin-berri
4732647de2
feat(ui): make LiteAdmin enterprise-only (#44399) 2026-10-03 11:07:54 -07:00
ishaan-berri
f0abe1bea1
feat(lens): live dot field timeline and full-screen traces view (#44390)
* feat(lens): add dot field layout and live status helpers for the traces timeline

* test(lens): cover dot field layout, agent colors and live status

* feat(lens): draw the traces timeline as a live dot field

* feat(lens): add the sweep animation for the live traces timeline

* feat(lens): put the lens tabs in a compact header and fill the screen with traces
2026-10-03 10:56:48 -07:00
moe-berri
ad8babae33
fix(lens): paginate trace reads within ClickHouse limits (#44384) 2026-10-03 10:38:40 -07:00
devin-ai-integration[bot]
8b1990b4bc
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register typesafe as a provider so Jev deployments load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move provider endpoints under llms and validate proxy bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add Cloudflare Clef and Strands Decider backends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register decisions routes for managed agents and gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(decisions): use raw regex for cloudflare missing account match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): avoid cast in Cloudflare response unwrapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): default model, evaluation health probe, short Cloudflare names

The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.

* fix(decisions): let health_check_params override the evaluation probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit the decisions endpoint across providers, limits, health and chaos

Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).

The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.

* fix(decisions): send env API keys to a configured api_base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register the routes through the lazy feature registry

The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.

The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.

* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell

* fix(proxy): let a config pass-through beat a lazily registered route in eager mode

With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:38:38 +00:00
devin-ai-integration[bot]
b024950353
feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop (#44391)
* feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(harness): name per-tool spec FunctionTool

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-03 17:33:01 +00:00
berriai-litellm-provider-info-sync[bot]
231a46e40b
feat(azure): add azure_ai/kimi-k2-thinking from Azure Kimi pricing page (#44382)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 10:04:59 -07:00
devin-ai-integration[bot]
a5e9f275ea
fix(proxy): keep the database error when the log_db_metrics failure hook raises (#44383)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 16:38:15 +00:00
devin-ai-integration[bot]
4c6c84afb7
perf(proxy): stop prompt-cache eligibility from tokenizing the whole conversation (#44221)
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them

Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
797353f13a
fix(otel): name postgres service spans by operation and table (#44240)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
564d236985
fix(otel): nest cache spans under their operation and name service spans by purpose (#44150)
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.

A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
devin-ai-integration[bot]
53d2ab6b05
fix(proxy): emit postgres service spans only on real DB reads in auth cache helpers (#44148)
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.

A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.

Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
berriai-litellm-provider-info-sync[bot]
66a422ea50
fix(azure): add MAI-Image max_input_tokens from models sold directly page (#44375)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 08:49:09 -07:00
devin-ai-integration[bot]
4ece6c9fb8
build(deps): suppress unfixed braces GHSA-vfj7-8cjw-p6xm to clear osv-scan (#44347)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 07:55:14 -07:00
devin-ai-integration[bot]
e768ad55ce
refactor(types): replace Any with proven types in 4 files (#44370)
* refactor(types): replace Any with proven types in 8 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert unproven email logger protocol and iterator annotation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert unproven deepagents and http handler annotations

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 03:52:28 -07:00
devin-ai-integration[bot]
e340e546e2
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* wip

* wip

* wip

* chore(trace): checkpoint ongoing Rust migration

* refactor(trace): group Python bridge under trace package

* refactor(traces): read span conventions through a Convention trait

Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.

The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.

* feat(trace): export Rust-owned wire schemas and enforce contract bounds

* fix(trace): bound quoted counts in ClickHouse wire schemas

* feat(trace): generate Python wire contracts with datamodel-code-generator

* test(trace): validate migrated callers and generated contracts at the native boundary

* refactor(traces): rename normalization convention to format

* fix(traces): reconcile spend evidence and preserve unknown costs

* feat(traces): normalize additional telemetry formats

* test(traces): cover captured normalization fixtures

* refactor(traces): isolate SDK normalization rules

* feat(tracing): seed all trace exports for local dashboard

* fix(clickhouse): preserve custom LiteLLM request metadata

* docs(traces): define normalization module boundaries

* docs(traces): define resolution and OTLP boundaries

* fix(ui): normalize nullable trace message names

* refactor(traces): split resolver modules and cover resolution behavior

* test(traces): replace normalization snapshots with behavior assertions

* fix(ui): align dashboard API contracts with generated types

* refactor(traces): type normalization and storage boundaries

* fix(traces): seed captured SDK spend and preserve provider identities

* wip

* test(traces): verify guide discovery and content ordering

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:59:51 +00:00
devin-ai-integration[bot]
d260765652
refactor: clean up fresh tech debt from 2026-10-02 (#44362)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 02:12:03 -07:00
devin-ai-integration[bot]
e92bd50de1
refactor(ui): inject Lens backends as services instead of a fake HTTP client (#44369)
The interactive Lens demo used to build a fetch shim that encoded in-memory
fixtures as HTTP responses so the shared ApiClient could decode them again,
and every trace view branched on demo vs live to pick a URL. Lens and the
trace views now depend on two small service interfaces, LensApi and
TracesApi, with named operations. The live layer wraps the existing HTTP
calls, the demo layer reads fixtures directly, and a React context provides
whichever one the session runs on. Without a provider the hooks fall back to
the live implementation, so the live app and existing tests are unchanged.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 09:11:08 +00:00
devin-ai-integration[bot]
09313bcf1b
refactor(ui): reorganize Lens dashboard components (#44354)
* refactor(ui): reorganize Lens dashboard components

* refactor(ui): align Lens forms with react-hook-form

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep Lens submit errors out of form validity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): reduce Lens lint budget usage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): use a query key factory for Lens queries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 08:29:16 +00:00
moe-berri
cb00dbecd7
feat(ui): improve trace inspection and ROI estimation (#44351)
* feat(lens): simplify trace inspection in the gateway drawer

* feat(lens): add a conversation view for full traces

* test(lens): keep normalized message fixtures type safe

* refactor(lens): make conversation view read like a chat

* refactor(lens): use a quiet trace view menu

* refactor(lens): make trace view tabs explicit

* refactor(lens): restore compact trace view switch

* fix(lens): preserve complete conversation history and tool types

* feat(lens): add trace full-screen and close controls

* fix(lens): show forwarded answers and agent errors once

* fix(lens): reset full screen when closing a trace

* fix(roi): estimate linked authors and clarify model selection

* fix: preserve trace errors and ROI results across partial failures

* fix(roi): correct pagination variable typing

* fix(ui): place loaded root failures in conversation order

* fix(roi): read estimator recommendations from model catalog

* revert: remove catalog-driven ROI recommendations
2026-10-03 01:18:16 -07:00
devin-ai-integration[bot]
5724117116
fix(responses): merge bridged tool calls into the same choice as the text (#44346)
* revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295)

This reverts commit ca1994e403.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): merge bridged tool calls into the same choice as the text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 07:00:52 +00:00
tin-berri
8e32d4568c
feat(ui): move LiteAdmin into the header with a docked side panel (#44293)
The floating bottom-right LiteAdmin button covered page controls such as
the Logs pagination buttons, and Playground had to hide it entirely.
Render the trigger as a pill in the header tools ahead of Docs and open
LiteAdmin as a panel docked beside the content column, which narrows the
page instead of covering it. Add a Cmd/Ctrl+J toggle and drop the
Playground override.

The Logs and trace drawers treated Cmd+J as a plain J and advanced the
selection, so they now share RunDrawer's rule that letter shortcuts
yield to modified presses and typing.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 00:00:10 -07:00
devin-ai-integration[bot]
0c86d6bfc9
refactor(dashboard): migrate to zod 4 and openai 6 (#44345)
* refactor(dashboard): migrate to zod 4 and openai 6

Bump the dashboard to real zod 4.6.5 and openai 6.49.0 so every module
imports from bare "zod" instead of "zod/v4". Ports the ten files that
still used the zod 3 API (error params, record, passthrough, strict,
email/date validators, union discriminator codes) and adapts the
LiteAdmin tool schemas and form plumbing where openai 6 and zod 4
changed behaviour.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(dashboard): prettier format zod 4 schema files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(dashboard): restore system one missing state message under zod 4

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:53:41 +00:00
devin-ai-integration[bot]
d96e56c76f
revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295) (#44344)
This reverts commit ca1994e403.

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:47:11 +00:00
devin-ai-integration[bot]
7b432d78d2
fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model (#44341)
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): streamed alias matching a capability rule bills the deployment price

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): assert every streamed chunk carries the client alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): log the client alias on the priced streamed response, the same as non-streamed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:44:52 +00:00
devin-ai-integration[bot]
a76ba8c01e
revert(cost): revert "fix(cost): price rule-only model names at the deployment's rate" (#44144) (#44335)
This reverts commit f9a32ffcb5.

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:09:00 +00:00
devin-ai-integration[bot]
0fce5bccbc
fix(model-prices): mark azure us/eu responses-only models as mode responses (#44323)
* fix(model-prices): mark azure us/eu responses-only models as mode responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): keep azure us/eu o3-deep-research on chat, which Azure lists as chat capable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 23:05:55 -07:00
moe-berri
caee45fed4
feat(roi): add GitLab sources and branch cost attribution (#44324)
* feat(roi): support GitLab and tagged branch costs

* fix(roi): count tagged branches independently of estimation status

* test(roi): capture live GitHub and GitLab report validation

* fix(roi): open estimate details at the start

* fix(roi): clarify cost views and unify report layout

* feat(roi): showcase per-PR costs in the sample report

* fix(roi): separate report tabs and preserve branch cost attribution

* fix(roi): preserve demo previews and align progress spacing

* fix(roi): isolate demo loading and parallelize fork lookups

Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression

* fix(roi): separate demo and live loading states

Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers

* fix(roi): ignore refreshes from a previous source

* fix: trust gateway context for ROI estimator exclusion

* fix: preserve historical ROI estimator exclusion
2026-10-03 06:01:43 +00:00
yucheng-berri
12d4b75b7a
fix(search_tools): encrypt search tool litellm_params at rest (#43631)
* fix(search_tools): encrypt search tool litellm_params at rest

Encrypt every string value of a search tool's litellm_params on create and
update and decrypt on every DB read, so legacy plaintext rows load unchanged.
Include the table in master key rotation, LITELLM_MIGRATE_FROM_MASTER_KEY and
the migrate-encryption scan.

* fix(search_tools): keep edits made while the master key rotates

Write each rotated search tool row only if it still holds the litellm_params
that were read, and re-read and rotate it again if it was edited in between,
so a PUT that lands during /key/regenerate is not overwritten.

* fix(search_tools): retry rotation writes until the row stops changing

Rotate a search tool row again for as long as it keeps being edited instead of
giving up after five attempts, and stop with a warning only when the conditional
write fails on an unchanged row. Build the decrypted read result without
mutating it in place.

* refactor(search_tools): rotate edited rows in a loop, drop the step comment

Retry the conditional rotation write in a loop instead of recursion so sustained
edits cannot deepen the call stack, drop the step comment on the rotation call,
and stop mutating local state in the rotation tests.

* test(search_tools): drop the rotation test docstring

* Store search tool params as written when no encryption key is configured

* Rotate search tools under the salt key, keep non-ciphertext values and loaded tools that do not decrypt

* Treat a search tool as undecryptable only when its provider is ciphertext-length

* Drop suppressions the type discipline gate on main now reports as unused

* Show the loaded search tool in the admin list and info views when its DB params do not decrypt

* Keep the DB row's other fields when the admin views substitute loaded params
2026-10-02 22:53:18 -07:00
yucheng-berri
5d42cb7cfa
fix(guardrails): encrypt guardrail litellm_params secrets at rest (#43627)
* fix(guardrails): encrypt guardrail litellm_params secrets at rest

* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation

- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals

* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor

- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes

* test(guardrails): drive the real guardrail rotator from the master key rotation test

- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key

* Annotate guardrail param encryption collections for type-discipline gate

* Type guardrail param recursion through validated JSON containers

* Type guardrail registry test helpers and drop section comment

* Reject client-supplied encrypted values in guardrail litellm_params

* Allow depth-bounded contains_encrypted_marker in the recursion detector

* Keep a loaded guardrail when its DB params do not decrypt with the current key

* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models

* Keep the loaded guardrail when an undecryptable param has no loaded value

* Drop suppressions the type discipline gate on main now reports as unused

* Assert what the reinitialized guardrail holds after an edit to an undecryptable one

* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize

* Type the rotation test helpers and drop the new test docstrings

* fix(guardrails): refuse to approve a submission whose params do not decrypt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:49:41 -07:00
devin-ai-integration[bot]
6d8434f940
fix(proxy): return 4xx instead of 500 for missing required params, invalid pagination and unknown ids (#43787)
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run search_endpoints tests in proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(llm_http_handler): keep provider error text when re-raising mapped errors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): allow promptless image edits and default search models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): default missing image edit image to None

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): build image edit defaults without mutating request data

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): inject a fake router for the search default model test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on vector store and file lookup handlers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): keep the lookup handler raise block to a single statement

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover missing required body params and provider lookup status codes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bind spend-row request id with partial to satisfy B023

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): only reject non-positive page_size on vector store list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): remove unreachable fine-tuning body validation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover streaming anthropic messages reaching the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): count only provider calls when asserting missing params never reach the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve merge-base request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve interaction completion model defaults

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-02 22:48:53 -07:00
devin-ai-integration[bot]
8efb4a21f6
fix(vector_stores): return managed file ids from vector store file list (#43800)
* fix(vector_stores): return managed file ids from vector store file list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(vector_stores): cover managed file list route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vector_stores): only map round-trippable managed ids and index flat file ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy-extras): build managed file gin index concurrently

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(proxy-extras): move the managed file gin index migration after main's newest

* fix(vector_stores): satisfy lint gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(proxy): drop the stale no-index note on the raw file-id guard

* test(vector_stores): cover managed file ids on the vector store file list end to end

Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 21:28:50 -07:00
Ankit Jha
0d5ea45c86
feat(scaleway): add rerank support (#44160)
* feat(scaleway): add rerank support

Fixes #43856

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>

* fix(scaleway): caller headers cannot replace the provider key; inject the async client in tests

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>

* refactor(utils): fold Scaleway into the Jina rerank branch to stay inside the complexity budget

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>

---------

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>
2026-10-02 21:20:05 -07:00
michelligabriele
119179942e
fix(guardrails): run unified guardrails on /v1/images/edits (#44195) 2026-10-02 21:14:35 -07:00
Koh Jun Hao
5ce81631fe
fix(bedrock): send the anthropic-workspace-id header to Bedrock Mantle (#44173)
aws_bedrock_project_id reached Bedrock Mantle as anthropic-workspace, a
header AWS ignores, so requests ran under account defaults and models
that require a project data-retention mode failed with a 400. Send the
header AWS documents, anthropic-workspace-id, on both Mantle routes.

Re-lands #31994 by @gunjanjaswal, merged into the since-deleted
litellm_oss_staging branch on 2026-08-22, which never reached main.

Fixes #31947

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 21:12:33 -07:00
joeym82956
2394fc3682
fix(streaming): keep the finish_reason of the provider's last content chunk in the logged response (#44318)
* fix(streaming): keep the finish_reason of the provider's last content chunk in the logged response

When a provider sends the last content or tool-call delta and the finish_reason in one chunk (vLLM does this when speculative decoding finishes a reply in one step), the wrapper strips the finish_reason from that chunk and re-emits it on a terminal chunk from finish_reason_handler(). The terminal chunk reached the caller but was never added to self.chunks, so the response built for callbacks and SpendLogs defaulted to finish_reason "stop": tool calls and replies truncated by max_tokens were logged as plain stops. Append the terminal chunk in both the sync and async end-of-stream paths.

* fix(streaming): only log the terminal chunk when the provider sent a finish_reason

Review feedback: when a stream ends before any finish_reason arrives (e.g. an Anthropic stream cut after message_start), finish_reason_handler() synthesizes "stop"; adding that chunk to self.chunks made the response builder take the placeholder usage instead of estimating it. Append the terminal chunk only when received_finish_reason or intermittent_finish_reason is set, and test that case.

---------

Co-authored-by: joeym82956 <244252723+joeym82956@users.noreply.github.com>
2026-10-02 21:11:40 -07:00
devin-ai-integration[bot]
8596fe954d
feat(interactions): durable cross-pod settlement for background interaction billing (#41955)
* feat(interactions): durable cross-pod settlement for background interaction billing

Background interaction billing lived only in the creating replica's memory, so a
DELETE routed to another replica, or a restart of the creating one, never billed
the completed provider work and the budget reservation was refunded at the poll
timeout. The create now registers the billing context in a settlement store
before returning, the proxy installs a Prisma-backed store at boot
(LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica
claims the row once through a conditional update before billing or releasing,
startup resumes every unclaimed row with its remaining timeout, and a give-up
records an unsettled outcome instead of silently reconciling to zero. The SDK
keeps an in-memory store and behaves as before.

* fix(interactions): survive a settlement install failure at boot and stop carrying request headers

* fix(interactions): drop the stored request context once a settlement row is settled

* fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline

* fix(interactions): carry a missing model through the settlement context for agent-only background creates

An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it.

* fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere

The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials.

* fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill

* fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed

A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler.

A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere.

* fix(interactions): settle an unverified registration through the durable claim

A create whose settlement-store write raised no longer bills through a
private in-memory gate that a later boot's resume cannot see. The claim
asks the durable store first and falls back to the local gate only when
the store answers that no row exists, and a missing settlement table reads
as no rows so a replica without the migration still settles in process.

* test(proxy): keep the settlement test where the proxy-infra shard collects it

The merge of main moved test_background_interaction_settlement.py under
tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy
files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it
and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected
whole by the proxy-infra shard, which is where the test ran before the merge.

* fix(interactions): raise on a non-2xx Gemini interaction fetch

AsyncHTTPHandler.get never raises for status and the Gemini GET transform
only raised when the body was not JSON, so a 500 or 404 carrying Gemini's
JSON error body parsed as an interaction with no status. A delete on a
replica other than the creator then claimed the settlement as released and
forwarded the delete instead of failing closed, and the bill was lost. The
transform now raises GeminiError with the vendor's status, as the delete
transform already does; the in-process poll already retries a fetch that
raises

* test(integration): audit durable background interaction settlement across replicas

Twenty-six deterministic cells drive a one-worker creator and a two-worker
settler against an owned scripted Gemini upstream: cross-replica deletes
bill once, failed and cancelled interactions release, a later replica
resumes unclaimed rows, custom deployment pricing bills at the deployment
rate, a fetch the settler cannot make fails the delete closed, odd ids are
refused, a missing settlement table keeps in-process billing, polling
disabled registers nothing, the budget reservation is released by the
settler, an upstream outage mid-burst fails closed and recovers, killed
workers hand their polls to the respawned ones, and concurrent deletes on a
slow upstream settle exactly once. The support upstream gains a scripted
interaction store with per-id GET status and delay, and the process helper
gains an owned upstream a test can stop and restart

* test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate

* chore(ui): regenerate dashboard API types after merging main

* test(integration): accept the 422 budget refusal and a respawned worker's resume

The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged

* test(integration): delete the pinned key's interaction with a second key

A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 21:08:38 -07:00
Matthew Lapointe
ad566c90dd
fix(proxy): use OpenAI workload identity federation tokens on /openai_passthrough (#44140)
* fix(proxy): use OpenAI workload identity federation tokens on /openai_passthrough

The OpenAI passthrough routes (HTTP and websocket) only looked up a static
OpenAI API key, so proxies authenticating to OpenAI through workload identity
federation failed with "Required 'OPENAI_API_KEY'". When no non-empty static
key is configured, exchange the workload identity subject token for a bearer
token, scoped to the passthrough's own OPENAI_API_BASE so the token is never
sent to a non-OpenAI host

* fix(proxy): close the OpenAI websocket passthrough cleanly when the workload identity exchange fails

A rejected, unreachable or unreadable workload identity exchange raised out of
the websocket route before the handshake was accepted, so clients saw a bare
handshake failure. Log the cause and close with 1011 and a fixed reason instead

* refactor(openai): resolve workload identity bearer tokens for an api base in the OpenAI provider module

The passthrough route only owns the static key lookup now and asks the OpenAI
provider module for a workload identity bearer token scoped to its api base

---------

Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
2026-10-03 04:07:09 +00:00
ishaan-berri
80a2f4d8a8
feat(lens): add preset watch-for checks to investigation setup (#44313)
* feat(lens): add preset watch-for checks for common agent failures

* feat(lens): add keyboard-driven watch-for picker with lens dot animation

* feat(lens): use the watch-for picker in investigation setup

* feat(lens): show preset checks by name in the criteria tab

* test(lens): cover saving and editing watch-for presets

* feat(lens): shorten watch-for summaries and start with three presets on

* feat(lens): lay out watch-for presets as toggle tiles with a clear add-your-own button

* feat(lens): open a custom check from the watch-for picker

* test(lens): cover watch-for tiles and the add-your-own button

* fix(lens): draw the selected tile border inside the tile so the dialog edge cannot clip it
2026-10-02 21:06:04 -07:00
Aasif Multani
54260bec88
fix(bedrock): honor unsupported reasoning effort levels for OpenAI GPT models (#44183)
Bedrock GPT-5.6 rejects reasoning effort minimal with 400 unsupported_value. The model map already
sets supports_minimal_reasoning_effort=false for these models, but neither the Converse path nor the
native Responses path read it, so the value was forwarded. Drop it under drop_params and raise
UnsupportedParamsError otherwise, matching how the OpenAI GPT-5 path treats explicitly disabled
effort levels

Co-authored-by: Aasif-Multani <20943280+Aasif-Multani@users.noreply.github.com>
2026-10-02 21:03:58 -07:00
devin-ai-integration[bot]
6b6222578e
test(integration): let run.py select cells by pytest node id (#44319)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 21:01:57 -07:00
tin-berri
db14931401
fix(auto-router): attribute router-day savings without spend logs or a session (#44226)
Per-router auto-router savings were missing on proxies with
disable_spend_logs and for requests without a session id under
missing_session_id: omit. The router-day row is written by the
auto-router turn, whose enqueue sat behind the spend-logs flag and whose
builder dropped sessionless requests, while LiteLLM_DailyUserSpend has
neither gate. That money then showed only as unattributed savings.

Enqueue the turn whether or not spend logs are kept, and keep a
sessionless turn with an empty session id that writes the router-day row
while the session upserts skip it. The day and session rows still commit
in one statement, so late baseline corrections keep their ordering.

Without spend logs, write only the router-day aggregate: the turn drops
its session id, so no per-session row is stored, and baseline capture is
skipped, since a baseline observation can only publish once its
request's spend log exists. The flag keeps its meaning of no per-request
or per-session data, while the daily per-router money matches
LiteLLM_DailyUserSpend.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 20:56:45 -07:00
devin-ai-integration[bot]
f9a32ffcb5
fix(cost): price rule-only model names at the deployment's rate (#44144)
* fix(cost): price streamed aliases that only match a capability rule from the deployment model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): satisfy basedpyright delta after the cost-candidate sort

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(cost): skip the capability-rule check for exact cost-map keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): price streamed alias rows from the deployment's own rates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): isolate streamed alias deployments with per-run model names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(streaming): keep the unpriceable-stamp case on a truly unmapped model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): treat capability-rule matches as unmapped in every cost lookup

Route every model-info lookup on cost paths through get_priced_model_info
and _cached_get_priced_model_info_helper, which raise ModelNotMappedError
when the only match is a pricing-free fallback-generalizations rule. An
alias that matches a capability rule now falls through to the deployment's
real model instead of billing 0, and the earlier candidate-sorting fix is
reverted since the priced helper is the single choke point.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): assert the rule alias bills the same as the plain alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): treat router-registered rule-only model_cost entries as unmapped

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): count string rates and check pricing before the capability-rule match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost): repoint cost-path patches at get_priced_model_info

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(cost): give the model_cost row cast a cast-ok reason

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): type the lazy get_priced_model_info export

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): narrow the fix to ordering rule-only cost candidates last

Drops get_priced_model_info and its call-site swaps, the lazy-import entry,
and the cost-path code check. Only the candidate sort in completion_cost and
pricing_entry_for_cost_calc stays, with its tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover rule-only base_model billing on every endpoint, client and failure mode

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 03:18:22 +00:00
devin-ai-integration[bot]
ca1994e403
fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice (#44295)
* fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): preserve deferred logging bridge coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover gpt-6 family bridge routing and merged tool calls in integration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 20:17:55 -07:00
ishaan-berri
2ccb7ed06e
feat(lens): issue briefs with problem, user goal, outcome and test cases (#44311)
* refactor(lens): move analysis prompts into markdown files

* feat(lens): ask the investigator for a scoped agent fix brief with two options

* test(lens): cover the agent fix brief through investigation and merges

* chore(ui): regenerate api types for the lens fix brief

* feat(ui): build copyable lens fix prompts

* feat(ui): show the lens fix brief with copy buttons for claude code and codex

* test(ui): cover copying a lens fix option

* refactor(lens): replace the fix options with a plain issue brief

* feat(lens): ask for problem, user goal, outcome and test cases without prescribing code changes

* test(lens): cover the issue brief through investigation and merges

* chore(ui): regenerate api types for the lens issue brief

* refactor(ui): drop the lens fix prompt builders

* feat(ui): add a lens issue brief panel

* feat(ui): show the lens issue brief in the finding drawer

* test(ui): cover the lens issue brief and the legacy fallback

* feat(ui): render a lens issue brief as a markdown document

* test(ui): pin the lens issue brief markdown layout

* feat(ui): show the issue brief as a copyable file with claude code and codex buttons

* feat(ui): pass the finding title into the issue brief

* test(ui): cover copying the issue brief for claude code and codex

* feat(ui): bold the input and expected labels in lens test cases

* test(ui): pin the bold test case labels in the issue brief

* feat(ui): render the issue brief as formatted markdown

* test(ui): cover the rendered issue brief sections and raw markdown copy
2026-10-03 03:00:29 +00:00
devin-ai-integration[bot]
233db9f285
feat(bedrock): serve gpt-5.6+ chat completions natively by default, with chat_completions/ opt-in for gpt-oss and grok (#44307)
* feat(bedrock): send grok chat completions through runtime openai path

Unspecified bedrock grok was rewritten to Converse. Chat completions now hit bedrock-runtime /openai/v1/chat/completions, and converse/ still uses Converse

* feat(bedrock): serve gpt-oss and gpt-5.6 chat completions on runtime's native openai path

* fix(bedrock): route gpt-oss response_format to Converse and decide the route once from the raw request

* fix(bedrock): serve region-path and GovCloud gpt-oss ids on native Chat Completions

The cost-map parity tests require every regional variant of a flagged id to carry the same supports_ flags, so the six us-gov gpt-oss entries now carry the native-route flags too. A region path in the model name (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model: the route is looked up on the id after the path, the path's region picks the endpoint and the SigV4 scope, an explicit aws_region_name still wins, and the body carries the bare id AWS expects

* fix(bedrock): keep params AWS refuses natively off the chat completions route

Drop the params each family 400s or 503s on runtime Chat Completions from the native config's supported list (GPT-5.6 penalties, stop, and logprobs, Grok penalties, gpt-oss logit_bias) so drop_params drops them as Converse did, gate legacy functions on GPT-5.6 the same way as tools, and send an Anthropic-style thinking block to Converse, the only route that forwards it

* fix(bedrock): keep schema-less json_object on Converse for the chat completions models

* fix(bedrock): keep every json_object response_format on Converse for the chat completions models

* fix(rust): declare the bedrock runtime chat completions flags on ModelInfo

* fix(bedrock): opt into the native chat completions route through supported_endpoints

* docs(cost-map): describe the bedrock native chat completions capability flags

* revert: docs(cost-map): describe the bedrock native chat completions capability flags

This reverts commit 4101c0ceb2.

cost-map-guard runs main's schema generator under pull_request_target and compares
its output to the PR's committed schema, so a PR that changes the generator's output
cannot pass that required check until the generator change lands on main first. The
descriptions move to a follow-up that lands the generator change ahead of the schema

* test(bedrock): move the native chat completions tests under tests/unit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): drop reasoning_effort none for grok on the native chat completions route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): keep converse extension params on the converse route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): inline http image urls and keep stop on converse for native chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(bedrock): share the sync remote media inliner

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(image-handling): infer the image mime type when the server sends a generic content type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): stop sending aws_bedrock_project_id as OpenAI-Project on the runtime chat completions route

* feat(bedrock): make native chat completions an opt-in bedrock/chat_completions/ route

Bare Bedrock OpenAI and Grok model ids stay on Converse as on main. The
bedrock/chat_completions/<model> prefix opts a deployment into bedrock-runtime's
/openai/v1/chat/completions, and a request carrying a Converse-only param
still falls back to Converse. The cost map no longer decides the route.

* fix(bedrock): keep chat_completions/-prefixed deployments on the native Responses surface

* fix(bedrock): keep provider response headers on the runtime chat completions route

* feat(bedrock): serve gpt-5.6 and newer on runtime chat completions by default

Unprefixed bedrock/<gpt-5.6+> models whose cost-map row lists /v1/chat/completions
now route to the native OpenAI-compatible endpoint; converse/ pins Converse and
chat_completions/ still opts gpt-oss and Grok in. Guardrails, application inference
profile ARNs, and tools with reasoning keep falling back to Converse per request.
Hoist the remote-media url comprehension into a single-clause helper.

* fix(bedrock): refuse temperature and top_p natively on GPT 5.6 and newer like Converse does

AWS answers temperature and top_p with a 400 on the native Chat Completions endpoint for the GPT 5.6+ models, the same models whose Converse route already dropped both under drop_params via supports_sampling_params: false. The native config now honors that price-map flag, the gpt-6 and gpt-6.1 rows carry it, and the gpt-6 family joins gpt-5 in refusing frequency_penalty, presence_penalty, logprobs, and top_logprobs before the request reaches AWS.

* fix(bedrock): refuse GPT sampling and logprob params natively only while reasoning is on

On bedrock-runtime's native chat completions endpoint, GPT-5.x and GPT-6.x
accept temperature, top_p, frequency_penalty, presence_penalty, logprobs,
and top_logprobs once reasoning_effort is "none", and refuse them with any
other effort or when the effort is unset. The previous commit refused the
sampling params unconditionally from the cost map's supports_sampling_params
flag, which lost the reasoning-off case and never covered the penalties or
logprobs. The refusal now keys on the model being a GPT id and reasoning
being active, raises a 400 UnsupportedParamsError naming the params unless
drop_params drops them, and lets everything through under "none". Grok and
gpt-oss keep their unconditional family refusals.

* refactor(bedrock): keep the Converse route-prefix strip inside the bedrock llms module

* fix(bedrock): forward a non-string reasoning_effort on the native route instead of crashing

A list or dict reasoning_effort hit a frozenset membership test in
without_refused_reasoning_effort and raised TypeError, which the proxy
surfaced as a 500 APIConnectionError with no upstream call. The value is
now left alone unless it is a string Bedrock's native endpoint refuses,
so AWS answers the malformed value with its own 400 like it does for an int

* fix(bedrock): route overlong GPT version digits to Converse and send native chat completions to the runtime endpoint

A model id with more than 4300 version digits raised ValueError in the route check; the digits are now bounded so such ids fall back to Converse. The native chat completions URL now follows Converse's precedence: aws_bedrock_runtime_endpoint (or AWS_BEDROCK_RUNTIME_ENDPOINT) wins over api_base, so a deployment that sets both keeps sending to the same host

* fix(bedrock): route model_id overrides to Converse and never send an empty bearer natively

A deployment whose litellm_params carry model_id (an application inference profile or provisioned throughput ARN) went to the native Chat Completions route with the base model in the URL and model_id left in the body. It now takes Converse like the bedrock/arn:... model form, which encodes the override into the request URL

A blank api_key on a SigV4 deployment became an Authorization header reading Bearer with nothing after it on the native route, since the OpenAI-like header builder writes any non-None key and the signer keeps a non-AWS4 Authorization header. validate_environment now resolves the key through bedrock_bearer_token, so a blank key is signed with SigV4 the way Converse signs it

* test(bedrock): audit the native GPT chat completions route on the integration rig

Adds the /audit cells for the runtime chat completions route: the scripted Bedrock runtime peer, the happy and fallback wire tests, the sad-path and regex worst-case tests, the chaos burst tests, the Messages adapter tests, and the Responses native-route tests. Tests only, no product diff.

* test(bedrock): harden the runtime chat completions audit cells

The chaos peer's shared counter and process now come from the same spawn context, since a fork-context Value handed to a spawn-context process raises on Linux. The peer-kill test waits for the first six answers to reach the client before killing the peer instead of counting accepted requests. The Responses wire tests look the spend row up under both the ciphertext id the caller received and the issued id behind it, matching the chaos file's rule for the pre-encryption row

* fix(bedrock): refuse or drop a non-string reasoning_effort before the native chat completions call

A reasoning_effort sent as an int, a list, or an object on a GPT 5.6+ deployment the native
route serves now answers 400 from litellm before any wire request, naming the type and the
drop_params way out, and is dropped under drop_params so AWS applies its default effort, the
way Converse dropped it on main. The tip since a0cef91f0b forwarded it for AWS to refuse

* test(bedrock): pin router retries off and give the chaos bursts config deployments on an owned proxy

* test(bedrock): wait for the replacement worker before tearing down the sigkill chaos proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo <mateo@berri.ai>
2026-10-02 19:55:54 -07:00
devin-ai-integration[bot]
5dbe4f95e8
feat(bedrock): drop lookaround regex patterns from tool schemas for Converse models that reject them (#44138)
* feat(bedrock): drop lookaround regex patterns from tool schemas for Converse models that reject them

* fix(bedrock): rename the lookaround flag to supports_regex_lookaround and keep dropped patternProperties names allowed

The cost-map flag becomes a generic supports_regex_lookaround capability, which the
cost-map schema admits as a supports_* boolean, and the Converse transform now owns
the drop decision instead of the shared tools factory. A patternProperties key dropped
from an object closed by additionalProperties: false leaves its value schema as that
object's additionalProperties, so the names it allowed stay allowed, on the OpenAI
non-Python-regex drop too. tool_with_sanitized_parameters also cleans Anthropic-shape
tools (input_schema).

* fix(router): keep a deployment's supports_regex_lookaround off the shared cost-map entry

A deployment's model_info.supports_regex_lookaround was written to the shared
bedrock/<model> cost-map key, so every sibling deployment of that model id
inherited one deployment's choice. The flag now stays under the deployment's
own id, which is what the Converse lookaround check reads first, and the
shared entry keeps the cost map's value

* test(bedrock): audit the Converse lookaround drop on the wire across endpoints, SDKs, flags and chaos

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 19:47:25 -07:00