Commit graph

41664 commits

Author SHA1 Message Date
yuneng-jiang
02dd853bed
Merge pull request #36332 from BerriAI/litellm_backport_1_95_x_bp-195x-0809sec
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
chore(release): backport proxy request-handling maintenance and refresh runtime deps for 1.95.1
2026-08-08 19:07:43 -07:00
Yuneng Jiang
286062541c
chore(deps): raise the published aiohttp floor to 3.14.2 2026-08-08 18:58:49 -07:00
Yuneng Jiang
c47b4d6a29
chore(deps): bump msal to 1.37.0 2026-08-08 18:58:11 -07:00
Yuneng Jiang
31c49afb10
chore: refresh uv.lock for 1.95.1 2026-08-08 16:38:58 -07:00
Yuneng Jiang
7b3a9ed645
bump: version 1.95.0 → 1.95.1 2026-08-08 16:38:37 -07:00
Yuneng Jiang
5d328a387b
chore: update Next.js build artifacts (2026-08-08 23:38 UTC, node v20.20.2) 2026-08-08 16:38:04 -07:00
Yuneng Jiang
865fc20b7a
chore(deps): bump cryptography to 50.0.0 2026-08-08 16:34:06 -07:00
Yuneng Jiang
7b1c95b4fb
chore(deps): bump h2 to 4.4.1 2026-08-08 16:32:40 -07:00
Yuneng Jiang
1a748a245f
chore(deps): bump gitpython to 3.1.58 2026-08-08 16:32:15 -07:00
Yuneng Jiang
defd3a3714
chore(deps): bump aiohttp to 3.14.3 2026-08-08 16:31:56 -07:00
yuneng-jiang
47d20d9df5
Merge pull request #36011 from BerriAI/litellm_maint_batch_2026_07
fix(proxy)!: apply request-parameter checks consistently across body, path and form inputs

(cherry picked from commit c898d341c0)
2026-08-08 16:29:15 -07:00
yuneng-jiang
f4d17d5c0b
Merge pull request #35844 from BerriAI/litellm_/terraform-provider-dep-bump-5feb4a
chore(deps): bump grpc and golang.org/x modules in the terraform provider

(cherry picked from commit 2e255191ab)
2026-08-08 16:27:38 -07:00
yuneng-jiang
ecf07e1307
Merge pull request #35835 from BerriAI/litellm_/elated-margulis-7f300f
refactor(ui): route MCP session tokens through the shared storage helper

(cherry picked from commit e4fd790f1c)
2026-08-08 16:27:26 -07:00
yuneng-jiang
72a4a55f43
Merge pull request #35552 from BerriAI/litellm_/backport-35523-rc-1-95-0-f65d9e
fix(ui): land general login on the keys dashboard, send MCP consent to /ui/connect (backport #35523 to rc/1.95.0)
2026-08-01 17:35:25 -07:00
Yuneng Jiang
7d0963a296
chore: update Next.js build artifacts (2026-08-02 00:27 UTC, node v20.20.2) 2026-08-01 17:27:37 -07:00
yuneng-jiang
69f5fe35d8
Merge pull request #35523 from BerriAI/litellm_ui_login_no_mcp_landing
fix(ui): land general login on the keys dashboard, send MCP consent to /ui/connect

(cherry picked from commit ceaf556b2e)
2026-08-01 17:24:14 -07:00
yuneng-jiang
6753639325
Merge pull request #35414 from BerriAI/litellm_sync_rc_1_95_0
chore(release): sync rc/1.95.0 with the v1.95.0-rc.1 main SHA
2026-07-31 14:28:40 -07:00
Yuneng Jiang
27c6a4c4ca
chore(release): sync rc/1.95.0 with the v1.95.0-rc.1 main SHA 2026-07-31 14:15:18 -07:00
yuneng-jiang
9439174d9d
Merge pull request #35299 from BerriAI/litellm_backport_35271_rc_1_95_0
chore(release): backport #35271 to rc/1.95.0
2026-07-30 19:37:45 -07:00
yuneng-jiang
8b92b36573
revert(proxy)!: stop enforcing user budget on team keys (#35271)
Reverts #32005. Team-scoped keys are governed by the team and team-member
budgets only; the key owner personal max_budget no longer applies to them,
restoring the hierarchy that existed before that PR.

The skip_user_budget_on_team_key opt-out existed solely to turn the new
behavior back off, so it is removed along with the behavior: the
ConfigGeneralSettings field, the /config/list allowed_args entry that
surfaced it as an Admin UI toggle, and the argument threaded through
reserve_budget_for_request and _get_budget_counters.

Regression tests cover both enforcement points in the restored direction:
test_common_checks_personal_user_budget_skipped_for_team_key for the
read-time check and test_should_not_reserve_user_budget_counter_for_team_key
for the optimistic reservation path.

(cherry picked from commit 6f1625d23b)
2026-07-30 17:38:51 -07:00
yuneng-jiang
cad32fd9bc
Merge pull request #35049 from BerriAI/litellm_hotfix_35047_mcp_e2e_poll
Some checks are pending
CodeQL / Analyze (actions) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Helm unit test / unit-test (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
test(e2e): poll MCP tools across multi-worker lag (#35047)
2026-07-29 09:06:21 -07:00
mubashir1osmani
82fa66908b
test(e2e): poll MCP tools across multi-worker lag (#35047)
* fix(mcp): resolve call_tool by registry without requiring tool map

Multi-worker reloads put MCP servers in the registry from the DB but do
not re-run tools/list on every process. Gating call_tool on
tool_name_to_mcp_server_name_mapping made cold workers 500 with Tool not
found after another worker had already listed the tool. Treat a registry
match on server id/name/alias as enough; upstream rejects unknown tools

* test(e2e): poll MCP register, tools/list, and tools/call across multi-worker lag

Stage multi-worker gateways only load MCP servers and tool maps on the
process that handled the request. Poll until the server is listed, the
tool appears on tools/list, and tools/call is not a cold-worker 500 so
key-access and Datadog MCP e2e stop racing the LB

* Revert "fix(mcp): resolve call_tool by registry without requiring tool map"

This reverts commit 8b56e51e39.

* test(e2e): tighten MCP multi-worker lag classifier

Only retry tools/call on gateway shapes Tool <name> not found and
server_not_found, not any 500 that mentions tool/server not found, so
upstream failures are not retried until the poll deadline

* test(e2e): drop unit file for MCP lag classifier

The live await_call_tool polls already cover multi-worker lag; a separate
string-match unit module is not worth keeping

(cherry picked from commit c274cf321c)
2026-07-28 22:24:21 -07:00
yuneng-jiang
2cd62cfb83
Merge pull request #35020 from BerriAI/litellm_hotfix_e2e_model_servable_timeout
test(e2e): bound the post-/model/new servable wait at 40s
2026-07-28 17:58:14 -07:00
mubashir1osmani
87be33f935
fix(e2e): reject first listing that returns after the 40s deadline
A poll may start with remaining budget and still return after started+timeout
if the transport overruns its clamp. Recheck the first-listing deadline after
the response so a late listing does not open the continuous DB-sync phase

(cherry picked from commit 7ff2bcbf14)
2026-07-28 17:47:19 -07:00
mubashir1osmani
38d03fd341
fix(e2e): never skip the final deadline-clamped model-servable poll
When less than one full poll interval remained in the first-listing budget,
the pre-sleep check returned NotServable without another /v1/models call.
Sleep only min(interval, time left) so a model that becomes listable in the
last seconds of the timeout still gets a clamped final poll

(cherry picked from commit 8439195922)
2026-07-28 17:47:19 -07:00
mubashir1osmani
5953a66eab
test(e2e): drop proxy_client model-servable unit tests
Keep the create_model DB-sync wait in the harness; the pure-function unit
file is not needed for this PR

(cherry picked from commit 89204651d1)
2026-07-28 17:47:19 -07:00
mubashir1osmani
5aa66ea33e
fix(e2e): wait one default DB reload interval of continuous listing
create_model returned after the first /v1/models hit that listed the model,
so chat could still land on a cold gateway worker (numWorkers>1 / peer pod)
and 400 Invalid model name. Require continuous listing for the product
default add_deployment interval (30s) after first sight so every worker has
synced from the DB; first listing still bounded at 40s

(cherry picked from commit 7d1ee2ff86)
2026-07-28 17:47:19 -07:00
mubashir1osmani
e1afe2e29c
test(e2e): bound the post-/model/new servable wait at 40s
_await_model_servable used poll_timeout (120s), the spend/log read-back
budget. A stuck model reload therefore stalled every suite that creates a
deployment for two minutes before failing

Give create_model a fixed harness middle ground: model_servable_timeout=40s,
polled every 2s, with each /v1/models call capped at 5s and clamped to the
remaining deadline so one slow GET cannot overrun the wait. Happy path still
returns on the first listing. Not derived from proxy general_settings or env

Transport.get accepts an optional per-call timeout for that clamp. Unit tests
cover the deadline arithmetic and clamp without a live proxy

(cherry picked from commit c082a0e648)
2026-07-28 17:47:19 -07:00
yuneng-jiang
9ead580272
Merge pull request #34864 from BerriAI/litellm_internal_staging
Some checks failed
CodeQL / Analyze (actions) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
Helm unit test / unit-test (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
CodSpeed Benchmarks / benchmarks (push) Has been cancelled
chore(ci): promote internal staging to main
2026-07-28 16:05:20 -07:00
mubashir1osmani
7cd009caf7
fix(proxy): avoid DB outage during planned RDS IAM rotation (#34749)
* fix(proxy): warm rotate Prisma client for IAM refresh

* fix(proxy): drain Prisma operations during IAM rotation

* fix(proxy): bound the drain wait when retiring a replaced prisma engine

A replaced engine waited indefinitely for its drain tracker to empty.
Hung queries self-release via prisma's 30s default HTTP timeout, but a
transaction whose owner is hard-cancelled before commit/rollback leaks
its drain count forever, keeping the retired engine and its DB
connection pool alive indefinitely; at one rotation per 12 minutes such
engines accumulate. Cap the wait at 90 seconds, which exceeds every
legitimate operation bound (30s HTTP timeout, 60s max interactive
transaction timeout in this codebase), then kill the engine anyway.
Work killed at the deadline degrades to the pre-drain behavior and is
retried by the existing reconnect/backoff layers.

---------

Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-07-28 13:27:50 -07:00
yuneng-jiang
f4a68a75ff
feat(ui): mark Cost Optimization as beta in the left nav (#34984) 2026-07-28 12:04:45 -07:00
ryan-crabbe-berri
51ad1b0a57
test(e2e): skip passthrough headers test until stage can route custom paths to provider creds (#34980) 2026-07-28 11:36:09 -07:00
ryan-crabbe-berri
01ffd1296b
fix(e2e): poll for both spend rows before asserting the cache-hit contract (#34968)
The cache-hit and paid rows for the two driver calls flush from different
pods on independent update_spend timers, so waiting only for the cache-hit
row can return a half-arrived result set where the paid-row assertion then
fails on an empty list. Requiring both row kinds in the poll predicate lets
the existing deadline absorb the slower flush without weakening any assertion
2026-07-28 11:21:36 -07:00
ryan-crabbe-berri
b930e2fc2b
fix(gateway): route /a2a through the gateway component (#34958)
* fix(gateway): route /a2a through the gateway component

A2A message-send runs the completion bridge, an outbound LLM call, but the
ingress only listed /v1/a2a so the serving routes at /a2a/{agent_id} fell to
the backend catch-all. Backend pods hold no provider credentials, so every
invocation died with a missing-provider-key auth error while the same call
succeeds on the gateway fleet. Adds /a2a to the ingress gateway prefixes and
the gateway route allowlist, plus a parity test so an ingress prefix that the
gateway trims can never reappear

* revert(test): drop the allowlist parity tests

---------

Co-authored-by: yuneng-jiang <yuneng@berri.ai>
2026-07-28 10:22:49 -07:00
tin-berri
d91fd084f7
Fix cache leakage card layout to keep date picker on right (#34885)
* Fix cache leakage card layout to keep date picker on right and prevent content overlap

Removes flex-wrap and mt-3 to ensure date picker stays pinned to the right side of the card header regardless of zoom level, preventing it from covering card content below

* Remove overflow-hidden from Card to allow dropdowns and overlays to display fully

Fixes date picker dropdown being clipped when opened in cards like the Cache Leakage Card. By removing overflow-hidden from the Card container, popovers, dropdowns, and other overflow content can now display properly without being clipped by the card boundaries.

* Make cache leakage card descriptions consistent with line clamping

Adds line-clamp-2 to ensure both 'by model' and 'by virtual key' cards maintain consistent height. Removes conditional anthropic-specific text that caused height variations between dimensions.
2026-07-28 10:12:00 -07:00
ryan-crabbe-berri
daf22ec871
test(e2e): make MCP and prometheus e2e tests robust to data-plane sync lag (#34854)
* test(e2e): harden harness and tests against data-plane pod churn

A stage autoscaler scale-down produced a 2s window of ALB 502s that killed six
budget tests on their first management call, and a freshly scaled-up pod that
had not run its 30s DB object sync yet failed two MCP tests and one prometheus
cardinality test. Retry transient gateway errors (502/503/504, connection
errors) once at the shared e2e_http dispatch seam, poll MCP server registration
to the poll deadline instead of asserting a single-shot listing, anchor the MCP
guardrail full-sync wait to the later of the guardrail and server writes, and
turn the prometheus alias poll into a drive-and-scrape convergence loop that
re-sends traffic for missing aliases and unions results across scrapes

* test(e2e): drain request body in retry stub handler so keep-alive reuse cannot misparse leftovers as requests

* revert(e2e): drop the transient-502 retry seam

A raw 502 during a pod scale-down is what a real client sees, so the suite
retrying past it hides an availability gap instead of flagging it. The
gateway-side fix is graceful drain on the deployment; until then the failures
are signal

* test(e2e): cap per-alias driver re-drives in the prometheus cardinality poll

Bounds worst-case provider spend to 4 completions per alias while scrapes keep
polling to the deadline; counters persist on whichever pod served them, so the
cap costs no convergence unless that pod dies

* test(e2e): drop driver re-drives from the prometheus cardinality poll

The per-key cardinality contract is process-local and counters persist on
whichever pod served the driver call, so unioning aliases across free scrape
polls converges without re-sending billable traffic. The residual gap, a pod
dying inside the poll window, is deferred to direct per-pod scraping
2026-07-27 19:22:52 -07:00
mubashir1osmani
328e41b1f9
test(e2e): unblock the ui suite, fix the mcp registration race, park two known product bugs (#34853)
* test(e2e): let the ui suite run from a read-only cwd

The playwright suite never executed on stage. It died in globalSetup before a
single test ran, and the reported error was a red herring.

/app/e2e/ui is a read-only filesystem in the packaged e2e image (the image
runner already redirects playwright's own artifacts to TMPDIR for this reason),
but the suite wrote three things relative to cwd: the per-role storageState
files, the failure-screenshot directory, and the html report. Reproduced in the
pod: storageState raises EROFS, mkdir test-results raises ENOENT.

Worse, the catch block that exists to capture a screenshot threw its own ENOENT
while handling a failure, so the real login error was replaced by a filesystem
error. That is why the run looked like a missing directory rather than whatever
actually went wrong.

Route every artifact through ARTIFACT_DIR (E2E_UI_ARTIFACT_DIR, default "." to
keep run_e2e.sh behavior unchanged), make the diagnostic screenshot best-effort
so it can never mask the underlying failure, and point playwright's reporter and
outputDir at the same place so a bare `npx playwright test` works there too.

fixtures/users.ts had its own copy of the five storageState filenames; it now
re-exports the ones from constants so the paths have a single definition.

Verified in the read-only pod: both writes fail before, both succeed after.
85 tests enumerate and tsc --noEmit is clean.

Refs LIT-4821

* fix(e2e): create the ui artifact root before writing into it

storageState() does not create missing parents, and nothing created ARTIFACT_DIR
itself. Pointing E2E_UI_ARTIFACT_DIR at a writable path that did not exist yet
therefore failed with ENOENT on the very first role's snapshot, before any UI
test ran; the same class of failure the artifact-dir change was meant to remove,
just moved one level up.

Reproduced: writing admin.storageState.json into a missing directory raises
ENOENT. My earlier pod verification masked this because the probe called
mkdirSync itself, which the real code path never did.

mkdir the root once at the top of globalSetup, before the login loop. recursive
makes it idempotent, handles nested paths, and keeps the default "." a no-op.
Playwright creates its own outputDir lazily, so globalSetup is the only place
that needs this, and migration.serverRootPath.globalSetup delegates here so it is
covered too.

* test(e2e): skip the mid-conversation cache checks pending LIT-4873

A mid-conversation role="system" reminder invalidates the prompt cache on the
vertex_ai, azure_ai and bedrock_invoke Messages paths. Measured on the reminder
turn, same conversation shape throughout:

  direct to api.anthropic.com            7013 read  cache preserved
  litellm -> anthropic/claude-opus-4-8   7013 read  cache preserved
  litellm -> vertex_ai/claude-opus-4-8      0 read  cache destroyed

and the Vertex control with the same added assistant/user turns but no reminder
reads 7013, so it is the reminder on the non-first-party paths and not the extra
turns. Anthropic keeping the cache rules out provider behavior; litellm's
first-party anthropic path keeping it rules out the shared Messages transform.

That makes these assertions correct and the failure a real billing bug, so the
tests are skipped rather than weakened; the bodies stay intact and must be
restored unchanged with the fix. Registry rows are left in place, so the three
mid_conversation_system.nonstream.cache_hit cells report as uncovered gaps.

Skips are decorators rather than a pytest.skip() inside the shared helper: a
mid-function skip fires only after setup has already registered a real
deployment via /model/new and left the rest of the body unreachable.

Only Vertex was measured end to end. Azure Foundry and Bedrock Invoke are
inferred from matching nightly failures and should be confirmed with the fix.

Refs LIT-4821, LIT-4873
2026-07-27 19:20:57 -07:00
yuneng-jiang
3c0b1db633
test(e2e): realign Admin UI specs with the MCP dialog and keyless landing (#34870)
Both specs assert against UI that has since moved, so they fail on selectors
rather than on behavior.

The MCP discovery modal became a shadcn/Base UI dialog when mcp-servers
migrated off antd, so `.ant-modal` no longer matches it; locate it by its
dialog role instead. The create form below it is still an antd Modal and keeps
its existing locator.

The no-team internal user has no keys, and a keyless non-admin is now sent to
/ui/connect on the post-login landing, which has no sidebar. Wait for that
redirect to settle, then navigate to the keys page explicitly; the redirect is
gated on the ?login=success marker that the fresh navigation drops, so the
dashboard sticks and the rest of the test is unchanged.
2026-07-27 18:11:18 -07:00
yuneng-jiang
5c95017bc1
test: unstale the reasoning-effort grid count and the responses bridge test (#34868)
* test(reasoning-effort-grid): bump cell-count assertion for claude-opus-5

The claude-opus-5 grid entry added in ae81625ee6 raised the Anthropic direct
route to 31 model combos, but test_grid_cell_count still expected 30, so the
suite went red on the tripwire rather than on any behavior change.

* test(openai): swap the retired deep-research model out of the bridge test

OpenAI shut down o3-deep-research and o4-mini-deep-research on 2026-07-23, so
the live call in this test now comes back as a 400 'Model not found'. The test
was never about deep research specifically; the bridge fires on any model whose
cost-map mode is "responses", so it now uses gpt-5.5-pro, the newest
responses-only OpenAI model, and is renamed to say that.

gpt-5.5-pro was confirmed present on the CI account with an authenticated
GET /v1/models before being picked.
2026-07-27 18:07:04 -07:00
yuneng-jiang
f2cda740f7
chore: update Next.js build artifacts (2026-07-28 00:06 UTC, node v20.20.2) (#34859) 2026-07-27 17:20:51 -07:00
devin-ai-integration[bot]
bdf8f8c309
fix(guardrails): classify all 4xx HTTPException guardrail blocks as intervened (#33821)
Some checks are pending
CodSpeed Benchmarks / benchmarks (push) Waiting to run
UI Unit Tests / ui-unit-tests (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* fix(guardrails): classify all 4xx HTTPException guardrail blocks as intervened

* fix(guardrails): narrow HTTPException block classification to 400/403/422

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-27 16:50:13 -07:00
tin-berri
8f86c87f8e
Merge pull request #34564 from BerriAI/litellm_fix_router_registry_leak_on_edit
fix(router): release the pre-routing strategy slot when a deployment is replaced or deleted
2026-07-27 16:46:40 -07:00
tin-berri
09856a40cd
Merge pull request #34586 from BerriAI/litellm_lit4795_headroom_anthropic
fix(guardrails): compress content-parts messages in headroom guardrail (Anthropic traffic)
2026-07-27 16:12:45 -07:00
Mateo Wang
2a7885aee7
Merge pull request #34539 from BerriAI/litellm_fix_responses_bridge_streaming_contract
fix(responses_bridge): keep one chat completion id per stream and always stream completed responses
2026-07-27 15:41:06 -07:00
Tin Chi Lo
50bdf250f6 fix(router): repair deployment indices before releasing strategies on delete
delete_deployment resolved the outgoing deployment through get_deployment before
popping it, and ran the strategy release before repairing the index maps. Both
halves of that ordering could leave the router inconsistent. A resolution failure
meant the entry left the model_list with its registry slots still held, so the
alias stayed routable and the name could not be reused; a failure inside the
release meant the outer handler returned None with the entry already popped and
model_id_to_deployment_index_map never repaired, breaking every later lookup and
delete until a restart.

upsert_deployment already had this right: it pops, repairs the caches and indices,
and only then releases the slot. delete_deployment now follows the same sequence
and resolves the deployment from the item it just popped rather than through a
lookup that can fail. Releasing the slot is secondary to structural integrity, so
it runs last and a failure there is logged instead of abandoning a removal that has
already happened.
2026-07-27 15:32:27 -07:00
tin-berri
9bb75d67af
Merge pull request #34675 from BerriAI/litellm_tool_spend_rollup
fix(proxy): roll up tool spend daily instead of scanning SpendLogs
2026-07-27 15:31:19 -07:00
yuneng-jiang
38ea85b4bb
Merge pull request #32583 from BerriAI/litellm_/redact-langsmith-api-key-c92cc3
fix(proxy): sanitize per-key callback config out of logged metadata
2026-07-27 15:25:03 -07:00
ryan-crabbe-berri
0171170fc7
fix(ui): validate default team values in Default User Settings (#34815)
* fix(ui): validate default team values in Default User Settings

The Default User Settings form accepted any free-text team id, and the
proxy persisted it without checking the team exists. New users were then
silently never added to the default team because the consume-time 404
from team_member_add was swallowed at debug level.

Backend: PATCH /update/internal_user_settings now rejects unknown and
duplicate team ids with a 400 naming them, before any persistence or
team budget side effects. Team-add failures in _add_user_to_team now log
at ERROR with user and team ids.

UI: DefaultUserSettings rewritten as a shadcn + react-hook-form + zod
form following the org-settings pattern. The team id free-text input is
replaced with a searchable server-backed team picker, so only existing
teams can be selected; zod blocks empty and duplicate rows. The shared
deriveErrorMessage helper now unwraps the HTTPException detail.error
shape so backend validation errors surface readably in toasts.

* fix(ui): restore read-only view with Edit Settings toggle on default user settings

Parity with the pre-migration form: the tab renders a read-only summary
of the saved defaults, Edit Settings opens the RHF form, Cancel discards
pending edits and returns to the summary, and a successful save returns
to the summary showing the new values. Model sentinel labels in the
summary are derived from ModelSelect's now-exported special values
instead of duplicating the strings.

* refactor(ui): rename MODEL_SELECT_SPECIAL_VALUES_ARRAY to MODEL_SENTINEL_OPTIONS

* fix(ui): move Edit Settings into the card header action slot
2026-07-27 15:05:51 -07:00
Tin Chi Lo
47a0c22f64 fix(router): rebuild the adaptive companion when an upserted complexity router participates in adaptive routing
The finalize re-run in upsert_deployment keyed off the auto_router/adaptive_router
prefix only, so editing a complexity router with adaptive enabled released its
adaptive_routers entry (and post-call hook) without rebuilding it: complexity
routing kept serving while bandit recording, DB persistence and
/adaptive_router/state went silently dark until the next full reload. Gate the
re-run on a participation predicate that mirrors both arms of the finalize pass,
drop the import that pass no longer uses, and pin the registry helpers with
direct contract tests
2026-07-27 14:54:39 -07:00
mateo-berri
4299c6d191 fix(responses-bridge): return CustomStreamWrapper from the completed-response stream helper 2026-07-27 14:48:18 -07:00