Commit graph

53373 commits

Author SHA1 Message Date
devin-ai-integration[bot]
54ae4c5bbf
fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (#43082)
Some checks are pending
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rerun integrations shard after unrelated gitlab prompt manager timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit azure_storage client reuse against a local Data Lake sink

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): cover the exact TTL expiry boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): wait for a rejected write before flipping the sink back

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): add azure_storage log delivery cells behind an opt-in lane

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): restart the proxy mid burst and bound the loss to the unflushed queue

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): drop the redundant stop after the owned proxy exits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): read azure_storage objects at the auth-mode-dependent layout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): install the datalake sdk in the e2e lint environment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure_storage): drop the opt-in real Azure e2e cells and their e2e-dev dependency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-30 18:50:15 -07:00
yuneng-jiang
431ecd8920
chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (#43961)
gitpython 3.1.62 (2026-09-07) and tornado 6.5.10 (2026-09-15) are past the
3-day uv cooldown. diskcache still has no fixed release, so its ignore moves
from 2026-10-01 to 2026-11-01
2026-09-30 18:27:06 -07:00
devin-ai-integration[bot]
a3a7650569
fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (#43956)
* fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): patch the shared proxy logger directly in the straiker api_version test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): keep the straiker stray-version block marker separate from the shared block marker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): drop redundant comments on the straiker api_version integration tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 18:20:57 -07:00
devin-ai-integration[bot]
2c3866ebb4
fix(azure_storage): name Data Lake objects without base64 padding or slashes (#43914)
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 18:17:13 -07:00
devin-ai-integration[bot]
6997223068
fix(grayswan): send request conversation and tool calls to post-call monitor (#43770)
* fix(grayswan): send request conversation and tool calls to post-call monitor

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(grayswan): tighten post-call context typing and wire test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): resolve post-call surface from request route before call_type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): omit tools from post-call monitor when request context is empty

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(grayswan): apply ruff format to post-call context changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): merge response text and tool calls into one assistant monitor message

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): only merge tool calls into the response text for single-choice responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): audit post-call context across endpoints, modes and outages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): share the upstream model probe reply across audit responders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): assert the full generic guardrail body and kill a real serving worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): normalize the client user agent in the generic body assert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): normalize accept-encoding in generic body assertion

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): keep volatile header placeholders only when the header is present

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): capture monitor calls immutably in the unit test client

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): type the test helper parameters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 17:52:53 -07:00
berriai-litellm-provider-info-sync[bot]
8a1f3568ba
chore(cost-map): sync openrouter prices from the models API (#43950)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 17:48:48 -07:00
berriai-litellm-provider-info-sync[bot]
38b0762992
fix(wandb): set supports_vision true on GLM-5.3-Flash (#43951)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 17:35:36 -07:00
yuneng-jiang
f553c80cd3
test(proxy): scope user_api_key_auth overrides in proxy_server tests (#43952)
TestPriceDataReloadAPI, TestPriceDataReloadIntegration and TestInvitationEndpoints set
app.dependency_overrides[user_api_key_auth] and never removed it. Under xdist the shared
proxy app kept the override, disabling auth for later tests on the same worker and failing
test_harness_smoke.py::test_auth_as_cleans_up_on_exit. Set it through monkeypatch.setitem
so it is undone at teardown

A per-test leak check over every tests/test_litellm/proxy file that touches
dependency_overrides found these 20 tests as the only leakers; it reports none after this change
2026-10-01 00:27:53 +00:00
devin-ai-integration[bot]
ed4caebb65
fix(anthropic): forward the dangerous-tool-use beta to Azure AI Foundry (#43934)
Map dangerous-tool-use-2026-09-03 for azure_ai in the beta header config so the
Claude Code auto mode beta reaches Foundry instead of being stripped, matching
the anthropic, bedrock, and vertex_ai entries

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-30 16:40:11 -07:00
yuneng-jiang
f8f05767da
test(ci): refresh qualified retired OpenAI fixtures (#43938)
* test(ci): refresh qualified retired OpenAI fixtures

* test(ci): compare fallback input usage instead of provider wording
2026-09-30 16:10:18 -07:00
devin-ai-integration[bot]
c42d06fb80
fix(router): bill service tiers at catalog rates for custom-priced deployments (#43890)
* fix(router): inherit catalog service-tier rates for custom-priced deployments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): apply tier-suffixed long-context rates when only tier thresholds are set

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): cover canonical cost-map backend model resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): use descriptive names for service-tier pricing fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 16:06:37 -07:00
moe-berri
6fd9334751
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main

Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.

* feat(lens): add trace analysis and standalone worker

* fix(lens): clarify review limits and finalize main integration

* fix(lens): simplify worker setup and show the next check

* fix(lens): simplify analyzer setup and resolve integration failures

* fix(lens): preserve durations and evidence from later trace reads

* fix(lens): trust server context for internal analysis exclusion

* fix(lens): pin reviewed analyzer image and verify request inclusion

* test(lens): select time units before entering custom duration

* test(lens): allow the standalone analyzer lifetime HTTP client

* test(lens): run analyzer tests in active proxy coverage shard
2026-09-30 22:42:09 +00:00
devin-ai-integration[bot]
c51d5b12ac
test(bedrock): restore the AWS env after a failed live call in the auth tests (#43921)
* test(bedrock): restore the AWS env after a failed live call in the auth tests

* test(bedrock): assert the regression test's failing call actually ran

* test(bedrock): drop the test that tests the auth tests

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-30 22:36:57 +00:00
ryan-crabbe-berri
632b69b5c8
refactor(proxy): answer every team access check with TeamAccess.allows (#43364)
* refactor(proxy): route every team-admin decision through auth/team_access.py

Move the six team-admin helpers out of common_utils, team_endpoints and
key_management_endpoints into litellm/proxy/auth/team_access.py under public
names, and point every management route and helper at them. The key routes
keep checking team admin before org admin, so a team admin whose user row is
gone still passes as before. Status codes and bodies are unchanged, which the
223-case team-admin matrix confirms at the merge base and at the tip

common_utils keeps `_is_user_team_admin` as an alias because the published
litellm-enterprise 0.1.71 wheel still imports it from there

* refactor(proxy): answer every team access check with TeamAccess.allows

Replace the six helpers in auth/team_access.py with one resolver in
litellm/proxy/management/teams/access.py. Each route passes the roles it
accepts (TEAM_OR_ORG_ADMIN or TEAM_ADMIN_ONLY), and /team/update and
/team/info rank roles through strongest_role so org admin still outranks
team admin there

The org lookup moves behind an OrgRoles protocol, implemented by
PrismaOrgRoles in management/users/service.py, and get_team_access in
management/teams/dependencies.py is the only place that reads proxy_server
globals. _check_key_admin_access keeps its name and body from main

Routes that checked org admin first now read the roster first, so a team
admin whose org lookup errors now passes on /team/delete, /team/block,
/team/unblock, member reset_spend and reset_budget, and the team callback
routes. No allowed caller is denied
2026-09-30 15:27:33 -07:00
joshua-berri
52b9fa2ba1
feat(agents): add identity registration and dashboard controls (#43723)
* feat(agents): identity registration and dashboard

* fix(agents): preserve retired identity ownership

* fix(agents): preserve configuration during identity updates

* fix(agents): retain intentional card edits in the dashboard

* fix: remove mutable agent identity registration constructions

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-30 15:21:21 -07:00
ryan-crabbe-berri
f39c811d34
fix(packaging): keep wheel paths under Windows MAX_PATH for Store Python (#43903)
* fix(packaging): keep wheel paths under Windows MAX_PATH for Store Python

pip install litellm fails on Microsoft Store Python because its user
site-packages is already 134 chars plus the profile name, and the content
filter guardrail ships YAML five directories deep under
litellm/proxy/guardrails/guardrail_hooks/litellm_content_filter/. The
existing wheel guard assumed a 100-char install prefix, so it never saw it.

Move categories/ and policy_templates/ to
litellm/proxy/guardrails/content_filter_data/ and drop the benchmark
fixtures from the wheel. Old category_file paths keep resolving because
the resolver only keys on the trailing categories/<file> or
policy_templates/<file> suffix.

Derive the guard's worst-case prefix from the Store Python site-packages
path with a 15-char profile name (149), fail files at 260 and directories
at 248 (CreateDirectoryW), and fix the off-by-one that let a 260-char
path through.

Fixes #43851

* ci: run the Windows wheel install guard on pull requests

The two Windows jobs live in CircleCI, which never runs on pull requests,
so nothing installs the wheel on Windows before merge. Add a GitHub Actions
job on windows-latest that builds the wheel and runs the guard.

Two things make the run deterministic instead of image dependent. The job
turns the LongPathsEnabled registry key off first, because runner images
ship with it on and python.exe is long-path aware, so a 300-char path would
install fine. The guard installs with pip instead of uv, because uv writes
files from Rust, which switches to extended-length paths on its own and can
never hit MAX_PATH.

* fix(guardrails): keep the old content filter package dir as a category search root

Deployments that copied their own category YAML into
guardrail_hooks/litellm_content_filter/ before the data move would have had
that file rejected by the new directory jail and missing from by-name loads,
inherit_from lookups, the UI category listing and the category YAML endpoint.
Every lookup now searches the bundled data dir first and the old package dir
second, with the bundled copy winning on a name clash.

* fix(guardrails): resolve category files through safe_join

By-name category lookups and the suffix search in the category_file resolver now go through safe_join, so a name or suffix that would escape its data root never reaches the filesystem. The LITELLM_CONTENT_FILTER_ALLOW_EXTERNAL_PATHS opt-out keeps its unjailed search. Clears the two CodeQL path-injection findings on the new lookup code.

* fix(guardrails): keep symlinked category files loadable by name

By-name category lookups resolved symlinks through safe_join, so a category file symlinked into the categories folder from elsewhere stopped loading. Those lookups now only reject names that leave the folder lexically and return the link untouched, matching how by-name loads behaved before the data move. The category_file resolver keeps its realpath jail as before.

* fix(guardrails): keep the category viewer inside the category folders

GET /guardrails/ui/category_yaml/{name} hands raw file contents to any valid key, and on main it refused a symlink whose target left the categories folder. The previous commit let by-name lookups follow symlinks again, which also let the viewer read whatever a symlink in a legacy categories folder pointed at. The viewer now checks the found file's real path against every categories folder it searches and answers 400 as before, while the guardrail's own by-name loads keep following symlinks

The roots come in through a FastAPI dependency so the check is testable against a temp folder, and the content filter's realpath containment moves to path_utils.is_within so both surfaces share it. The test that patched os.path.commonpath covered a branch that no longer exists and goes with it

* ci: drop the Windows wheel install job from pull requests

The job took about 13 minutes on every PR to guard an edge case. The
guard still runs its path-length check on Linux in base_sdk_install and
on Windows in the CircleCI windows_release_wheel job.
2026-09-30 22:20:36 +00:00
yujonglee
629c2b5808
feat(tracing): store spend in ClickHouse automatically (#43928) 2026-09-30 22:17:43 +00:00
devin-ai-integration[bot]
b41715b0c2
fix(transcription): honor base_url alias for Groq Whisper and report it as the api base (#43917)
* fix(transcription): honor base_url alias for Groq Whisper and report it as the api base

transcription() and speech() only accepted api_base, so a deployment configured with base_url leaked the alias into the provider params, which Groq rejected as an unknown param, and the request never reached the internal gateway. get_api_base() now reads the same alias so response headers and logs show the configured endpoint instead of the provider default

Resolves LIT-9071

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(speech): keep base_url after existing audio params, route Vertex speech to it, skip empty alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 15:03:23 -07:00
devin-ai-integration[bot]
e662772ad1
fix(proxy): register a UI-configured arize callback next to otel under OTel v2 (#43906)
Callbacks saved from the Admin UI reach litellm through _add_custom_logger_callback_to_specific_event, which skipped the freshly built logger whenever a callback of the same exact class was already registered. With LITELLM_OTEL_V2 enabled every OTel preset (otel, arize, ...) is an OpenTelemetryV2, so a UI-added arize was treated as a duplicate of the yaml otel callback and never attached, and no trace ever reached Arize

The exists check now compares the exact class and the logger's callback_name, so presets that share a class register side by side while a true re-registration of the same preset is still skipped

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 14:58:40 -07:00
devin-ai-integration[bot]
fc8f3a26bb
fix(proxy): relay Azure passthrough body model groups through the router (#43896)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 14:57:31 -07:00
yuneng-jiang
a6f6c64b6e
feat(ui): adopt the new LiteLLM logo and monogram (#43913)
* feat(ui): adopt the new LiteLLM logo and monogram

Swap the bundled admin UI logos for the new brand assets: the primary
logo in blue for light mode and white for dark mode, and the monogram for
the collapsed sidebar, favicons, and the built-in guardrail cards.

/get_image gains a variant=monogram query parameter so the collapsed
sidebar can request the monogram while admin-configured UI_LOGO_PATH /
UI_LOGO_PATH_DARK logos still take precedence. The bundled light logo
moves from JPEG to a transparent PNG.

* test(ui): query collapsed sidebar logos by role to stay within the lint budget

* fix: point remaining logo consumers at the new bundled assets

The Rust gateway UI served /get_image from the removed litellm_logo.jpg,
the non-root get_image tests pinned logo.jpg, and a cookbook script read
litellm/proxy/logo.jpg. Point them at the monogram and logo.png.

* fix(mcp): serve the BYOK OAuth page logo from /get_image

The page pointed at /ui/assets/logos/litellm_logo.jpg, which the rebrand
removes from the dashboard sources, so the next UI build would drop it.
/get_image?variant=monogram is always served by the proxy and follows any
admin-configured logo.

* fix(ui): invert the LiteLLM monogram on dark guardrail cards

The blue monogram has a transparent train cut-out, so on a dark card it
read as a muddy blue block. Inverting it yields the brand's white mark,
which the logo guidelines prescribe for dark backgrounds.

* fix(gateway-ui): serve theme and variant aware logos from the dashboard export

The Rust gateway served one monogram for every /get_image request, and the
committed export lacked it, so /get_image returned 404 until the next UI
release build. Pick the full or monogram logo in light or dark from the
query, ship those assets in the dashboard's public dir and the committed
export, and drop two comments that restated asserted paths.
2026-09-30 14:51:06 -07:00
devin-ai-integration[bot]
cbe69723b1
test(s3_v2): pin async 5xx retry through the production AsyncHTTPHandler (#43080)
* test(s3_v2): pin async 5xx retry through the production AsyncHTTPHandler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Update tests/unit/integrations/test_s3_v2.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: yucheng-berri <yucheng@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-30 14:30:27 -07:00
yuneng-jiang
e61733b170
test(e2e): align completion, SAIL, and spend-log fixtures with supported contracts (#43902) 2026-09-30 14:21:22 -07:00
devin-ai-integration[bot]
6b9766fa0c
feat(proxy): add native ROI calculator for gateway spend vs merged PRs (#43669)
* feat(proxy): add native ROI calculator for gateway spend vs merged PRs

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): serialize ROI Prisma inputs with builtin containers

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* style(proxy): format ROI calculator backend files

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix: parse fenced ROI estimates and retain completed reports

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi-calculator): correct estimator and dashboard behavior

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* chore(ui): drop next dev generated AGENTS.md block

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): chunk ROI spend user lookup

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ui): show reused ROI estimates after sync

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(security): address ROI CodeQL alerts

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(proxy): make ROI calculator unit tests discoverable

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(ci): run ROI calculator tests in proxy infra shard

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi): page repository search and recover polling errors

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* feat(roi): bring scheduled analysis and guided setup into the gateway

* fix(roi): recover interrupted syncs and resolve review findings

* fix(roi): preserve cached estimates across report scope changes

* fix(roi): normalize scheduler timestamps to UTC

* fix(roi): fence cancelled syncs and read reports from writer

* fix(roi): preserve reports during metadata outages

* refactor(roi): isolate outage validation and verify uncached retry

* fix(roi): make scheduled job registration repeatable

* style(roi): format scheduler import

* fix(roi): continue syncing accessible repositories

* fix(roi): preserve reports and identity during upstream outages

* fix(roi): persist refreshed identities for reused estimates

* perf(roi): skip writes for unchanged cached identities

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-authored-by: moe-berri <moe@berri.ai>
2026-09-30 14:16:41 -07:00
yujonglee
268eb4d6e6
feat(tracing): add OTLP trace ingestion and reads (#43915) 2026-09-30 21:12:29 +00:00
devin-ai-integration[bot]
f285229b51
fix(proxy): delete large teams without per-member transaction fan-out (#42998)
* fix(proxy): delete large teams without per-member transaction fan-out

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): evict email-only member caches and reset team members metric on delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep new delete-team literals within the LIT002 ceiling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve deleted-team member ids before the locked delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup

`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.

* fix(repositories): slice find_by_emails into bounded IN statements

The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.

* fix(proxy): delete a team once when /team/delete repeats its id

The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.

* test(integration): audit cells for /team/delete on large, legacy and concurrent teams

Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.

Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.

Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.

* test(integration): pin each chaos outage to a live /team/delete

The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 13:49:18 -07:00
devin-ai-integration[bot]
657bb777fa
fix(hosted_vllm): keep reasoning_content on replayed assistant messages (#43599)
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages

vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm
chat templates consume it, so popping it made reasoning models lose
earlier reasoning across tool loops. thinking_blocks is still removed
for vLLM compatibility.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(hosted_vllm): forward replayed reasoning_content only when it is a string

* test(integration): cover hosted_vllm reasoning_content replay across endpoints

* test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-09-30 13:12:28 -07:00
joshua-berri
405ed414cb
feat(agents): authenticate Entra identities and delegated requests (#43722)
* feat(agents): authenticate Entra identities and delegated requests

* fix(agents): enforce target policy and preserve trusted authentication

* fix(proxy): make inference model selection exhaustive

* fix(agents): authorize targets against current database policy

* fix(proxy): make exhaustive model resolution return explicit

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-30 13:05:22 -07:00
shrey-berri
2bf0cddcc1
fix(bedrock): keep applicable beta headers (#43829) 2026-09-30 13:04:49 -07:00
devin-ai-integration[bot]
1fa3cde6a2
fix(traces): correct ClickHouse rollup partitioning, dedupe keys, and retention changes (#43901)
* fix(traces): correct ClickHouse rollup partitioning, dedupe keys, and retention changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(traces): pin spend dedupe timestamps within one second

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 19:45:11 +00:00
devin-ai-integration[bot]
50f5cc9bbb
feat(otel v2): excluded_services opt-out for datastore spans on tenant destinations (#43278)
* feat(otel): excluded_services opt-out for datastore spans on tenant destinations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep upstream support unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): excluded_services resolves from the otel callback config only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): name the otel callback logger so excluded_services owner lookup matches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): assert no aux datastore traces reach the tenant sink

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): read bogus-start proxy log from the results dir

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): read only this invocation's bogus-start proxy log

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): assert operator kept db spans over the whole recorded window

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): split operator db-span asserts by trace scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): build the otel logger after preset callbacks and validate the exclusion env at boot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): tolerate a bogus exclusion env when callback config wins

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep bogus exclusion env fatal when a preset parses it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): hoist the preset check out of the callback loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): use a rule-scoped pyright suppression

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): log and drop unknown excluded_services instead of failing boot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for the operator spend-writer span before checking the tenant for postgres

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): leave callback init and boot untouched when excluded_services is unset

Read callback_settings.otel.excluded_services directly instead of making the otel callback build its own logger, and drop the new boot-time parse of callback_settings.otel, so a proxy without the setting behaves exactly as on main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): normalize callback_settings excluded_services without rereading OTel env vars

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): log and ignore malformed excluded_services instead of failing startup

Lowercase and trim names, drop non-string items, and add an integration matrix over endpoints, clients, cache hits, destination outages and setting shapes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): cover failed upstream calls in the excluded_services matrix

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): pin operator Langfuse credentials in preset-only excluded_services tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
2026-09-30 12:20:27 -07:00
ishaan-berri
e78845afcf
feat(ui): agent traces tab on logs with timeline and otel setup guide (#43891)
* feat(ui): add agent trace api calls

* feat(ui): add agent trace types

* feat(ui): add span tree row types

* feat(ui): build span tree rows, grouped per agent

* test(ui): cover span tree rows, grouping and previews

* test(ui): add research trace fixture

* test(ui): add swarm trace fixture

* test(ui): add deep agent trace fixture

* test(ui): add trace list fixture

* feat(ui): fetch agent traces with a rolling live window

* feat(ui): look up a span's request log at its own time

* test(ui): cover rolling trace window and span log lookup

* feat(ui): add status mark for spans

* feat(ui): add duration bar for span timeline

* feat(ui): add copy button

* feat(ui): render span content as messages

* feat(ui): open a span's request log in place

* feat(ui): show raw span attributes

* feat(ui): add span detail pane

* test(ui): cover span detail pane

* feat(ui): span tree with timeline and keyboard nav

* feat(ui): run view with back out of load errors

* test(ui): cover the run view and initial selection

* feat(ui): add runs table toolbar

* feat(ui): add runs table

* feat(ui): add runs timeline with drag to zoom

* test(ui): cover runs timeline bucketing and drag

* feat(ui): add time range and live controls

* feat(ui): agent traces section with timeline and range

* test(ui): cover agent traces section

* feat(ui): add agent traces page

* feat(ui): otel setup guide for agent traces

* test(ui): cover agent traces setup guide

* feat(ui): add agent traces preview image

* feat(ui): add langgraph logo

* feat(ui): add langchain logo

* feat(ui): add openai agents logo

* feat(ui): add crewai logo

* feat(ui): add pydantic ai logo

* feat(ui): add llamaindex logo

* feat(ui): add opentelemetry logo

* feat(ui): add agent traces tab to logs

* test(ui): cover agent traces tab on logs

* feat(ui): collapse sidebar on logs for a full screen view

* test(ui): cover sidebar collapse on logs
2026-09-30 19:10:27 +00:00
yujonglee
41df8cf4d0
feat(traces): add Rust storage foundation (#43819)
* wip

* feat(traces): establish shared Rust storage foundation

* fix(traces): escape ClickHouse text parameters

* test(traces): exercise response cap with bounded strings

* fix(traces): remove unnecessary lint expectation

* fix(traces): encode ClickHouse timestamp units in Rust

* test(traces): mark exception match as a regex

* refactor(traces): execute schema setup in Rust

* refactor(traces): use shared logging execution wrapper

* docs(traces): replace foundation README with boundary rules

* fix(traces): use current bridge execution facade

* fix(traces): account for protocol cast in lint budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 12:00:00 -07:00
devin-ai-integration[bot]
264b09ac8d
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails

The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.

Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.

Resolves LIT-8931

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): reject empty guardrail rewrites instead of forwarding raw input

An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the texts-replacing guardrail helper explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover skip_system_message_in_guardrail on the live proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system

Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): annotate new guardrail tests with return types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the guardrail test doubles explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:44:35 -07:00
devin-ai-integration[bot]
253627f484
fix(ui): render access group MCP and agent selections as wrapping chips (#41228)
* fix(ui): render access group MCP and agent selections as wrapping chips

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): cover 20 selected MCP servers rendering as separate chips

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): cover MCP and agent chip selection in access group create dialog

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 18:43:32 +00:00
devin-ai-integration[bot]
3930c5bab6
fix(proxy): strip caller credentials from websocket passthrough (#43855)
* fix(proxy): strip caller credentials from websocket passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover configured x-api-key in websocket passthrough credential test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:22:25 -07:00
joshua-berri
79756cbb9b
feat(agents): enforce authoritative agent permissions (#43721)
* feat(agents): authoritative permissions

* fix: enforce authoritative managed agent permissions

* fix(agents): only consult the identity store for managed targets

is_agent_allowed entered the identity-store path whenever a prisma client
was configured, so an ordinary agent paired with an internal user returned
503 instead of 200. Classify the target from the registry first and fall
back to the store only when the registry has no entry, so an unmanaged
target never depends on the store being reachable.

* fix(agents): gate the managed path on an admitted policy object

Ten call sites branched on `managed_agent_policy is not None`, which any
MagicMock attribute satisfies, so the managed path fired on unmanaged
subjects and died in Pydantic validation as a 503. Route every check
through a shared helper that requires a real AgentResponse.

* test(mcp): stub the writer replica the fresh-policy reads use

reload_admitted_user now passes check_db_only through to get_user_object,
so the user row is read from writer_db. Point the mocks at the replica the
code actually reads and give each parametrized case its own user id.

* fix(agents): cap a managed agent at the invoking team's agents

resolve_agent_access returned the managed policy's grants before the
agent_caller ceiling was applied, so a managed agent acting on behalf of a
user reached agents that user's team was never granted. Intersect with the
caller ceiling the unmanaged path already honours.

* fix(agents): restore token narrowing and scope the private-access suppressions

The managed-model check lost its valid_token narrowing when it moved to the
shared helper. Make the caller-access resolver public rather than reaching
into it from module scope, and give each remaining private access a reason.

* docs(agents): drop the comment claiming admins skip the A2A permission check

The check has never had an admin bypass on this path, so the comment
described behaviour the code does not implement.

* test(proxy): stub the writer reads and restore the MCP manager singleton

Fresh-policy user lookups read writer_db, so the team and rest-endpoint
mocks stubbed a replica the code no longer reads, and the dashboard
session fake still had the pre-kwarg signature. The manager reload also
rebound global_mcp_server_manager in every MCP module without restoring
it, leaking an empty manager into later files.

* style: sort imports under the litellm package ruff config

* fix(mcp): cap a managed agent's servers and tools at the invoking caller

managed_agent_servers and managed_agent_tools returned the agent's own
grants without the agent_caller ceiling the unmanaged resolvers apply, so
a managed agent reached MCP servers and tools the echoed caller could not.
Call the existing ceiling helpers on both axes.

* refactor(mcp): return the caller-capped tools without an interim list

The ceiling helper already returns a sequence, so materializing it into a
list added a mutable collection for nothing. Sort at the return sites
instead, which also makes the tool order stable across both branches.

* fix(agents): preserve actor ceilings during managed target checks

* fix(agents): keep managed permission ceilings authoritative

* fix(mcp): fail closed on authoritative caller team outages

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-30 11:11:37 -07:00
ryan-crabbe-berri
0736a14326
fix(ui): right-align money and count columns across tables (#37889)
* fix(ui): right-align money and count columns across tables

Make numeric the one alignment token for the three shared table
wrappers (DataTable, MemberTable, SimpleTable) via NUMERIC_CELL_CLASS,
and flag every money, cost, and bare-count column that was still
left-aligned. Raw ui/table usages that render spend or budgets get the
same class on their header and cell.

* test(ui): render the organizations alignment test with providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): query alignment cells by role instead of DOM traversal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 10:59:53 -07:00
yucheng-berri
7204942756
fix(proxy): keep request-body credentials out of stored spend-log requests (#43635)
* fix(proxy): keep request-body aws credentials out of stored spend-log requests

* fix(proxy): redact every credential-named request-body field in stored spend-log requests

Replace the hard-coded AWS key check in the spend-log request-body sanitizer with
SensitiveDataMasker's key classification, so Azure, Vertex, watsonx, OCI, GigaChat,
Gemini and header credentials are redacted too. Proxy-stamped key identity metadata
is kept.

* fix(proxy): keep request identifiers named like keys in stored spend-log requests

* refactor(proxy): drop the AWS-only snapshot exclusion now that spend-log redaction is name-based

* refactor(proxy): use SensitiveDataMasker's key classification without an exclusion list

* refactor(proxy): always redact credential-named fields in stored spend-log payloads
2026-09-30 17:31:48 +00:00
berriai-litellm-provider-info-sync[bot]
025292e75b
feat(pricing): add vertex_ai gemini-3.8 flash tts rows (#43876)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 10:20:57 -07:00
yucheng-berri
e5c74cb2a6
fix(cost_calculator): stop copying optional_params into response hidden params (#43637)
* fix(cost_calculator): stop copying optional_params into response hidden params

* test(cost_calculator): assert the stored spend-log request and logging payload carry no forwarded credentials
2026-09-30 10:17:21 -07:00
devin-ai-integration[bot]
2ed9761921
fix(router): strip encrypted reasoning the pinned deployment cannot decrypt (#43781)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 09:56:27 -07:00
devin-ai-integration[bot]
82d8b3797c
fix(proxy): attribute completed batch cost rows to /batches in daily activity (#43870)
* fix(proxy): attribute completed batch cost rows to /batches in daily activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): wait for priced batch tokens before asserting team endpoint activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 16:16:52 +00:00
berriai-litellm-provider-info-sync[bot]
efdccd8811
chore(cost-map): add fireworks priority prices for ember-1, nemotron and glm 5.3 us rows (#43811)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 09:08:16 -07:00
berriai-litellm-provider-info-sync[bot]
971e60660b
chore(cost-map): add openai gpt-image-2.5 batch prices from the pricing page (#43869)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 09:06:31 -07:00
devin-ai-integration[bot]
b80052839e
refactor(rust): centralize bridge execution wrappers (#43871)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 15:54:27 +00:00
devin-ai-integration[bot]
9dda4d895f
fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (#43764)
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
2026-09-30 07:32:16 -07:00
devin-ai-integration[bot]
b781d157d7
chore(model_prices): add Gemini Veo, Mistral and Azure Claude 4.5 deprecation dates (#43857) 2026-09-30 07:26:17 -07:00
devin-ai-integration[bot]
b370996b9d
test(router): settle the shared logging worker before recording shadow callbacks (#43847)
* test(router): settle the shared logging worker before recording shadow callbacks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): always stop the shared logging worker after settling it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 04:44:43 -07:00
devin-ai-integration[bot]
6bc17f98d7
refactor: clean up fresh tech debt from 2026-09-29 (#43830)
* refactor: clean up fresh tech debt from 2026-09-29

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(routing): pin usage-based routing Redis reads through the proxy and SDK

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 02:40:37 -07:00