The .npmrc file (ignore-scripts=true, min-release-age=3d) is temporarily
removed during the Docker build since lifecycle scripts are needed by
npm ci. However, the unconditional `mv` fails when the build context
doesn't include .npmrc (e.g. when LiteLLM is vendored in a subdirectory).
Make all .npmrc mv operations conditional. This is safe because npm ci
already installs from package-lock.json with pinned versions and
integrity hashes.
PR #25258 changed _cleanup_stale_managed_objects from update_many to
execute_raw via _expire_stale_rows, but the tests were not updated.
The tests now mock _expire_stale_rows on the instance and assert
update_many calls only for job completion, not stale cleanup.
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
- team-admin: assert Admin Settings is not visible (role-specific check)
- proxy-admin: use users[Role.ProxyAdmin].password from constants instead of duplicating the env var fallback inline
Pin all cosign public key references to the immutable commit hash
(0112e53) that first introduced the key, instead of fetching it from
the release tag. This addresses the concern that an attacker with push
access could replace the key on main/tags and re-sign tampered images.
Docs now show two verification methods: commit hash (recommended) and
release tag (convenience), with explanation of why the hash is stronger.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: batch-limit stale managed object cleanup to prevent 300K row UPDATE (#25257)
* Add STALE_OBJECT_CLEANUP_BATCH_SIZE constant
Configurable batch limit (default 1000) for stale managed object cleanup,
preventing unbounded UPDATE queries from hitting 300K+ rows at once.
* Batch-limit stale managed object cleanup with single bounded SQL query
Two fixes to _cleanup_stale_managed_objects:
1. Replace unbounded update_many with a single execute_raw using a
subquery LIMIT, capping each poll cycle to STALE_OBJECT_CLEANUP_BATCH_SIZE
rows. Zero rows loaded into Python memory — everything stays in Postgres.
Uses the same PostgreSQL raw-SQL pattern as spend_log_cleanup.py
(the proxy requires PostgreSQL per schema.prisma).
2. Extract _expire_stale_rows as a separate method for testability.
Keeps the file_purpose='response' filter to avoid incorrectly expiring
long-running batch or fine-tune jobs that legitimately exceed the
staleness cutoff.
* docs: add STALE_OBJECT_CLEANUP_BATCH_SIZE to env vars reference
* test: remove deprecated embed-english-v2.0 cohere embedding tests
Adds a new endpoint to bulk-update team_member_permissions across
teams. Supports apply_to_all_teams (with cursor-based pagination)
or a specific list of team_ids. Merges new permissions into each
team's existing set rather than overwriting.
Also fixes test isolation bug in test_get_prompt_info_by_base_id
where leaked prisma_client state from other tests caused a
TypeError on await.
* Remove redundant matrix unit test workflow
All test paths in test-litellm-matrix.yml are fully covered by the
newer semantic unit test workflows (test-unit-*.yml), making the
matrix workflow redundant CI spend.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Add Codecov coverage reporting to semantic unit test workflows
Add coverage collection (--cov) and Codecov OIDC upload to both
reusable base workflows and all 12 caller workflows, replacing the
coverage reporting that was previously only in the matrix workflow.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Move id-token/pull-requests permissions to job level for multi-job workflows
For workflows with multiple jobs (llm-providers, proxy-db), move
id-token: write and pull-requests: write from workflow level to job
level so permissions are scoped to only the jobs that need them.
Removes zizmor inline suppressions that were masking the issue.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* [Docs] Enforce Black formatting in contributor docs
Black formatting is now enforced in CI. Update CLAUDE.md, AGENTS.md,
and CONTRIBUTING.md to instruct contributors and AI agents to run
`poetry run black .` before committing, and add VS Code setup guidance.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: fixes
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The proxy_e2e_azure_batches_tests workflow is consistently flaky and
does not provide reliable signal on whether changes break anything.
Remove the workflow from both CircleCI and GitHub Actions, along with
the test directory it exclusively used.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: multiple concurrent budget windows per API key and team (#24883)
* feat(proxy): add BudgetLimitEntry type and wire budget_limits into key/team models
* feat(schema): add budget_limits Json column to VerificationToken and TeamTable
* feat(migrations): add migration for budget_limits column on keys and teams
* feat(keys): initialize budget_limits windows with reset_at on key create/update
* feat(teams): initialize budget_limits windows with reset_at on team create/update
* feat(auth): add _virtual_key_multi_budget_check and _team_multi_budget_check
* feat(auth): call multi-budget checks from common_checks for keys and teams
* feat(proxy): increment per-window Redis spend counters after each request
* feat(budget): reset individual budget windows on schedule via reset_budget_job
* feat(ui): add hourly option to BudgetDurationDropdown
* feat(ui): add budget_limits field to KeyResponse type
* feat(ui): add Budget Windows editor to key edit view
* feat(ui): add Budget Windows editor to create key form
* fix(proxy): strip budget_limits=None before Prisma upsert to fix login 500
Prisma rejects nullable JSON fields (Json? without @default) when passed as
Python None — it needs the field omitted entirely so the DB stores NULL via
the column's nullable constraint. This was breaking /v2/login because the UI
session key creation path hit the upsert with budget_limits=None.
* ui(key-edit): use antd InputNumber+Button for budget windows, add reset hints
* ui(create-key): use antd InputNumber+Button for budget windows, add reset hints
* docs(users): add multiple budget windows section with API + dashboard walkthrough
* fix: BudgetExceededError returns HTTP 429 instead of 400
- Add status_code=429 to BudgetExceededError class
- auth_exception_handler hardcoded code=400 → code=429
* fix: no-op else branch in multi-budget auth checks causes KeyError
- BudgetLimitEntry objects must be coerced via model_dump() not left as-is
- Move _virtual_key_multi_budget_check into common_checks (was asymmetric
with _team_multi_budget_check which already lived there)
* fix: len() on JSON string returns char count not window count
Guard with isinstance check + json.loads() before iterating per-window
Redis counters in increment_spend_counters
* fix: silent except:pass hides Redis reset failures in reset_budget_windows
Log Redis counter reset failures as warnings so they are observable
* test: add unit tests for multi-budget window enforcement
5 tests covering: no budget_limits passes, under budget passes,
over hourly window raises 429, over monthly window raises 429,
BudgetLimitEntry objects coerced without KeyError
* fix: key per-window counters stable across reorders (duration key, not index)
* fix: team+key per-window spend increments use duration key, not index
* fix: budget window reset uses duration key; log failures instead of swallowing
* refactor: extract BudgetWindowsEditor to shared component
* refactor: key_edit_view imports BudgetWindowsEditor from shared component
* refactor: create_key_button imports BudgetWindowsEditor from shared component
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
* fix(reset_budget_job): extract _reset_expired_window helper to fix PLR0915 too many statements
* feat(skills): Skills Registry & Hub — register skills, browse in AI Hub, public skill hub (#25118)
* feat(skills): add domain and namespace fields to plugin types
* feat(skills): store and return domain/namespace inside manifest_json
* feat(skills): add /public/skill_hub endpoint for unauthenticated access
* feat(skills): whitelist /public/skill_hub from auth requirements
* feat(skills): add domain, namespace to Plugin and RegisterPluginRequest types
* feat(skills): smart URL parser — paste github URL, auto-detect source type and name
* feat(skills): replace enable toggle with Public badge, make rows clickable
* feat(skills): add skill detail view with Overview and How to Use tabs
* feat(skills): add MakeSkillPublicForm modal for publishing skills to the hub
* feat(skills): rename panel to Skills, wire in skill detail view on row click
* feat(skills): add skill hub table columns — name, description, domain, source, status
* feat(skills): add SkillHubDashboard with stats row, domain dropdown filter, and table
* feat(skills): add Skill Hub tab to AI Hub with Select Skills to Make Public button
* feat(skills): move Skills to top-level nav item directly under MCP Servers
* feat(skills): add skillHubPublicCall and NEXT_PUBLIC_BASE_URL support
* feat(skills): add Skill Hub tab to public AI Hub page
* feat(skills): add skills page routing in main app router
* feat(skills): add /skills page route
* chore: update package-lock after npm install
* docs(skills): add Skills Gateway doc page with mermaid architecture diagram
* docs(skills): add Skills Gateway to sidebar under Agent & MCP Gateway
* docs(skills): add loom walkthrough video to Skills Gateway doc
* chore: fixes
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
Co-authored-by: Yuneng Jiang <yuneng@berri.ai>
* fix(bedrock): strip [1m]/[200k] context window suffixes before cost lookup
* test(bedrock): add test for [1m] context window suffix stripping in cost lookup
* schema: add allowed_models to BudgetTable, default_team_member_models to TeamTable
* migration: add allowed_models and default_team_member_models columns
* types: add allowed_models to TeamMemberAddRequest, TeamMemberUpdateRequest, UpdateTeamRequest
* utils: add allowed_models param to add_new_member, persist to budget table
* common_utils: add allowed_models to _upsert_budget_and_membership
* team endpoints: seed allowed_models on member_add, persist on member_update and team/update
* auth: enforce per-member allowed_models at request time
* networking: add allowed_models to Member type and teamMemberUpdateCall
* TeamMemberTab: add Model Scope column showing per-member allowed_models
* EditMembership: add Allowed Models multi-select field
* TeamInfo: add default_team_member_models field in Settings tab
* chore: sync schema.prisma copies from root
* fix(team_member_update): update existing budget in-place instead of creating new one
When a member already has a budget_id, patch only the fields the caller
provided rather than always creating a fresh budget record. The old
code ignored existing_budget_id entirely, so updating only allowed_models
silently dropped the stored max_budget / tpm_limit / rpm_limit values.
* fix(auth): pass llm_router to _check_team_member_model_access
Without the router, _can_object_call_model cannot resolve wildcard model
names (e.g. openai/*) or access-group names in allowed_models, causing
legitimate requests to be denied. Thread the existing llm_router from
_run_common_checks through to the new member-scope check.
* feat(ui): add Team Member Settings accordion to Create Team modal
Groups default_team_member_models, member budget/key duration, and
tpm/rpm defaults into a single collapsible section. The model picker
is filtered to only show the models selected for the team, and the
copy distinguishes it from the team-level Models field.
* feat(ui): consolidate Team Member Settings into accordion in edit team form
Moves default_team_member_models + per-member budget/key/tpm/rpm fields
into a collapsible "Team Member Settings" panel. Keeps the top-level
form focused on team-wide settings (team models, team budget, tpm/rpm).
* fix(ui): use tremor Accordion for Team Member Settings in edit team form
* fix(ui): move Team Member Settings accordion above budget fields in Create Team
* chore: fixes
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Yuneng Jiang <yuneng@berri.ai>
* Add STALE_OBJECT_CLEANUP_BATCH_SIZE constant
Configurable batch limit (default 1000) for stale managed object cleanup,
preventing unbounded UPDATE queries from hitting 300K+ rows at once.
* Batch-limit stale managed object cleanup with single bounded SQL query
Two fixes to _cleanup_stale_managed_objects:
1. Replace unbounded update_many with a single execute_raw using a
subquery LIMIT, capping each poll cycle to STALE_OBJECT_CLEANUP_BATCH_SIZE
rows. Zero rows loaded into Python memory — everything stays in Postgres.
Uses the same PostgreSQL raw-SQL pattern as spend_log_cleanup.py
(the proxy requires PostgreSQL per schema.prisma).
2. Extract _expire_stale_rows as a separate method for testability.
Keeps the file_purpose='response' filter to avoid incorrectly expiring
long-running batch or fine-tune jobs that legitimately exceed the
staleness cutoff.
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
* fix(ui): resolve login redirect loop when reverse proxy adds HttpOnly to cookies
When LiteLLM is behind nginx-ingress or similar with security-hardened
configs, the reverse proxy adds HttpOnly to all Set-Cookie headers. This
makes the JWT token unreadable by JavaScript, causing an infinite login
redirect loop. Fix by returning the JWT token in the /v2/login response
body so the frontend can set a JS-accessible cookie directly.
Fixes#19663
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: address Greptile review feedback
- Add window guard to setTokenCookie for SSR consistency with clearTokenCookies
- Add SSR test for window undefined case
- Add code comment explaining why JWT is included in response body
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: address second round of Greptile review feedback
- Add loginCall integration tests verifying setTokenCookie is called with
token and skipped when absent (backward-compatibility path)
- Use encodeURIComponent/decodeURIComponent in setTokenCookie/getCookie
for defense-in-depth against non-standard token formats
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* Update ui/litellm-dashboard/src/utils/cookieUtils.ts
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Update ui/litellm-dashboard/src/utils/cookieUtils.ts
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix(ui): use sessionStorage instead of cookie for login token storage
Replace setTokenCookie (which is a no-op when reverse proxy adds HttpOnly)
with storeLoginToken using sessionStorage. Add sessionStorage fallback to
getCookie so the token is found even when the cookie is HttpOnly. Also handle
'=' in cookie values with .slice(1).join("=") and clear sessionStorage on
logout.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ui): use shared getCookie in page.tsx and user_dashboard.tsx
Replace local getCookie functions in page.tsx and user_dashboard.tsx
with the shared one from cookieUtils that has the sessionStorage
fallback. Without this, the HttpOnly cookie fix was incomplete —
page.tsx (the dashboard entry point) could not read the token,
causing the redirect loop to persist.
Also scope the sessionStorage fallback to the "token" key only,
and clear sessionStorage in page.tsx deleteCookie.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ui): scope deleteCookie sessionStorage cleanup to token key only
Also document the sessionStorage cross-tab trade-off: per-tab scope
means users behind an HttpOnly proxy must log in once per tab, but
this is intentional to avoid localStorage XSS exposure.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* Update ui/litellm-dashboard/src/utils/cookieUtils.ts
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* style: remove stray double blank line in user_dashboard.tsx
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ui): guard storeLoginToken against empty/whitespace-only tokens
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ui): preserve sessionStorage token across beforeunload clear
The existing beforeunload handler calls sessionStorage.clear() to
flush cached UI data on page refresh. This also wiped the token
stored by storeLoginToken, re-introducing the redirect loop after
any page refresh in the HttpOnly proxy scenario. Now the token is
saved and restored across the clear.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ui): set JS-accessible cookie at /ui path as HttpOnly workaround
sessionStorage alone is unreliable. Also set the token via
document.cookie at path=/ui — nginx only adds HttpOnly to server-set
Set-Cookie headers, so a JS-set cookie is always readable.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ui): use dynamic cookie path based on server_root_path
Hardcoded path=/ui breaks when LiteLLM is deployed with a custom
server_root_path. Now derives the cookie path from serverRootPath
so it works at /ui, /myapp/ui, etc.
Also reuse clearTokenCookies() in deleteCookie() to avoid duplication.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* refactor(ui): remove circular dependency in cookieUtils.ts
Derive the UI cookie path from window.location.pathname instead of
importing serverRootPath from networking.tsx. This breaks the
cookieUtils → networking → cookieUtils cycle that could cause
serverRootPath to be undefined under certain bundler configurations.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ui): harden getUiCookiePath regex and add missing tests
- Use regex /\/ui(?=\/|$)/ to match "/ui" only as a full path segment,
preventing false matches on paths like "/my-ui-tool/login".
- Add unit tests for storeLoginToken empty/whitespace guard and
cookie-at-/ui-path behavior.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: fix Black formatting in audit_logs.py
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix CI: formatting, test params, remove token from login JSON
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: reformat with Black 23.x to match CI
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: keep token in login JSON body for UI storeLoginToken flow
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: use storeLoginToken in exchangeLoginCode, add credentials include
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* revert: remove unrelated changes from HttpOnly cookie fix branch
Reset files not related to the login cookie fix back to main:
- prometheus.py, bedrock converse, guardrail handler
- auth_checks.py, reset_budget_job.py, audit_logs.py
- test_user_api_key_auth.py
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Revert "revert: remove unrelated changes from HttpOnly cookie fix branch"
This reverts commit 0684a1e275.
* Revert "fix: use storeLoginToken in exchangeLoginCode, add credentials include"
This reverts commit 866405f443.
* Revert "fix: keep token in login JSON body for UI storeLoginToken flow"
This reverts commit 086c41640c.
* Revert "fix: reformat with Black 23.x to match CI"
This reverts commit b2c3334c88.
* Revert "fix CI: formatting, test params, remove token from login JSON"
This reverts commit 2905d47bd4.
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
* docs(blog): add cosign Docker image verification instructions
Add steps for verifying Docker images with cosign to three security blog posts:
CI/CD v2, Security Townhall, and Security Update.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs(proxy): add cosign verification to Docker/Helm/Terraform deploy page
Add image signature verification steps to the main deployment doc so
users pulling Docker images know how to verify them with cosign.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: fixes
* Update index.md
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* [Docs] Scope cosign signing docs to GHCR and specify starting version
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* [Docs] Add starting version callout to ci_cd_v2 blog post
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Extend LATENCY_BUCKETS beyond 5 minutes so request/LLM latency metrics
can distinguish long runs up to the typical default LLM request timeout.
Made-with: Cursor
* fix(presidio): use correct text positions in anonymize_text (#24160)
The Presidio anonymizer endpoint returns items with start/end positions
that reference the *anonymized output* text, not the original input.
anonymize_text() was applying these positions to the original text,
causing garbled output with remnants of un-masked PII data.
When output_parse_pii is False, return redacted_text["text"] directly
from the anonymizer response instead of manually splicing.
When output_parse_pii is True, use analyze_results positions (which
correctly reference the original text) to build numbered replacement
tokens and the pii_tokens mapping.
* address review: remove dead code, fix token numbering order
- Remove unused `anon_item_by_entity` dict (Greptile P2)
- Number tokens left-to-right (<PERSON_1> first in text, not last)
- Add assertion for token numbering order in test
## Problem
When `get_cache_key(**kwargs)` is called with kwargs that already
contains `preset_cache_key` (which can happen when cache key is
recomputed in certain code paths), the call to
`_set_preset_cache_key_in_kwargs()` fails with:
```
TypeError: _set_preset_cache_key_in_kwargs() got multiple values
for keyword argument 'preset_cache_key'
```
This is because `preset_cache_key` is passed both explicitly:
```python
self._set_preset_cache_key_in_kwargs(
preset_cache_key=hashed_cache_key, **kwargs
)
```
And implicitly via `**kwargs` unpacking when `kwargs["preset_cache_key"]`
exists.
## Solution
Filter out `preset_cache_key` from kwargs before passing to
`_set_preset_cache_key_in_kwargs()`:
```python
kwargs_for_preset = {k: v for k, v in kwargs.items() if k != "preset_cache_key"}
self._set_preset_cache_key_in_kwargs(
preset_cache_key=hashed_cache_key, **kwargs_for_preset
)
```
## Testing
Added unit tests covering:
- kwargs with existing preset_cache_key (the bug case)
- kwargs without preset_cache_key (regression test)
- Verification that preset_cache_key is correctly set in litellm_params
Fixes#25081.
is_tool_name_prefixed() checked for the presence of MCP_TOOL_PREFIX_SEPARATOR
(default '-') anywhere in the tool name. Any non-MCP tool whose name
contains a hyphen (e.g. 'text-to-speech', 'code-review') was silently
misclassified as an MCP-prefixed tool. When the semantic tool filter is
enabled, these tools would be routed through semantic matching and
potentially dropped.
Fix: accept an optional known_server_prefixes set. When supplied, the
function extracts the candidate prefix (text before the first separator)
and checks it against the normalised set of registered server prefixes.
Only a genuine match returns True. Without the set, legacy behaviour is
preserved for backward compatibility.
Updated _get_mcp_server_from_tool_name() to build the prefix set from
the live registry and pass it through.
9 new tests.
Co-authored-by: d 🔹 <258577966+voidborne-d@users.noreply.github.com>
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run