Commit graph

38954 commits

Author SHA1 Message Date
Vikas Sharma
51a67bb2a0 refactor(tickerr): slim to 103-line dumb pipe, server owns all normalization
- Remove _PROVIDER_MAP (25 entries), _ERROR_TYPE_MAP, _normalize_provider,
  _get_provider, _get_status_code, _latency_ms, _fire_and_forget, _parse_sample_rate
- Provider read directly from custom_llm_provider, passed as-is
- Model passed as-is, no prefix stripping
- status_code sent raw, server derives error_type
- Inline fire-and-forget into _report, remove semaphore
- Add TICKERR_DISABLED kill switch
- Add opt-in success sampling (TICKERR_SAMPLE_RATE, default 0)
- event_type: "failure"/"success" replaces is_resolution boolean
- Remove client_tier (trust concern for first merge)
- Tests import only TickerrLogger, no internal symbols
- 103 lines source, 19 tests, 55 lines docs
2026-05-29 01:31:20 +05:30
imviky-ctrl
e279076193 test(tickerr): move tests to integrations folder so CI picks them up for coverage 2026-05-06 17:19:15 +05:30
imviky-ctrl
6123020b52 test(tickerr): add 23 tests to reach 100% coverage on tickerr.py
Covers:
- log_failure_event sync hook
- _normalize_provider: vertex_ai, vertex_ai_anthropic, azure, bedrock,
  bedrock_converse, command-*, mixtral-*, deepseek, openrouter,
  fireworks_ai, cerebras, xai
- _report: no exception, client_tier/region payload fields,
  empty model → None, timeout/auth error_type mapping
- _fire_and_forget: success path + semaphore release via finally
- _extract_status_code: non-digit string, 529
- _latency_ms: zero, large value
2026-05-06 16:51:32 +05:30
imviky-ctrl
8047392b21 fix: add missing provider aliases, fix test semaphore accounting
- Add vertex_ai_anthropic → anthropic, bedrock_converse → aws,
  azure_ai → azure to _PROVIDER_MAP so composite providers report
  the canonical slug instead of their raw internal key
- Fix test_fire_and_forget_respects_semaphore_cap to release only
  as many slots as were actually acquired, preventing semaphore
  count from exceeding _MAX_INFLIGHT if an assert fires mid-loop
2026-05-05 15:48:28 +05:30
imviky-ctrl
c2c2f8285d style: apply Black formatting to tickerr.py 2026-05-05 15:34:15 +05:30
imviky-ctrl
a39ba5c52f test(tickerr): fix race condition and private API usage in tests
- test_fire_and_forget_silent_on_network_error: mock threading.Thread
  entirely so no real thread or network call can escape the test boundary
- test_semaphore_released_on_thread_start_failure: replace _inflight._value
  (private CPython attr) with a non-blocking acquire/release assertion
2026-05-05 15:28:54 +05:30
imviky-ctrl
94a1638630 fix(tickerr): fix semaphore leak, remove 500 from error map, fix semaphore test
- Semaphore leak: wrap t.start() in try/except and release _inflight on
  RuntimeError so the slot is not permanently lost if the OS thread limit
  is hit
- Remove 500 from _ERROR_TYPE_MAP: 500 is a generic internal server error
  (crash/bug/misconfiguration), not a capacity/overload condition — sending
  'overloaded' for 500s would corrupt crowd-sourced signal
- Rewrite semaphore cap test to actually invoke real _fire_and_forget with
  exhausted semaphore and assert threading.Thread is never called
- Add test for semaphore release on thread start failure
- Update error-type test to explicitly assert 500 is excluded
2026-05-05 15:20:18 +05:30
imviky-ctrl
72d1a93ea1 test(tickerr): add unit tests for TickerrLogger callback
Covers: provider normalization, status code extraction, latency
calculation with both datetime and float types, error type mapping,
payload construction, semaphore cap, silent network failure, and
async hook wiring.
2026-05-05 15:04:08 +05:30
imviky-ctrl
f4b0d84a98 fix(tickerr): fix TypeError, error-type fallback, and thread cap
- _latency_ms(): handle both datetime objects and floats — LiteLLM
  passes datetime, not float, causing TypeError on every invocation
- _ERROR_TYPE_MAP fallback: removed default "overloaded" for unmapped
  codes (400, 404, 502, etc.) — now returns None for unrecognised codes
- _fire_and_forget(): added semaphore cap (_MAX_INFLIGHT=5) to prevent
  unbounded thread spawning during burst failures
- Removed dead is_resolution parameter from _report() — only called
  from failure hooks so it was always False
2026-05-05 15:04:08 +05:30
imviky-ctrl
64aab2de56 feat(integrations): add Tickerr callback for LLM failure reporting
Adds Tickerr as a built-in LiteLLM callback. Tickerr is a crowd-sourced
outage radar for AI agents — when an agent hits a 5xx error, it reports
anonymously to Tickerr and receives live signal from other agents hitting
the same issue (including a fallback model recommendation).

Changes:
- litellm/integrations/tickerr.py — TickerrLogger (stdlib-only, no new deps)
- litellm/__init__.py — adds "tickerr" to _custom_logger_compatible_callbacks_literal
- litellm/litellm_core_utils/custom_logger_registry.py — registers TickerrLogger
- litellm/integrations/callback_configs.json — UI config entry
- docs/my-website/docs/observability/tickerr.md — integration docs

Usage:
    litellm.callbacks = ["tickerr"]

No API key required. Anonymous. Non-blocking (daemon thread, 5s timeout).

https://tickerr.ai
2026-05-05 15:04:08 +05:30
yuneng-jiang
f318ef03bd
Merge pull request #27170 from BerriAI/litellm_/unruffled-mcclintock-62a296
[Fix] Docker: Remove Hardcoded Prisma Binary Target For Multi-Arch Builds
2026-05-04 21:43:14 -07:00
shin-berri
43c78057d4
Merge pull request #27169 from BerriAI/litellm_/vigorous-albattani-2b7480
[Perf] CI: Skip Redundant Playwright Apt Install in E2E UI Job
2026-05-04 21:39:43 -07:00
Yuneng Jiang
4ee586a321
[Fix] Docker: Remove Hardcoded Prisma Binary Target For Multi-Arch Builds
PRISMA_CLI_BINARY_TARGETS="debian-openssl-3.0.x" was hardcoded in
docker/Dockerfile.non_root by #17695. On a buildx linux/arm64 leg this
forces prisma to download the amd64 schema-engine into an arm64 image,
so 'prisma migrate deploy' fails at startup with 'Could not find
schema-engine binary'.

Removing the env lets prisma auto-detect per build platform: amd64
builds still resolve to debian-openssl-3.0.x (Wolfi falls back to
debian, same binary as before), and arm64 builds now correctly fetch
linux-arm64-openssl-3.0.x. The offline-cache pre-warm goal of #17695 is
preserved — only which binaries fill the cache changes.

Fixes #19458
2026-05-04 21:37:16 -07:00
Yuneng Jiang
19ad964c4a
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/vigorous-albattani-2b7480 2026-05-04 21:19:34 -07:00
Yuneng Jiang
c1c0506d2c
[Perf] CI: Skip Redundant Playwright Apt Install in E2E UI Job
The cimg/python:3.12-browsers base image already ships every Chromium
system dependency Playwright needs (libnss3, libatk-bridge2.0-0,
libcups2, etc. — the install log shows them all as "already the newest
version"). Passing --with-deps to `npx playwright install` therefore
runs an apt-get update + install for nothing, but pays the full cost of
hitting Ubuntu mirrors. On a recent run those mirrors stalled hard:
apt-get update alone took 6m53s at 81.5 kB/s with several archives
returning connection refused.

Drop --with-deps and persist ~/.cache/ms-playwright alongside
node_modules so the Chromium binary is also reused across runs. Bump
the cache key to v2 so the existing v1 entry (which only contained
node_modules) is not loaded and skipped over the new browser path.
2026-05-04 21:19:31 -07:00
yuneng-jiang
cd38ecd532
Merge pull request #27156 from BerriAI/yj_build_may4
[Infra] Build UI
2026-05-04 21:12:20 -07:00
shin-berri
ff3d089ab8
Merge pull request #27160 from BerriAI/litellm_/peaceful-gates-6e46e7
[Fix] Proxy: Break managed-resources import cycle on Python 3.13
2026-05-04 21:11:53 -07:00
yuneng-jiang
f2969ca78a
Merge pull request #27165 from BerriAI/litellm_/friendly-lichterman-35cf02
Some checks are pending
Unit Tests: Caching (Redis) / caching-redis (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / schema-migration (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Security / security (push) Waiting to run
[Fix] CI: Enable VCR replay for test_azure_o_series
2026-05-04 20:59:46 -07:00
Yuneng Jiang
0976fbc6c4
[Fix] Tests: Restore /metrics access for prometheus test suite
/metrics now requires auth by default; tests/otel_tests/test_prometheus.py
makes 4+ unauthenticated GETs against http://0.0.0.0:4000/metrics, so
every prometheus test in CI now fails the metric assertion.

Set require_auth_for_metrics_endpoint: false in otel_test_config.yaml
to opt out for this test job, which scrapes /metrics directly. Verified
locally: 8/8 prometheus tests green (one flaky retry on
test_proxy_success_metrics that pre-dates this PR).

Also drop the -x stop-on-first-failure flag from the otel test command
so all failures in the job surface in a single CI run rather than
hiding behind whichever one trips first.
2026-05-04 20:54:54 -07:00
Yuneng Jiang
6a6c79d992
[Fix] CI: Enable VCR replay for test_azure_o_series
The Azure o-series tests were excluded from the conftest's VCR auto-marker
because of a respx/vcrpy transport-patching conflict, but the only respx
reference in the file was an unused `MockRouter` import. Drop the dead
import and remove the file from the conflict set so cassettes record on
first run and replay thereafter, eliminating the 60-95s live Azure latency
that was crashing xdist workers under --timeout=120 thread-mode timeouts.
2026-05-04 20:48:26 -07:00
Sameer Kankute
b0edffb883
Merge pull request #27103 from BerriAI/litellm_azure-deployment-image-body
fix(azure): omit model from deployment image gen and image edit bodies
2026-05-05 09:09:45 +05:30
Yuneng Jiang
e6f524f951
[Fix] Tests: Pick chat-completion OTEL trace by content, not recency
The /otel-spans endpoint returns process-wide spans and tags
most_recent_parent by max start_time. After tightening that route to
proxy_admin (sk-1234), the GET /otel-spans request itself emits auth
spans that beat the chat-completion spans on start_time, so
most_recent_parent now points at the request's own auth trace
(['postgres', 'postgres']) and the >=5-span assertion fails.

Pick the chat-completion trace by content: it is the only trace whose
span list is a superset of {postgres, redis, raw_gen_ai_request,
batch_write_to_db}. Verified locally end-to-end against
otel_test_config.yaml + OTEL_EXPORTER=in_memory: 3/3 runs green.
2026-05-04 20:35:09 -07:00
Sameer Kankute
4487d8352f
Merge pull request #27115 from Sameerlite/litellm_health_check_reasoning_effort
feat(proxy): add health_check_reasoning_effort for model health checks
2026-05-05 09:00:09 +05:30
Yuneng Jiang
8a1b6635fa
[Fix] Tests: Use master key for /otel-spans in test_chat_completion_check_otel_spans
/otel-spans now requires proxy admin (returns 401 'Only proxy admin
can be used to generate, delete, update info for new keys/users/teams.
Route=/otel-spans' for non-admin callers). Switch the GET call to use
the master key sk-1234 while keeping the generated key for the
chat-completion request that produces the spans.
2026-05-04 20:23:11 -07:00
Sameer Kankute
b4ee6a2355
test(proxy): cover health_check_reasoning_effort for completion mode
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-05 08:52:57 +05:30
Sameer Kankute
bb0e4168ad
refactor(azure): move image gen JSON helper; rename image edit finalize hook
- Add image_generation/http_utils.azure_deployment_image_generation_json_body; call
  from azure.py (keeps AzureChatCompletion focused on chat).
- Rename finalize_image_edit_multipart_data to finalize_image_edit_request_data with
  docstring covering multipart and JSON POST payloads (review feedback).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-05 08:49:46 +05:30
Yuneng Jiang
193907a4a3
[Fix] Lint: Mark _user_has_admin_view re-export in common_utils
Ruff F401 flagged the aliased import as unused within common_utils.py
because the name is consumed only by external modules (~15 callers
across guardrails, spend tracking, MCP, agents, management endpoints).
Add `# noqa: F401  re-exported` so the alias survives lint while
keeping a single source of truth in litellm.proxy._types.
2026-05-04 20:16:59 -07:00
Yuneng Jiang
8cac6c5bff
[Fix] Proxy: Address Greptile feedback on hook-cycle PR
- Move _user_has_admin_view to litellm.proxy._types as
  user_api_key_has_admin_view (single source of truth). common_utils.py
  and isolation.py both import from there now, removing the duplicated
  role-check that could silently diverge if new admin roles are added.
- Add pytest.importorskip("litellm_enterprise") to the two regression
  tests that assert managed_files / managed_vector_stores are registered;
  those keys come from ENTERPRISE_PROXY_HOOKS so the tests would fail
  unconditionally in a checkout without the enterprise extra installed.
2026-05-04 20:13:31 -07:00
Yuneng Jiang
727ab8dcc4
[Fix] Proxy: Break managed-resources import cycle on Python 3.13
The Python 3.13 CCI smoke matrix surfaces a partially-initialized-module
ImportError when loading the managed files hook chain:

  litellm.proxy.hooks/__init__ (mid-import)
    -> enterprise.enterprise_hooks
    -> litellm_enterprise.proxy.hooks.managed_files
    -> litellm.llms.base_llm.managed_resources.isolation
    -> litellm.proxy.management_endpoints.common_utils
    -> litellm.proxy.utils  (re-enters litellm.proxy.hooks)

The except ImportError block in hooks/__init__.py silently swallowed the
failure, leaving managed_files unregistered and POST /files returning
500 "Managed files hook not found".

Two-layer fix:
- Inline the 3-line _user_has_admin_view check in isolation.py instead
  of importing it from litellm.proxy.management_endpoints.common_utils.
  litellm.llms.* should not depend on litellm.proxy.* — removing this
  layering violation breaks the cycle at its root.
- Define PROXY_HOOKS and get_proxy_hook before the conditional
  enterprise import in litellm/proxy/hooks/__init__.py, so any future
  re-entry resolves the public names instead of hitting an
  ImportError on a partially-initialized module.

Also fold in two unrelated CCI repairs surfaced in the same staging run:
- tests/otel_tests/test_key_logging_callbacks.py: per-key
  gcs_bucket_name / gcs_path_service_account are now stripped by
  initialize_dynamic_callback_params, so the GCS client falls through
  to the env-only branch. Update the assertion to match the new
  "GCS_BUCKET_NAME is not set" message.
- .circleci/config.yml: tests/pass_through_tests now resolves
  google-auth-library@10.x via the @google-cloud/vertexai 1.12.0 bump,
  which uses dynamic ESM imports Jest 29 cannot load without
  --experimental-vm-modules. Pass that flag in the Vertex JS test step.

Adds tests/test_litellm/proxy/hooks/test_proxy_hooks_init.py as a
regression guard: managed_files / managed_vector_stores must register,
and isolation.py must not transitively import litellm.proxy.utils.
2026-05-04 20:05:24 -07:00
Yuneng Jiang
7c8409d013
chore: update Next.js build artifacts (2026-05-05 02:13 UTC, node v20.20.2) 2026-05-04 19:13:25 -07:00
yuneng-jiang
9ea824d5bf
Merge pull request #27143 from BerriAI/cursor/fix-secret-fields-in-spend-logs-a532
fix(security): prevent secret_fields from leaking into spend logs
2026-05-04 19:07:54 -07:00
yuneng-jiang
be5f217aaf
Merge pull request #26861 from BerriAI/litellm_fix_scim_virtual_key_deactivation
fix(scim): revoke virtual keys when SCIM deprovisions a user
2026-05-04 19:03:55 -07:00
Cursor Agent
5923c3209b
fix(security): prevent secret_fields from leaking into spend logs
secret_fields (containing raw HTTP headers including Authorization
Bearer tokens) was being included in proxy_server_request['body']
because the body snapshot was a copy.copy(data) of the full request
dict. This body gets serialized and persisted in the LiteLLM_SpendLogs
table, exposing user credentials in the database.

Root cause: data['secret_fields'] was set before the body snapshot at
data['proxy_server_request']['body'] = copy.copy(data), so the full
raw headers (including auth tokens) ended up in the snapshot.

Fix (defense in depth):
1. Exclude 'secret_fields' when creating the body snapshot in
   litellm_pre_call_utils.py (primary fix)
2. Strip 'secret_fields' in _sanitize_request_body_for_spend_logs_payload
   as a secondary safeguard

secret_fields remains available on the live data dict for legitimate
downstream consumers (MCP, Responses API).

Co-authored-by: Krrish Dholakia <krrish-berri-2@users.noreply.github.com>
2026-05-05 02:01:41 +00:00
yuneng-jiang
555a8131fe
Merge pull request #26951 from stuxf/codex/skills-containers-tenant-guard
chore(proxy): tighten resource ownership checks
2026-05-04 18:47:17 -07:00
yuneng-jiang
2f305050ce
Merge pull request #27004 from stuxf/fix/managed-resource-service-account-isolation
fix(proxy): isolate managed resources for service-account API keys
2026-05-04 18:45:55 -07:00
user
3dcb6bd3f9
Merge remote-tracking branch 'upstream/litellm_internal_staging' into codex/skills-containers-tenant-guard
# Conflicts:
#	litellm/proxy/auth/auth_utils.py
2026-05-05 01:41:25 +00:00
user
7faba9656f
Merge remote-tracking branch 'upstream/litellm_internal_staging' into fix/managed-resource-service-account-isolation 2026-05-05 01:38:11 +00:00
yuneng-jiang
281296f9cf
Merge pull request #27151 from BerriAI/litellm_yj_may4
[Infra] Merge dev branch
2026-05-04 18:29:52 -07:00
user
aee064ad37
Merge remote-tracking branch 'upstream/litellm_internal_staging' into fix/managed-resource-service-account-isolation 2026-05-05 01:29:05 +00:00
yuneng-jiang
dcb357ee2d
Merge pull request #27149 from BerriAI/litellm_/peaceful-bell-ba8ca5
[Fix] Tests: Replace deprecated openrouter/claude-3.7-sonnet with claude-sonnet-4.5
2026-05-04 18:27:45 -07:00
yuneng-jiang
efca16ccfa
Merge pull request #27043 from stuxf/fix/ssti-prompt-managers
fix(security): sandbox jinja2 in gitlab/arize/bitbucket prompt managers
2026-05-04 18:23:41 -07:00
Yuneng Jiang
e35cd5af76
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_yj_may4 2026-05-04 18:22:47 -07:00
Yuneng Jiang
7f550a5d67
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/peaceful-bell-ba8ca5 2026-05-04 18:21:33 -07:00
Yassin Kortam
db2a3cafb6
Merge pull request #27131 from BerriAI/litellm_fix/routing-groups-ui
feat: routing groups ui
2026-05-04 18:16:49 -07:00
mateo-berri
4179159f0f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_azure-deployment-image-body 2026-05-04 18:16:46 -07:00
Yassin Kortam
a56256e5ee feat: routing groups ui 2026-05-04 18:09:14 -07:00
yuneng-jiang
42cd9493e9
Merge pull request #27071 from stuxf/fix/strip-pricing-fields
chore(proxy): drop client-supplied pricing fields from request bodies
2026-05-04 18:08:41 -07:00
yuneng-jiang
68c120a68f
Merge pull request #26957 from stuxf/chore/guardrail-coverage
chore(guardrails): cover multimodal + Responses-API content shapes
2026-05-04 18:01:27 -07:00
Yuneng Jiang
00d0c3e745
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/peaceful-bell-ba8ca5 2026-05-04 17:51:54 -07:00
Yuneng Jiang
22782f3c3f
[Fix] Tests: Replace deprecated openrouter/claude-3.7-sonnet with claude-sonnet-4.5
OpenRouter has dropped active endpoints for anthropic/claude-3.7-sonnet,
causing test_reasoning_content_completion to fail with a 404 "No endpoints
found" error. Switch to anthropic/claude-sonnet-4.5, which is current and
supports reasoning streaming.
2026-05-04 17:51:50 -07:00