Commit graph

33508 commits

Author SHA1 Message Date
ryan-crabbe
253792a1d1
Merge pull request #21611 from BerriAI/litellm_perf_skip_throwaway_usage
perf: skip throwaway Usage() construction in ModelResponse.__init__
2026-02-24 14:53:17 -08:00
Ryan Crabbe
6a0dd0a45f Merge remote-tracking branch 'origin/main' into litellm_perf_skip_throwaway_usage
# Conflicts:
#	tests/llm_translation/test_llm_response_utils/test_convert_dict_to_chat_completion.py
2026-02-24 14:51:57 -08:00
yuneng-jiang
80ebe722d9
Merge pull request #22029 from BerriAI/litellm_spend_tracking_logging
[Infra] Add Spend Tracking Lifecycle Logging
2026-02-24 12:52:32 -08:00
yuneng-jiang
2cabbccf6f address greptile review feedback (greploop iteration 3)
- Add missing traceback to team member spend enqueue error log
2026-02-24 12:49:48 -08:00
yuneng-jiang
8b56e1d969 trigger review 2026-02-24 12:37:43 -08:00
Ishaan Jaff
33719e6b38
docs: update v1.81.12-stable release notes to point to v1.81.12-stable.1 (#22036) 2026-02-24 12:30:18 -08:00
yuneng-jiang
70ef4d0d69 address greptile review feedback (greploop iteration 2)
- Remove re-raise in _store_transactions_in_redis so one Redis
  push failure doesn't drop remaining transaction types
- Downgrade per-push success log from info to debug to reduce noise
- Fix misleading error message in update_database — entity spend
  updates run as independent tasks and are not affected by this catch
2026-02-24 12:10:19 -08:00
Sameer Kankute
5219b1d0c3
Merge pull request #22035 from BerriAI/litellm_openai_codex_day_0_codex_5.3
[Feat] OpenAI codex 5.3 day 0 support
2026-02-25 01:29:27 +05:30
yuneng-jiang
235d60eb88 address greptile review feedback (greploop iteration 1)
- Add traceback to cache update warning logs (user, end_user, team, tag)
- Remove duplicate info log in non-redis commit path
2026-02-24 11:58:59 -08:00
Ishaan Jaff
c343bfffda
fix(router): emit x-litellm-overhead-duration-ms header for streaming requests (#22027)
* fix(router): preserve _hidden_params in FallbackStreamWrapper so x-litellm-overhead-duration-ms is emitted for streaming requests

* test(router): add regression test for FallbackStreamWrapper _hidden_params preservation
2026-02-24 11:56:16 -08:00
Ishaan Jaff
e44b9b6b35
feat(prometheus): add opt-in stream label to litellm_proxy_total_requests_metric (#22023)
Set prometheus_emit_stream_label: true in litellm_settings to emit a
stream label (True/False/None) on litellm_proxy_total_requests_metric.

Opt-in to avoid breaking cardinality on existing deployments.
2026-02-24 11:51:42 -08:00
Sameer Kankute
5d291c739f Fix phase docs link 2026-02-25 01:21:38 +05:30
Sameer Kankute
74abf0c8e6 Fix phase docs link 2026-02-25 01:19:10 +05:30
Sameer Kankute
aded14a55a Fix release version for gpt-5.3-codex 2026-02-25 01:04:12 +05:30
yuneng-jiang
c43a8dc842 feat(proxy): add warning/error level logging throughout spend tracking lifecycle
Elevate silent debug-level and bare except:pass error paths to
warning/error so spend tracking failures are visible in production logs.

All new log messages are prefixed with "Spend tracking -" for easy
filtering. Changes cover the full request-to-DB lifecycle:
enqueue, in-memory flush, Redis buffer push/pop, DB commit,
cache updates, spend log writes, and pod lock management.

Also fixes a copy-paste bug in _update_team_cache that logged
"end user" instead of "team".
2026-02-24 10:17:35 -08:00
yuneng-jiang
4321bc9285
Merge pull request #21985 from BerriAI/litellm_ui_testing_coverage_00
[Fix] UI - Virtual Keys: restrict Edit Settings to key owners
2026-02-24 10:16:40 -08:00
Sameer Kankute
1c48d8fda7 Add gpt-5.3-codex in model cost map 2026-02-24 23:37:09 +05:30
Ishaan Jaff
5e9f24f74c
fix(bedrock): pass timeout param to bedrock rerank http client (#22021)
* fix(bedrock): pass timeout to bedrock rerank http client

* refactor: extract large functions to fix PLR0915 ruff lint errors
2026-02-24 09:32:11 -08:00
Sean Marsh Glover
4652c73259
feat(proxy): limit concurrent health checks with health_check_concurrency (#20584)
* staged first pass

* black

* Update litellm/proxy/health_check.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* simpler

* restore cached logo

* fix tests for perform_health_check max_concurrency arg

* implement pr suggestion

* and the helm chart

* add configureable resources and probes to the deployment in the helm chart

* more helm chart unittests

* move some background healthcheck loggin to debug

---------

Co-authored-by: Sean Glover <sglover@athenahealth.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-24 08:16:59 -08:00
Harshit Jain
1fa0aad3f2
Merge pull request #22008 from BerriAI/litellm_fix_CVE
security: fix critical/high CVEs in OS-level libs and NPM transitive
2026-02-24 21:44:19 +05:30
Harshit28j
132e2ed671 Merge branch 'main' of https://github.com/BerriAI/litellm into litellm_fix_CVE
# Please enter a commit message to explain why this merge is necessary,
# especially if it merges an updated upstream into a topic branch.
#
# Lines starting with '#' will be ignored, and an empty message aborts
# the commit.
2026-02-24 21:09:16 +05:30
Harshit28j
3e6c10a071 security: fix critical/high CVEs in OS-level libs and NPM transitive 2026-02-24 19:40:09 +05:30
Sameer Kankute
a37cd0fe7c
Merge pull request #22005 from BerriAI/litellm_mcp_server_ui_fix
Fix: Transport Type for OpenAPI Spec on UI
2026-02-24 19:38:34 +05:30
Sameer Kankute
c17caf4cc7
Merge pull request #21992 from BerriAI/litellm_fix_oauth_mcp
fix: Missing OAuth session state
2026-02-24 19:37:09 +05:30
Sameer Kankute
6531d01959
Merge pull request #21982 from BerriAI/litellm_fix_pat_token_mcp
Fix: skip health check for MCP integration with passthrough token auth
2026-02-24 19:36:08 +05:30
Sameer Kankute
7a8499e89f
Merge pull request #21940 from BerriAI/litellm_oss_staging_02_23_2026
litellm oss staging 02 23 2026
2026-02-24 19:32:56 +05:30
Sameer Kankute
b38059b014
Merge branch 'main' into litellm_oss_staging_02_23_2026 2026-02-24 19:32:48 +05:30
Sameer Kankute
816f9052ff Fix: Transport Type for OpenAPI Spec on UI 2026-02-24 19:27:12 +05:30
Julio Quinteros Pro
d0330aa4e3
Merge pull request #22001 from jquinter/revert/pr-21957
Revert PR #21957: atomic RPM rate limiting
2026-02-24 09:53:25 -03:00
Julio Quinteros Pro
737a04b3ea Revert "Merge pull request #21957 from jquinter/fix/flaky-rpm-limit-test"
This reverts commit 77453ada2a, reversing
changes made to 7622f26918.
2026-02-24 09:52:50 -03:00
Julio Quinteros Pro
77453ada2a
Merge pull request #21957 from jquinter/fix/flaky-rpm-limit-test
fix: atomic RPM rate limiting in model rate limit check
2026-02-24 09:51:26 -03:00
Sameer Kankute
ac720defc3 Add documentation related to phase 2026-02-24 17:50:38 +05:30
Sameer Kankute
ef67b6b533 Add support for phase param 2026-02-24 17:48:55 +05:30
Shivam Rawat
7622f26918
Merge pull request #21997 from BerriAI/doc_fix_remove_harcoded_api_key
[Doc] replaced azure openai key with mock key
2026-02-24 03:32:37 -08:00
shivam
c86b174642 replaced with mock key 2026-02-24 03:28:28 -08:00
Sameer Kankute
12f37cea43 fix: Missing OAuth session state. Please retry 2026-02-24 14:22:38 +05:30
yuneng-jiang
c119adb6dc [Fix] UI - Virtual Keys: restrict Edit Settings button to key owners
Non-owner Internal Users could see and interact with the "Edit Settings"
button in the key Settings tab for keys they don't own. The button was
gated by `rolesWithWriteAccess.includes(userRole)` (role-only check)
instead of `canModifyKey` (ownership-aware), unlike the Regenerate and
Delete buttons which already used the correct check.

Replace the condition with `canModifyKey` so the Edit Settings button
follows the same proxy-admin / team-admin / key-owner logic as the
other action buttons. Add tests covering all permission paths.
2026-02-23 23:06:27 -08:00
Sameer Kankute
46ed7fc706 Add Additonal header field on UI for testing passthrough 2026-02-24 12:06:07 +05:30
Sameer Kankute
62d5d96e12 Fix: skip health check for MCP integration with passthrough token authentication 2026-02-24 12:05:37 +05:30
yuneng-jiang
f55fe7afdc
Merge pull request #21980 from BerriAI/litellm_ui_testing_coverage_00
[Refactor] UI - Onboarding: Extract testable view components
2026-02-23 22:03:19 -08:00
yuneng-jiang
36f7722b0f fix: add QueryClientProvider, remove stale file, use should naming in tests
- Wrap onboarding page with QueryClientProvider to prevent runtime crash
  (mirrors the same pattern used in LoginPage)
- Stage deletion of stale litellm/ui/litellm-dashboard/src/app/onboarding/page.tsx
  committed at the wrong path
- Rename all 16 test names to start with "should" per AGENTS.md convention
2026-02-23 21:53:42 -08:00
yuneng-jiang
a3491490f9 refactor onboarding 2026-02-23 21:30:08 -08:00
yuneng-jiang
02a53989cf fix: narrow onSubmit values and use semantic loading assertion 2026-02-23 21:20:22 -08:00
yuneng-jiang
0873494270 feat: extract OnboardingFormBody component with tests
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-23 21:08:19 -08:00
₳Ⱡ₥Ø₲
d0bcafacf0
fix(aiohttp): only set enable_cleanup_closed when required (#21897)
* fix(aiohttp): only set enable_cleanup_closed when required

* add tests
2026-02-23 21:06:29 -08:00
yuneng-jiang
a347cf0a33 feat: extract OnboardingErrorView component with tests
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-23 21:06:17 -08:00
yuneng-jiang
a3d4a8752f feat: extract OnboardingLoadingView component with tests
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-23 21:04:48 -08:00
Ishaan Jaff
c79d94fd16
feat(realtime): guardrail hook for voice transcription (#21976)
* feat(realtime): add guardrail hook for voice transcription in Realtime API

Adds a new `realtime_input_transcription` guardrail event hook that fires
after Whisper transcription completes, before the LLM generates a response.

When a guardrail blocks, a synthetic warning is sent to the client and
`response.create` is never forwarded — the LLM never responds.

Also rewrites `create_response: true` → `false` in client `session.update`
so the proxy controls when responses are triggered.

* feat(realtime): speak guardrail block message as audio via TTS

Instead of sending synthetic text events when a guardrail blocks,
send response.create with forced instructions so OpenAI's TTS speaks
the warning message — user hears the block instead of just seeing text.

* fix(realtime): speak exact content filter error message via TTS

Extract the human-readable error string from HTTPException.detail
so the spoken warning says e.g. "Content blocked: keyword 'system update'
detected" instead of the raw str(e) repr.

* fix(realtime): reliably enforce create_response=false for guardrails

- Proxy now injects session.update with create_response=false immediately
  on session.created (when guardrails are active), instead of rewriting
  the client's session.update — works regardless of what the client sends
- Add response.cancel before the warning response.create to kill any
  in-flight LLM response that snuck through before the guardrail fired

* refactor(realtime): call apply_guardrail directly, remove dedicated hook method

The async_realtime_input_transcription_hook in CustomGuardrail and
ContentFilterGuardrail was just a thin wrapper that called apply_guardrail —
the same interface used by /chat and /messages. Remove the wrapper and call
apply_guardrail directly from run_realtime_guardrails, keeping the pattern
consistent across all endpoints.

* docs: add Realtime API guardrails tutorial and flow diagram

* fix: address Greptile review comments

- Forward user_api_key_dict through realtime_api/main.py (_arealtime) so
  it actually reaches RealTimeStreaming instead of always being None
- Run guardrail interception in provider_config path too (e.g. Gemini),
  not only the OpenAI direct path
- Narrow exception catch to HTTPException/ValueError only; re-raise
  unexpected errors so programming bugs surface in logs rather than
  silently appearing as guardrail blocks
- Update tests: mock apply_guardrail directly (hook method was removed),
  replace session.update client-rewrite test with session.created
  injection test matching the new server-side approach

* fix: address latest Greptile review comments

- Remove fastapi import from SDK-layer file; check for status_code/detail
  attrs instead to identify guardrail-block exceptions vs programming errors
- Add store_message() before continue in transcription interception so
  transcription events are logged in the non-provider_config path
- Inject create_response=false on session.created in provider_config path
  (Gemini etc.) to match the OpenAI path — prevents LLM auto-responding
  before guardrail runs on VAD-detected turns
2026-02-23 21:04:40 -08:00
Ishaan Jaff
22fb39ab1e
feat(content-filter): add employment discrimination topic blockers for 5 protected classes (#21962)
Adds YAML topic category files for military_status, disability, age_discrimination,
religion, and gender_sexual_orientation to block employment discrimination prompts
like "Do not hire veterans because they may have mental health issues."

Previously these were not blocked because:
- The prebuilt regex patterns used strict \b word boundaries that didn't match
  plurals (veterans, disabilities, Muslims)
- gender_sexual_orientation pattern was LGBTQ+-focused and missed women/female
- age_discrimination pattern missed "over 50" phrasing
- No conditional (identifier + discriminatory intent) detection existed for these
  protected classes

Each new YAML file uses the bias_racial.yaml pattern: identifier_words (protected
class terms) + additional_block_words (discriminatory employment actions), plus
always_block_keywords for explicit discriminatory phrases. Exceptions prevent false
positives for legitimate diversity programs, accommodation discussions, etc.

Also fixes regex plurals in patterns.json: veterans?, disabilit(y|ies), muslims?,
adds wom[ae]n?/females? to gender pattern, and over\s+\d+ to age pattern.

Evals: 100% precision/recall/F1/accuracy on all 5 new categories (89 total cases,
0 FP, 0 FN). Existing insults and investment evals unaffected.
2026-02-23 21:03:26 -08:00
Nicolò Pignatelli
b8dddab311
feat: add groq/openai/gpt-oss-safeguard-20b model pricing (#21951)
* feat: add groq/openai/gpt-oss-safeguard-20b model pricing

Add pricing and context window data for OpenAI's GPT-OSS-Safeguard-20B
model on Groq, a reasoning model trained for safety classification tasks.

- Input: $0.075/1M tokens
- Cached input: $0.037/1M tokens
- Output: $0.30/1M tokens
- Context window: 131,072 tokens
- Max output: 65,536 tokens

Reference: https://console.groq.com/docs/model/openai/gpt-oss-safeguard-20b

* docs: add gpt-oss-safeguard-20b to Groq provider docs
2026-02-23 21:03:18 -08:00