Commit graph

42355 commits

Author SHA1 Message Date
yuneng-jiang
3284dee734
Merge pull request #25737 from joereyna/fix/remove-missing-mcps-coverage-path
fix: remove non-existent litellm_mcps_tests_coverage from coverage combine
2026-04-14 19:54:13 -07:00
yuneng-jiang
50786007fc
Merge pull request #25736 from BerriAI/docs_visual_guide_for_guardrail_fallbacks
docs update
2026-04-14 19:51:54 -07:00
yuneng-jiang
ffc3a9736d
Merge pull request #25739 from BerriAI/litellm_flakyBedrockGptOssToolCall
[Test] Replace flaky bedrock gpt-oss tool-call live test with request-body mock
2026-04-14 19:49:04 -07:00
joereyna
ccbdaa9187
fix(ci): increase test-server-root-path timeout to 30m 2026-04-14 19:42:10 -07:00
Yuneng Jiang
e2043e11f1
[Test] add request-body mock test for bedrock gpt-oss tool schema
Complements the stubbed-out live integration test by verifying the
outgoing Bedrock Converse request body for GPT-OSS is well-formed when
the caller supplies a tool schema with OpenAI-style metadata
($id, $schema, additionalProperties, strict):
- correct converse URL for bedrock/converse/openai.gpt-oss-20b-1:0
- toolConfig.tools[0].toolSpec has the expected name/description
- inputSchema.json keeps type/properties/required and strips fields
  Bedrock does not accept
2026-04-14 19:36:57 -07:00
Yuneng Jiang
8e44a02a22
[Test] stub flaky bedrock gpt-oss function-calling stream test
GPT-OSS on Bedrock intermittently emits truncated toolUse.input deltas
(e.g. accumulated args of '{"":"'), causing
test_function_calling_with_tool_response to hard-fail on json.loads.
The model flakiness is not a litellm regression: the same base test
passes for Anthropic in the same CI run, and the streaming delta path
at invoke_handler.py has not changed recently.

Follow the existing override pattern in TestBedrockGPTOSS
(test_prompt_caching, test_completion_cost, test_tool_call_no_arguments)
and stub the test to pass. The underlying bedrock converse streaming
tool-call path is already covered by Claude/Nova/Llama Converse suites
in test_bedrock_completion.py and test_bedrock_llama.py, so removing
the live GPT-OSS check loses no unique litellm-side signal.
2026-04-14 19:13:42 -07:00
Yuneng Jiang
d6a69b9c81
[Test] mark bedrock gpt-oss function-calling stream test flaky
Bedrock GPT-OSS occasionally emits truncated toolUse.input deltas
(e.g. accumulated args of '{"":"'), which causes
test_function_calling_with_tool_response to hard-fail on json.loads.
Other overrides in TestBedrockGPTOSS already handle similar
model-side flakiness; apply retries=6 delay=5 scoped to this subclass
so other providers keep strict behavior.
2026-04-14 19:10:55 -07:00
Ishaan Jaffer
84c507bc14
fix(mypy): use explicit None check for cache rate values to satisfy type checker 2026-04-14 19:03:56 -07:00
joereyna
a01cf44c35
fix: remove non-existent litellm_mcps_tests_coverage from coverage combine 2026-04-14 18:59:25 -07:00
Yuneng Jiang
38f8d7a008
Point contributors toward litellm_oss_branch in guard error messages 2026-04-14 18:41:59 -07:00
Yuneng Jiang
ab71d3d700
Also reject PRs from forks, not just non-allowlisted branches 2026-04-14 18:39:54 -07:00
shivam
fd110cd5cf
docs update 2026-04-14 18:33:42 -07:00
Ishaan Jaffer
e20d9df1b6
remove test file 2026-04-14 18:29:15 -07:00
Ishaan Jaffer
e0a988e39a
feat(ui/log-details): pass rawInputTokens, cacheReadTokens, cacheCreationTokens to CostBreakdownViewer from SpendLogs 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
0148effd6e
feat(ui/cost-breakdown): show separate Input / Cache Read / Cache Write line items in cost breakdown drawer 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
6b1dc1156e
fix(ui/usage): subtract cache tokens from Input Tokens summary card to avoid double-counting 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
781fc6311b
feat(cost-calculator): compute and store per-type cache costs in CostBreakdown (cache_read_cost, cache_creation_cost) 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
b5a4c26248
feat(logging): pass cache_read_cost and cache_creation_cost through set_cost_breakdown 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
c84597ecd0
fix(bedrock/converse): capture raw input_tokens as text_tokens before cache inflation in PromptTokensDetailsWrapper 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
5c056cae9f
fix(anthropic): store raw text_tokens in PromptTokensDetailsWrapper before cache inflation 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
c1dcfa70c9
feat(types): add cache_read_cost and cache_creation_cost fields to CostBreakdown TypedDict 2026-04-14 18:29:12 -07:00
Ishaan Jaffer
d805fe7103
test(bedrock): add unit tests for cache token billing with prompt caching
Adds TestBedrockInvokeCacheTokenBilling covering the Bedrock InvokeModel path:
- baseline: no cache tokens, prompt_tokens equals input_tokens
- cache_read: prompt_tokens inflated by design, prompt_tokens_details carries breakdown
- cache_creation: same pattern for write tokens
- cost_calculation_correct_with_cache_read: core billing regression test
- cost_calculation_correct_with_cache_creation: write-rate billing regression test
- back_to_back_requests_cost: full end-to-end scenario (cache write then read)

These lock in the fix from PR #25517 - cache tokens were being double-counted
in AnthropicConfig.calculate_usage causing 10-50x inflated cost on cache reads.
2026-04-14 18:29:12 -07:00
Yuneng Jiang
45d1e1b341
[Infra] Guard main branch with PR source-branch check
Adds a GHA that fails PRs to main unless the head branch is
'litellm_internal_staging' or 'litellm_hotfix_*'. Also fails merge_group
events since merge queue is not in use.
2026-04-14 18:19:14 -07:00
yuneng-jiang
5c1f7d99bf
Merge pull request #25731 from BerriAI/docs_guardrail
fallbacks image
2026-04-14 18:13:12 -07:00
shivam
65ce89dc67
update 2026-04-14 18:02:41 -07:00
yuneng-jiang
ec0953f7b0
Merge pull request #25728 from BerriAI/litellm_fix_together_ai_test_model
[Fix] Test - Together AI: replace deprecated Mixtral with serverless Qwen3.5-9B
2026-04-14 18:01:56 -07:00
shivam
19629004f5
fallbacks image 2026-04-14 17:58:11 -07:00
harish876
d20c70f24c Optimize database query which fetches latest model_id, model_name pairs and dedupes them in memory.
Current fix includes
 - Updates test case
 - Optimized query with docstring. The change leverages deduplication and sorting logic from SQL
 - Added a bench script to differentiate peak memory usage before and after
2026-04-15 00:54:37 +00:00
Yuneng Jiang
045d32a242
bump: version 1.83.7 → 1.83.8 2026-04-14 17:47:24 -07:00
Yuneng Jiang
a9c6156137
[Fix] Test - Together AI: replace deprecated Mixtral with serverless Qwen3.5-9B
Mixtral-8x7B-Instruct-v0.1 is no longer on Together AI's serverless tier
and now requires a dedicated endpoint, causing multiple tests to fail in CI:

  - test_together_ai.py::TestTogetherAI::test_empty_tools
  - test_completion.py::test_completion_together_ai_stream
  - test_completion.py::test_customprompt_together_ai
  - test_completion.py::test_completion_custom_provider_model_name
  - test_text_completion.py::test_async_text_completion_together_ai

Qwen/Qwen3.5-9B is currently serverless on Together AI and supports
function calling, satisfying BaseLLMChatTest capability requirements.
2026-04-14 17:43:35 -07:00
yuneng-jiang
2af0768ab9
Merge pull request #25727 from BerriAI/litellm_removeChatUiSwaggerLink
Some checks are pending
CodeQL / Analyze (actions) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
CodSpeed Benchmarks / benchmarks (push) Waiting to run
Helm unit test / unit-test (push) Waiting to run
Read Version from pyproject.toml / read-version (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
Unit Tests: Security / security (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
[Refactor] Remove Chat UI link from Swagger docs message
2026-04-14 17:36:52 -07:00
yuneng-jiang
bb91f3ace9
Merge pull request #25726 from BerriAI/docs_yj_apr14
[Docs] Regenerate v1.83.3-stable release notes from previous stable
2026-04-14 17:29:34 -07:00
user
2911d99d77
test(gemini): stub API key for format param tests 2026-04-15 00:28:39 +00:00
Yuneng Jiang
58ce769092
Remove Chat UI link from Swagger docs message 2026-04-14 17:27:25 -07:00
Yuneng Jiang
05ad48236f
[Docs] Regenerate v1.83.3-stable release notes from v1.82.3-stable baseline
The previous v1.83.3 changelog was generated against v1.83.0-nightly and
missed ~3 weeks of work. This regenerates it against the previous stable
release and restructures the LLM API Endpoints section to group by API
type (Responses, Batch, Count Tokens, Video Generation, Pass-Through,
etc.) matching the convention used in v1.82.3, v1.82.0, and v1.81.14.
Adds ~25 previously uncited PRs, cross-section duplications for
cross-cutting changes, and a verified first-time-contributors list.
2026-04-14 17:19:42 -07:00
user
b1bc3c166d
test(prompts): isolate in-memory version tests 2026-04-14 23:37:13 +00:00
ryan-crabbe-berri
6d2b7b76ad
Merge pull request #25721 from BerriAI/litellm_fix-invite-user-default-role
fix: default invite user modal global role to least-privilege
2026-04-14 16:34:04 -07:00
yuneng-jiang
260d51ad67
Merge pull request #25723 from BerriAI/litellm_release_notes_v1_83_3_v1_83_7rc1
[Docs] Add release notes for v1.83.3-stable and v1.83.7.rc.1
2026-04-14 16:32:17 -07:00
user
f521e27371
test(gemini): align API key expectations 2026-04-14 23:28:13 +00:00
Ryan Crabbe
3aae15f5d8
[Docs] Use GitHub avatar for Ryan Crabbe in release notes
Replace the expiring LinkedIn CDN image URL with a stable GitHub
avatar URL for v1.83.3 and v1.83.7.rc.1 release notes.
2026-04-14 16:22:07 -07:00
Yuneng Jiang
966be2982a
[Docs] Add missed content PRs to v1.83.7.rc.1 and update runbook
- Add 8 content PRs that merged directly to the release branch outside the listed staging PRs: #23769 (Ramp callback), #25252 (JWT OAuth2 override), #25254 (AWS GovCloud mode), #25258 (batch-limit cleanup), #25334 (router custom_llm_provider), #25345 (Triton embeddings), #25347 (tag-based routing), #25358 (Baseten pricing attribution)
- Add @kedarthakkar to new contributors (first-ever PR via #23769)
- Update RELEASE_NOTES_GENERATION_INSTRUCTIONS: require walking git log range between release tags in addition to staging PRs, and verify new-contributor status per author rather than trusting the GH release body floor
2026-04-14 16:13:09 -07:00
user
abb8d8d454
ci: trigger CI re-run for codecov
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 23:09:17 +00:00
user
b16d0b1d5e
test: add coverage for credential leak prevention changes
Add 50 tests across 3 files covering the new MaskedHTTPStatusError,
safe response helpers, _redact_string in error paths, Gemini
interactions x-goog-api-key header auth, and RAG ingestion header
usage.

Fix missing early-validation for Gemini API key in _get_token_and_url()
which caused TypeError when key was None (headers got None value).
Harmonize error messages between the two validation sites.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 23:09:17 +00:00
user
74f55b0671
fix: apply mask_sensitive_info to async streaming error body
Addresses Greptile review — the async streaming path was missing
mask_sensitive_info() on the response body, while the sync path had it.
2026-04-14 23:09:17 +00:00
user
25f93bed91
security: prevent API key leaks in error tracebacks, logs, and alerts
Gemini API keys embedded in URLs as ?key= query parameters leak through
httpx error tracebacks, which are then captured by traceback.format_exc()
and forwarded to logging callbacks, Slack/Teams alerts, and HTTP client
responses.

Short-term: all httpx.HTTPStatusError handlers now raise
MaskedHTTPStatusError(...) from None, which masks the URL and breaks
exception chaining so the original error never appears in tracebacks.

Long-term: moved all Gemini/Vertex URL constructions from ?key={api_key}
to x-goog-api-key header (Google's documented auth method), so the key
is never in the URL at all. WebSocket realtime is the only exception
since WS clients cannot use custom headers.

Additionally hardened all outbound credential paths:
- WebSocket close reasons now pass through _redact_string()
- Callback pipeline (failure_handler) redacts traceback_exception and
  error_str before forwarding to integrations (Langfuse, Datadog, etc.)
- Slack/Teams alert messages redacted in send_llm_exception_alert,
  ProxyLogging.failure_handler, and post_call_failure_hook
- HTTP error responses in proxy SSE and health endpoints redacted
- Exception messages in exception_mapping_utils redacted
- print_verbose() stdout output redacted when set_verbose=True
- HTTPHandler.put() now has MaskedHTTPStatusError (was missing)
2026-04-14 23:09:17 +00:00
Yuneng Jiang
4a1da629fa
[Fix] Correct pip install versions for v1.83.3-stable and v1.83.7.rc.1 docs
PyPI publishes 1.83.3 and 1.83.7 (no .post1 / rc1 suffixes) — align the pip install commands with the actual published versions.
2026-04-14 16:00:27 -07:00
Yuneng Jiang
8eec2c69b7
[Docs] Add release notes for v1.83.3-stable and v1.83.7.rc.1
- Retitle existing v1.83.3 preview file to v1.83.3-stable (same commit)
- Add new v1.83.7.rc.1 preview release notes
- Update RELEASE_NOTES_GENERATION_INSTRUCTIONS runbook with guidance on resolving staging PRs to their underlying commits
2026-04-14 15:58:13 -07:00
Ryan Crabbe
a428ae7599
fix: default invite user modal global role to least-privilege
Pre-select "Internal User Viewer" in the Global Proxy Role dropdown
on both the standalone and embedded Invite User forms so admins don't
have to remember to pick a role, and the default lands on the least
privileged option rather than silently posting an undefined role.
2026-04-14 15:45:35 -07:00
yuneng-jiang
7b36cfc0de
Merge pull request #25719 from BerriAI/litellm_fix-ui-unit-tests-get-cookie
test(ui): add getCookie to cookieUtils mock in user_dashboard test
2026-04-14 15:32:13 -07:00
Ryan Crabbe
bdaaa5c187
test(ui): add getCookie to cookieUtils mock in user_dashboard test
user_dashboard.tsx imports getCookie from @/utils/cookieUtils, but the
vi.mock factory in user_dashboard.test.tsx only exports clearTokenCookies.
Vitest throws `No "getCookie" export is defined on the "@/utils/cookieUtils"
mock`, breaking all three beforeunload-listener tests.

Add getCookie to the mock factory so it matches the current imports.
2026-04-14 15:18:10 -07:00