Commit graph

36905 commits

Author SHA1 Message Date
Yuneng Jiang
d4288b4ff4
ci: fix LITELLM_LICENSE interpolation in e2e_ui_testing
Remove LITELLM_LICENSE from the run step's environment block — YAML
environment maps may pass the literal string "${LITELLM_LICENSE}" instead
of interpolating the project env var, overriding it with a value that
fails license validation. The project-level env var is inherited
automatically by the proxy process.
2026-04-09 21:45:34 -07:00
Ryan Crabbe
26e99f22b3
refactor: consolidate route auth for UI and API tokens
Unify UI and API token authorization through the shared RBAC path
and backfill missing routes in role-based route lists.
2026-04-09 21:36:35 -07:00
yuneng-jiang
9e6d2d2069
Merge pull request #25468 from BerriAI/litellm_/compassionate-shannon
[Test] UI - Unit tests: raise global vitest timeout and remove per-test overrides
2026-04-09 21:35:43 -07:00
Yuneng Jiang
9071dbba12
fix(ui): remove leftover setIsOwnKey call after state removal 2026-04-09 21:24:22 -07:00
yuneng-jiang
cc43d09d79
Potential fix for pull request finding 'CodeQL / Unused variable, import, function or class'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-04-09 21:19:02 -07:00
Yuneng Jiang
c7f610c57e
Merge remote-tracking branch 'origin/main' into litellm_/compassionate-shannon 2026-04-09 21:15:03 -07:00
Yuneng Jiang
71640f062c
Merge remote-tracking branch 'origin/main' into litellm_regen_key_modal_antd 2026-04-09 21:14:57 -07:00
yuneng-jiang
42e5583788
Merge pull request #25471 from BerriAI/litellm_doc_mcp_per_user_token_env_vars
[Docs] Add missing MCP per-user token env vars to config_settings
2026-04-09 21:13:10 -07:00
Yuneng Jiang
ee374c4884
ci: pass LITELLM_LICENSE to e2e_ui_testing proxy
Key regeneration is an enterprise feature — without LITELLM_LICENSE the
endpoint returns a 403 and the Playwright test for "Regenerate key"
never sees the success view. Other CircleCI jobs already pass this
secret; the e2e_ui_testing job was missing it.
2026-04-09 21:09:29 -07:00
Yuneng Jiang
ce0b57b4ff
[Docs] Add missing MCP per-user token env vars to config_settings
MCP_PER_USER_TOKEN_DEFAULT_TTL and MCP_PER_USER_TOKEN_EXPIRY_BUFFER_SECONDS
were added in #25441 but not documented, causing test_env_keys.py to fail.
2026-04-09 21:04:34 -07:00
yuneng-jiang
1214dc93b3
Merge pull request #25467 from jaydns/fix/parameterize-combined-view-query
fix(proxy): use parameterized query for combined_view token lookup
2026-04-09 21:03:24 -07:00
yuneng-jiang
9b33d9d427
Merge pull request #25445 from jaydns/fix/proxy-input-validation
fix(proxy): improve input validation on management endpoints
2026-04-09 21:02:52 -07:00
Yuneng Jiang
d0168bcff1
ci: retrigger e2e 2026-04-09 20:57:03 -07:00
Yuneng Jiang
92dbd2c491
address greptile review feedback (greploop iteration 1)
Remove leftover 10000ms per-test timeout in add_model_tab.test.tsx that was
missed in the initial sweep. The test now inherits the 30000ms global.
2026-04-09 20:54:36 -07:00
yuneng-jiang
aa0fa104ba
Merge pull request #25437 from joereyna/litellm_fix_responses_websocket_model_query_param
fix(responses-ws): append ?model= to backend WebSocket URL
2026-04-09 20:46:47 -07:00
Yuneng Jiang
3a316b9131
[Test] UI - Unit tests: raise global vitest timeout and remove per-test overrides
Raise vitest testTimeout from 10s to 30s and drop per-test timeout overrides
across UI unit tests. Group CreateUserButton and TeamInfo tests under nested
describe blocks to make the most flaky suites easier to scan.
2026-04-09 20:42:30 -07:00
Yuneng Jiang
1d50f774e2
fix(ui): support all duration suffixes in regenerate expiry preview
calculateNewExpiryTime only handled s/h/d, but the grace-period
validation and backend accept m, w, and mo as well. Entering any of
those in the Expire Key field caused the function to return null,
which then propagated as expires: null in the onKeyUpdate payload —
the parent UI would then render the expiry as "Never" even though
the backend had correctly applied the new expiry.

Extend the suffix check to cover s/m/h/d/w/mo, matching "mo" before
"m" so "1mo" isn't misread as minutes. Also nullish-coalesce the
call site so an unparseable duration falls back to the previous
expiry instead of null. Add parametric tests for each supported
suffix plus a regression test for the null fallback.
2026-04-09 20:31:35 -07:00
Yuneng Jiang
f95ef935ef
fix(ui): prefer form values over API echo in regenerate update payload
The regenerate endpoint returns a GenerateKeyResponse that inherits
max_budget/tpm_limit/rpm_limit from KeyRequestBase, so the API echoes
the existing values back. The previous updatedKeyData layout spread
...response *after* the explicit formValues assignments, which meant
the user's just-submitted edits were silently overwritten by the API
echo before being propagated to the parent via onKeyUpdate.

Reorder so the response spread comes first and the formValues-derived
fields override it, and add a regression test that mocks a response
with stale limits to lock the behavior in. Also drop the two leftover
debug console.log statements.
2026-04-09 20:25:10 -07:00
Yuneng Jiang
839d9bd5f3
refactor(ui): polish regenerate key success view
- Label the key block with a small "Virtual Key" caption so the gray
  box is clearly the key container.
- Move the Copy Key action to the modal footer as a primary button
  with icon; inline copy icon next to the key is removed.
- Swap the button to "Copied" with a check icon on success instead of
  firing a notification — less noisy and keeps feedback in place.
- Disable clicking outside the modal to close (maskClosable=false) so
  users must explicitly dismiss via Close or X.
- Enlarge the key text and let its container span the full modal
  width.
- Tests updated accordingly, including a new test for the copied-state
  swap and the "Virtual Key" label.
2026-04-09 20:14:01 -07:00
jayden
4dc416ee74
fix(proxy): use parameterized query for combined_view token lookup 2026-04-09 20:10:40 -07:00
yuneng-jiang
ce75598f97
Merge pull request #25384 from BerriAI/litellm_/bold-pare
[Fix] UI: improve storage handling and Dockerfile consistency
2026-04-09 19:55:24 -07:00
shivam
5c4915ad0d
fix(proxy): pass-through multipart uploads and Bedrock custom body
- Route multipart forwarding on forward_multipart instead of empty _parsed_body
  so litellm_logging_obj no longer forces json= for file uploads.
- Remove custom_body from pass-through endpoint signatures; FastAPI treated it
  as a JSON body and rejected multipart before the handler ran. Bedrock passes
  JSON via request.state (LITELLM_PASS_THROUGH_CUSTOM_BODY_STATE_KEY).
- Use build_request + send(stream=True) for streaming multipart; httpx 0.28
  AsyncClient.request does not accept stream=.
- Add regression test for non-empty _parsed_body multipart path; update Bedrock
  custom-body test and query-params test for forward_multipart.

Made-with: Cursor
2026-04-09 19:43:57 -07:00
ishaan-berri
2c0d20b327
Merge pull request #25441 from csoni-cweave/per-user-mcp-oauth-token
feat(mcp): add per-user OAuth token storage for interactive MCP flows
2026-04-09 19:03:59 -07:00
Yuneng Jiang
15f7cc9134
refactor(ui): replace success-view divs in regenerate key modal with antd
Use Flex, Typography.Paragraph (with copyable), and Typography.Text
instead of raw divs + code block + CopyToClipboard wrapper. Drops the
direct react-copy-to-clipboard dependency in this component in favor
of antd's native copyable support.

Also fixes two test issues surfaced when running the e2e locally:
- RegenerateKeyModal.test.tsx no longer mocks react-copy-to-clipboard
  (the component no longer imports it), removing the CJS require()
  inside an ESM mock factory flagged by Greptile.
- keys.spec.ts scopes the Regenerate and Copy lookups to the modal.
  The Regenerate button has an icon whose aria-label ("sync") is
  concatenated into the button's accessible name, so an exact-match
  lookup on "Regenerate" failed; and the new Paragraph copyable
  renders a generic "Copy" button that collided with the other
  copyable fields on the key info view.
2026-04-09 18:43:28 -07:00
shivam
b6357cd986
fix(proxy): reject non-admin spend log detail when DB is unavailable
Non-admins previously skipped RBAC when prisma_client was None but could
still read payloads from custom loggers. Return 403 unless admin view.
Add test_ui_view_request_response_forbids_non_admin_without_db.

Made-with: Cursor
2026-04-09 18:39:33 -07:00
shivam
1f474d5bb3
fix(proxy): spend logs RBAC—avoid common_utils cycle, tighten ownership check
- Drop module-level common_utils import; import team helpers inside callers.
- Inline admin-view role check in _is_admin_view_safe to break import cycle.
- Require non-null row.user before treating spend log as owned by the key
  (fixes None==None bypass for service keys).
- Document deferred proxy_server imports in _get_permitted_team_ids_for_spend_logs.
- Update tests (common_utils patches, regression test, ruff cleanups).

Made-with: Cursor
2026-04-09 18:34:23 -07:00
shivam
288ccb39c0
resolved greptile comments 2026-04-09 18:20:05 -07:00
shivam
31f750146b
added option to allow team user to see logs of team 2026-04-09 17:43:20 -07:00
yuneng-jiang
e7551a1d43
Merge pull request #25444 from joereyna/litellm_fix_vertex_fine_tuned_model_test
fix(test): mock headers in test_completion_fine_tuned_model
2026-04-09 17:36:48 -07:00
harish876
13108039c8 remove conftest patch. TODO: make a different PR for this 2026-04-09 22:29:49 +00:00
joereyna
afd46e7a16
format vertex test file 2026-04-09 15:28:34 -07:00
harish876
7ebc144c18 Add file content streaming support for OpenAI and related utilities
- Introduced `afile_content_streaming` and `file_content_streaming` functions in `litellm/files/main.py` to handle asynchronous and synchronous file content streaming.
- Added `FileContentStreamingResponse` class in `litellm/files/streaming.py` to manage streaming responses with logging capabilities.
- Updated OpenAI API integration in `litellm/llms/openai/openai.py` to support new streaming methods.
- Enhanced file content retrieval in `litellm/proxy/openai_files_endpoints/files_endpoints.py` to route requests for streaming.
- Added unit tests for the new streaming functionality in `tests/test_litellm/llms/openai/test_openai_file_content_streaming.py` and `tests/test_litellm/proxy/openai_files_endpoint/test_files_endpoint.py`.
- Refactored type hints and imports for better clarity and organization across modified files.
2026-04-09 22:14:46 +00:00
jayden
d910a95661
fix(proxy): improve input validation on management endpoints 2026-04-09 14:14:53 -07:00
joereyna
1571f5e45a
fix(test): mock headers in test_completion_fine_tuned_model 2026-04-09 13:18:35 -07:00
Chetan Soni
ce2add3b16 feat(mcp): add per-user OAuth token storage for interactive MCP flows 2026-04-09 12:42:42 -07:00
Krrish Dholakia
3a6db708ce
docs: add Docker Image Security Guide for cosign verification and deployment best practices (#25439)
Some checks are pending
CodeQL / Analyze (actions) (push) Waiting to run
CodeQL / Analyze (javascript-typescript) (push) Waiting to run
CodeQL / Analyze (python) (push) Waiting to run
CodSpeed Benchmarks / benchmarks (push) Waiting to run
Helm unit test / unit-test (push) Waiting to run
Read Version from pyproject.toml / read-version (push) Waiting to run
Scorecard supply-chain security / Scorecard analysis (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Security / security (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
- New doc page covering all signed image variants, verification commands,
  CI/CD enforcement (K8s Sigstore Policy Controller, GCP Binary Authorization,
  AWS/EKS, GitHub Actions), digest pinning, and safe upgrade patterns
- Added to sidebar under Setup & Deployment
- Cross-linked from the existing deploy.md cosign section

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Krrish Dholakia <krrish-berri-2@users.noreply.github.com>
2026-04-09 11:50:15 -07:00
Krrish Dholakia
f8243eee88
Revert "fix(proxy): set key_alias=user_id in JWT auth for Prometheus metrics …" (#25438)
This reverts commit 8d945c86b7.
2026-04-09 11:32:14 -07:00
joereyna
3ac4333be1
fix(responses-ws): use urllib.parse to append model param, fix test mocking 2026-04-09 11:14:34 -07:00
joereyna
f6dde296fa
fix(responses-ws): append ?model= to backend WebSocket URL 2026-04-09 10:33:20 -07:00
Abhijoy Sarkar
c688d9d6bc
Add PromptGuard guardrail integration (#24268)
* Add PromptGuard guardrail integration

Add PromptGuard as a first-class guardrail vendor in LiteLLM's proxy,
supporting prompt injection detection, PII redaction, topic filtering,
entity blocklists, and hallucination detection via PromptGuard's
/api/v1/guard API endpoint.

Backend:
- Add PROMPTGUARD to SupportedGuardrailIntegrations enum
- Implement PromptGuardGuardrail (CustomGuardrail subclass) with
  apply_guardrail handling allow/block/redact decisions
- Add Pydantic config model with api_key, api_base, ui_friendly_name
- Auto-discovered via guardrail_hooks/promptguard/__init__.py registries

Frontend:
- Add PromptGuard partner card to Guardrail Garden with eval scores
- Add preset configuration for quick setup
- Add logo to guardrailLogoMap

Tests:
- 30 unit tests covering configuration, allow/block/redact actions,
  request payload construction, error handling, config model, and
  registry wiring

* Fix redact path and init ordering per review feedback

- P1: Update structured_messages (not just texts) when PromptGuard
  returns a redact decision, so PII redaction is effective for the
  primary LLM message path
- P2: Validate credentials before allocating the HTTPX client so
  resources aren't acquired if PromptGuardMissingCredentials is raised
- Add tests for structured_messages redaction and texts-only redaction

* Harden PromptGuard integration: fail-open, event hooks, images, docs

- Add block_on_error config (default fail-closed, configurable fail-open)
- Declare supported_event_hooks (pre_call, post_call) like other vendors
- Forward images from GenericGuardrailAPIInputs to PromptGuard API
- Wrap API call in try/except for resilient error handling
- Add comprehensive documentation page with config examples
- Register docs page in sidebar alongside other guardrail providers
- Expand test suite from 32 to 40 tests covering new functionality

* Fix dict[str, Any] -> Dict[str, Any] for Python 3.8 compat

* Address remaining Greptile feedback: timeout, redact guard

- Add explicit 10s timeout to async_handler.post() to prevent
  indefinite hangs when PromptGuard API is unresponsive
- Guard redact path: only update inputs["texts"] when the key
  was originally present, avoiding phantom key injection
- Add test: redact with structured_messages only does not create
  texts key (41 tests total)

* Fix CI lint: black formatting, add PromptGuardConfigModel to LitellmParams

- Reformat promptguard.py to match CI black version (parenthesization)
- Add PromptGuardConfigModel as base class of LitellmParams for proper
  Pydantic schema validation, consistent with all other guardrail vendors
- Use litellm_params.block_on_error directly (now a typed field)

* Address Greptile review: redact path, null decision, error context

- P1: Filter _extract_texts_from_messages to user-role messages only,
  preventing system/assistant content from being injected into texts
- P1: Strengthen test_redact_updates_structured_messages assertion from
  weak `in` check to strict equality, catching the injection bug
- P2: Use `result.get("decision") or "allow"` to handle explicit null
  decision values (not just absent keys)
- P2: Wrap bare exception re-raise in GuardrailRaisedException so the
  caller knows which guardrail failed (block_on_error=True path)
- P2: Add static Promptguard entry in guardrail_provider_map so the
  preset works before populateGuardrailProviderMap is called
- Add test for explicit null decision treated as allow

* Fix black formatting: collapse f-string in error message
2026-04-09 08:12:24 -07:00
michelligabriele
cd9c511df6
feat(proxy): add credential overrides per team/project via model_config metadata (#24438) 2026-04-09 07:22:27 -07:00
Sameer Kankute
cb057ad44b
fix(websearch_interception): ensure spend/cost logging runs when stream=True
The deployment hook now converts stream=True→False in wrapper_async's
scope so the streaming early-return path is skipped and logging executes.
logging_obj.stream is synced after the hook, and the original stream
intent is recovered for the short-circuit path.

Made-with: Cursor
2026-04-09 18:48:38 +05:30
Yuneng Jiang
e42baeb5ab
[Refactor] UI - Virtual Keys: migrate regenerate key modal to AntD
Replace Tremor components in the regenerate key modal with Ant Design
equivalents and move the component to a new PascalCase file. The form
layout now uses Row/Col to place Max Budget, TPM Limit, and RPM Limit
on one row and Expire Key with Grace Period on another, reducing the
vertical footprint. The success view shows an Alert banner, the key
alias as secondary context, and the regenerated key in a monospace
block with an inline primary Copy button.

Also adds unit tests for the new component and updates the existing
Playwright spec to match the new banner and button text.
2026-04-08 23:52:25 -07:00
Yuneng Jiang
20ed120d1a
[Fix] Let setSecureItem propagate storage errors to callers
Remove the silent try/catch from setSecureItem so OAuth hooks can
surface actionable "enable storage" guidance instead of a cryptic
"state lost" error after the round-trip. Add a local try/catch in
ChatUI where the storage write is non-critical.
2026-04-08 22:13:55 -07:00
Sameer Kankute
97f722f558
feat(cost): add baseten model api pricing entries (#25358)
Add Baseten Model API pricing entries for Nemotron, GLM, Kimi, GPT OSS, and DeepSeek models with validated model slugs. Include a focused regression test to assert provider and per-token pricing values.

Made-with: Cursor
2026-04-08 21:39:58 -07:00
Krrish Dholakia
f42ffed2bd
Litellm oss staging 04 02 2026 p1 (#25055)
* fix(vertex_ai): support pluggable (executable) credential_source for WIF auth (#24700)

The WIF credential dispatch in load_auth() only handled identity_pool and
aws credential types. When credential_source.executable was present (used
for Azure Managed Identity via Workload Identity Federation), it fell
through to identity_pool.Credentials which rejected it with MalformedError.

Add dispatch to google.auth.pluggable.Credentials for executable-type
credential sources, following the same pattern as the existing identity_pool
and aws helpers.

Fixes authentication for Azure Container Apps → GCP Vertex AI via WIF
with executable credential sources.

* feat(logging): add component and logger fields to JSON logs for 3rd p… (#24447)

* feat(logging): add component and logger fields to JSON logs for 3rd party filtering

* Let user-supplied extra fields win over auto-generated component/logger, tighten test assertions

* Feat - Add organization into the metrics metadata for org_id & org_alias (#24440)

* Add org_id and org_alias label names to Prometheus metric definitions

* Add user_api_key_org_alias to StandardLoggingUserAPIKeyMetadata

* Populate user_api_key_org_alias in pre-call metadata

* Pass org_id and org_alias into per-request Prometheus metric labels

* Add test for org labels on per-request Prometheus metrics

* chore: resolve test mockdata

* Address review: populate org_alias from DB view, add feature flag, use .get() for org metadata

* Add org labels to failure path and verify flag behavior in test

* Fix test: build flag-off enum_values without org fields

* Gate org labels behind feature flag in get_labels() instead of static metric lists

* Scope org label injection to metrics that carry team context, remove orphaned budget label defs, add test teardown

* Use explicit metric allowlist for org label injection instead of team heuristic

* Fix duplicate org label guard, move _org_label_metrics to class constant

* Reset custom_prometheus_metadata_labels after duplicate label assertion

* fix: emit org labels by default, remove flag, fix missing org_alias in all metadata paths

* fix: emit org labels by default, no opt-in flag required

* fix: write org_alias to metadata unconditionally in proxy_server.py

* fix: 429s from batch creation being converted to 500 (#24703)

* add us gov models (#24660)

* add us gov models

* added max tokens

* Litellm dev 04 02 2026 p1 (#25052)

* fix: replace hardcoded url

* fix: Anthropic web search cost not tracked for Chat Completions

The ModelResponse branch in response_object_includes_web_search_call()
only checked url_citation annotations and prompt_tokens_details, missing
Anthropic's server_tool_use.web_search_requests field. This caused
_handle_web_search_cost() to never fire for Anthropic Claude models.

Also routes vertex_ai/claude-* models to the Anthropic cost calculator
instead of the Gemini one, since Claude on Vertex uses the same
server_tool_use billing structure as the direct Anthropic API.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* fix(anthropic): pass logging_obj to client.post for litellm_overhead_time_ms (#24071)

When LITELLM_DETAILED_TIMING=true, litellm_overhead_time_ms was null for
Anthropic because the handler did not pass logging_obj to client.post(),
so track_llm_api_timing could not set llm_api_duration_ms. Pass
logging_obj=logging_obj at all four post() call sites (make_call,
make_sync_call, acompletion, completion). Add test to ensure make_call
passes logging_obj to client.post.

Made-with: Cursor

* sap - add additional parameters for grounding

- additional parameter for grounding added for the sap provider

* sap - fix models

* (sap) add filtering, masking, translation SAP GEN AI Hub modules

* (sap) add tests and docs for new SAP modules

* (sap) add support of multiple modules config

* (sap) code refactoring

* (sap) rename file

* test(): add safeguard tests

* (sap) update tests

* (sap) update docs, solve merge conflict in transformation.py

* (sap) linter fix

* (sap) Align embedding request transformation with current API

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) mock commit

* (sap) run black formater

* (sap) add literals to models, add negative tests, fix test for tool transformation

* (sap) fix formating

* (sap) fix models

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) commit for rerun bot review

* (sap) minor improve

* (sap) fix after bot review

* (sap) lint fix

* docs(sap): update documentation

* fix(sap): change creds priority

* fix(sap): change creds priority

* fix(sap): fix sap creds unit test

* fix(sap): linter fix

* fix(sap): linter fix

* linter fix

* (sap) update logic of fetching creds, add additional tests

* (sap) clean up code

* (sap) fix after review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) add a possibility to put the service key by both variants

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) update test

* (sap) update service key resolve function

* (sap) run black formater

* (sap) fix validate credentials, add negative tests for credential fetching

* (sap) fix validate credentials, add negative tests for credential fetching

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) fix after bot review

* (sap) lint fix

* (sap) lint fix

* feat: support service_tier in gemini

* chore: add a service_tier field mapping from openai to gemini

* fix: use x-gemini-service-tier header in response

* docs: add service_tier to gemini docs

* chore: add defaut/standard mapping, and some tests

* chore: tidying up some case insensitivity

* chore: remove unnecessary guard

* fix: remove redundant test file

* fix: handle 'auto' case-insensitively

* fix: return service_tier on final steamed chunk

* chore: black

* feat: enable supports_service_tier to gemini models

* Fix get_standard_logging_metadata tests

* Fix test_get_model_info_bedrock_models

* Fix test_get_model_info_bedrock_models

* Fix remaining tests

* Fix mypy issues

* Fix tests

* Fix merge conflicts

* Fix code qa

* Fix code qa

* Fix code qa

* Fix greptile review

---------

Co-authored-by: michelligabriele <gabriele.michelli@icloud.com>
Co-authored-by: Josh <36064836+J-Byron@users.noreply.github.com>
Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: milan-berri <milan@berri.ai>
Co-authored-by: Alperen Kömürcü <alperen.koemuercue@sap.com>
Co-authored-by: Vasilisa Parshikova <vasilisa.parshikova@sap.com>
Co-authored-by: Lin Xu <lin.xu03@sap.com>
Co-authored-by: Mark McDonald <macd@google.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>
2026-04-08 21:37:10 -07:00
Austin Varga
541e81de2f
fix: expose reasoning effort fields in get_model_info + add together_ai/gpt-oss-120b (#25263)
* fix: expose reasoning effort fields in get_model_info and add together_ai/gpt-oss-120b

- litellm/utils.py: pass supports_none_reasoning_effort and
  supports_xhigh_reasoning_effort through _get_model_info_helper so
  get_model_info() returns them (previously silently dropped). Fixes #25096.

- model_prices_and_context_window.json: add together_ai/openai/gpt-oss-120b
  with supports_reasoning: true so reasoning_effort is accepted for this
  model without requiring drop_params. Fixes #25132.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: consolidate duplicate together_ai/openai/gpt-oss-120b entry and sync backup file

* fix: link commit to GitHub account for CLA verification

---------

Co-authored-by: Austin Varga <austin@knowmi.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-08 21:34:03 -07:00
kejunleng
4e32479e7d
feat(dashscope): preserve cache_control for explicit prompt caching (#25331)
DashScope inherits OpenAIGPTConfig which strips cache_control from
messages and tools by default. Override remove_cache_control_flag_from_messages_and_tools()
to preserve cache_control, following the same pattern used by ZAI, MiniMax, and Databricks.

Verified through 10-round multi-turn conversation tests:
- Explicit caching works correctly: cached_tokens grows each round from R4 onwards,
  with cache_creation_tokens reported on first cache build.
- Implicit caching is not affected: models that rely on implicit prefix-matching caching
  produce identical cached_tokens with and without this change, confirmed by comparing
  results against both the reverted codebase and direct API calls bypassing litellm.
- No errors or regressions observed on any model, including those that do not support
  explicit caching — the DashScope API silently ignores unrecognized cache_control fields.

Fixes #25330

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 21:32:04 -07:00
milan-berri
e0a578fbdd
fix: remove leading space from license public_key.pem (#25339)
* fix: remove leading space from license public_key.pem

PEM must begin with -----BEGIN; a leading ASCII space breaks
cryptography.load_pem_public_key on older cryptography (e.g. 41.x),
causing OpenSSL no start line / deserialize errors.

Made-with: Cursor

* test: assert license public_key.pem loads as valid PEM

Regression guard for leading whitespace before -----BEGIN, which breaks
load_pem_public_key on older cryptography (e.g. 41.x).

Made-with: Cursor
2026-04-08 21:30:38 -07:00
Sameer Kankute
6a0e0ce061
fix(router): pass custom_llm_provider to get_llm_provider for unprefixed model names (#25334)
Fixes 'LLM Provider NOT provided' errors when models are configured with
custom_llm_provider but model names lack provider prefix (e.g., 'gpt-4.1-mini'
instead of 'azure/gpt-4.1-mini').

Changes:
- Router now passes deployment's custom_llm_provider to get_llm_provider()
- Fixes 6 code paths: file creation, file content, batch operations, vector store
- Adds regression tests for file creation and file content operations

Made-with: Cursor
2026-04-08 21:27:13 -07:00