Commit graph

41449 commits

Author SHA1 Message Date
Sameer Kankute
728e5b13f7
Merge pull request #22922 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-06 00:43:28 +05:30
Sameer Kankute
f06e9e6368 Fix doc 2026-03-06 00:42:45 +05:30
Varad Khonde
6d4a281ba0 fix(gemini): handle 'minimal' reasoning_effort param for gemini-3.1-flash-lite-preview 2026-03-06 00:26:45 +05:30
Sameer Kankute
7aff1dc0d3
Merge pull request #22919 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-06 00:26:34 +05:30
Sameer Kankute
04f38332de Fix doc 2026-03-06 00:25:31 +05:30
Spencer Burridge
c919031ff0
feat(proxy): include user_email in jwt upsert user creation (#22915)
* Include user_email in new user creation within get_user_object

Enhance the get_user_object function to include user_email in the parameters when creating a new user. This change is accompanied by a new test to verify that user_email is correctly included during the upsert process.

* Improve error handling in test_get_user_object by logging exceptions

Updated the test_get_user_object_upsert_includes_user_email function to log exceptions when they occur, enhancing the visibility of potential issues during testing. This change helps in diagnosing failures related to the mock LiteLLM_UserTable.
2026-03-05 10:55:11 -08:00
Sameer Kankute
46fa9a33da
Merge pull request #22918 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-05 23:53:34 +05:30
Sameer Kankute
cae1f5fbae Fix doc 2026-03-05 23:52:56 +05:30
Sameer Kankute
9df686e044
Merge pull request #22917 from BerriAI/litellm_gpt-4.5_fix
Fix doc
2026-03-05 23:51:00 +05:30
Sameer Kankute
cf376d2c0e Fix doc 2026-03-05 23:50:20 +05:30
Ishaan Jaff
a42132f329
fix(passthrough): propagate Azure 429/5xx errors in async streaming instead of silent HTTP 200 (#22913)
* fix(passthrough): raise_for_status in _async_streaming to propagate Azure 429s

* address greptile review feedback (greploop iteration 1)

Guard data/json args when content is provided to avoid httpx ValueError

* address greptile review feedback (greploop iteration 2)

Use bare raise to preserve original traceback in _async_streaming exception handler

* address greptile review feedback (greploop iteration 3)

Close httpx streaming response on error to prevent connection pool exhaustion

* address greptile review feedback (greploop iteration 4)

Guard aclose() call to prevent masking original exception; add explicit test for content param forwarding

* address greptile review feedback (greploop iteration 5)

Pass content to sign_request so AWS body-hash signing is correct when content is the sole body source

* revert sign_request content change - request_data expects dict, not bytes

Bedrock's sign_request calls json.dumps(request_data) — passing content bytes
would TypeError. sign_request should only receive data/json (dict), not raw bytes.
2026-03-05 10:12:43 -08:00
Sameer Kankute
8dca085640
Merge pull request #22916 from BerriAI/litellm_gpt-5.4_day_0
Add day 0 support for gpt-5.4
2026-03-05 23:41:35 +05:30
Sameer Kankute
3b457b5d8e Add day 0 support for gpt-5.4 2026-03-05 23:40:24 +05:30
giulio-leone
d6310ff36e fix: downgrade WebSearch logs from info to debug to reduce production noise
All operational/diagnostic messages in WebSearchInterceptionLogger are now
debug-level to avoid flooding production logs while still remaining available
when verbose logging is enabled.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 18:58:40 +01:00
Sameer Kankute
b9a8d42882 Add day 0 support for gpt-5.4 2026-03-05 23:26:24 +05:30
giulio-leone
7b0ed0ff91 fix: replace sk-fake with safe test key to avoid secret scanner
Replace 'sk-fake' with 'fake-key-for-testing' in websearch interception
tests to prevent false-positive secret scanner triggers.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 18:29:28 +01:00
Julio Quinteros Pro
6db3f2f668
Merge pull request #22892 from BerriAI/fix/greptile-type-safety-improvements
fix(types): address type-safety issues from mypy PR review
2026-03-05 14:06:17 -03:00
Julio Quinteros
023794ba62 fix(merge): resolve conflict with main in cost_tracking_settings
PR #22890 used cast(str, ...) / cast(Optional[str], ...) for the return
statements; this PR's approach uses str() for explicit runtime coercion
(addressing Greptile's concern). Keep the str() version and drop the
now-unused cast import.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 14:01:07 -03:00
Giulio Leone
6b7d767637
feat(anthropic): support top-level cache_control for automatic prompt caching (#22442) 2026-03-05 08:34:56 -08:00
giulio-leone
660de94493 fix: change all verbose_logger.warning → info in websearch handler
Per Sameerlite's review: warning-level logs trigger Slack alerts.
All 6 remaining .warning() calls were operational/fallback messages,
not actual errors. Changed to .info() to match the first fix at L510.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 17:14:30 +01:00
Cesar Garcia
5c418d8de6
Merge pull request #22909 from Chesars/fix/greptile-reasoning-summary-followup 2026-03-05 12:41:39 -03:00
Chesars
5b904f6054 fix(anthropic): align translate_thinking_for_model with default summary injection + add docs
- Update translate_thinking_for_model (3rd code path) to inject
  summary="detailed" by default, consistent with the other two paths
- Add disable_default_reasoning_summary flag check via shared helper
- Add tests for flag enabled/disabled and user-provided summary
- Document disable_default_reasoning_summary in reasoning_content.md
2026-03-05 12:22:18 -03:00
Cesar Garcia
c9e60d9909
Merge pull request #22904 from Chesars/claude/add-claude-param-default-M4Yic
feat(anthropic): add opt-out flag for default reasoning summary
2026-03-05 11:59:28 -03:00
Chesars
3e9ea6f49b refactor: extract summary_disabled logic into shared helper and add missing env var test
- Extract duplicated summary_disabled evaluation from handler.py and
  transformation.py into a shared is_default_reasoning_summary_disabled()
  helper in utils.py to prevent future divergence.
- Add test_summary_excluded_when_env_var_set to handler test class to
  close env-var test coverage gap flagged by Greptile.
2026-03-05 11:58:25 -03:00
Chesars
607a9683a4 feat(anthropic): add opt-out flag for default reasoning summary
Add `litellm.disable_default_reasoning_summary` flag (default False) and
env var `LITELLM_DISABLE_DEFAULT_REASONING_SUMMARY` to allow users to
opt out of the automatic `summary="detailed"` injection when routing
Anthropic thinking requests to OpenAI's Responses API.

Default behavior is preserved (summary="detailed" is always added),
but users who don't want to pay for summary tokens can now disable it.

https://claude.ai/code/session_01VJU9EwVvgvmeCe3Yu1aULa
2026-03-05 11:58:25 -03:00
giulio-leone
b04ba60e6e fix(websearch): downgrade max_tokens adjustment log level
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-03-05 14:19:15 +01:00
Sameer Kankute
5183a6e850
Merge pull request #22866 from mubashir1osmani/feat/bedrock-mantle-provider-clean
feat: bedrock mantle provider
2026-03-05 18:24:00 +05:30
Sameer Kankute
a282bf9726
Merge pull request #22893 from BerriAI/litellm_messages-to-responses-mapping-docs
docs(anthropic): add v1/messages → /responses parameter mapping reference
2026-03-05 18:21:37 +05:30
Sameer Kankute
0620f99fa4
Merge pull request #22867 from BerriAI/litellm_bedrock-azure-cache-control-scope
fix(bedrock,azure_ai): strip scope from cache_control for Anthropic messages
2026-03-05 18:20:59 +05:30
Sameer Kankute
c04c120df2
Merge pull request #22884 from BerriAI/litellm_vertex-output-config-drop
fix(vertex_ai): drop unsupported output_config parameter from all requests
2026-03-05 18:20:47 +05:30
Sameer Kankute
501671aa43 fix(agents): PUT update_agent_in_db clears static_headers and extra_headers when omitted
For full-replace PUT semantics, always include static_headers and extra_headers
in update_data, defaulting to {} and [] when not supplied. Previously,
omitting these fields left stale DB values intact (e.g. auth headers).

Made-with: Cursor
2026-03-05 16:16:21 +05:30
Sameer Kankute
4fda3e8351
Merge pull request #22896 from BerriAI/litellm_mistral-document-ai-2512-cost-map
feat(cost): add azure_ai/mistral-document-ai-2512 to model cost map
2026-03-05 16:14:27 +05:30
Sameer Kankute
bb1297fe1b feat(cost): add azure_ai/mistral-document-ai-2512 to model cost map
Made-with: Cursor
2026-03-05 16:07:46 +05:30
Sameer Kankute
dd0ccdd8b2 fix(o-series): generalize is_model_o_series_model to match any o+digit prefix
Replace the hardcoded startswith(("o1", "o3", "o4")) check with a broader
o+digit pattern, ensuring future o-series models (e.g. o2, o5) are
automatically recognized once added to open_ai_chat_completion_models.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 15:54:53 +05:30
Sameer Kankute
6a8adf8bdf docs(anthropic): add v1/messages → /responses parameter mapping reference
Documents exactly how every request and response field gets translated
when LiteLLM routes an Anthropic /v1/messages call through the OpenAI
Responses API path (for OpenAI/Azure targets). Covers messages content
block mapping, tools, tool_choice, thinking→reasoning, context_management,
and the reverse response translation. Wired into the /v1/messages sidebar.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 15:51:19 +05:30
Julio Quinteros Pro
b52bd43740
Merge pull request #22822 from BerriAI/fix/plr0915-too-many-statements
fix(lint): resolve PLR0915 too-many-statements in 4 files
2026-03-05 07:13:10 -03:00
Julio Quinteros
c1076de5bd fix(types): address type-safety issues from mypy PR review
- CreateBatchRequest.output_expires_after: drop Optional since total=False
  already makes the key absent-or-present; Optional[T] incorrectly allowed
  the key to exist with value None, which is incompatible with the OpenAI
  SDK's OutputExpiresAfter | NotGiven expectation on batches.create()
- cost_tracking_settings._resolve_model_for_cost_lookup: replace implicit
  object-to-str returns with explicit str() calls so the function is safe
  even if the surrounding truthiness guards are later weakened

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 07:11:31 -03:00
Julio Quinteros Pro
ad2969badb
Merge pull request #22890 from BerriAI/fix/mypy-type-errors
fix(mypy): resolve type errors across 9 files
2026-03-05 07:05:23 -03:00
Julio Quinteros Pro
44498da62a
Merge pull request #22887 from BerriAI/fix/schema-add-realtime-mode
fix(test): add 'realtime' to model mode enum in schema validation
2026-03-05 07:03:27 -03:00
Julio Quinteros Pro
de18b47f83
Merge pull request #22891 from BerriAI/fix/prisma-schema-duplicate-spec-path
fix(schema): remove duplicate spec_path field in LiteLLM_MCPServerTable
2026-03-05 07:02:33 -03:00
Julio Quinteros
16f415ad74 fix(schema): remove duplicate spec_path field in LiteLLM_MCPServerTable
PR #22850 (BYOK MCP servers) accidentally re-declared spec_path which was
already added by PR #22820, causing Prisma schema validation to fail with
error P1012 "Field is already defined".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 07:00:11 -03:00
Julio Quinteros
f45a9df52d fix(mypy): resolve type errors across 9 files
- batches/main.py: import FileExpiresAfter, cast output_expires_after on assignment
- openai/openai.py, azure/batches/handler.py: add # type: ignore[arg-type] on
  batches.create / batches.retrieve TypedDict unpacking calls
- searchapi/transformation.py: cast optional_params["country"] to str before .lower()
- openrouter/image_edit/transformation.py: cast iterated value to str for size/quality params
- spend_log_cleanup.py: narrow bool | None to bool with `or False`
- cost_tracking_settings.py: cast base_model/resolved_model to str and
  custom_llm_provider to Optional[str] in return statements
- text_moderation.py: suppress misc TypedDict ** expansion error; use cast for response
- prompt_shield.py: use cast instead of TypedDict(**response_json) construction

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:58:20 -03:00
Sameer Kankute
04c904f4d6 fix(agents): fix header isolation, streaming trace headers, and patch clearing
- create_a2a_client already uses httpx.AsyncClient directly (no shared cache);
  test_create_a2a_client_uses_fresh_httpx_client will now pass in CI
- asend_message_streaming now injects X-LiteLLM-Trace-Id and
  X-LiteLLM-Agent-Id headers, matching the non-streaming path
- patch_agent_in_db uses key-presence check ("in agent") instead of
  is not None so PATCH with static_headers=None/[] correctly clears the field

Made-with: Cursor
2026-03-05 15:26:22 +05:30
Julio Quinteros Pro
1c7f93f90c
Update litellm/fine_tuning/main.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-05 06:45:31 -03:00
Sameer Kankute
594499e806 Add tests 2026-03-05 15:14:45 +05:30
Julio Quinteros
db8e909ef2 fix(test): add 'realtime' to model mode enum in schema validation
gemini/gemini-live-2.5-flash-preview-native-audio-09-2025 uses mode='realtime'
but the schema in test_aaamodel_prices_and_context_window_json_is_valid did
not include 'realtime' as a valid enum value, causing a ValidationError.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:41:51 -03:00
Sameer Kankute
175d0905b7 docs(agents): add A2A agent authentication headers guide
New page: docs/a2a_agent_headers.md

Covers all three header forwarding methods:
- Static headers (admin-configured, always sent)
- Forward client headers (extra_headers — admin lists names, client supplies values)
- Convention-based x-a2a-{agent_name/id}-{header} (no admin config needed)

Documents merge precedence (static wins), header isolation guarantee,
combining all three methods, and API reference for static_headers / extra_headers fields.

Registered in sidebars.js under /a2a - A2A Agent Gateway.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 15:11:18 +05:30
Julio Quinteros
0133d11eb2 fix: replace assert with RuntimeError and fix return type annotation
- a2a_protocol/main.py: replace bare assert with descriptive RuntimeError
  in _execute_a2a_send_with_retry so retry exhaustion gives a clear message
- fine_tuning/main.py: fix _resolve_fine_tuning_timeout return type from
  float to Union[float, httpx.Timeout] to accurately reflect the passthrough path

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:38:38 -03:00
Sameer Kankute
9a13c76e2f
Merge pull request #22553 from dsteeley/fix/streaming-multi-tool-call-premature-finish
fix(streaming): output_item.done for function_call must not emit finish_reason
2026-03-05 15:05:43 +05:30
Julio Quinteros
b3bbcd3955 fix(lint): remove unreachable None check in _resolve_fine_tuning_timeout
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-05 06:29:57 -03:00