Commit graph

290 commits

Author SHA1 Message Date
yuneng-jiang
72fba093c8 Merge remote-tracking branch 'origin/main' into litellm_dev_sameer_16_march_week 2026-03-21 15:11:29 -07:00
yuneng-jiang
10b0139bf8
Merge branch 'main' into litellm_oss_staging_03_05_2026 2026-03-21 14:58:11 -07:00
Sameer Kankute
4f1e484a9b Merge branch 'main' into litellm_dev_sameer_16_march_week
Resolve conflicts in common_request_processing.py (keep main streaming,
post_call_success_hook try/finally, deferred logging; retain skip_pre_call_logic)
and utils.py (defer + internal-call skip + sync success callbacks for all calls).

Tighten _has_post_call_guardrails for event_hook=None; align deferred
guardrail test. Sync model_prices_and_context_window_backup.json.

Pyright: narrow ignores for passthrough StreamingResponse and post_call hook.
Made-with: Cursor
2026-03-22 00:29:38 +05:30
Krish Dholakia
a5b7e49713
Merge branch 'main' into litellm_oss_staging_03_17_2026 2026-03-21 10:40:48 -07:00
Krish Dholakia
c350d08d66
Merge branch 'main' into litellm_oss_staging_03_05_2026 2026-03-21 10:31:50 -07:00
Cesar Garcia
a4f091c025
Merge pull request #24073 from Chesars/feat/gemini-context-circulation
feat(gemini): support context circulation for server-side tool combination
2026-03-20 23:29:30 -03:00
Sameer Kankute
ecfcf241c6
Merge pull request #24119 from BerriAI/main
merge main
2026-03-19 15:53:32 +05:30
Chesars
6f4b4d3c42 feat(gemini): support context circulation for server-side tool combination
Enables Gemini 3+ models to combine built-in tools (Google Search, etc.)
with custom functions via `include_server_side_tool_invocations=True`.
Server-side invocations are surfaced in provider_specific_fields and
automatically re-injected on subsequent turns for multi-turn coherence.

Closes #24047
2026-03-18 22:33:01 -03:00
Sameer Kankute
694cf22c9e
Update tests/test_litellm/llms/vertex_ai/test_vertex_ai_batch_transformation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-18 11:09:20 +05:30
Sameer Kankute
c4d27cb239 fix(vertex-ai): address greptile review – proxy retrieve URL, timeout forwarding, sync logging
- Fix retrieve_api_base derivation to handle custom proxies with
  path-based routing (not just :cancel suffix)
- Forward timeout to POST calls in cancel_batch (sync + async)
- Add try/except error logging to sync cancel path (parity with async)
- Add tests for timeout forwarding and custom proxy retrieve URL

Made-with: Cursor
2026-03-18 10:30:05 +05:30
Sameer Kankute
d0d593beb8
Update tests/test_litellm/llms/vertex_ai/test_vertex_ai_batch_transformation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-03-18 09:48:24 +05:30
Sameer Kankute
74ae17d153 greptile comments 2026-03-18 09:41:46 +05:30
Cesar Garcia
3c7e37799a
Merge pull request #23928 from Chesars/fix/gemini-context-caching-custom-api-base
fix(gemini): pass model to context caching URL builder for custom api_base
2026-03-18 00:43:16 -03:00
Sameer Kankute
018ccff23f fix(vertex-ai): address greptile review feedback on batch cancel
- Add try/except httpx.HTTPStatusError blocks in _async_cancel_batch for
  both POST cancel and GET retrieve calls, with verbose_logger error logging
- Fix endpoint extraction inconsistency: compute endpoint from URL without
  :cancel suffix so it matches behaviour of create_batch/retrieve_batch
- Add explicit validation that api_base ends with ':cancel' before
  stripping it, raising a descriptive error for unsupported custom proxy
  URL rewriting scenarios
- Use string-based patch() in test instead of patch.object() for robustness
  against import order changes

Made-with: Cursor
2026-03-18 09:11:20 +05:30
Cesar Garcia
4947074aac
Merge pull request #23907 from Chesars/fix/vertex-count-tokens-location-override
fix(vertex): respect vertex_count_tokens_location for Claude count_tokens
2026-03-18 00:30:55 -03:00
Chesars
8828f002be fix(gemini): pass model to context caching URL builder for custom api_base
_get_token_and_url_context_caching() was hardcoding model=None when
calling _check_custom_proxy(), which raises ValueError when api_base
is set because Gemini proxy URLs need the model name:
{api_base}/models/{model}:cachedContents

Fixes #23846
2026-03-17 23:21:24 -03:00
Chesars
8f015e2db2 fix(vertex): respect vertex_count_tokens_location for Claude count_tokens
The count_tokens handler unconditionally overrode vertex_location to
us-central1 for Claude models, ignoring the user-configured
vertex_count_tokens_location parameter. Also, us-central1 is no longer
a supported region — Google now supports us-east5, europe-west1, and
asia-southeast1.

Now vertex_count_tokens_location takes precedence, vertex_location is
used as fallback, and us-east5 is the default only when neither is set.

Fixes #23872
2026-03-17 19:14:08 -03:00
Chesars
0c28b47057 fix(vertex): streaming finish_reason="stop" instead of "tool_calls" for gemini-3.1-flash-lite-preview
Models like gemini-3.1-flash-lite-preview send the final streaming chunk
with empty content (text:"") alongside finishReason:"STOP", instead of
omitting content entirely. The existing fix (PR #21577) only handled
chunks without content, so this case was missed.

Now, after processing candidates, if tool_calls were seen in earlier
chunks and a choice has finish_reason="stop", it is overridden to
"tool_calls" to match the OpenAI spec.

Fixes #22900
2026-03-17 17:47:01 -03:00
Sameer Kankute
0bc609affd fix(vertex-ai): support batch cancel via Vertex API
Add Vertex batch cancellation support in LiteLLM batch APIs, route proxy cancel fallback using request provider headers, and return post-cancel batch state via retrieve to keep response shape compatible.

Made-with: Cursor
2026-03-17 11:23:47 +05:30
Sameer Kankute
22b333cae6 Fix downloading vertex ai files 2026-03-16 12:08:06 +05:30
Awais Qureshi
c7ba7948bc
PR #22867 added _remove_scope_from_cache_control for Bedrock and Azur… (#23183)
* PR #22867 added _remove_scope_from_cache_control for Bedrock and Azure AI but omitted Vertex AI. This applies the same pattern to VertexAIPartnerModelsAnthropicMessagesConfig."

* PR #22867 added _remove_scope_from_cache_control for Bedrock and Azure AI but omitted Vertex AI. This applies the same pattern to VertexAIPartnerModelsAnthropicMessagesConfig."

* PR #22867 added _remove_scope_from_cache_control to AzureAnthropicMessagesConfig
 but missed VertexAIPartnerModelsAnthropicMessagesConfi Rather than duplicating the method again, moved it up to the base AnthropicMessagesConfig so all providers
  inherit it, and removed the now-redundant copy from the Azure AI subclass.

* PR #22867 added _remove_scope_from_cache_control to AzureAnthropicMessagesConfig
 but missed VertexAIPartnerModelsAnthropicMessagesConfi Rather than duplicating the method again, moved it up to the base AnthropicMessagesConfig so all providers
  inherit it, and removed the now-redundant copy from the Azure AI subclass.

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2026-03-13 22:41:25 -07:00
Chesars
5c1e5c2510 Merge main into litellm_oss_staging_03_05_2026 2026-03-14 00:42:39 -03:00
Chesars
2d33d6496b Merge branch 'upstream/main' into HEAD
# Conflicts:
#	tests/test_litellm/llms/vertex_ai/gemini/test_vertex_and_google_ai_studio_gemini.py
2026-03-13 22:56:08 -03:00
yuneng-jiang
61cee53200 fix(tests): fix flaky qwen global endpoint test by mocking AsyncHTTPHandler at class level
The test was creating a real AsyncHTTPHandler instance and patching its
post method, but the internal code creates its own handler, bypassing
the mock. This caused real API calls to Vertex AI, resulting in 401
auth errors in CI. Switched to patching AsyncHTTPHandler at the class
level, matching the pattern used by the passing GPT-OSS test.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 09:07:48 -07:00
Cursor Agent
9a356644bf
fix(tests): stabilize 3 failing CI tests
1. Add missing __init__.py files in tests/test_litellm/llms/gemini/ and
   subdirectories (realtime/, image_edit/) to fix ModuleNotFoundError
   with pytest-xdist parallel workers.

2. Update test_transform_request_uses_dynamic_max_tokens to use
   claude-3-7-sonnet-20250219 (max_output_tokens=64000) since
   claude-3-5-sonnet-20241022 was removed from model_prices JSON
   during deprecated model cleanup. The test assertion was outdated.

3. Update context caching TTL tests to use gemini-2.5-pro instead of
   gemini-1.5-pro. The old model was removed from model_prices JSON,
   causing supports_system_messages to return False, which prevented
   system_instruction from appearing in the transformation output.

Co-authored-by: yuneng-jiang <yuneng-jiang@users.noreply.github.com>
2026-03-13 00:26:31 +00:00
Cursor Agent
e242356570
fix(ci): fix ruff lint errors and 9 failing unit tests on main
Lint fixes (check_code_and_doc_quality job):
- Remove unused variable reasoning_effort in gpt_5_transformation.py (F841)
- Remove unused timezone imports in mcp_server rest_endpoints.py and server.py (F401)
- Remove unused ProxyBaseLLMRequestProcessing import in realtime endpoints.py (F401)
- Add BaseRealtimeHTTPConfig to TYPE_CHECKING block in utils.py (F821)
- Add PLR0915 per-file-ignore for mcp_server/rest_endpoints.py in ruff.toml

Test fixes (litellm_mapped_tests_llms job):
- Gemini video cost tests: pass explicit model_info to video_generation_cost()
  instead of relying on gemini/veo-3.0-generate-preview being in model_prices JSON
- Anthropic max_tokens tests: mock get_max_tokens() to return expected values
  instead of depending on claude-3-5-sonnet-20241022 being in model_prices JSON
- Vertex AI pydantic obj test: update from removed gemini-1.5-pro to gemini-2.5-flash,
  update expected request body to use response_json_schema format
- Vertex AI/Bedrock file_content integration tests: update mocks to target
  base_llm_http_handler.retrieve_file_content (the new code path via
  ProviderConfigManager) instead of the old vertex_ai_files_instance/
  bedrock_files_instance paths

Co-authored-by: yuneng-jiang <yuneng-jiang@users.noreply.github.com>
2026-03-12 19:58:43 +00:00
Chesars
4e6e1d8de8 merge: resolve conflicts with upstream staging (bedrock + mcp tests)
Keep both sets of tests: upstream's OAuth2 token injection test and
our case-insensitive tool matching tests. Use upstream's version of
the bedrock output_config test (more comprehensive).
2026-03-12 13:40:16 -03:00
Chesars
feed274aa3 Reapply "feat: add model_cost aliases expansion support"
This reverts commit 3d2df7e8b5.
2026-03-12 13:36:57 -03:00
Cesar Garcia
6bd7cd7573
Merge branch 'main' into litellm_oss_staging_03_11_2026 2026-03-12 10:43:08 -03:00
Sameer Kankute
412a283569 Revert "fix(vertex): skip harmful schema transforms for Gemini 2.0+ tool parameters"
This reverts commit a9c3095cc5.
2026-03-12 18:26:11 +05:30
Chesars
1be6b31e2f merge: resolve conflicts between main and litellm_oss_staging_03_11_2026 2026-03-12 09:38:31 -03:00
Chesars
3948513a4c fix(vertex-ai): warn on region override and remove dead is_global_only_vertex_model
Add verbose_logger.warning when user-specified region is overridden by
supported_regions. Remove now-unused is_global_only_vertex_model function
and its tests since get_vertex_region handles all region logic directly.
2026-03-11 15:12:53 -03:00
Chesars
f3ceb69e9f fix(vertex-ai): override unsupported user region for models with supported_regions
- get_vertex_region now overrides user-specified region when it's not in
  the model's supported_regions list (prevents 404 for users with a
  global VERTEXAI_LOCATION default hitting global-only models)
- Add supported_regions: ["global"] to glm-5-maas in both JSON files
- Update tests to cover the override behavior
2026-03-11 14:55:51 -03:00
Chesars
689cbaa6c1 fix(vertex-ai): update tests to match new get_vertex_region model_cost lookup
- Remove redundant get_vertex_region() call in partner models main.py
  (already called inside get_complete_vertex_url)
- Rewrite test mocks to use patch.dict(litellm.model_cost) instead of
  patching the removed is_global_only_vertex_model symbol
- Align test assertions with new behavior: user-specified region is
  preserved (not overridden) for global-only models
2026-03-11 14:25:08 -03:00
Sameer Kankute
f243e5615f
Merge branch 'main' into litellm_oss_staging_03_10_2026 2026-03-11 18:50:03 +05:30
Cesar Garcia
3d2df7e8b5
Revert "feat: add model_cost aliases expansion support" 2026-03-10 22:39:19 -03:00
Cesar Garcia
5f5e47fc24
Merge pull request #22138 from Chesars/fix/unify-finish-reason-mapping
fix(completion): unify finish_reason mapping to OpenAI-compatible values
2026-03-10 19:29:04 -03:00
Cesar Garcia
6bca746d23
Merge pull request #21601 from Chesars/feat/model-cost-aliases
feat: add model_cost aliases expansion support
2026-03-10 18:07:23 -03:00
Chesars
2315d4b73c fix: resolve merge conflicts with staging branch
Keep unified _FINISH_REASON_MAP dict approach, discard upstream's
inconsistent _VALID_OPENAI_FINISH_REASONS frozenset that mapped to
values not in the OpenAIChatCompletionFinishReason Literal.
2026-03-10 17:27:57 -03:00
Chesars
a9c3095cc5 fix(vertex): skip harmful schema transforms for Gemini 2.0+ tool parameters
Gemini 2.0+ natively accepts JSON Schema in tool parameters, including
bare {} (TYPE_UNSPECIFIED), anyOf with null, and lowercase types. The
existing _build_vertex_schema pipeline was coercing {} to {"type": "object"},
breaking JsonValue/Any field semantics (issue #22391).

Add _build_vertex_schema_for_gemini_2() that only resolves $ref (which
Gemini doesn't support in tools) and filters unsupported fields. Use it
for Gemini 2.0+ models, keeping the full transform for Gemini 1.5.
2026-03-10 11:15:27 -03:00
Chesars
a6cb510703 merge: resolve conflicts between main and litellm_oss_staging_03_04_2026
Resolved 14 file conflicts:
- image_edits.md: combined OpenRouter + Black Forest Labs providers
- utils.py: kept staging's message-level cache_control check
- networking.tsx: kept export on 4 tool interfaces
- tool_management_endpoints.py: kept ToolOutputPolicy import
- Accepted main's version for: schema.prisma, a2a_protocol, mcp_server,
  _types.py, auth_checks.py, db_spend_update_writer, endpoints.py,
  spend_tracking_utils, a2a_endpoints, model_prices backup
2026-03-10 10:45:04 -03:00
Sameer Kankute
4dc277e427 fix(vertex_ai): strip LiteLLM-internal keys from extra_body before merging to Gemini request
PR #20950 added extra_body forwarding to Vertex AI Gemini. LiteLLM-internal
keys (cache, tags) were being merged into the request body, causing Vertex AI
to reject with 400: 'Unknown name "cache": Cannot find field.'

- Add _LITELLM_INTERNAL_EXTRA_BODY_KEYS frozenset (cache, tags)
- Skip these keys in _pop_and_merge_extra_body before merging
- Add regression tests for cache and tags stripping

Fixes regression from 1.79.3 → 1.81.12 when using proxy cache with
extra_body={"cache": {"use-cache": True, "ttl": 86400}}

Made-with: Cursor
2026-03-09 10:21:29 +05:30
Ishaan Jaff
fc81edc4c4
revert: undo PR #22589 and follow-up vertex anyOf fixes (#23083)
* Revert "fix(vertex): drop bare {} schemas from anyOf before adding nullable=True (#23060)"

This reverts commit 3ad9a536d3.

* Revert "Merge pull request #22589 from Chesars/fix/vertex-preserve-any-type-schema"

This reverts commit da941e4261, reversing
changes made to f77f28a5f8.
2026-03-07 17:49:49 -08:00
Ishaan Jaff
3ad9a536d3
fix(vertex): drop bare {} schemas from anyOf before adding nullable=True (#23060)
When anyOf contains a mix of concrete types, bare {} (any-type), and null,
convert_anyof_null_to_nullable was adding nullable=True to the {} entry,
producing {nullable: True} with no type field. Gemini rejects this as an
anyOf entry without a concrete type, breaking tool calls that use
Optional[List[...]] or similar union types (common in LangChain/Pydantic).

Fix: strip any-type schemas from anyOf before the nullable=True pass.
If only any-type schemas remain after null removal (anyOf: [{}, null]),
collapse the anyOf entirely and set nullable=True on the parent schema
instead — correctly representing 'any nullable value' for Gemini.

Regression introduced by da941e4261.
2026-03-07 15:59:10 -08:00
michelligabriele
5e34fdce77
feat(vertex_ai): support explicit AWS credentials for WIF auth (#21472)
* feat(vertex_ai): support explicit AWS credentials for WIF auth

The current Vertex AI AWS Workload Identity Federation implementation
exclusively uses google.auth.aws.Credentials.from_info(), which requires
EC2 instance metadata access to obtain AWS credentials. In environments
where the metadata service is blocked for security reasons, this makes
WIF unusable.

Add support for explicit AWS credentials by implementing a custom
AwsSecurityCredentialsSupplier (google-auth >= 2.29.0). When aws_* keys
(e.g. aws_role_name, aws_region_name) are present in the WIF credential
JSON, LiteLLM uses BaseAWSLLM.get_credentials() to obtain AWS creds via
STS AssumeRole (or any other supported AWS auth flow), wraps them in the
custom supplier, and passes them to aws.Credentials() — bypassing the
metadata service entirely.

When no aws_* keys are present, the existing from_info() flow is used
unchanged, preserving full backward compatibility.

* refactor(vertex_ai): extract AWS WIF auth to own class + add docs

Address PR review feedback:
- Move _AWS_CREDENTIAL_KEYS, _extract_aws_params(), and
  _credentials_from_aws_with_explicit_auth() from VertexBase into
  new VertexAIAwsWifAuth class in vertex_ai_aws_wif.py
- Add documentation for explicit AWS credentials WIF auth method
  in vertex.md (supported params, JSON example, SDK/Proxy tabs)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(vertex_ai): use lazy credentials provider to prevent stale STS tokens

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 09:27:20 -08:00
Sameer Kankute
a2c11d431a fix(vertex_ai): drop unsupported output_config parameter from all requests
Vertex AI does not support the output_config parameter in its API.
This parameter is being added by Anthropic/Gemini transformations but needs
to be removed before sending requests to Vertex AI endpoints.

This fix addresses the "Extra inputs are not permitted" error (issue #22312)
when using Claude models with structured outputs on Vertex AI.

Changes:
- Drop output_config in Gemini model transformation
- Drop output_config in Anthropic partner model transformation
- Drop output_config in Anthropic experimental pass-through transformation
- Add comprehensive tests to verify output_config is dropped

Fixes: #22312
Made-with: Cursor
2026-03-05 13:02:17 +05:30
Gustavo Martin Alvarez
b3a17596fe
fix(gemini): resolve image token undercounting in usage metadata (#22608)
* fix(gemini): ensure image token accumulation in usage metadata

Fixed an issue where image tokens were being overwritten instead of accumulated in Gemini responses. Added support for both camelCase and snake_case token count keys. Fixes #22082.

* test: add regression test for image token accumulation and cleanup files

* fix(gemini): ensure consistent accumulation for responseTokensDetails

* fix(gemini): harden token count parsing and add vertex accumulation test

Parse tokenCount/token_count as int-safe values to satisfy mypy and avoid None/object arithmetic. Add regression test for duplicate modality accumulation in Vertex _calculate_usage.
2026-03-04 18:52:08 -08:00
Cesar Garcia
d346f5cfab
Merge pull request #17550 from Chesars/fix/gemini-async-streaming-custom-client-17148
Fix: User specified async client ignored with Gemini streaming+async
2026-03-04 19:46:51 -03:00
Chesars
872554df42 Fix: User specified async client ignored with Gemini streaming+async
The user-specified async client was being overwritten by
`litellm.module_level_aclient` in `streaming_handler.py` when using
async+streaming with Gemini.

This fix adds a `gemini_client` parameter to `make_call()` (matching
the existing pattern in `make_sync_call()`) so the user's custom client
is preserved and not overwritten.

Fixes #17148
2026-03-04 19:38:08 -03:00
Chesars
c1a8bdd164 fix(gemini): support detail parameter for image resolution on Gemini 2.x models
Add global media_resolution support for Gemini 2.x models (2.0, 2.5) when
using OpenAI's detail parameter on images. Previously, the detail parameter
was only working for Gemini 3+ models (per-part) and was silently ignored
for older Gemini models.

- Add _get_highest_media_resolution() and _extract_max_media_resolution_from_messages()
  to extract highest detail from all images/files in a request
- Update _transform_request_body() to add mediaResolution to generationConfig
  for Gemini 2.x models only (not 1.x which doesn't support it, not 3+ which
  uses per-part)
- Add mediaResolution field to GenerationConfig TypedDict
- Support detail extraction from both image_url and file content types
- Add comprehensive unit tests and update documentation
2026-03-04 17:32:19 -03:00