Commit graph

28789 commits

Author SHA1 Message Date
fzowl
5bb7394c93
Merge branch 'BerriAI:main' into main 2025-12-14 19:21:00 +01:00
dongbin-lunark
0be6dd57c3
fix: pass credentials to PredictionServiceClient for Vertex AI custom endpoints (#17757)
When calling Vertex AI Model Garden custom endpoints with service account
credentials (instead of ADC), the credentials were not passed to the
PredictionServiceClient, resulting in "default credentials not found" error.

- Move credential loading before cache check to ensure availability
- Pass credentials to sync/async PredictionServiceClient
- Add vertex_credentials param to async_completion and async_streaming

Closes #8597
2025-12-14 08:41:43 +05:30
Cesar Garcia
c892c2c83d
fix(anthropic): use dynamic max_tokens based on model (#17900)
* fix(anthropic): use dynamic max_tokens based on model

When users don't specify max_tokens in requests to Anthropic models,
LiteLLM now uses the correct max_output_tokens value from the model
pricing JSON instead of a hardcoded 4096.

This fixes truncated responses for Claude 3.5+ models which support
higher output limits (8192 for Claude 3.5, 128k for Claude 3.7, etc.)

Fixes #8835

* fix(anthropic): restore env var support for backwards compatibility

Keep DEFAULT_ANTHROPIC_CHAT_MAX_TOKENS as fallback when model is not
found in JSON, allowing users to configure via environment variable.
2025-12-14 08:31:27 +05:30
Cesar Garcia
bd1a075a89
feat(stability): add Stability AI image generation support (#17894)
Add direct Stability AI REST API support for image generation endpoints.
This enables using Stability's SD3, SD3.5, and Stable Image models via
LiteLLM's OpenAI-compatible interface.

Changes:
- Add STABILITY provider to LlmProviders enum
- Create StabilityImageGenerationConfig with multipart/form-data support
- Add OpenAI size to Stability aspect_ratio mapping
- Register provider in ProviderConfigManager
- Add 9 Stability models to model_prices_and_context_window.json
- Add documentation at docs/providers/stability.md
- Add 25 unit tests

Supported models:
- stability/sd3, sd3-large, sd3-large-turbo, sd3-medium
- stability/sd3.5-large, sd3.5-large-turbo, sd3.5-medium
- stability/stable-image-ultra, stable-image-core
2025-12-14 08:29:45 +05:30
Cesar Garcia
5262896d62
fix(perplexity): use API-provided cost instead of manual calculation (#17887)
Fixes #15337

Perplexity API returns pre-calculated costs in `usage.cost.total_cost`
that include the `request_cost` (fixed per-request fee). LiteLLM was
ignoring this and calculating costs manually, resulting in ~27x
underreporting (e.g., $0.0002 vs actual $0.006).

Changes:
- Use `usage.cost.total_cost` from Perplexity response when available
- Fall back to manual calculation if cost object not present
- Add tests for both behaviors
2025-12-14 08:24:44 +05:30
Kerem Turgutlu
1da0bdd33d
fix gemini web search requests count (#17921)
* fix gemini web search requests count

* filter queries
2025-12-14 08:18:26 +05:30
Ishaan Jaffer
88d5efb2b7 docs v1.80.10.rc.1 2025-12-13 18:02:06 -08:00
Ishaan Jaffer
bd0ae49c74 ui new build 2025-12-13 17:28:04 -08:00
Ishaan Jaffer
efa9f69991 TestRunwaymlImageGeneration 2025-12-13 17:21:20 -08:00
Ishaan Jaff
2641f58be5
Litellm 1 80 10 (#17945)
* update providers

* v0

* docs fix

* docs fix

* docs fix
2025-12-13 17:19:58 -08:00
Ishaan Jaffer
accebc49a2 TestNvidiaNim 2025-12-13 16:38:11 -08:00
Ishaan Jaffer
f6c4ad92e4 async def test_update_team_guardrails_with_org_id(): 2025-12-13 16:24:58 -08:00
Ishaan Jaffer
95e818fdc4 bump litellm-proxy-extras 2025-12-13 16:22:26 -08:00
Ishaan Jaffer
d9794f0811 fix bedrock embeddings - validate env 2025-12-13 16:16:30 -08:00
Ishaan Jaffer
92b72fc759 test_runwayml_tts_async 2025-12-13 16:10:48 -08:00
Ishaan Jaffer
2fd8621b38 test recraft 2025-12-13 16:10:34 -08:00
Ishaan Jaffer
050264f7d7 test_recraft_image_edit_api 2025-12-13 16:09:52 -08:00
Ishaan Jaff
14eed8aff7
[Fixes] A2a Gateway - ensure azure foundry agents work (#17943)
* add agents  v2 fixes azure

* fix auth

* get_azure_ad_token fix

* docs foundry
2025-12-13 16:08:03 -08:00
yuneng-jiang
232bb33fec
Merge pull request #17942 from BerriAI/litellm_ui_notification
[Feature] Show progress and pause on hover for Notifications
2025-12-13 15:37:55 -08:00
yuneng-jiang
6567e43560
Merge pull request #17940 from BerriAI/litellm_ui_mcp_headers
[Fix] Add extra_headers and allowed_tools to UpdateMCPServerRequest
2025-12-13 15:20:55 -08:00
yuneng-jiang
db7454c6ee Show progress and pause on hover for notifications 2025-12-13 15:16:13 -08:00
Alexsander Hamir
6635325629
fix: filter internal params in fallback code and fix test issues (#17941)
- Filter skip_mcp_handler and other internal params in fallback_utils.py before calling acompletion
  Fixes issue where internal parameters were being passed to provider APIs causing errors
- Remove deployment field from GCS bucket logger test metadata
  Fixes model name mismatch where deployment field was overriding the model in logging
- Update Bedrock Titan test to use non-deprecated model (titan-text-express-v1)
  Fixes test failure due to deprecated amazon.titan-text-lite-v1 model
2025-12-13 15:05:26 -08:00
Ishaan Jaff
ed356fdfc0
[Docs] Cursor Integration (#17939)
* docs cursor

* remove bloat

* stash changes

* docs fix

* simpler docs

* docs

* docs cursor

* add cursor/chat/completions
2025-12-13 14:44:40 -08:00
yuneng-jiang
61767779f8 Adding tests 2025-12-13 14:44:26 -08:00
yuneng-jiang
5ee2338f9e Adding extra_headers and allowed_tools in UpdateMCPServerRequest 2025-12-13 14:36:37 -08:00
Alexsander Hamir
fab1b81b7f
fix: add agent_id field to GCS PubSub spend_logs_payload.json test expectation (#17938)
- Add agent_id: null to expected JSON to match actual payload structure
- Fixes test_async_gcs_pub_sub_v1 test failure
- agent_id is an optional field in SpendLogsPayload that is always included (as null when not provided)
2025-12-13 13:35:20 -08:00
yuneng-jiang
b681729e0c
Merge pull request #17937 from BerriAI/litellm_a2a_doc_fix
[Docs] Add Import Image to A2A Docs
2025-12-13 13:30:06 -08:00
yuneng-jiang
8b0dd58a64 Importing Image in A2A Doc 2025-12-13 13:29:01 -08:00
Alexsander Hamir
425b8400b0
fix: add storage_backend and storage_url columns to schema.prisma files (#17936)
- Updated litellm/proxy/schema.prisma to include storage_backend and storage_url columns
- Updated litellm-proxy-extras/litellm_proxy_extras/schema.prisma to include storage_backend and storage_url columns
- Fixes database schema mismatch causing 'storage_backend column does not exist' errors
- Keeps all schema files in sync with root schema.prisma
2025-12-13 13:28:34 -08:00
yuneng-jiang
1b28ea7728
Merge pull request #17935 from BerriAI/litellm_fix_agent_docs_2
[Docs] Fixing links
2025-12-13 13:10:25 -08:00
yuneng-jiang
d18ed525b8 Fixing links 2025-12-13 13:09:39 -08:00
yuneng-jiang
af7f50d42b
Merge pull request #17934 from BerriAI/litellm_merge_agent_docs
[Docs] Merge Agent Usage with A2A Cost Tracking
2025-12-13 13:05:23 -08:00
yuneng-jiang
773c4d08b4 Merge Agent Usage with A2A Cost Tracking 2025-12-13 13:04:28 -08:00
yuneng-jiang
6745c81800
Merge pull request #17932 from BerriAI/litellm_doc_fix
[Docs] Agent Usage doc fix
2025-12-13 12:54:50 -08:00
yuneng-jiang
b692e87836 Agent doc fix 2025-12-13 12:54:00 -08:00
Ishaan Jaff
24d6ec67c7
[QA] Cursor Integration x LiteLLM (#17855)
* fix utils.py

* ValidUserMessageContentTypesLiteral

* add _transform_tool_choice

* _transform_responses_api_content_to_chat_completion_content

* TestContentTypeTransformation

* test_map_tool_choice_string_auto

* fix validate_chat_completion_user_messages

* fix _is_input_item_tool_call_output

* fix LiteLLMCompletionResponsesConfig
2025-12-13 12:49:45 -08:00
yuneng-jiang
2d75875ea6
Merge pull request #17931 from BerriAI/litellm_agent_usage_md
[Docs] Agent Usage Doc
2025-12-13 12:49:06 -08:00
yuneng-jiang
39b8acae55 Agent Usage Doc 2025-12-13 12:46:56 -08:00
Alexsander Hamir
892d7e8d70
[Fix] CI/CD - Fix Bedrock tool calling test failures with non-serializable objects and internal parameters (#17930)
* fix(bedrock): filter non-serializable objects from request params

- Enhanced filter_exceptions_from_params() to filter callable objects (functions) and Logging objects
- Applied filtering in Bedrock's _prepare_request_params() before deepcopy
- Applied filtering to additional_request_params before JSON serialization
- Prevents TypeError during deepcopy (APIConnectionError objects) and JSON serialization (functions, Logging objects)
- Fixes test_bedrock_tool_calling test failures

Root cause: MCP-related functions (handle_chat_completion_with_mcp, completion_callable) and litellm_logging_obj were incorrectly added to optional_params via add_provider_specific_params_to_optional_params(), which then ended up in additional_request_params. These objects should be in litellm_params, not optional_params.

* fix(bedrock): filter internal MCP parameters from API requests

Filter out LiteLLM internal/MCP-related parameters (skip_mcp_handler,
mcp_handler_context, _skip_mcp_handler) from additional_request_params
before sending to Bedrock API to prevent 'extraneous key' errors.

- Added filter_internal_params() helper function in core_helpers.py
- Applied filtering in Bedrock's _prepare_request_params() method
- Fixes test_bedrock_completion.py::test_bedrock_tool_calling

* fix: mypy type error

* fix: add filter_exceptions_from_params to recursive function ignore list

- Add filter_exceptions_from_params to IGNORE_FUNCTIONS in recursive_detector.py
- Function is safe: has max_depth parameter (default 20) to prevent infinite recursion
2025-12-13 12:38:07 -08:00
yuneng-jiang
5a3cf8b171
Merge pull request #17929 from BerriAI/litellm_v18010_doc
[Docs] v1.80.10 draft
2025-12-13 12:05:47 -08:00
yuneng-jiang
e8bbb561b9 Adding image 2025-12-13 12:03:24 -08:00
yuneng-jiang
261437623a Agent usage docs WIP 2025-12-13 11:35:04 -08:00
yoshi-p27
a154671320
Regex guardrails update (#17915)
* Update patterns.json

* regex filtering update
2025-12-13 11:12:24 -08:00
yuneng-jiang
952e555ed3
Merge pull request #17928 from BerriAI/litellm_ui_logs_fix
[Fix] UI - Request and Response in Logs
2025-12-13 10:52:02 -08:00
yuneng-jiang
67a9d61e74 Fixing regression on logs 2025-12-13 10:43:04 -08:00
Alexsander Hamir
32fdb9e60e
fix: Add headers to Request scope in JWT tests to fix KeyError (#17927)
- Add 'headers': [] to all Request(scope={'type': 'http'}) instances in test_jwt.py
- Fixes KeyError: 'headers' when accessing request.headers in user_api_key_auth
- All 7 previously failing tests now pass:
  - test_allow_access_by_email (2 variants)
  - test_allowed_routes_admin (4 variants)
  - test_team_token_output (2 variants)

The Starlette Request object requires 'headers' key in scope dictionary
when accessing request.headers property.
2025-12-13 10:36:13 -08:00
Alexsander Hamir
1393c76578
[Fix] CI/CD - Fix failing proxy and core integration tests (#17926)
* Fix: Add prisma generate to proxy tests CI job

- Add prisma generate command before pytest in litellm_mapped_tests_proxy job
- Fixes 3 test failures: test_health_liveliness_endpoint, test_health_liveness_endpoint, test_health_readiness
- Matches pattern used in litellm_mapped_enterprise_tests job

* Fix: Add missing prompt_spec parameter to TestCustomPromptManagement

- Add prompt_spec parameter to get_chat_completion_prompt() method signature
- Fixes 2 test failures: test_custom_prompt_management_with_prompt_id and test_custom_prompt_management_with_prompt_id_and_prompt_variables
- Aligns test mock with base class method signature from CustomPromptManagement

* Fix: Handle string datetime values in OpenTelemetry timestamp conversion

- Add _to_timestamp helper to handle datetime, float, and string inputs
- Fixes test_handle_success_spans_and_metrics failure
- Handles string datetime format from JSON deserialization (e.g. '2025-06-22 10:59:08.399523')
- Applied to all timestamp conversions in OpenTelemetry metrics methods

* Fix OpenTelemetry timestamp parsing to handle datetime strings with/without microseconds
2025-12-13 10:11:01 -08:00
Alexsander Hamir
5b6b613561
[Fix] CI/CD - Fix failing proxy unit test and langfuse trace_id test (#17924)
* fix: correct Request headers format in JWT auth test

Fix test_jwt_non_admin_team_route_access by converting headers to bytes
format as required by Starlette's ASGI specification. Headers must be
bytes tuples with lowercase header names.

This allows dict(request.headers) to work correctly and enables the
authorization check to run, producing the expected error message.

* fix: ignore UUID trace_id from standard_logging_object, use litellm_call_id

The issue was that standard_logging_object.trace_id contains a UUID
(from litellm_trace_id default), which was being used instead of
falling back to litellm_call_id. This caused the test to fail because
it expected 'my-unique-call-id' but got a UUID.

Now we properly detect UUIDs (36 chars with 4 hyphens in specific positions)
and ignore them, allowing the fallback to litellm_call_id to work correctly.
This ensures we use litellm_call_id when no explicit trace_id is provided,
which gets stored in the cache and returned by _get_trace_id().

* fix: use existing_trace_id when provided instead of litellm_call_id

When existing_trace_id is provided in metadata, it should be used as the
trace_id to return (and store in cache), not litellm_call_id. This fixes
the test case where existing_trace_id is set and should be returned by
_get_trace_id().
2025-12-13 09:32:43 -08:00
yuneng-jiang
7633d84d51
Merge pull request #17922 from BerriAI/litellm_ui_build_3a
[Infra] UI build
2025-12-13 09:08:46 -08:00
yuneng-jiang
26b8e7f889 UI build 2025-12-13 09:07:43 -08:00