mateo-berri
43c838f4b9
Merge branch 'litellm_internal_staging' into litellm_decrease_anys_opus5
2026-08-29 21:41:04 -07:00
mateo-berri
b418ccd738
fix(azure): flatten top-level tool schema combinators on Azure chat completions
...
Azure's chat completions validator rejects tool parameters carrying a
top-level anyOf/oneOf/allOf for every model family. AzureOpenAIConfig and
the o-series config now flatten them via the shared helper moved to
prompt_templates common_utils. Requests bridged to the Responses API for
gpt-5.4+ with reasoning active keep the union, which that surface accepts
2026-08-29 21:27:57 -07:00
mateo-berri
0e27e09fae
Merge branch 'litellm_internal_staging' into litellm_fix_chat_anyof_tool_schema
2026-08-29 20:57:35 -07:00
yucheng-berri
d44d281d1d
fix(proxy): emit timing headers and overhead for /v1/messages and /v1/responses ( #38840 )
2026-08-29 18:11:58 -07:00
Mateo Wang
2a79a81b46
Merge pull request #38837 from BerriAI/litellm_fix_azure_responses_anyof_tool_schema
...
fix(azure): flatten top-level tool schema combinators for Azure Responses GPT-4-family deployments
2026-08-29 16:46:22 -07:00
Mateo Wang
1f5e76155b
Merge pull request #38836 from BerriAI/litellm_fix_messages_effort_budget_cap
...
fix(anthropic): cap reasoning_effort thinking budget below max_tokens on /v1/messages
2026-08-29 16:45:18 -07:00
Mateo Wang
ecd42ea77a
Merge pull request #38792 from BerriAI/litellm_fix_responses_anyof_tool_schema
...
fix(openai): flatten top-level anyOf/oneOf/allOf in Responses API tool schemas
2026-08-29 16:45:11 -07:00
Mateo Wang
a979c89b88
Merge pull request #38804 from BerriAI/litellm_registry_audit_rolling_38693
...
fix(models): registry audit: new Together/Fireworks/Gemini/Mistral/xAI models, xai retirement repricing, bedrock grok-4.6 caching, deprecation dates
2026-08-29 16:44:45 -07:00
mateo-berri
1c4674441c
docs(openai): trim tool-flattening docstrings to upstream facts
2026-08-29 16:43:05 -07:00
mateo-berri
a85e16f731
Merge remote-tracking branch 'origin/litellm_fix_audio_speech_content_type' into litellm_fix_gemini_tts_container
2026-08-29 16:27:27 -07:00
mateo-berri
d804b9d4fe
fix(vertex_ai): skip non-dict property values in set_schema_property_ordering
...
The typed rewrite made the properties recursion call .get on every child,
so a malformed schema with a string or list property value raised
AttributeError where it previously passed through untouched.
2026-08-29 16:26:47 -07:00
mateo-berri
855f56fa94
fix(openai): flatten top-level tool schema combinators on chat completions
2026-08-29 16:25:28 -07:00
mateo-berri
af186eaaf3
fix(azure): flatten top-level tool schema combinators for Azure Responses GPT-4-family deployments
2026-08-29 16:23:01 -07:00
Mateo Wang
6bc8dafa99
Merge pull request #38740 from BerriAI/litellm_vertex_gemini_35_transcribe
...
feat(vertex_ai): support gemini-3.5-transcribe on /v1/audio/transcriptions
2026-08-29 16:19:56 -07:00
mateo-berri
71a951691a
fix(anthropic): cap reasoning_effort thinking budget below max_tokens on /v1/messages
...
A deployment carrying reasoning_effort in its litellm_params on the
/v1/messages passthrough mapped the effort to a legacy thinking block
whose budget_tokens was forwarded as is, so any request whose max_tokens
sat at or below that budget was rejected upstream with a 400. The mapped
budget now runs through the same cap the adaptive-to-legacy branch and
the chat path already use: it is clamped to max_tokens - 1, and dropped
with a warning when even the minimum budget cannot fit.
The cap helper becomes public since three call sites outside
AnthropicConfig use it.
2026-08-29 15:24:05 -07:00
tin-berri
36ea28b092
fix(anthropic): emit signature-only thinking blocks on the /v1/messages bridge ( #38809 )
2026-08-29 15:04:22 -07:00
mateo-berri
9448293903
fix(openai): flatten tool schema unions only for models whose validator rejects them
...
GPT-5 and later accept a top-level anyOf natively and call tools better with it intact, so the flattening now runs only for the gpt-4, gpt-3.5, chatgpt-4o, o1, o3, and o4 families. Non-dict tool entries pass through untouched, a typeless root that carries properties counts as an object, and the bounded $ref walker is listed in the recursion detector allowlist.
2026-08-29 14:38:08 -07:00
mateo-berri
c251d6d609
fix(vertex_ai): label TTS audio bytes with their real content-type
2026-08-29 14:11:51 -07:00
Mateo Wang
8dd9c4acb1
Merge pull request #30782 from emerzon/litellm_veo_31_lite
...
feat(vertex-ai): add veo 3.1 lite model metadata
2026-08-29 13:36:02 -07:00
Mateo Wang
306daf13b5
Merge pull request #38752 from BerriAI/litellm_deflake_20260829
...
fix: bound Hugging Face config fetch and keep embedding tests off the network
2026-08-29 13:33:06 -07:00
mateo-berri
2bd7b58640
fix(registry): correct xai retired slug pricing, bedrock grok caching, and unsourced entries
...
Reprice ten more retired xAI slugs (grok-3 and grok-3-mini families,
grok-4-1-fast) to the grok-4.3 rates they now bill at, with family-correct
deprecation dates. Restore cache_read_input_token_cost on the Bedrock Grok 4.6
entries so implicit cache hits bill at the cache-read rate while explicit
cachePoint stays unsupported. Drop the unsourced 1080p video rate and the
gemini/ live native-audio entry the Gemini API 404s on. Add Groq qwen3.8-27b
tool-use flags per Groq docs. Extend the xai and gemini tests to lock all of
this in
2026-08-29 13:24:09 -07:00
mateo-berri
9b8ad46f37
fix(openai): flatten top-level anyOf/oneOf/allOf in Responses API tool schemas
...
OpenAI's function-calling validator rejects tool parameters carrying
oneOf/anyOf/allOf/enum/const/not at the top level, while the ChatGPT
backend Codex talks to natively accepts them, so an MCP tool declaring a
top-level union 400s through the proxy. Merge the branches into the
object schema for OpenAI itself only, walking the namespace-nested tools
current Codex builds send, on both /v1/responses and /v1/responses/compact
2026-08-29 13:11:04 -07:00
Mateo Wang
9ed7de6c02
Merge pull request #38670 from BerriAI/devin_ai_38659_cohere_embed_dispatch
...
fix(bedrock): route all cohere.embed models to the cohere embedding config
2026-08-29 12:55:32 -07:00
Mateo Wang
817bbe1dc6
Merge pull request #34440 from dan2k3k4/litellm_soniox_srt_cue_grouping
...
fix(soniox): align synthesized SRT/VTT cues to real speech timing
2026-08-29 12:49:57 -07:00
Mateo Wang
c453920f7a
Merge pull request #38285 from BerriAI/litellm_azure_v1_image_routes
...
fix(azure): use /openai/v1 image routes for v1, preview and latest api versions
2026-08-29 12:45:37 -07:00
mateo-berri
886d39c3a2
test(bedrock): expect cohere embed base64 encoding_format to normalize to float
2026-08-29 12:44:01 -07:00
mateo-berri
a007fa49e5
Merge branch 'litellm_internal_staging' into litellm_veo_31_lite
2026-08-29 12:04:42 -07:00
mateo-berri
4c42c01cb2
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_soniox_srt_cue_grouping
...
# Conflicts:
# litellm/llms/soniox/common_utils.py
2026-08-29 12:02:21 -07:00
Devin AI
23703a5341
Merge remote-tracking branch 'origin/litellm_internal_staging' into devin/1787944648-registry-audit-rolling
2026-08-29 19:02:03 +00:00
mateo-berri
9e01bd1441
fix(azure): send the deployment name as the body model on v1 image routes
2026-08-29 11:41:58 -07:00
mateo-berri
68f891fd2b
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_keyless_key_managed_resource_owner
2026-08-29 11:33:08 -07:00
Mateo Wang
c42ac262d3
Merge branch 'litellm_internal_staging' into fix_databricks_oauth_url
2026-08-29 07:00:58 -07:00
mateo-berri
163c7033d6
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5
2026-08-29 06:38:33 -07:00
mateo-berri
c37260a2bd
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5
...
# Conflicts:
# basedpyright-code-budget.json
# enterprise/litellm_enterprise/proxy/audit_logging_endpoints.py
# litellm/_lazy_imports.py
# litellm/a2a_protocol/litellm_completion_bridge/transformation.py
# litellm/integrations/bitbucket/bitbucket_client.py
# litellm/integrations/compression_interception/handler.py
# litellm/integrations/prometheus_helpers/prometheus_api.py
# litellm/litellm_core_utils/model_response_utils.py
# litellm/litellm_core_utils/url_utils.py
# litellm/llms/anthropic/experimental_pass_through/context_management/dispatcher.py
# litellm/llms/anthropic/experimental_pass_through/responses_adapters/handler.py
# litellm/llms/anthropic/skills/transformation.py
# litellm/llms/azure/files/handler.py
# litellm/llms/bedrock/realtime/handler.py
# litellm/llms/chatgpt/chat/streaming_utils.py
# litellm/llms/compactifai/chat/transformation.py
# litellm/llms/oci/chat/cohere.py
# litellm/llms/vertex_ai/vector_stores/rag_api/transformation.py
# litellm/proxy/agent_endpoints/agent_registry.py
# litellm/proxy/client/cli/commands/credentials.py
# litellm/proxy/client/cli/commands/teams.py
# litellm/proxy/common_utils/get_routes.py
# litellm/proxy/db/routing_prisma_wrapper.py
# litellm/proxy/guardrails/guardrail_hooks/custom_code/sandbox.py
# litellm/proxy/guardrails/guardrail_hooks/hiddenlayer/hiddenlayer.py
# litellm/proxy/guardrails/guardrail_hooks/llm_as_a_judge/__init__.py
# litellm/proxy/guardrails/guardrail_hooks/promptguard/promptguard.py
# litellm/rust_bridge/responses_websocket.py
# litellm/secret_managers/secret_manager_handler.py
# ruff-strict-budget.json
# type-discipline-budget.json
2026-08-29 06:37:10 -07:00
Mateo Wang
e48f8f016f
Merge pull request #38148 from mubashir1osmani/litellm_hosted_vllm_videos
...
feat(hosted_vllm): add vLLM-Omni videos API
2026-08-29 06:30:35 -07:00
Devin AI
f0849eb0c9
fix(models): xai retirement repricing, bedrock grok-4.6 caching, openai/gemini deprecation dates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 13:11:26 +00:00
mateo-berri
8d4620649f
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_decrease_anys_opus5
...
# Conflicts:
# basedpyright-code-budget.json
# litellm/caching/valkey_semantic_cache.py
# litellm/integrations/compression_interception/handler.py
# litellm/integrations/custom_logger.py
# litellm/llms/custom_httpx/container_handler.py
# litellm/llms/infinity/rerank/transformation.py
# litellm/proxy/agent_endpoints/agent_registry.py
# litellm/repositories/base_repository.py
# litellm/repositories/credentials_repository.py
# litellm/repositories/team_repository.py
# ruff-strict-budget.json
# type-discipline-budget.json
2026-08-29 06:03:33 -07:00
Devin AI
401b12e64b
Merge remote-tracking branch 'origin/litellm_internal_staging' into devin/1787944648-registry-audit-rolling
2026-08-29 13:02:37 +00:00
Mateo Wang
6a3333d3c8
Merge pull request #38747 from BerriAI/litellm_aws_partition_helper
...
fix(aws): build every AWS endpoint and ARN from the region's partition (aws-cn, aws-us-gov)
2026-08-29 04:10:59 -07:00
Mateo Wang
39e4f1ae13
Merge pull request #38727 from BerriAI/litellm_aws_external_id_embed_sagemaker
...
fix(aws): forward aws_external_id in Bedrock embeddings and SageMaker credential loading
2026-08-29 03:30:32 -07:00
Devin AI
db1b1e2195
fix: bound Hugging Face config fetch and keep embedding tests off the network
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-29 09:45:52 +00:00
mateo-berri
1947c65081
test(aws): type the new partition test parameters
2026-08-29 02:27:18 -07:00
mateo-berri
7fbcd9c3ed
fix(bedrock): treat partial record counts as unknown on batch retrieve
2026-08-29 01:40:07 -07:00
mateo-berri
ad8c1457d1
fix(aws): build every AWS endpoint and ARN from the region partition
...
Adds litellm/litellm_core_utils/aws_partition.py mapping a region to its
AWS partition (aws, aws-cn, aws-us-gov, and the iso partitions), its DNS
suffix, and its ARN prefix, and uses it at every AWS host and ARN build
site: bedrock (runtime, agent, agentcore, legacy client, batches, files,
realtime), sagemaker, polly, secrets manager, s3 log uploads, bedrock
passthrough routes, and rag ingestion. ARN detection now accepts
arn:aws-cn: and arn:aws-us-gov: prefixes.
STS region resolution now falls back to the configured aws_region_name
after the aws_sts_endpoint host and the AWS_REGION/AWS_DEFAULT_REGION env
vars, so cn and gov role assumption no longer silently signs against
us-west-2.
A partition sweep test walks every endpoint builder with cn regions and
asserts no amazonaws.com host or arn:aws: prefix comes out, plus an AST
guard that fails on any new f-string hardcoding either literal.
2026-08-29 01:21:59 -07:00
mateo-berri
5d34fb20ff
fix(bedrock): map real batch record counts and guard zero-count retire
2026-08-29 01:05:36 -07:00
mateo-berri
1a26608769
feat(vertex_ai): route gemini transcribe models to generateContent on /v1/audio/transcriptions
2026-08-29 00:46:03 -07:00
mateo-berri
44a5d7a47a
fix(aws): forward aws_external_id in bedrock embeddings and sagemaker credential loading
2026-08-28 18:25:14 -07:00
mateo-berri
1a5a856e3e
fix(guardrails): defer native /v1/messages stream logging until post_call scans finish
2026-08-28 16:05:45 -07:00
Mateo Wang
27c09248e4
Merge pull request #38593 from BerriAI/litellm_gpt5_default_reasoning_effort
...
fix(gpt-5): stop forwarding temperature and top_p to reasoning models that reject them
2026-08-28 15:20:41 -07:00
Mateo Wang
d4f6b4491d
Merge pull request #38656 from aaaaaandrew/litellm_preserve_stream_selected_model
...
fix(streaming): preserve provider model for cost calculation
2026-08-28 13:00:35 -07:00