Commit graph

51971 commits

Author SHA1 Message Date
albertbausili
c9c268fa90 chore: merge upstream main into neuraltrust guardrail PR 2026-09-28 11:54:52 +02:00
albertbausili
5c437f7efb fix(guardrails): apply buffered tool argument rewrites 2026-09-21 18:06:10 +02:00
kerry-berri
aa2219d466
Merge pull request #42250 from BerriAI/litellm-providers/price-sync-aws-bedrock
chore(prices): sync AWS Bedrock prices: 25 models [enrichment failed: AWS Bedrock, 66 held]
2026-09-21 08:59:45 -07:00
kerry-berri
37fb6e72d4
Merge pull request #42249 from BerriAI/litellm-providers/price-sync-azure
chore(prices): sync Azure prices: 1 model, 1 deprecated
2026-09-21 08:59:08 -07:00
kerry-berri
e270ba5aa6
Merge pull request #42251 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 3 models
2026-09-21 08:41:41 -07:00
berriai-litellm-provider-info-sync[bot]
739227fefc
chore(prices): sync OpenRouter prices: 3 models
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
2026-09-21 15:31:20 +00:00
berriai-litellm-provider-info-sync[bot]
ade39978a9
chore(prices): sync AWS Bedrock prices: 25 models [enrichment failed: AWS Bedrock, 66 held]
anthropic.claude-fable-5: 
anthropic.claude-fable-5-1: 
anthropic.claude-opus-4-7: 
anthropic.claude-opus-4-8: 
anthropic.claude-opus-5: 
anthropic.claude-sonnet-4-6: 
anthropic.claude-sonnet-5: 
global.anthropic.claude-fable-5: 
global.anthropic.claude-fable-5-1: 
global.anthropic.claude-opus-4-7: 
global.anthropic.claude-opus-4-8: 
global.anthropic.claude-opus-5: 
global.anthropic.claude-sonnet-4-6: 
global.anthropic.claude-sonnet-5: 
us-gov.anthropic.claude-fable-5-1: 
us-gov.anthropic.claude-opus-4-8: 
us-gov.anthropic.claude-opus-5: 
us-gov.anthropic.claude-sonnet-5: 
us.anthropic.claude-fable-5: 
us.anthropic.claude-fable-5-1: 
us.anthropic.claude-opus-4-7: 
us.anthropic.claude-opus-4-8: 
us.anthropic.claude-opus-5: 
us.anthropic.claude-sonnet-4-6: 
us.anthropic.claude-sonnet-5:
2026-09-21 15:31:16 +00:00
berriai-litellm-provider-info-sync[bot]
a3dcec463b
chore(prices): sync Azure prices: 1 model, 1 deprecated
azure_ai/MAI-Image-2.5-Pro: deprecation_date
2026-09-21 15:31:08 +00:00
kerry-berri
0e252c2de3
Merge pull request #42246 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 4 models
2026-09-21 08:22:06 -07:00
kerry
97fc8220df fix(prices): align deepseek-v4-pro-0813 off-peak and cache-hit rates with the base cache read rate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 15:14:19 +00:00
Yassin Kortam
e912ebe999
Merge pull request #42152 from BerriAI/litellm_claude_code_safeguards_passthrough
fix(anthropic): forward safeguards and anthropic-beta unchanged on native /v1/messages
2026-09-21 10:12:44 -05:00
kerry-berri
701c2b7256
Merge pull request #34941 from BerriAI/litellm_azure_ai_mai_image_2_5_pro 2026-09-21 08:07:36 -07:00
berriai-litellm-provider-info-sync[bot]
33e64e53f9
chore(prices): sync OpenRouter prices: 4 models
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
2026-09-21 15:00:53 +00:00
Yassin Kortam
1ac4d7ae04 fix(anthropic): type safeguards and safeguard_results as the arrays Anthropic sends
Driving a real Claude Code 2.1.278 through the proxy, and a direct call to
api.anthropic.com, both show these two fields are JSON arrays on the wire rather
than objects. The request carries safeguards as
[{"type": "dangerous_tool_use", "classifier_context": {...}}] under beta
dangerous-tool-use-2026-09-03, and the 200 comes back with safeguard_results as
[{"type": "dangerous_tool_use", "status": {"type": "available", "tool_uses": {...}}}].

No runtime change: the request filter matches on TypedDict keys and never inspects
the value. The test fixtures move to the captured shapes so the regression tests
pin what the client and the provider actually exchange.
2026-09-21 09:55:09 -05:00
Yassin Kortam
32133a329c
Merge pull request #42239 from BerriAI/litellm_agentcore_a2a_message_stream
fix(a2a): send message/stream for Bedrock AgentCore streaming requests
2026-09-21 09:53:41 -05:00
kerry-berri
f008c018a3
Merge pull request #42243 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 4 models
2026-09-21 07:49:01 -07:00
kerry
7b2d3b36b6 fix(prices): align deepseek-v4-pro-0813 cache hit cost with cache read cost
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 14:40:46 +00:00
berriai-litellm-provider-info-sync[bot]
346ad002c8
chore(prices): sync OpenRouter prices: 4 models
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/meta-llama/llama-3.1-70b-instruct: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token
2026-09-21 14:30:57 +00:00
kerry-berri
0d45883312
Merge pull request #42240 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 07:09:26 -07:00
berriai-litellm-provider-info-sync[bot]
18ca95c3a7
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 14:00:51 +00:00
yassin
6786eb0131 fix(a2a): send message/stream for Bedrock AgentCore streaming requests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 13:39:32 +00:00
kerry-berri
dfbdcf1c9d
Merge pull request #42234 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 06:39:30 -07:00
albertbausili
080b5a486a fix(guardrails): finish stream checks before releasing tool calls 2026-09-21 15:35:31 +02:00
Yassin Kortam
def37c6532
Merge pull request #42097 from BerriAI/litellm_budget_exceeded_422
fix(proxy): return 422 instead of 429 for BudgetExceededError
2026-09-21 08:35:30 -05:00
Yassin Kortam
502de6bed9
Merge pull request #42207 from BerriAI/litellm_helm_componentized_replica_count
fix(helm): render a fixed replicaCount on componentized deployments when HPA is disabled
2026-09-21 08:31:06 -05:00
berriai-litellm-provider-info-sync[bot]
6db2bce43c
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 13:30:42 +00:00
Devin AI
2888b4f5f4 fix(bedrock): whitelist regional qwen3-next keys for converse routing check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 13:30:35 +00:00
Devin AI
d9a97d74db fix(model_prices): add groq qwen3.6-27b deprecation date and bedrock qwen3-next regional pricing
Groq lists qwen/qwen3.6-27b for shutdown on 2026-09-14. Adds the six regional
Bedrock qwen.qwen3-next-80b-a3b entries priced per AWS's published regional
rates (absorbs #42191) with a regression test that the regional entry is used
instead of the US rate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 13:21:50 +00:00
Devin AI
30adc1b2b9 Merge remote-tracking branch 'origin/main' into litellm_azure_ai_mai_image_2_5_pro 2026-09-21 13:20:20 +00:00
kerry-berri
0c01d297b9
Merge pull request #42227 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 2 models
2026-09-21 06:10:10 -07:00
Mateo Wang
0c1dd2cee4
Merge pull request #42220 from BerriAI/litellm_any_sweep_20260921
refactor(types): replace Any with proven types in 32 files
2026-09-21 06:07:47 -07:00
berriai-litellm-provider-info-sync[bot]
85e72a2e2f
chore(prices): sync OpenRouter prices: 2 models
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 13:00:50 +00:00
albertbausili
e80c20aab7 fix(guardrails): hold streamed tool calls until inspection succeeds 2026-09-21 14:48:45 +02:00
Devin AI
8795be0a65 refactor(types): replace Any with proven types in 32 files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 10:51:46 +00:00
albertbausili
4d88cb480d test(guardrails): pin the streamed tool-call block contract for NeuralTrust
LiteLLM forwards streamed tool-call deltas as they arrive and only scans
the assembled call once the stream ends, in every streaming mode. Under
incremental_diff the answer text and the turn's finish_reason stay
withheld, so a client cannot treat the turn as complete and the block
surfaces as a 400 instead of trailing a finished-looking stream.

Cover that in tests and say so in the hook README, so the remaining
exposure is documented rather than implied by the new mode.
2026-09-21 11:41:43 +02:00
albertbausili
ae61e703d2 feat(guardrails): let NeuralTrust transform verdicts reach streaming clients
The native hook never plumbed streaming_transform_mode, so a TrustGuard
transform was silently dropped on streamed tokens while the generic HTTP
path could turn incremental_diff on. Redaction only worked if a customer
kept the adapter the native guardrail is meant to replace.

Expose streaming_transform_mode on the guardrail card and wire it through
the initializer. Under incremental_diff each reply scan withholds the whole
reply via stream_holdback_chars: TrustGuard re-reads the full reply every
round and its redaction spans move as the reply grows, so releasing tokens
early lets a later scan try to rewrite bytes already on the wire, which the
engine can only reject as stream_transform_underflow. Holding them also
means a block lands with nothing streamed.

A scalar transform payload no longer rewrites the request conversation that
the framework attaches to reply scans for context; it rewrites the scanned
reply instead.
2026-09-21 11:15:37 +02:00
Mateo Wang
1cac8bd9ab
Merge pull request #42193 from BerriAI/litellm_responses_bridge_safety_identifier
fix(responses): forward safety_identifier through the chat completion bridge
2026-09-21 01:36:26 -07:00
albertbausili
c955057c65 chore: merge origin/main into the NeuralTrust guardrail branch 2026-09-21 10:23:45 +02:00
Devin AI
b38504b5a6 test(e2e): assert the forwarded Converse body without requiring the model to accept safety_identifier
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 07:37:28 +00:00
Devin AI
0317903a44 ci(e2e-changed): surface failed test ids from the pytest log
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 07:16:02 +00:00
yassin
d266d7324b fix(helm): leave spec.replicas unset unless replicaCount is explicitly configured
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 06:35:05 +00:00
yassin
333fadad6c fix(helm): render a fixed replicaCount on componentized deployments when HPA is disabled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 06:25:43 +00:00
Devin AI
7e0fa40fe3 test(e2e): gate the Bedrock edge capture behind a provider_edge_host opt-in
The Buildkite ephemeral stack runs the gateway in another pod, so it cannot reach the pytest host's provider edge. The GitHub changed-e2e lane runs gateways on the runner and sets E2E_PROVIDER_EDGE_HOST_REACHABLE

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 06:17:03 +00:00
Devin AI
ba93c7402a test(e2e): tolerate provider retries in safety_identifier capture assertion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 04:53:10 +00:00
Devin AI
83223885e6 fix(responses): forward safety_identifier through the chat completion bridge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 04:09:59 +00:00
kerry-berri
e484a7c89c
Merge pull request #42192 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 2 models
2026-09-20 20:39:41 -07:00
berriai-litellm-provider-info-sync[bot]
d5921713a8
chore(prices): sync OpenRouter prices: 2 models
openrouter/~moonshotai/kimi-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/moonshotai/kimi-k3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 03:30:42 +00:00
kerry-berri
0ece1cd426
Merge pull request #42187 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 2 models
2026-09-20 20:09:22 -07:00
berriai-litellm-provider-info-sync[bot]
4258bd366c
chore(prices): sync OpenRouter prices: 2 models
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 03:00:39 +00:00
kerry-berri
946da34260
Merge pull request #42184 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 2 models
2026-09-20 19:38:29 -07:00