kerry
7761d04450
test(vertex): cover malformed batch usage details
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 02:37:51 +00:00
kerry
27a486e4d3
test(cost): cover modality guards and image detection fallbacks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:34:17 +00:00
kerry
4a8ec7b9d8
fix(cost): bill gemini-embedding-2-preview per token like the GA entries
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:18:50 +00:00
kerry
6a18105275
fix(vertex): only bill image rate without modality details when every input is an image
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:16:00 +00:00
kerry
ac8e1a355c
test(vertex): load local pricing in embedding billing tests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 01:00:32 +00:00
kerry
e5845c17ff
fix(vertex): bill image inputs at the image rate when usage lacks modality details
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 00:51:49 +00:00
kerry
0c91d9157c
refactor(vertex): drop unused resolved_files from embed response parsing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 00:45:08 +00:00
kerry
d4f2119b03
fix(cost): bill gemini-embedding-2 per token and stop double charging audio
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 00:41:46 +00:00
tin-berri
c626ff098b
Merge pull request #40877 from BerriAI/litellm_lit7658_cache_cost_v0_fresh
...
feat(proxy): predict prompt-cache costs across deployments
2026-09-14 16:25:41 -07:00
yassin
fbc1011d27
fix(bedrock): end the realtime session when the client disconnects instead of waiting for Nova Sonic
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:42:42 +00:00
yassin
b4d0f4ad26
refactor(realtime): move session ownership marker keys into constants
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:33:48 +00:00
yassin
d2342f06ce
fix(bedrock): stamp the realtime success ownership marker when Nova Sonic spend is logged
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 21:13:29 +00:00
yassin
771430df7b
fix(bedrock/realtime): keep partial spend on cancelled output task and make the committed-session refusal non-retryable
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 19:07:21 +00:00
yassin
5c04ec2b93
fix(bedrock/realtime): keep the pending session.update until a provider stream is committed
...
Peek at the pending session.update instead of popping it, so an eager fallback failure before the bridge starts does not lose the replay for the next attempt. Move the websocket scope keys to constants.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 19:07:20 +00:00
yassin
cb5d901774
fix(bedrock/realtime): propagate deferred Nova Sonic stream failures to the router
...
Bedrock realtime caught every exception inside both forwarding tasks and
gathered them with return_exceptions=True, so a provider failure surfacing
after the websocket handshake (lazy duplex stream: 503/429/validation only
show up on await_output or the input publisher) made async_realtime return
normally and the router recorded a success instead of running fallbacks and
cooldown accounting. session.updated is now acked only after Bedrock is
ready, provider failures escape as BedrockError with the AWS status code,
a failure after the client disconnected is not reported as a provider
failure, and a fallback attempt on the same websocket replays the pending
session.update instead of emitting a second session.created.
Resolves LIT-6484
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 19:07:20 +00:00
Mateo Wang
c2c2a623c0
Merge pull request #39846 from BerriAI/litellm_bedrock_mantle_govcloud_cost_row
...
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
fix(bedrock_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names
2026-09-12 21:13:58 -07:00
devin-ai-integration[bot]
77dc1a6c03
fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events ( #33352 )
...
* fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* style(anthropic-adapter): drop added comments per repo convention
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-09-12 21:13:35 -07:00
Mateo Wang
70e3f5a02e
Merge pull request #39836 from BerriAI/litellm_lit_6975_bedrock_files_delete_list
...
feat(bedrock): support file delete and list for S3-backed managed files
2026-09-12 21:13:27 -07:00
kerry-berri
9ae727bc8e
Merge pull request #40929 from BerriAI/litellm_fireworks_short_key_lookup
...
fix(fireworks): resolve short model names to long cost map keys
2026-09-12 20:49:44 -07:00
Devin AI
a519d805bb
refactor(fireworks): resolve cost map key through a provider config hook
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-13 03:29:11 +00:00
Devin AI
0904051fda
refactor(fireworks): move cost map key construction under llms/
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-13 03:24:23 +00:00
Mateo Wang
9d984371fd
Merge pull request #40909 from BerriAI/litellm_databricks_reasoning_effort_thinking
...
fix(databricks): translate reasoning_effort to thinking for Gemini 2.5
2026-09-12 17:46:44 -07:00
ryan-crabbe-berri
c134fb7a38
Merge pull request #39395 from seyeong-han/litellm_meta_muse_voice_realtime
...
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
feat(realtime): add Meta Muse Voice transcription
2026-09-12 15:58:24 -07:00
Mateo Wang
fe5ff9d3b0
Merge pull request #40771 from BerriAI/litellm_regression_coverage_followup
...
test: tighten regression tests added in #37974
2026-09-12 15:29:36 -07:00
Mateo Wang
c040061f15
Merge pull request #40766 from BerriAI/litellm_fix_v1_messages_disconnect_partial_cost
...
fix(anthropic): price recovered tokens when a /v1/messages client disconnects mid-stream
2026-09-12 15:21:43 -07:00
ryan-crabbe-berri
4647cd1215
fix(realtime): keep a Muse turn active for turnless partials after speechEnd
...
Muse partials carry no turnId and belong to the most recent speechStart,
and the docs say the model may keep post processing a turn after speechEnd
until speechComplete. Releasing the active turn on speechEnd made any
partial arriving in that window raise and get dropped in ENDPOINTING mode.
The turn now stays active until its speechComplete or final transcript.
2026-09-12 15:13:44 -07:00
Tin Chi Lo
cba843cc16
feat(proxy): predict prompt-cache costs across deployments
2026-09-12 14:02:45 -07:00
mateo-berri
c7b607c46e
fix(databricks): keep the Claude fallback when gating the anthropic thinking payload
...
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Gate the reasoning_effort translation on the cost-map flag or the model name containing
claude, so unmapped Claude serving endpoints keep translating. Flag the newer Claude
entries that were missing it. Expose supports_anthropic_thinking_payload as a public
helper next to the other supports_* wrappers instead of importing the private factory.
Drop the adaptive-only guard, since the adaptive flags only ever match Claude ids, and
add regression tests for an unmapped Claude endpoint and an adaptive Claude model
2026-09-12 13:13:38 -07:00
mateo-berri
10a0da7a32
Merge litellm_internal_staging into devin/1784568628-databricks-gemini-reasoning-effort
2026-09-12 13:08:22 -07:00
ryan-crabbe-berri
a2e383a1a5
fix(realtime): ignore a late speechStart for a finished Muse turn
...
A duplicate speechStart for a turn that already stopped used to make that
closed turn active again, so the next turnless PUSH_TO_TALK transcript was
routed to the finished item and dropped.
2026-09-12 12:53:03 -07:00
ryan-crabbe-berri
6b78438c99
fix(realtime): close every Muse turn on its own terminal signal
...
Turns no longer wait behind each other in a FIFO queue, so an empty
server_vad turn (speechStart then speechEnd with no transcript) cannot
stall every later turn, and a PUSH_TO_TALK speechComplete now closes its
turn without waiting for a speechEnd that never arrives. Each turn keeps
its own idempotent emit state, so late or duplicate speechEnd,
speechComplete and transcript frames are no-ops, and finished turns are
remembered in a bounded map instead of a separate tombstone deque.
The session.created ack and the sanitized error frame are now typed as
members of OpenAIRealtimeEvents, which removes the typing.cast calls
that the strict ruff budget flagged.
2026-09-12 12:39:59 -07:00
yujonglee
347b642bdd
refactor(ocr): complete native lifecycle and preserve Azure auth ( #40734 )
...
* refactor(ocr): extract call completion boundary
* fix(ocr): release completion state after dispatch
* test(ocr): prove wrapper completion handoff
* test(ocr): narrow mapped failure assertion
* fix(ocr): preserve wrapper invocation kwargs
* fix(ocr): retain completion through finalization
* fix(ocr): make completion ownership explicit
* refactor(ocr): resolve logging executor explicitly
* fix(callbacks): preserve completion lifecycle behavior
* refactor(ocr): move public OCR into native lifecycle
* refactor(ocr): remove unused rust bridge capability
* wip
* wip
* refactor
* wip
* fix(ocr): preserve reducto native compatibility
* wip
* fix(ocr): document native callable casts
* perf(ocr): bound responses and reduce native scheduling overhead
* refactor(python-bridge): organize placeholder routes
* refactor test
* fix(ocr): normalize DeepSeek document content
* perf(ocr): skip unused callback work and benchmark callback overhead
* fix(ocr): align conversion contracts
* test(ocr): cover official provider response shapes
* fix(ocr): restore Python fallback and honor Rust opt-out
* fixes and refactor
* fix(ocr): preserve Azure Document Intelligence authentication
* fix(rust): enforce OCR response limits and lint contracts
* test(rust): align native OCR contract coverage
* test(ocr): isolate Azure auth precedence coverage
2026-09-12 11:56:49 -07:00
devin-ai-integration[bot]
eddfb5fb20
fix(responses): preserve hosted web search calls ( #40828 )
...
* fix(responses): preserve hosted web search calls
Co-Authored-By: Claude Code <noreply@anthropic.com>
(cherry picked from commit 09183b3346 )
* chore: remove unrelated generated schema documentation changes
(cherry picked from commit ca6a860757 )
* fix(responses): preserve hosted search context during replay
(cherry picked from commit 425f1e9b3a )
* chore: regenerate dashboard API types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Tin Chi Lo <tin@berri.ai>
Co-authored-by: Claude Code <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 09:52:19 -07:00
Mateo Wang
99e14fc2e5
Merge pull request #40798 from BerriAI/litellm_lit7523_mantle_reasoning_summary
...
fix(bedrock_mantle): gate reasoning.summary on the OpenAI Responses path
2026-09-11 20:30:49 -07:00
ryan-crabbe-berri
17fde7a261
refactor(realtime): move Meta Muse Voice onto BaseRealtimeConfig
...
Replace the hand-rolled Meta realtime handler with a MetaRealtimeConfig
that plugs into the shared realtime handler and RealTimeStreaming relay.
Clients keep speaking the OpenAI realtime wire: session.update,
input_audio_buffer.append/commit and the OpenAI transcription events.
Unsupported transcription settings are logged and dropped, matching the
Gemini realtime precedent, and the Meta-specific session.mode, keywords,
language_bias, DIARIZATION and speaker extensions are removed.
Drop the MODEL_API_KEY env var in favor of the standard META_API_KEY,
remove the private-logging flag so spend logs record the transcript the
same way other realtime models do, and add per-second pricing for
muse-voice-transcribe-1.0.
The relay now sends raw bytes from transform_realtime_request straight to
the backend after pace_backend_send, and transcription sessions never
trigger response.create.
2026-09-11 20:12:15 -07:00
Young Han
1acb994998
fix(realtime): bound Muse audio before decoding
2026-09-11 20:10:53 -07:00
Young Han
b82b31a44f
feat(realtime): add Meta Muse Voice transcription
2026-09-11 20:10:53 -07:00
devin-ai-integration[bot]
1fde15c1ec
fix(shadow-eval): skip hosted web search samples ( #40827 )
...
(cherry picked from commit a78cd2fe02 )
Co-authored-by: Tin Chi Lo <tin@berri.ai>
2026-09-12 02:11:48 +00:00
devin-ai-integration[bot]
7057b2f6c4
fix(fireworks_ai): keep reasoning_content on replayed assistant messages ( #40682 )
...
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 18:47:37 -07:00
shivam
37a1b859d7
fix(bedrock_mantle): harden reasoning summary validation
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 23:34:14 +00:00
ryan-crabbe-berri
4e9d414603
Merge pull request #40748 from BerriAI/litellm_fix_gemini_reasoning_effort_400
...
fix(vertex_ai): return 400 for invalid reasoning_effort instead of 500
2026-09-11 16:31:16 -07:00
yujonglee
0dd5e6e289
feat(ocr): add Reducto legacy and v3 adapters ( #40535 )
...
* feat(ocr): add Reducto adapters
* fix(ocr): decline missing Reducto credentials
* fix(ocr): map Reducto credentials in gateway errors
* test(ocr): keep Reducto coverage at SDK boundary
* test(ocr): remove stale gateway Reducto cases
* fix(ocr): stop retaining Reducto responses by default
* refactor(ocr): preserve Reducto extra params
* refactor(ocr): adopt request preparation contract
* fix(ocr): preserve provider model passthrough
* fix(ocr): reject unknown Reducto models
* fix(ocr): preserve Reducto provider options
2026-09-11 16:22:56 -07:00
shivam
bf9e8ea9ef
fix(bedrock_mantle): gate reasoning.summary on the OpenAI Responses path
...
Mantle's /openai/v1/responses rejects reasoning.summary values other than "auto" with 400 unsupported_parameter. Drop it with a warning under drop_params, otherwise raise UnsupportedParamsError naming the remedy.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 23:19:17 +00:00
Mateo Wang
f90b5cad8a
Merge pull request #40740 from BerriAI/litellm_bedrock_openai_xhigh_flags
...
fix(cost-map): bedrock reasoning effort flags, registry audit fixes for vertex/openai/together/openrouter, absorb cerebras and inception rows
2026-09-11 14:41:18 -07:00
yuneng-jiang
d51a7af655
fix(search): propagate GET provider HTTP errors ( #40779 )
2026-09-11 14:07:32 -07:00
mateo
e3130a87bc
test(cost-map): type the monkeypatch fixture in cerebras and inception registry tests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 19:50:27 +00:00
mateo
4ffd4ecc83
fix(model_prices): absorb cerebras/inception PRs, fix vertex/openai/together/openrouter pricing
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-11 19:37:45 +00:00
yujonglee
89f1f9567d
refactor(ocr): route native requests through core ( #40532 )
...
* refactor(ocr): route native Mistral through core
* fix(ocr): preserve Azure API base resolution
* chore(ocr): document bridge boundary casts
* fix(ocr): keep Azure environment resolution in Rust
* fix(ocr): centralize native execution and isolate request logging
* refactor(ocr): narrow native migration to bridge routing
---------
Co-authored-by: Stack Plan <stack-plan@example.invalid>
2026-09-11 12:37:19 -07:00
Mateo Wang
0fe9de8550
Merge pull request #39507 from BerriAI/litellm_fix_oci_streaming_chunk_ids
...
fix(oci): pin one response id per streamed completion, skip the [DONE] sentinel
2026-09-11 11:46:35 -07:00
ryan-crabbe-berri
f72b117b21
fix(vertex_ai): return 400 for invalid reasoning_effort instead of 500
...
Both reasoning_effort mappers ended their if/elif chain in a bare ValueError.
exception_type() has no branch for ValueError, so it fell through to the shared
APIConnectionError fallback and the proxy answered a malformed client request
with a retryable HTTP 500 carrying no hint of the accepted values.
Raise UnsupportedParamsError (400) instead, listing the supported set, matching
what the Anthropic and Bedrock transforms already do and what this same file
already does at its five other param-validation sites.
This also covers 'xhigh' and 'max', which are members of litellm's own
REASONING_EFFORT literal but have no Gemini mapping, so callers bridging from
OpenAI-shaped code were hitting the 500 without typing anything wrong.
Fixes #40474
Claude-Session: https://claude.ai/code/session_01XT1qsbjLwnhiN5sQ2hNUxr
2026-09-11 10:00:35 -07:00