mateo-berri
c0fd8f6012
fix(anthropic): merge a case-variant Anthropic-Beta client header instead of clobbering it
2026-09-15 11:03:10 -07:00
ryan-crabbe-berri
08b433267e
Merge pull request #40917 from BerriAI/litellm_credential_conflict_409
...
fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in
2026-09-15 10:58:26 -07:00
ryan-crabbe-berri
6dcca8c4ae
Merge pull request #41028 from BerriAI/litellm_bulk_new_user
...
feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation
2026-09-15 10:57:46 -07:00
Devin AI
a426df108a
feat(http): opt-in outbound HTTP/2 for httpx clients
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:54:03 +00:00
Mateo Wang
e5547a56a9
Merge pull request #41257 from BerriAI/litellm_pr_template_coding_tool_proof
...
docs(github): ask for interactive coding-tool proof in the PR template
2026-09-15 10:44:53 -07:00
berriai-litellm-provider-info-sync[bot]
363b7835a3
chore(prices): sync AWS Bedrock prices: 4 models
...
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
a8dd1394f9
chore(prices): drop source from us-gov rows again
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
1b69cfcf3c
chore(prices): sync AWS Bedrock prices: 4 models
...
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
kerry
68f97321dd
chore(prices): drop source from us-gov rows to match usgov pricing test
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
8a9a08caf0
chore(prices): sync prices for 2 providers: 14 models
...
chatgpt-image-latest: input_cost_per_image_token_batches
gemini-2.0-flash: input_cost_per_audio_token_batches
gemini-2.0-flash-lite: input_cost_per_audio_token_batches
gemini-2.5-flash: input_cost_per_audio_token_batches
gemini-2.5-flash-lite: input_cost_per_audio_token_batches
gemini-3-flash-preview: input_cost_per_audio_token_batches
vertex_ai/gemini-3-flash-preview: input_cost_per_audio_token_batches
gemini-3.1-flash-lite: input_cost_per_audio_token_batches
vertex_ai/gemini-3.1-flash-lite: input_cost_per_audio_token_batches
gpt-image-1: input_cost_per_image_token_batches
gpt-image-1-mini: input_cost_per_image_token_batches
gpt-image-1.5: input_cost_per_image_token_batches
gpt-image-1.5-2025-12-16: input_cost_per_image_token_batches
gpt-image-2: input_cost_per_image_token_batches
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
01b069eb6c
chore(prices): sync AWS Bedrock prices: 25 models
...
anthropic.claude-fable-5:
anthropic.claude-fable-5-1:
anthropic.claude-opus-4-7:
anthropic.claude-opus-4-8:
anthropic.claude-opus-5:
anthropic.claude-sonnet-4-6:
anthropic.claude-sonnet-5:
global.anthropic.claude-fable-5:
global.anthropic.claude-fable-5-1:
global.anthropic.claude-opus-4-7:
global.anthropic.claude-opus-4-8:
global.anthropic.claude-opus-5:
global.anthropic.claude-sonnet-4-6:
global.anthropic.claude-sonnet-5:
us-gov.anthropic.claude-fable-5-1:
us-gov.anthropic.claude-opus-4-8:
us-gov.anthropic.claude-opus-5:
us-gov.anthropic.claude-sonnet-5:
us.anthropic.claude-fable-5:
us.anthropic.claude-fable-5-1:
us.anthropic.claude-opus-4-7:
us.anthropic.claude-opus-4-8:
us.anthropic.claude-opus-5:
us.anthropic.claude-sonnet-4-6:
us.anthropic.claude-sonnet-5:
2026-09-15 17:42:56 +00:00
berriai-litellm-provider-info-sync[bot]
dc3a0399f2
chore(prices): sync Vertex AI prices: 4 models
...
gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini/gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches
gemini-3.1-flash-live-preview: input_cost_per_second
gemini/gemini-3.1-flash-live-preview: input_cost_per_second
2026-09-15 17:42:56 +00:00
ryan-crabbe-berri
a7f180fdd8
feat(terraform): make credential adoption opt-in, escape names in request URLs
...
Terraform's convention is that create does not seize a resource the
configuration never made, and credential_values holds secrets that are never
read back into state, so a silent takeover overwrites values no plan showed.
A name collision now fails with the terraform import command that adopts the
existing credential explicitly, and adopt_existing = true opts into taking it
over during create. The provider detects the conflict by the proxy's 409 and
keeps the Prisma string match as a fallback for older proxies
Credential names and model_id went into URLs raw, so a name with a slash or a
question mark hit the wrong route. Every credential URL is now built from a
package const through fmt.Sprintf with url.PathEscape or url.QueryEscape,
which the endpoint audit can resolve. Toggling adopt_existing alone no longer
sends a PATCH, so it does not rewrite the stored secret
2026-09-15 10:41:44 -07:00
ryan-crabbe-berri
81806f33cf
fix(credentials): answer 409 on a name collision, let PATCH resolve values from model_id
...
POST /credentials let a duplicate name hit the unique index and handed back
Prisma's "Unique constraint failed" as a 500, so callers string-matched that
message to tell a caller mistake from a server fault. The unique violation now
maps to a 409 whose message names the PATCH route, two concurrent creates of
one name agree on it, and the detection lives in a repository helper the five
hand-rolled copies can move onto later
PATCH /credentials/{name} took a CredentialItem body, so the model_id the
Terraform adopt path sent was dropped. It now accepts UpdateCredentialItem and
shares the deployment lookup with create. Both handlers take the router as a
FastAPI dependency instead of reading the proxy global, which is what the
tests override
2026-09-15 10:41:43 -07:00
kerry-berri
b97bc10ec9
Merge pull request #41263 from BerriAI/litellm_usgov_rows_allow_price_list_source
...
test(pricing): let synced GovCloud Bedrock rows cite the AWS price list
2026-09-15 10:41:42 -07:00
Devin AI
9fce6e1989
test(pricing): let synced GovCloud Bedrock rows cite the AWS price list
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:09:12 +00:00
Devin AI
79450121f8
test(proxy): document access group test seam
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 17:00:11 +00:00
tin-berri
3ad9a7f336
Merge pull request #41174 from BerriAI/litellm_tier_model_affinity
...
fix(router): preserve session model choice within each complexity tier
2026-09-15 09:54:53 -07:00
mateo-berri
a502afe608
docs(github): ask for interactive coding-tool proof in the PR template
2026-09-15 09:34:45 -07:00
Devin AI
56d0f953f5
fix(proxy): list directly assigned team models in model access errors
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 16:33:41 +00:00
Devin AI
3000c00c50
fix(models): add supports_response_schema to fireworks deepseek-v4-flash-vision-exp short alias
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 14:03:24 +00:00
Devin AI
631ba7f9e3
fix(models): dedupe merged keys, price gemini *-latest aliases at live targets, add Nova cache read prices and Fireworks deprecation dates
...
Absorbs #41148 and #41152 into the rolling registry PR.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 13:36:11 +00:00
Mateo Wang
c8114ba41f
Merge pull request #40921 from BerriAI/litellm_unified_key_policy_hook
...
feat(proxy): unified custom_key_policy hook for key generate, update and regenerate
2026-09-15 06:34:28 -07:00
Devin AI
6c8ab2f22e
Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_14
2026-09-15 13:09:38 +00:00
Mateo Wang
9496f16f12
Merge pull request #41173 from BerriAI/litellm_realtime_health_check_credential_name
...
fix(health): resolve litellm_credential_name in realtime health checks
2026-09-15 05:24:29 -07:00
mateo-berri
1f2d050386
Merge remote-tracking branch 'origin/main' into litellm_unified_key_policy_hook
2026-09-15 05:11:29 -07:00
yassin
b64e430e93
test(proxy): record custom tokenizer loads with a mock instead of a mutable list
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 09:37:53 +00:00
mateo-berri
87b630429d
fix(realtime): send openai and xai health check keys as bearer tokens
2026-09-15 02:20:09 -07:00
yassin
0c611e63c8
fix(utils): cache custom HuggingFace tokenizers across /utils/token_counter requests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 09:12:49 +00:00
Mateo Wang
3ed6c19b8d
Merge pull request #40915 from BerriAI/litellm_internal_copy_37075
...
fix(vertex-live): bill Gemini Live sessions end to end (internal copy of #37075 )
2026-09-15 02:06:49 -07:00
Mateo Wang
c274fd8781
Merge pull request #41168 from BerriAI/litellm_bedrock_wif_session_policy_coverage
...
fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy
2026-09-15 01:49:41 -07:00
mateo-berri
5969fbb052
refactor(realtime): freeze the empty credential mapping in the health check
2026-09-15 01:44:46 -07:00
yucheng
5aa5c092d5
refactor(utils): set converted-stream logging flags inline instead of mutating a helper parameter
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 08:42:19 +00:00
mateo-berri
0770f663c1
fix(vertex-live): report repeated search queries once per grounded turn
...
The session usage collapsed duplicate query strings across turns while the
price was per turn, so two turns asking the same question paid two fees yet
reported web_search_requests 1. Sum each turn's grounding requests so the
counter matches the bill; duplicates within one turn still collapse.
2026-09-15 01:35:36 -07:00
yucheng
496c2a5513
fix(caching): replay agentic loop follow-up cache hits as plain objects
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 08:35:13 +00:00
mateo-berri
1771255b32
Merge remote-tracking branch 'origin/main' into litellm_realtime_health_check_credential_name
2026-09-15 01:17:31 -07:00
mateo-berri
d5938ff886
Merge remote-tracking branch 'origin/main' into litellm_bedrock_wif_session_policy_coverage
2026-09-15 01:16:45 -07:00
yucheng
1b474b075f
fix(proxy): attach litellm_call_id to client disconnect log record
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 08:15:14 +00:00
Mateo Wang
80b9ed4f2c
Merge pull request #41191 from BerriAI/litellm_router_test_cap_resets_per_fallback_hop
...
fix(router): count num_retries_per_request across fallback hops
2026-09-15 01:13:41 -07:00
Devin AI
0a8eb56ba4
fix(proxy): fall back to request data when logging object has no call id
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 08:03:45 +00:00
Mateo Wang
e5cb8b7534
Merge pull request #40984 from BerriAI/litellm_anthropic_guardrail_system_and_tool_use
...
fix(guardrails): scan the Anthropic top-level system prompt and tool_use arguments
2026-09-15 00:55:42 -07:00
yucheng
ce45d6a09d
style: drop explanatory docstrings from converted-stream helpers and tests
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 07:51:19 +00:00
yucheng
621db91d90
fix(caching): defer cache-hit callbacks by replayed result type, not request flags
...
A converted-stream request whose cache entry is a plain (non-stream) object is
replayed as that plain object, so nothing later fires the success callbacks.
Decide deferral from the replayed result's type instead of the request kwargs.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 07:51:19 +00:00
yucheng
8b86362703
fix(caching): replay cache hits for converted streams as streams
...
A deployment hook (Headroom, code interpreter, web search) can downgrade
kwargs["stream"] to False while the caller still expects to iterate the
result. The cache handler keyed stream replay and callback deferral off
the raw flag, so a cache hit returned a plain object to a caller that
iterates, and the Responses iterator never persisted the converted
stream in the first place. Key both off the conversion marker as well
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 07:51:19 +00:00
yucheng
95ef538789
fix(utils): log converted streams as streams so spend tracking works
...
Deployment hooks such as Headroom downgrade stream=True to a non-streaming provider call and the agentic loop then hands back a CustomStreamWrapper (or MockResponsesAPIStreamingIterator for Responses). wrapper_async still saw kwargs["stream"] is False, so it took the non-streaming success path with a lazy stream object: no standard_logging_object was built, the proxy cost callback raised failed_tracking_spend, and the wrapper's own end-of-stream dispatch was deduped away. Treat a lazy stream result as streaming for logging regardless of the downgraded kwarg. Regression in v1.99.0 via #35017
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 07:51:19 +00:00
Devin AI
bc17459548
fix(proxy): include litellm_call_id in LLM API exception logs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 07:46:41 +00:00
mateo-berri
1b040af414
test(router): type the retry-cap tests this PR adds or touches
2026-09-15 00:34:38 -07:00
mateo-berri
4f27573424
merge: origin/main into litellm_internal_copy_37075
2026-09-15 00:34:18 -07:00
yuneng-jiang
81863c1b17
Merge pull request #41188 from BerriAI/litellm_spend_reconciliation
...
test(spend): reconcile concurrent requests and daily activity
2026-09-15 00:32:20 -07:00
tin-berri
feab83aae1
Merge pull request #41186 from BerriAI/litellm_statusline_router_cost_label
...
fix(cli): label savings cost bars with the auto-router name
2026-09-15 00:32:08 -07:00