Commit graph

17245 commits

Author SHA1 Message Date
mateo-berri
3829ebdce5 fix(token_counter): estimate over-cap strings from evenly spaced samples instead of the prefix
A string above TOKEN_COUNTER_MAX_EXACT_CHARS was counted from its first cap characters and scaled, so a string whose start tokenizes unlike its end got a skewed count, and that count reaches fallback billing when the provider sends no usage. The estimate now tokenizes 16 evenly spaced samples that together total the cap and scales their sum by the string's length, keeping the same bound on work while tracking the whole string
2026-09-07 18:26:23 -07:00
mateo-berri
bb52fd44fa fix(cost_map): label the card's loaded_at as per-worker and cover the integrity-failure fallback 2026-09-07 18:20:41 -07:00
mateo-berri
86790a7723 fix(bedrock): bill Marengo embeddings per request instead of per estimated token
AWS prices Marengo 2.7 and 3.0 text and image embeddings per request, never per
token, and their responses carry no token count. The old transform estimated
prompt tokens from the vector length, which billed a text request at 128 tokens
times the per-token rate (0.00896 instead of 0.00007). Marengo responses now
report zero tokens with query_count and image_count derived from the request
batch, and all six Marengo cost-map entries price per request (with the video
and audio per-second and per-image rates on the base entries). query_count is a
new prompt_tokens_details field wired to input_cost_per_query in the cost
calculator.
2026-09-07 18:20:28 -07:00
yucheng-berri
9bc9104102
fix(proxy): log budget reservation notice once at config load (#40167)
* fix(proxy): log disable_budget_reservation notice once at config load

The disabled-budget-reservation reminder fired as a WARNING inside request
authentication, so every authenticated request on a proxy that deliberately
set the flag produced one warning line. The notice now runs once per worker
when general_settings loads, at INFO, and the request path only skips the
reservation. Reservation skipping and read-time budget checks are unchanged

* fix(proxy): keep budget notice sentinel with constants

* fix(proxy): expose shared budget notice state
2026-09-07 18:18:28 -07:00
tin-berri
1761fe236f
feat(complexity_router): add declarative custom dimensions to the heuristic scorer (#40156)
Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-07 18:17:30 -07:00
mateo-berri
3b199cd3da fix(azure_ai): price seven Foundry catalog names and charge the model router fee once
Add cost map entries for azure_ai/gpt-chat-latest, codex-mini, whisper,
model-router, cohere-command-a, grok-4-20-reasoning, and
grok-4-20-non-reasoning, priced from the live Azure AI Foundry and Azure
OpenAI pricing pages and the Azure Retail Prices API.

Skip the model router flat fee when the response model is the router
entry itself, since the generic cost already priced that fee. Before,
azure_ai/model_router charged it twice.

Resolves LIT-3157
2026-09-07 18:11:58 -07:00
mateo-berri
486af36963 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_5546_count_tokens_offload 2026-09-07 18:11:29 -07:00
mateo-berri
17b7003592 fix(azure_ai): cost embeddings, responses, images, and rerank relays instead of logging zero 2026-09-07 18:11:17 -07:00
mateo-berri
170fece7db fix(token_counter): release the GIL for HuggingFace counts and cap exact counting per string
Both proxy token counting endpoints already count in a worker thread, but the
HuggingFace tokenizer's encode holds the GIL for the whole call, so a 600k-token
count on a Claude model still froze the event loop for up to 0.8 s and every
other request with it. Count through encode_batch_fast, which releases the GIL,
and tokenize at most TOKEN_COUNTER_MAX_EXACT_CHARS characters of any one string
(default 4,000,000), scaling the exact count of that prefix by the string's
length above it so the largest payloads stay bounded.
2026-09-07 18:08:25 -07:00
mateo-berri
2f397fa128 fix(drop_params): honor string flags in litellm_settings and responses, and fail open on non-flag values 2026-09-07 18:06:29 -07:00
tin-berri
7da6fe54b5
fix: skip one-shot Claude Code cache injection (#40175) 2026-09-07 18:03:43 -07:00
Mateo Wang
26d589cd28
Merge pull request #39234 from BerriAI/litellm_fix_agent_mcp_grants
fix(mcp): clear error when an agent-bound key is denied a scoped MCP server + agent MCP grants in the UI
2026-09-07 18:02:16 -07:00
mateo-berri
c698ddeacb fix(azure): let the caller's api-version win over the deployment's on passthrough relays 2026-09-07 18:00:48 -07:00
mateo-berri
69608a29db Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_6975_bedrock_files_delete_list
# Conflicts:
#	litellm/proxy/openai_files_endpoints/common_utils.py
2026-09-07 17:58:02 -07:00
mateo-berri
6f99917b33 fix(router): rewrite multi-segment model groups as whole passthrough path segments 2026-09-07 17:57:54 -07:00
mateo-berri
192ea9ec80 fix(policy_engine): fail open on streaming shapes post_call pipelines cannot govern yet
A post_call pipeline now releases the original stream instead of refusing the
request on every shape it has no handler for: a background request, a pipeline
guardrail without the unified apply_guardrail interface, a route with no
endpoint translation, a buffered stream no translation resolves, and a rewrite
the translation cannot write back (tool-call edits, text edits on translations
without write-back, n>1 chat, an unended Anthropic stream, a Responses dump
with no event envelope). Each case logs a warning naming the policy and
guardrail. Real blocks and writable text masks are unchanged.
2026-09-07 17:55:37 -07:00
mateo-berri
9041768fb4 feat(cost_map): derive source_revision from the loaded bytes instead of a _metadata stamp
The revision an operator checks is now the git blob id of the exact bytes the process
loaded, the same id git rev-parse <commit>:model_prices_and_context_window.json prints,
so it is always present, never goes stale between bot writes, and needs no stamp in the
JSON that every PR touching the file would have to regenerate. The _metadata block, the
generated_at field, the schema and guard changes, and the bot stamping are dropped
2026-09-07 17:47:51 -07:00
mateo-berri
1d7e81cf5d fix(streaming): guard empty choices and missing role when assembling stream chunks 2026-09-07 17:47:06 -07:00
mateo-berri
fb7d06da4b test(budget_reservation): type the tiny-budget reservation helper 2026-09-07 17:45:53 -07:00
mateo-berri
22d7616fe7 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit_7022_azure_ai_passthrough_config
# Conflicts:
#	tests/test_litellm/proxy/pass_through_endpoints/test_llm_pass_through_endpoints.py
2026-09-07 17:44:44 -07:00
mateo-berri
aa1c76bc3b test(cost_map): skip every reserved top-level key in the price map schema test 2026-09-07 17:28:59 -07:00
mateo-berri
6e93d23e1e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit6552_fix_build_base_response_empty_choices 2026-09-07 17:28:18 -07:00
mateo-berri
dc09d9e7cf feat(bedrock): add TwelveLabs Marengo Embed 3.0 embeddings 2026-09-07 17:24:17 -07:00
mateo-berri
0d5ea553da Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_post_call_policy_pipeline 2026-09-07 17:24:10 -07:00
yucheng-berri
1009976c49
fix(bedrock): keep x-amzn-RequestId on chat error responses (#40089)
* fix(bedrock): keep x-amzn-RequestId on chat error responses

Bedrock chat error paths built BedrockError from only a status code and a
message, so the provider response headers were gone before exception mapping
ran and the proxy had nothing to forward. AWS support needs x-amzn-RequestId
to investigate a server-side error.

- converse and invoke chat handlers pass the real headers and response when
  they turn an httpx.HTTPStatusError into a BedrockError, and read the body
  through error_response_text so a streamed body nobody read does not throw
- every bedrock chat get_error_class honors the headers it is already handed:
  invoke, moonshot, bedrock-hosted openai, agentcore and the invoke agent
- BedrockError carries those headers into the response it synthesizes when a
  caller has headers but no response, skipping values httpx cannot carry
- the bedrock 500 mapping forwards the provider response like its 4xx and 503
  siblings instead of fabricating a blank one

The proxy now returns llm_provider-x-amzn-requestid on Bedrock chat errors.

* fix(bedrock): keep request-id on text-classified errors

The context-window and image branches of _map_bedrock_exception built their
litellm exception without the provider response, so a Bedrock 400 classified
by its body text lost x-amzn-RequestId while the sibling branches kept it.

Also narrows the new BedrockError types and trims its docstrings.

* chore(bedrock): drop the docstrings on the new error helpers

* fix(bedrock): keep request-id on every error path that has one

The ticket's root cause is that every BedrockError raise site under
litellm/llms/bedrock/ was built from status and message alone. The first
commits covered the chat and invoke handlers; this covers the rest.

Embeddings, rerank, image generation, image edit, count tokens, search and
the transformation layers now hand on the provider response or its headers,
and both bedrock_mantle configs return a BedrockError instead of the
OpenAI error that drops them.

Two blockers surfaced while verifying the streaming path. The trailing
`except Exception` in make_call and make_sync_call swallowed the BedrockError
raised a few lines above, relabelling a provider status as a 500, and the
non-200 branch read an unread streamed body, which throws.

The raise sites left alone have no provider response to carry: timeouts,
credential and config errors, and mid-stream event frames.

* fix(bedrock): forward provider headers from the count tokens route

The count tokens route converts BedrockError into an HTTPException, and dropped
the headers the handler had just kept, so that route still lost the request id.

get_response_headers now takes a Mapping so an httpx.Headers can be handed to it
without a copy.

* fix(bedrock): classify every bedrock surface through BedrockError

Eleven bedrock configs still inherited a provider-agnostic get_error_class
that builds a blank response, so the request id was gone before the proxy
read it. Claude platform, bedrock anthropic-messages, both image edit
configs, passthrough, realtime, vector stores and agentcore search now
return BedrockError, and a parametrized audit drives all 36 configs.

* fix(proxy): keep provider headers on the httpx status error branch

_handle_llm_api_exception forwards safe_headers on every branch except the
httpx.HTTPStatusError one, which the bedrock passthrough route reaches, so
the request id was dropped before the client saw the response.

* fix(bedrock): keep the request id on the timeout mappings

Timeout takes no response argument, so the three bedrock timeout branches
dropped the provider headers even when the upstream answered 408 or 504
with an x-amzn-RequestId. They now ride on the exception, already
llm_provider-prefixed, which is the form the proxy emits.

* fix(bedrock): keep the provider response on mapped timeouts

The previous round attached llm_provider-prefixed headers directly to the
Timeout. That shadowed the raw upstream headers for _get_response_headers,
so router cooldown and fallback cooldown stopped honouring retry-after on
bedrock 408/504 replies.

Give Timeout an optional response instead, the way every other mapped
bedrock exception already carries one. Retry logic reads the raw
retry-after off the response, and the proxy prefixes those headers on the
way out, so clients still see llm_provider-x-amzn-requestid.

* chore(bedrock): drop the explanatory comment on Timeout.response
2026-09-07 17:16:47 -07:00
mateo-berri
31bd00dad9 fix(budget_reservation): merge staging and cover count_tokens routes in the mapped test 2026-09-07 17:11:33 -07:00
mateo-berri
4cc0180eab fix(router): keep unresolved drop_params strings so DB rows and env refs survive
The drop_params validator collapsed every string it did not recognize to None. A pre-fix DB row holds the flag as ciphertext, so a partial PATCH rebuilt the deployment without it and dropped the key from the stored row, and /model/new turned an os.environ/ reference into nothing before the loader could resolve it. The validator now returns the raw value when it is not a boolean flag, the field admits strings the way timeout already does, and the flag set follows pydantic's lax bool parsing instead of a hand-rolled true/false pair
2026-09-07 17:06:34 -07:00
Yuneng Jiang
7d3b68fea5
fix(files): preserve managed deletion routing and response identity 2026-09-07 16:59:53 -07:00
mateo-berri
0710231acc feat(cost_map): stamp and surface generated_at and source revision provenance
The cost map JSON now carries a top-level `_metadata` block with `generated_at` and `source_revision`, written by the two bot writers only when model data changed. The loader pops it before the map becomes `litellm.model_cost`, records it next to the fetch ETag, and `/reload/model_cost_map`, `/model/cost_map/source`, and the reload schedule status return it. The Price Data Reload card shows the stamp, the ETag, and when the pod loaded the map. The schema and the cost map guard treat `_metadata` as a non-model root key
2026-09-07 16:57:09 -07:00
ryan-crabbe-berri
43a02e2dbc fix(proxy): report the billed token rates in /cost/estimate
The rate fields reported base cost-map prices while the cost lines were billed at the token tier, off-peak window and regional multipliers the calculator picks for the request, so a line did not always equal tokens times its reported rate. get_billed_token_rates now resolves the rates once, the token-type breakdown and the endpoint both read from it, and a tiered-model test asserts every line equals its token count times the rate reported next to it

Claude-Session: https://claude.ai/code/session_011Tn3657NkV6ojLqewL64Kb
2026-09-07 16:53:41 -07:00
mateo-berri
5c23d296d3 fix(azure_ai): cost OCR relays per page so non-chat relays debit budgets 2026-09-07 16:51:56 -07:00
mateo-berri
614b151365 chore: merge origin/litellm_internal_staging into litellm_lit_7022_azure_ai_passthrough_config 2026-09-07 16:43:12 -07:00
mateo-berri
e01bb98960 merge: bring litellm_internal_staging into litellm_fix_agent_mcp_grants again
Staging moved by the auto-router classifier cost change (#40168) between the
first merge and its push; this merge picks it up so the PR merges cleanly
2026-09-07 16:36:55 -07:00
Yuneng Jiang
4ab5719ff9
test(batches): use immutable expectations with explicit test doubles 2026-09-07 16:36:27 -07:00
mateo-berri
5e4dec4b88 merge: bring litellm_internal_staging into litellm_fix_agent_mcp_grants
Take staging's test_bedrock_knowledgebase_hook.py, which drops the duplicate
embedding_executor parameter that turned the lint check red, and make the two
cross-module helpers this branch added public (raise_denied_scoped_mcp_access
and routes_through_gateway) so the private-usage budget stays at its base count
2026-09-07 16:35:40 -07:00
tin-berri
9d0c9b9382
feat(ui): itemize auto-router classification spend (#40168)
Resolves LIT-7141

Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-07 16:29:19 -07:00
Yuneng Jiang
7bff9bf9a2
fix(batches): handle provider cancellation and file cleanup gaps 2026-09-07 16:22:32 -07:00
mateo-berri
b827375e60 Merge branch 'litellm_internal_staging' into litellm_lit_4116_drop_params_string_coerce 2026-09-07 16:22:26 -07:00
mateo-berri
1a9611d1a0 Merge branch 'litellm_internal_staging' into litellm_lit_4116_drop_params_string_coerce
Resolve the conflicts in utils.py, types/router.py, and the tests, and collapse the 56 per-provider isinstance(drop_params, bool) gates to bool(drop_params) now that get_optional_params normalizes the flag once at the top
2026-09-07 16:20:36 -07:00
Mateo Wang
bb8fe4a32f
Merge pull request #39826 from BerriAI/litellm_lit_6348_fireworks_responses_api
feat(fireworks_ai): add native Responses API config
2026-09-07 16:15:58 -07:00
tin-berri
cd681a573f
fix(mcp): encrypt stored static headers and stdio environment (#40164)
Encrypt secret maps at the shared persistence boundary, preserve plaintext API/runtime views, and extend rotation and migration scanning to legacy rows.

Co-authored-by: Claude Code <noreply@anthropic.com>
2026-09-07 16:03:06 -07:00
ryan-crabbe-berri
c84131b81a feat(proxy): price cache and reasoning tokens in /cost/estimate
POST /cost/estimate now accepts cache_read_input_tokens,
cache_creation_input_tokens and reasoning_tokens, bills them at the
model's cache and reasoning rates, and reports each share per request,
per day and per month next to the rates it used.

Custom-priced deployments also get cache and reasoning lines in the cost
breakdown now, so the estimate and the spend logs reconcile with their
totals instead of showing zero for those tokens.

Requested by a customer (Pylon #7365).

Claude-Session: https://claude.ai/code/session_011Tn3657NkV6ojLqewL64Kb
2026-09-07 16:00:20 -07:00
yuneng-jiang
8af624b5f5
Merge pull request #40162 from BerriAI/litellm_default_branch_tooling
ci: follow the default branch in development tooling
2026-09-07 15:56:30 -07:00
mateo-berri
18aa52d5e1 chore: merge litellm_internal_staging into litellm_lit_6348_fireworks_responses_api 2026-09-07 15:43:08 -07:00
yuneng-jiang
1646901170
Merge pull request #40041 from BerriAI/litellm_presidio_ui_user_story_e2e
test(e2e/ui): automate the RC checklist's Presidio guardrail walk
2026-09-07 15:41:33 -07:00
yuneng-jiang
41e4f1a8c8
Merge pull request #40038 from BerriAI/litellm_/guardrail-automation-testing-3ecd3d
test(guardrails): pin the presidio spend-log record and the UI's masked-entity persistence
2026-09-07 15:39:51 -07:00
Yuneng Jiang
15ba07f20d
ci: avoid duplicate default branch fetches 2026-09-07 15:28:01 -07:00
ryan-crabbe-berri
38683643e0
Merge pull request #40170 from BerriAI/litellm_chatgpt_add_model_provider
feat(ui): list the ChatGPT subscription provider in the Add Model form
2026-09-07 15:21:58 -07:00
Mateo Wang
058d260509
Merge pull request #38914 from BerriAI/litellm_fix_skills_hook_import_side_effect
fix(proxy): register SkillsInjectionHook at proxy startup instead of import time
2026-09-07 15:18:54 -07:00
Mateo Wang
0e9e2c01f3
Merge pull request #38806 from BerriAI/litellm_fix_mcp_test_connection_oauth_bearer
fix(mcp): forward staged credentials on /mcp-rest/test/connection like /test/tools/list
2026-09-07 15:18:35 -07:00