Commit graph

1773 commits

Author SHA1 Message Date
yuneng-jiang
418c7c6012
Merge pull request #37721 from BerriAI/litellm_internal_staging
Some checks failed
Unit Tests / misc (push) Has been cancelled
Unit Tests / proxy-auth (push) Has been cancelled
Unit Tests / proxy-endpoints (push) Has been cancelled
Unit Tests / proxy-infra (push) Has been cancelled
Unit Tests / proxy-server (push) Has been cancelled
Unit Tests / responses-caching-types (push) Has been cancelled
GitHub Actions Security Analysis / zizmor (push) Has been cancelled
CI Coverage / assert-ci-coverage (push) Has been cancelled
CodeQL / Analyze (actions) (push) Has been cancelled
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (python) (push) Has been cancelled
CodSpeed Benchmarks / benchmarks (push) Has been cancelled
Helm unit test / unit-test (push) Has been cancelled
Scorecard supply-chain security / Scorecard analysis (push) Has been cancelled
Code Quality Checks / code-quality (push) Has been cancelled
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
Unit Tests: Documentation Validation / documentation (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
chore(ci): promote internal staging to main
2026-08-20 19:23:38 -07:00
Mateo Wang
66a6a09706
Merge pull request #37743 from BerriAI/litellm_cognition_provider_identity
feat(cognition): give Cognition its own provider identity
2026-08-20 18:38:23 -07:00
mateo-berri
ef1cde433e fix: add moonshot/kimi-k3 to the cost map
models.litellm.ai and released litellm versions read
model_prices_and_context_window.json from main at runtime, so Kimi K3 is
missing from the hosted catalog even though the entry is in review for
litellm_internal_staging in #37552. This copies that entry onto main so
the catalog picks it up on its next fetch.

Data only: the cost map and its backup copy, no code changes. Pricing
matches Moonshot's published rates ($3/M input, $0.30/M cache read,
$15/M output, 1,048,576-token context). The fireworks_ai and Azure
Foundry kimi-k3 variants are separate work in #37512 and #37658; neither
touches the native moonshot/kimi-k3 key.
2026-08-20 18:22:14 -07:00
mateo-berri
1a9e9951e1 fix(cognition): restore Lightning SWE pricing and declare responses support
The swe-1.7 rates were briefly lowered to the standard tier. The docs page
records the API-served swe-1.7 as the Cerebras-served Lightning tier, so put
the matching rates back rather than have the cost map and the docs disagree.

Cognition also answers /v1/responses through the chat-completions bridge, the
same as every other provider in the JSON registry, so the endpoints support
matrix should say so instead of under-declaring it.
2026-08-20 17:39:29 -07:00
Mateo Wang
211b399761
Merge pull request #37658 from BerriAI/litellm_model_registry_consolidated_20260820
fix(model_prices): consolidate eleven open registry audits into one changeset
2026-08-20 17:39:19 -07:00
mateo-berri
332a0f1b9b fix(cognition): price swe-1.7 from the published standard tier
The swe-1.7 rates were carried over from the closed prior attempt and
match SWE-1.7 Lightning, 5x the SWE-1.7 Max and Medium rates the vendor
publishes. swe-1.6 was already on the standard tier, so the two entries
disagreed with each other. Both now read 0.5 in, 2.5 out, 0.2 cached per
million tokens.

Also drops the redundant registry comment in constants.py.
2026-08-20 17:33:59 -07:00
Mateo Wang
286c75f69d
Merge pull request #37729 from BerriAI/devin_ai_fal_gpt_image_2
feat(fal_ai): add gpt-image-2 image generation support
2026-08-20 17:23:24 -07:00
Mateo Wang
fb3dd0fb98
Merge pull request #37722 from BerriAI/litellm_gpt56_max_input_tokens
fix(model-costs): correct gpt-5.6 max input tokens to 922k
2026-08-20 17:21:09 -07:00
mateo-berri
e00301703f feat(cognition): give Cognition its own provider identity
Cognition serves an OpenAI-compatible /v1/chat/completions endpoint, so it has been onboarded as
custom_llm_provider: openai. That books its traffic as OpenAI, which means OpenAI-specific cost
discounts and provider-level reporting apply to it.

Registers cognition through the JSON provider registry: a providers.json entry with
COGNITION_API_KEY and COGNITION_API_BASE, LlmProviders.COGNITION, the constants.py provider lists,
cost map entries for swe-1.6 and swe-1.7, the provider endpoints matrix, the dashboard provider
fields, and tests. JSON providers can now also be resolved from their base url alone, so an
api_base pointing at a known provider no longer falls through to an unresolved provider.
2026-08-20 17:16:53 -07:00
mateo-berri
6d66567915 fix(fal_ai): stop advertising /v1/images/edits for gpt-image-2 edit
The edit model is reached through the image generation path with fal's
image_urls param; /v1/images/edits is not wired for fal_ai and errors.
Point supported_endpoints at /v1/images/generations and say so in the
entry notes.
2026-08-20 17:02:07 -07:00
mateo-berri
5c9ed89301 fix: mark daybreak-blue-latest and gpt-5.6-sol as supporting computer use
OpenAI documents computer_use as a supported tool for Daybreak Blue and its
default snapshot gpt-5.6-sol, but neither entry carried supports_computer_use.
Sibling gpt-5.6-cyber and daybreak-red-latest already set it, so /model/info
and the capability gates reported blue as unable to use computer tools.

The gap came in with the source PR rather than the consolidation: #37029 sets
the flag on cyber and red only. Pinned by a new metadata test covering the
daybreak family and the blue alias agreeing with its snapshot.
2026-08-20 16:20:40 -07:00
Devin AI
618d907d5a fix(fal_ai): price gpt-image-2 unprefixed alias and edit endpoint
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-20 23:04:37 +00:00
Devin AI
60706d5f89 feat(fal_ai): add gpt-image-2 image generation support
Route fal.ai's openai/gpt-image-2 endpoints through a dedicated transformation that maps OpenAI image params (n, size, quality, output_format) into fal's schema, and register the model in the cost map.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-20 22:18:34 +00:00
Mateo Wang
d556fac56b
Merge pull request #37112 from mubashir1osmani/litellm_add_perplexity_agent_api_models
feat(perplexity): add Agent API third-party models
2026-08-20 14:49:12 -07:00
mateo
6bb677d30f fix(model-costs): correct gpt-5.6 input token cap
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-20 21:44:32 +00:00
mateo-berri
2996a18fa9 fix(model_prices): document 262k input limit on fireworks qwen3p8-max 2026-08-20 14:33:04 -07:00
mateo-berri
50a346da1c fix(model_prices): restore supports_vision on Mistral Small 4.0 entries 2026-08-20 13:54:03 -07:00
mateo-berri
0bdaa98b28 fix(perplexity): correct glm-5.2 cache read rate to the published catalog rate
The new perplexity/perplexity/glm-5.2 row carried 2.6e-07, which is glm-5.3's
cache read rate. api.perplexity.ai/v1/models publishes 0.14 usd per 1M cached
input tokens for glm-5.2, so the rate is 1.4e-07.
2026-08-20 12:48:41 -07:00
Devin AI
ba4a355afc fix(model_prices): add tpm/rpm to gemini live-translate preview entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-20 19:27:44 +00:00
Devin AI
bdd4c8e564 fix(model_prices): add Gemini live-translate, Voyage 4 series, Perplexity contextualized embeddings; absorb Fireworks + Bedrock batch registry PRs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-20 19:14:11 +00:00
Devin AI
a95a1d0232 fix(model_prices): drop duplicate zai-glm-5-2 entry superseded by staging
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-20 18:35:52 +00:00
Devin AI
b469294029 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_model_registry_consolidated_20260820 2026-08-20 18:35:08 +00:00
mateo-berri
059aec8887 Merge remote-tracking branch 'origin/litellm_internal_staging' into pr37112-head 2026-08-20 11:18:30 -07:00
Devin AI
b7017a7949 fix(model_prices): consolidate nine open registry audits into one changeset
Combines the model-cost-map data from #35911, #36017, #36080, #36113, #36188, #36444, #37029, #37252 and #37632 onto current litellm_internal_staging, merged per entry field so older branches no longer revert fields the base has gained since they were opened. Drops the Gemini deprecation dates from #36188 and the text-embedding-004 date from #36080 that the official docs contradict.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-20 18:13:39 +00:00
mateo-berri
10829fff04 Merge remote-tracking branch 'origin/litellm_internal_staging' into pr-37110-check 2026-08-20 11:00:02 -07:00
mateo-berri
4fac88790d fix(mistral): correct zai-glm-5-2 limits, add cached-input price and glm-5-2 alias
Mistral's live /v1/models reports max_context_length 1048576 and capabilities.reasoning
true for zai-glm-5-2, and its docs price cached input at $0.14/M. Without
cache_read_input_token_cost LiteLLM billed every cached prompt token at $0, so a repeat
request against a 21k-token cached prefix logged $0.0000135 instead of its real cost.

Mistral also serves the model under the short glm-5-2 name, which had no cost map entry
at all and therefore no pricing, so add it alongside.
2026-08-20 10:45:24 -07:00
Mateo Wang
e51addb802
Merge pull request #37628 from BerriAI/litellm_lit5876_openai_prompt_cache_breakpoint
feat(prompt-caching): map cache_control_injection_points to OpenAI prompt_cache_breakpoint on GPT-5.6+ targets
2026-08-20 10:24:43 -07:00
mateo-berri
d9aaa95978 Gate OpenAI prompt cache breakpoints on the real target and carry them through /v1/responses
The cache control hook also runs on litellm.responses() input. On a
GPT-5.6 deployment it wrapped a string-content item into a chat-shaped
{"type": "text"} part, which the Responses API rejects, and it never
marked input_text, input_image or input_file parts, so no breakpoint and
no prompt_cache_options reached the provider. Add the Responses part
types to the eligible block set and translate chat-shaped text parts on
non-assistant items to input_text in
ResponsesAPIRequestUtils.merge_prompt_management_input, which both the
async and the sync prompt management sites go through.

The dialect also fired for any GPT-5.6 name that resolved to provider
openai, including deployments pointed at a custom api_base that does not
understand prompt_cache_breakpoint. Decide it once per request from the
provider, the model map and the resolved api_base (request, then
litellm.api_base, then OPENAI_BASE_URL / OPENAI_API_BASE): only
api.openai.com and *.api.openai.com hosts speak the dialect, a top-level
prompt_cache_options opts a custom target in, and litellm_proxy/ targets
never get it. maybe_seed_default_injection_points takes api_base and
stamps the finished decision on the points as _litellm_openai_dialect so
the sync completion() path, whose hook params do not carry api_base,
honors it; maybe_inject_cache_control takes api_base from the
/v1/messages handler.

Eligibility now comes from a supports_prompt_cache_breakpoint model map
flag on the OpenAI gpt-5.6 entries, exposed through
litellm.utils.supports_prompt_cache_breakpoint, with the GPT version rule
kept only for models the map does not know. The OpenAI dialect no longer
reserves a slot for tool_config points, which OpenAI has no cache block
for, and with_prompt_cache_breakpoint plus the chat bridge helper return
a new block instead of mutating their input.
2026-08-20 05:17:13 -07:00
Mateo Wang
952c6d5675
Merge pull request #37517 from BerriAI/devin_ai_bedrock_grok_4_6_cost_map
feat: add bedrock grok 4.6 to model cost map
2026-08-20 03:03:44 -07:00
Mateo Wang
6fcdea03b0
Merge pull request #36969 from oneKn8/fix-cost-map-mid-conversation-flag
fix: add supports_mid_conversation_system to bare first-party Claude cost-map keys
2026-08-20 02:04:56 -07:00
Mateo Wang
a6163e0146
Merge pull request #37543 from BerriAI/litellm_lit_5785_vertex_regional_pricing
fix(vertex_ai): apply regional endpoint uplift to cost tracking
2026-08-19 17:56:34 -07:00
Mateo Wang
634e699555
Merge pull request #36331 from BerriAI/devin_ai_agentcore_search
feat(search): add Amazon Bedrock AgentCore web search provider
2026-08-19 15:59:17 -07:00
mateo-berri
a870d45a8a Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_lit_5785_vertex_regional_pricing
# Conflicts:
#	litellm/model_prices_and_context_window_backup.json
#	model_prices_and_context_window.json
2026-08-19 15:31:42 -07:00
mateo-berri
b39a339b7d fix(vertex_ai): apply regional endpoint uplift to cost tracking 2026-08-19 15:21:06 -07:00
Mateo Wang
d58b1c8558
Merge pull request #37516 from BerriAI/litellm_gemini_prompt_cache_min_tokens_4096
fix(model_prices): set prompt_cache_min_tokens=4096 for Gemini 3.5/3.6/3.7 Flash and 3.1 Pro Preview
2026-08-19 15:11:00 -07:00
mateo-berri
45884b9bd3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_gemini_prompt_cache_min_tokens_4096
# Conflicts:
#	litellm/model_prices_and_context_window_backup.json
#	model_prices_and_context_window.json
2026-08-19 14:55:44 -07:00
Mateo Wang
b8d5139701
Merge pull request #37473 from BerriAI/litellm_model_registry_audit_20260819
fix(model_prices): correct gemini 3.1 flash image and deepseek v4 pricing, add openai deprecation dates
2026-08-19 14:50:49 -07:00
Mateo Wang
f4b46c81da
Merge pull request #37283 from BerriAI/devin/1787058723-registry-deprecation-dates
fix(model_prices): add provider-announced deprecation_date to 205 registry entries
2026-08-19 14:43:45 -07:00
Devin AI
1f6bef79ca fix: drop source url from grok 4.6 cost map entries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-19 21:24:29 +00:00
Devin AI
087cdcff07 feat: add bedrock grok 4.6 to model cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-19 21:18:12 +00:00
mateo-berri
77716eeaed fix(model_prices): set prompt_cache_min_tokens=4096 for Gemini 3.5/3.6/3.7 Flash and 3.1 Pro Preview 2026-08-19 14:13:00 -07:00
Devin AI
5a0a8ffafe fix(model_prices): correct gemini and deepseek pricing and add deprecation dates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-19 13:17:15 +00:00
mateo-berri
24fc3f721c fix(databricks): match the claude 4.6 context limits to Anthropic's published values 2026-08-18 17:55:45 -07:00
mateo-berri
4285ffd82b fix(databricks): drop the minimal reasoning effort flag from the claude-opus-4-6 entry 2026-08-18 17:09:17 -07:00
Mateo Wang
f570af9fcf
Merge branch 'litellm_internal_staging' into add-databricks-model-pricing 2026-08-18 16:55:02 -07:00
yassin
cf2e50077c Merge branch 'litellm_internal_staging' into devin_ai_agentcore_search
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-18 21:53:22 +00:00
mateo-berri
be594f5984 feat(guardrails): count bedrock guardrail cost against spend and budgets
Price ApplyGuardrail usage units recorded by PR #37225 with a new
bedrock/guardrails entry in the model cost map (regional override via
bedrock/{region}/guardrails), add the per-request guardrail_cost to the
standard logging payload's response_cost and CostBreakdown, surface it in
the x-litellm-response-cost header, and bill blocked requests through the
failure hook so key and team budgets see what AWS bills
2026-08-18 14:16:07 -07:00
Devin AI
3d523d6d81 fix(model_prices): add provider-announced deprecation_date to 205 registry entries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-18 13:14:42 +00:00
mateo
94a29e0708 fix(gemini): price gemini 3.6 flash at Google's introductory rates on every service tier
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-17 18:42:57 +00:00
mubashir1osmani
539a61be08 feat(perplexity): add Agent API third-party models (DeepSeek V4 Flash, GLM 5.2, Kimi K3, Kimi K2.7 Code) 2026-08-16 14:45:22 -04:00