Commit graph

2557 commits

Author SHA1 Message Date
devin-ai-integration[bot]
671748067d
fix(model_prices): registry audit 2026-10-03, add cohere embed v5, grok-imagine-video-1.5-lite and openrouter gpt-image rows (#44376)
* fix(model_prices): registry audit 2026-10-03, add cohere embed v5 and absorb openrouter gpt-image rows

Co-authored-by: tinysolver <iam.tinysolver@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): add grok-imagine-video-1.5-lite and gemini 3.5 transcribe limits, drop deprecation-only date

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: tinysolver <iam.tinysolver@gmail.com>
2026-10-03 12:53:07 -07:00
berriai-litellm-provider-info-sync[bot]
9fe6442172
fix(azure): add azure_ai/flux.2-pro input token limit from the models sold directly page (#44405)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 11:40:44 -07:00
devin-ai-integration[bot]
6c32384d8c
fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3 (#44292)
* fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3

Bedrock prices Kimi K3 cache reads through implicit caching but Converse
rejects the explicit cachePoint marker ("This model doesn't support the
cachePoint field"), so any cache_control on the request answered 400.
Mark the three K3 rows supports_prompt_cache_breakpoint: false and have
bedrock_model_accepts_cache_points honor that flag before falling back to
supports_prompt_caching, keeping cached-token pricing intact.

* test(bedrock): assert cache points per request section

* fix(bedrock): honor a deployment's cache breakpoint flag for unmapped models

* fix(bedrock): read a converse-routed deployment's cache breakpoint flag

* refactor(bedrock): look up cache breakpoint flags by key

* test(bedrock): add the Kimi K3 cache point wire audit

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 18:11:46 +00:00
devin-ai-integration[bot]
8b1990b4bc
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register typesafe as a provider so Jev deployments load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move provider endpoints under llms and validate proxy bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add Cloudflare Clef and Strands Decider backends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register decisions routes for managed agents and gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(decisions): use raw regex for cloudflare missing account match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): avoid cast in Cloudflare response unwrapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): default model, evaluation health probe, short Cloudflare names

The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.

* fix(decisions): let health_check_params override the evaluation probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit the decisions endpoint across providers, limits, health and chaos

Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).

The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.

* fix(decisions): send env API keys to a configured api_base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register the routes through the lazy feature registry

The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.

The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.

* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell

* fix(proxy): let a config pass-through beat a lazily registered route in eager mode

With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:38:38 +00:00
berriai-litellm-provider-info-sync[bot]
231a46e40b
feat(azure): add azure_ai/kimi-k2-thinking from Azure Kimi pricing page (#44382)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 10:04:59 -07:00
berriai-litellm-provider-info-sync[bot]
66a422ea50
fix(azure): add MAI-Image max_input_tokens from models sold directly page (#44375)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 08:49:09 -07:00
devin-ai-integration[bot]
0fce5bccbc
fix(model-prices): mark azure us/eu responses-only models as mode responses (#44323)
* fix(model-prices): mark azure us/eu responses-only models as mode responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): keep azure us/eu o3-deep-research on chat, which Azure lists as chat capable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 23:05:55 -07:00
devin-ai-integration[bot]
233db9f285
feat(bedrock): serve gpt-5.6+ chat completions natively by default, with chat_completions/ opt-in for gpt-oss and grok (#44307)
* feat(bedrock): send grok chat completions through runtime openai path

Unspecified bedrock grok was rewritten to Converse. Chat completions now hit bedrock-runtime /openai/v1/chat/completions, and converse/ still uses Converse

* feat(bedrock): serve gpt-oss and gpt-5.6 chat completions on runtime's native openai path

* fix(bedrock): route gpt-oss response_format to Converse and decide the route once from the raw request

* fix(bedrock): serve region-path and GovCloud gpt-oss ids on native Chat Completions

The cost-map parity tests require every regional variant of a flagged id to carry the same supports_ flags, so the six us-gov gpt-oss entries now carry the native-route flags too. A region path in the model name (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model: the route is looked up on the id after the path, the path's region picks the endpoint and the SigV4 scope, an explicit aws_region_name still wins, and the body carries the bare id AWS expects

* fix(bedrock): keep params AWS refuses natively off the chat completions route

Drop the params each family 400s or 503s on runtime Chat Completions from the native config's supported list (GPT-5.6 penalties, stop, and logprobs, Grok penalties, gpt-oss logit_bias) so drop_params drops them as Converse did, gate legacy functions on GPT-5.6 the same way as tools, and send an Anthropic-style thinking block to Converse, the only route that forwards it

* fix(bedrock): keep schema-less json_object on Converse for the chat completions models

* fix(bedrock): keep every json_object response_format on Converse for the chat completions models

* fix(rust): declare the bedrock runtime chat completions flags on ModelInfo

* fix(bedrock): opt into the native chat completions route through supported_endpoints

* docs(cost-map): describe the bedrock native chat completions capability flags

* revert: docs(cost-map): describe the bedrock native chat completions capability flags

This reverts commit 4101c0ceb2.

cost-map-guard runs main's schema generator under pull_request_target and compares
its output to the PR's committed schema, so a PR that changes the generator's output
cannot pass that required check until the generator change lands on main first. The
descriptions move to a follow-up that lands the generator change ahead of the schema

* test(bedrock): move the native chat completions tests under tests/unit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): drop reasoning_effort none for grok on the native chat completions route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): keep converse extension params on the converse route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): inline http image urls and keep stop on converse for native chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(bedrock): share the sync remote media inliner

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(image-handling): infer the image mime type when the server sends a generic content type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): stop sending aws_bedrock_project_id as OpenAI-Project on the runtime chat completions route

* feat(bedrock): make native chat completions an opt-in bedrock/chat_completions/ route

Bare Bedrock OpenAI and Grok model ids stay on Converse as on main. The
bedrock/chat_completions/<model> prefix opts a deployment into bedrock-runtime's
/openai/v1/chat/completions, and a request carrying a Converse-only param
still falls back to Converse. The cost map no longer decides the route.

* fix(bedrock): keep chat_completions/-prefixed deployments on the native Responses surface

* fix(bedrock): keep provider response headers on the runtime chat completions route

* feat(bedrock): serve gpt-5.6 and newer on runtime chat completions by default

Unprefixed bedrock/<gpt-5.6+> models whose cost-map row lists /v1/chat/completions
now route to the native OpenAI-compatible endpoint; converse/ pins Converse and
chat_completions/ still opts gpt-oss and Grok in. Guardrails, application inference
profile ARNs, and tools with reasoning keep falling back to Converse per request.
Hoist the remote-media url comprehension into a single-clause helper.

* fix(bedrock): refuse temperature and top_p natively on GPT 5.6 and newer like Converse does

AWS answers temperature and top_p with a 400 on the native Chat Completions endpoint for the GPT 5.6+ models, the same models whose Converse route already dropped both under drop_params via supports_sampling_params: false. The native config now honors that price-map flag, the gpt-6 and gpt-6.1 rows carry it, and the gpt-6 family joins gpt-5 in refusing frequency_penalty, presence_penalty, logprobs, and top_logprobs before the request reaches AWS.

* fix(bedrock): refuse GPT sampling and logprob params natively only while reasoning is on

On bedrock-runtime's native chat completions endpoint, GPT-5.x and GPT-6.x
accept temperature, top_p, frequency_penalty, presence_penalty, logprobs,
and top_logprobs once reasoning_effort is "none", and refuse them with any
other effort or when the effort is unset. The previous commit refused the
sampling params unconditionally from the cost map's supports_sampling_params
flag, which lost the reasoning-off case and never covered the penalties or
logprobs. The refusal now keys on the model being a GPT id and reasoning
being active, raises a 400 UnsupportedParamsError naming the params unless
drop_params drops them, and lets everything through under "none". Grok and
gpt-oss keep their unconditional family refusals.

* refactor(bedrock): keep the Converse route-prefix strip inside the bedrock llms module

* fix(bedrock): forward a non-string reasoning_effort on the native route instead of crashing

A list or dict reasoning_effort hit a frozenset membership test in
without_refused_reasoning_effort and raised TypeError, which the proxy
surfaced as a 500 APIConnectionError with no upstream call. The value is
now left alone unless it is a string Bedrock's native endpoint refuses,
so AWS answers the malformed value with its own 400 like it does for an int

* fix(bedrock): route overlong GPT version digits to Converse and send native chat completions to the runtime endpoint

A model id with more than 4300 version digits raised ValueError in the route check; the digits are now bounded so such ids fall back to Converse. The native chat completions URL now follows Converse's precedence: aws_bedrock_runtime_endpoint (or AWS_BEDROCK_RUNTIME_ENDPOINT) wins over api_base, so a deployment that sets both keeps sending to the same host

* fix(bedrock): route model_id overrides to Converse and never send an empty bearer natively

A deployment whose litellm_params carry model_id (an application inference profile or provisioned throughput ARN) went to the native Chat Completions route with the base model in the URL and model_id left in the body. It now takes Converse like the bedrock/arn:... model form, which encodes the override into the request URL

A blank api_key on a SigV4 deployment became an Authorization header reading Bearer with nothing after it on the native route, since the OpenAI-like header builder writes any non-None key and the signer keeps a non-AWS4 Authorization header. validate_environment now resolves the key through bedrock_bearer_token, so a blank key is signed with SigV4 the way Converse signs it

* test(bedrock): audit the native GPT chat completions route on the integration rig

Adds the /audit cells for the runtime chat completions route: the scripted Bedrock runtime peer, the happy and fallback wire tests, the sad-path and regex worst-case tests, the chaos burst tests, the Messages adapter tests, and the Responses native-route tests. Tests only, no product diff.

* test(bedrock): harden the runtime chat completions audit cells

The chaos peer's shared counter and process now come from the same spawn context, since a fork-context Value handed to a spawn-context process raises on Linux. The peer-kill test waits for the first six answers to reach the client before killing the peer instead of counting accepted requests. The Responses wire tests look the spend row up under both the ciphertext id the caller received and the issued id behind it, matching the chaos file's rule for the pre-encryption row

* fix(bedrock): refuse or drop a non-string reasoning_effort before the native chat completions call

A reasoning_effort sent as an int, a list, or an object on a GPT 5.6+ deployment the native
route serves now answers 400 from litellm before any wire request, naming the type and the
drop_params way out, and is dropped under drop_params so AWS applies its default effort, the
way Converse dropped it on main. The tip since a0cef91f0b forwarded it for AWS to refuse

* test(bedrock): pin router retries off and give the chaos bursts config deployments on an owned proxy

* test(bedrock): wait for the replacement worker before tearing down the sigkill chaos proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo <mateo@berri.ai>
2026-10-02 19:55:54 -07:00
devin-ai-integration[bot]
5dbe4f95e8
feat(bedrock): drop lookaround regex patterns from tool schemas for Converse models that reject them (#44138)
* feat(bedrock): drop lookaround regex patterns from tool schemas for Converse models that reject them

* fix(bedrock): rename the lookaround flag to supports_regex_lookaround and keep dropped patternProperties names allowed

The cost-map flag becomes a generic supports_regex_lookaround capability, which the
cost-map schema admits as a supports_* boolean, and the Converse transform now owns
the drop decision instead of the shared tools factory. A patternProperties key dropped
from an object closed by additionalProperties: false leaves its value schema as that
object's additionalProperties, so the names it allowed stay allowed, on the OpenAI
non-Python-regex drop too. tool_with_sanitized_parameters also cleans Anthropic-shape
tools (input_schema).

* fix(router): keep a deployment's supports_regex_lookaround off the shared cost-map entry

A deployment's model_info.supports_regex_lookaround was written to the shared
bedrock/<model> cost-map key, so every sibling deployment of that model id
inherited one deployment's choice. The flag now stays under the deployment's
own id, which is what the Converse lookaround check reads first, and the
shared entry keeps the cost map's value

* test(bedrock): audit the Converse lookaround drop on the wire across endpoints, SDKs, flags and chaos

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 19:47:25 -07:00
tin-berri
f63d989ff9
feat: add Bespoke Nimble gateway and OSS classifier support (#44246)
* feat: add Bespoke Nimble gateway and OSS classifier support

* feat: accept Ollama's nimble model name for the Bespoke provider

* test: exempt the POST-only bespoke decisions route from the all-methods check

test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 19:37:08 -07:00
berriai-litellm-provider-info-sync[bot]
c5c0a48ef1
fix(bedrock): price amazon nova 2 pro preview at the standard tier (#44302)
Some checks are pending
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-infra-root (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-02 19:05:52 -07:00
berriai-litellm-provider-info-sync[bot]
9e31afbc6d
chore(openrouter): sync prices, limits and deprecation dates from the models API (#44287)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-02 17:31:44 -07:00
berriai-litellm-provider-info-sync[bot]
ac59ec6824
chore(prices): add xAI grok-voice-transcribe-1.0 deprecation date (#44237)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-02 13:30:09 -07:00
tin-berri
0dc23406eb
feat: add Laya gateway and OSS classifier providers (#43626)
* feat: add Laya gateway and classifier backend

* test: cover the pass-through model_group pin and repair the shard fakes

MockRequest in tests/pass_through_unit_tests gains an httpx.URL and an ASGI
scope, which get_request_route now reads inside
_init_kwargs_for_pass_through_endpoint, and the POST-only /laya/v1/systemone
route joins the protocol-constrained exemptions. A built-in pass-through pins
metadata.model_group to the resolved model so a client cannot choose its own
per-model budget key; test_pass_through_endpoints now proves that on a
non-Laya route and drops a duplicated assertion.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 11:29:36 -07:00
devin-ai-integration[bot]
a5fef4b4e6
fix(cost-map): add OpenAI TTS and GPT-5.x deprecation dates (#44175)
* fix(cost-map): add OpenAI TTS and GPT-5.x deprecation dates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(cost-map): credit OpenAI TTS deprecation dates from #44113

Co-authored-by: Ben Langfeld <210221+benlangfeld@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ben Langfeld <210221+benlangfeld@users.noreply.github.com>
2026-10-02 09:26:19 -07:00
berriai-litellm-provider-info-sync[bot]
58763b3021
chore(cost-map): take azure_ai claude-sonnet-4-5 retirement date from the Azure schedule (#44145) 2026-10-01 23:21:20 -07:00
berriai-litellm-provider-info-sync[bot]
086d76ab54
chore(cost-map): add azure_ai deprecation dates from the Azure retired models page (#44142)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-01 20:36:56 -07:00
berriai-litellm-provider-info-sync[bot]
6a8e0a270a
chore(cost-map): sync openrouter prices from the models API (#44105)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-01 17:24:51 -07:00
berriai-litellm-provider-info-sync[bot]
0737e26463
fix(cost-map): restore later azure Models API retirement dates and date gpt-6.1-sol (#44072)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry-berri <kerry@berri.ai>
2026-10-01 17:16:03 -07:00
devin-ai-integration[bot]
ac8c5aa4b9
fix(cost-map): add perplexity, openrouter, voyage and nebius models and fix registry metadata (#43907)
* fix(cost-map): add nebius qwen3.8-27b and correct nebius context limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): add and correct provider deprecation dates for deepseek, gemini and azure models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): add perplexity, openrouter and voyage models and correct gemini, nebius and perplexity metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 13:26:16 -07:00
berriai-litellm-provider-info-sync[bot]
ca05eca2d3
feat(vertex-ai): add vertex_ai/xai/grok-4.7 pricing (#44059)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-01 12:29:09 -07:00
berriai-litellm-provider-info-sync[bot]
3a11192f68
fix(cost-map): reprice fireworks deepseek v4.1 flash to the 2026-10-01 pricing update (#44024)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-01 08:16:30 -07:00
berriai-litellm-provider-info-sync[bot]
2b19ddb7a3
fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (#43916)
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 21:59:59 -07:00
berriai-litellm-provider-info-sync[bot]
ef6aa4ad66
chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (#43898)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 21:37:44 -07:00
berriai-litellm-provider-info-sync[bot]
9b8ddb0982
chore(cost-map): add fireworks inkling priority prices from the prices api (#43949)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 21:36:41 -07:00
berriai-litellm-provider-info-sync[bot]
8a1f3568ba
chore(cost-map): sync openrouter prices from the models API (#43950)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 17:48:48 -07:00
berriai-litellm-provider-info-sync[bot]
38b0762992
fix(wandb): set supports_vision true on GLM-5.3-Flash (#43951)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 17:35:36 -07:00
berriai-litellm-provider-info-sync[bot]
025292e75b
feat(pricing): add vertex_ai gemini-3.8 flash tts rows (#43876)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 10:20:57 -07:00
berriai-litellm-provider-info-sync[bot]
efdccd8811
chore(cost-map): add fireworks priority prices for ember-1, nemotron and glm 5.3 us rows (#43811)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 09:08:16 -07:00
berriai-litellm-provider-info-sync[bot]
971e60660b
chore(cost-map): add openai gpt-image-2.5 batch prices from the pricing page (#43869)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 09:06:31 -07:00
devin-ai-integration[bot]
b781d157d7
chore(model_prices): add Gemini Veo, Mistral and Azure Claude 4.5 deprecation dates (#43857) 2026-09-30 07:26:17 -07:00
berriai-litellm-provider-info-sync[bot]
cd0ac30881
fix(cost-map): add deprecation_date to two together_ai nvidia rows (#43809)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 21:24:04 -07:00
berriai-litellm-provider-info-sync[bot]
cae179e655
fix(bedrock): set gpt-6.1-sol max output tokens to 131072 (#43782)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 17:44:48 -07:00
berriai-litellm-provider-info-sync[bot]
e7460f1cff
chore(cost-map): take azure context limits from models-sold-directly (#43759)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 16:07:08 -07:00
berriai-litellm-provider-info-sync[bot]
56a63b4b29
feat(bedrock): add openai.gpt-6.1-sol us geo cris and Mantle rows (#43763)
Co-authored-by: kerry-berri <kerry@berri.ai>
2026-09-29 22:29:45 +00:00
berriai-litellm-provider-info-sync[bot]
27c110cb71
feat(bedrock): add openai gpt-6.1-sol global and base rows (#43758)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 13:00:01 -07:00
devin-ai-integration[bot]
1bfa3d4fa6
fix(model-prices): align Azure, Bedrock, Copilot, Gemini, Groq, OpenAI and OpenRouter entries with official docs (#43598)
* fix(model-prices): correct azure/eu/gpt-6-astra to Data Zone rates

Co-authored-by: rain <1504569896@qq.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): align groq, gemini and openai entries with official docs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): roll in verified Vertex, Gemini, OpenRouter and Azure AI registry fixes

Absorbs the fields from #43609, #43666, #43671 and #43644 that match the provider's own docs or price API today, and adds a cost test for the azure/eu/gpt-6-astra Data Zone tiers

Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): add Copilot, Bedrock Kimi K3, Gemini Robotics and OpenRouter values from official sources

Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: rain <1504569896@qq.com>
Co-authored-by: bunnysayzz <stfuazzo@gmail.com>
Co-authored-by: Michal Formanek <michal.formanek@generaliceska.cz>
2026-09-29 12:45:22 -07:00
berriai-litellm-provider-info-sync[bot]
f4a7c04d99
chore(cost-map): add openai gpt-6-astra ultrafast tier prices from the pricing page (#43745)
* chore(cost-map): add openai gpt-6-astra ultrafast tier prices from the pricing page

Price-Sync: litellm-providers

* feat(cost): support openai ultrafast tier fields in the model catalog

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: Kerry <kerry@berri.ai>
2026-09-29 19:05:57 +00:00
devin-ai-integration[bot]
abc85c2651
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing

Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins

Move Bedrock commitment rows to cost_per_second so they bill once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop legacy per-second fields from chat paths

Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost_calculator): recognize output-only per-second rates

Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:27:14 -07:00
berriai-litellm-provider-info-sync[bot]
0fe4028cd9
fix(cost-map): lower fireworks up-to-4b size tier to the pricing page price (#43740)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 11:14:14 -07:00
devin-ai-integration[bot]
bda2763f2c
chore(cost-map): add azure and openrouter gpt-6.1-sol rows (#43744)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 18:06:14 +00:00
berriai-litellm-provider-info-sync[bot]
a3552c451b
chore(cost-map): add openai gpt-6.1-sol from the pricing page (#43738)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 10:34:04 -07:00
berriai-litellm-provider-info-sync[bot]
d2cbc94fc6
feat(cost-map): add baseten DeepSeek-V4.1-Flash-Fast (#43735)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-29 10:09:11 -07:00
Flexomatic81
7b2cbf6e7f
fix(cost-map): add tool calling and reasoning flags, correct max output for nebius DeepSeek-V4.1-Flash (#43588) 2026-09-28 21:40:01 -07:00
devin-ai-integration[bot]
39d14bd855
feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers (#43641)
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): drive the router request test through an httpx MockTransport

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): assert tool definitions reach the router upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:08:16 -07:00
devin-ai-integration[bot]
b49661064f
feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5 (#43647)
* feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(cost-map): keep bedrock_mantle claude 5.5 change to cost map rows only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 01:44:13 +00:00
berriai-litellm-provider-info-sync[bot]
eae8ed7f3c
feat(bedrock): add xai grok-4.7 pricing and sync llama, mistral large 2407 and minimax m2.5 prices (#43623)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-28 16:31:15 -07:00
devin-ai-integration[bot]
5df502b360
feat(providers): add Prism provider (internal copy of #40914) (#41961)
* feat(providers): add Prism provider

* fix(providers): complete Prism registration

* feat(providers): expose Prism responses and messages

* feat(providers): add DeepSeek V4.1 Flash to Prism

* test(providers): exercise Prism endpoint requests

* fix(providers): align Prism pricing and limits with the live catalog

deepseek-v4.1-flash bills 0.17/0.63 USD per 1M input/output tokens and takes image input;
deepseek-v4-flash bills 0.17/0.21 and caps output at 384000 tokens, per GET /v1/models

* test(prism): assert cost-map invariants instead of pinning catalog facts

* test(prism): derive the asserted model list from the cost map instead of pinning it

* test(prism): capture requests through respx instead of appending to a list and swapping the client transport

---------

Co-authored-by: rajitkhanna <rajitskhanna@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
2026-09-28 16:07:09 -07:00
berriai-litellm-provider-info-sync[bot]
22cfc66af3
fix(cost-map): add Vertex batch cache prices to vertex_ai/claude-sonnet-5-5 (#43602)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-28 14:29:41 -07:00
berriai-litellm-provider-info-sync[bot]
9fd78ff6f4
fix(cost-map): add web search flag and model page source to anthropic claude-sonnet-5-5 (#43584)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-28 19:36:17 +00:00