Litellm oss staging 030626 (#29578)

* Fix incorrect agent API request example payload structure (#29556)

* fix(otel): add litellm_metadata fallback in _get_span_context and _end_proxy_span_from_kwargs (#29427)

* fix(otel): add litellm_metadata fallback in _get_span_context and _end_proxy_span_from_kwargs

On /v1/messages and other LITELLM_METADATA_ROUTES, the parent OTel span
is stored in litellm_params['litellm_metadata'] instead of
litellm_params['metadata']. When the request body contains a native
'metadata' field (e.g. Anthropic's {"user_id": "..."}),
litellm_params['metadata'] gets overwritten and the parent span is lost,
producing orphan root spans with a different trace_id.

Add fallback checks to litellm_metadata in:
- _get_span_context(): so child spans find the correct parent
- _end_proxy_span_from_kwargs(): so the proxy span gets closed

Fixes: https://github.com/BerriAI/litellm/issues/27934

* test(otel): tighten assertions per Greptile review

- test_span_context_metadata_takes_priority: assert litellm_metadata
  span is never accessed, proving metadata takes priority
- test_span_context_no_parent_when_neither_has_span: assert both ctx
  and detected_span are None

---------

Co-authored-by: shin-berri <shin-laptop@berri.ai>
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
Co-authored-by: Aneesh-Fiddler <aneeshfiddler@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>

* fix: remove premature end-user budget check from get_end_user_object (#29420)

* fix(proxy): remove premature end-user budget check from get_end_user_object

Problem:
- `_check_end_user_budget()` was called inside `get_end_user_object()`
- This caused budget checks to run BEFORE `skip_budget_checks` could be evaluated
- Zero-cost models (e.g., local vLLM) were incorrectly blocked when
  end-users exceeded their budget, even though they should bypass budget checks

Solution:
- Remove `_check_end_user_budget()` calls from `get_end_user_object()`
- Budget enforcement now happens exclusively in `common_checks()` where
  `skip_budget_checks` context is available
- `get_end_user_object()` keeps `route` as optional in function parameter for backwards compatibility and future implementation.

* refactor(tests): update budget enforcement tests to reflect changes in get_end_user_object

- test_get_end_user_object() verifies data fetching
- test_check_end_user_budget() verifies enforcement
- test_budget_enforcement_blocks_over_budget_users() integrates _check_end_user_budget()
- test_resolve_end_user_reraises_budget_exceeded() is now test_resolve_end_user since no budget exceeded is thrown in get_end_user_object()

* Gemini /images/generate and /images/edits billing fixes + add support for size and aspect ratio params (#29534)

* Fix Gemini image config mapping

* Address Gemini image config review

* Format Gemini image generation transform

* Fix Gemini image token usage logging

* Share Gemini image request helpers

* Fix Gemini Imagen model routing

* Fixes as per self code review

* Fixes per internal code review

* Stop gating Imagen imageSize forwarding

* Document Gemini image size mapping source

* chore: retrigger lint

* Clarify Gemini candidate count precedence

* Add Inception provider (#29522)

* add inception as provider (chat, fim)

* linting

* seperate test suite for chat and fim

* fix test coverage

* fix: model hub custom pricing model info (#29293)

* Opik user auth key metadata extractors (#28397)

* fix: enhance Opik metadata extraction to include user API key auth context fixed after refactoring to extractor logic

* test: add unit tests for OPik metadata extraction logic

* fix: enhance extract_opik_metadata function to prioritize metadata sources for improved accuracy

* fix(ci): clarified comments and edited unit tests

* test: add unit tests for OPik metadata extraction with auth and requester overrides

* fix(ui): replace fixed favicon.ico with current api get /get_favicon (#29532)

Signed-off-by: José Luis Di Biase <josx@interorganic.com.ar>

* fix(vertex/gemini): keep tool_call reference when a text-only assistant message follows (#29561)

`_gemini_convert_messages_with_history` tracks `last_message_with_tool_calls`
so a following tool result can be matched back to its tool call. The assignment
was inside a branch guarded by
`assistant_msg.get("tool_calls", []) is not None`, which is also True for a
text-only assistant message (an empty list is not None). As a result, an
assistant message with no tool calls that appears between a tool call and its
tool result overwrote the reference, and conversion failed with:

    Exception: Missing corresponding tool call for tool response message.

This shape is common: a model emits a short narration/assistant message after a
tool call before the tool result is appended.

Only update `last_message_with_tool_calls` when the assistant message actually
carries tool_calls (or a function_call). Adds a regression test.

Co-authored-by: shin-berri <shin-laptop@berri.ai>
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* Add 1-hour cache write pricing for EU/AU/JP Bedrock Anthropic models (#28572)

* fix(thinking): handle None thinking param in is_thinking_enabled (#28598)

Squash-merged by litellm-agent from Terrajlz's PR.

* feat(helm): support tpl rendering in podAnnotations (#28609)

Squash-merged by litellm-agent from devauxbr's PR.

* Forward custom_llm_provider through the Responses API bridge (Fixes #28505) (#28575)

* Forward custom_llm_provider through the Responses API bridge (Fixes #28505)

When a Chat Completions request to a GPT-5.4+ model contains both
`tools` and `reasoning_effort`, `completion()` auto-routes through
`responses_api_bridge`. The bridge handler called
`litellm.responses()` / `litellm.aresponses()` without forwarding the
already-resolved `custom_llm_provider`, so the downstream call
re-invoked `get_llm_provider()` with `custom_llm_provider=None` and
stripped a second provider prefix from a `provider/provider/model`
deployment string.

For a deployment configured as `openai/openai/openai/gpt-5.5`,
the bridge flow sent `openai/gpt-5.5` to the upstream API instead of
the correct `openai/openai/gpt-5.5`. Upstream APIs that enforce
model-name allow-lists rejected this as `key_model_access_denied`.

Fix: pass the locally-resolved `custom_llm_provider` into both the
sync `responses()` and async `aresponses()` calls so the downstream
`_resolve_model_provider_for_responses` sees an explicit provider
and skips the second prefix-strip.

New regression test
`tests/test_litellm/completion_extras/test_responses_bridge_provider_propagation.py`
pins both call sites: each must forward `custom_llm_provider`.

* fix(28505): set custom_llm_provider on request_data instead of as duplicate kwarg

Greptile flagged that the previous patch passed custom_llm_provider as an
explicit kwarg to responses()/aresponses() while request_data already
carried it via the spread of sanitized_litellm_params, which would raise
TypeError: got multiple values for keyword argument on every real bridge
call.

Switches to assigning request_data['custom_llm_provider'] before the call
so the resolved provider wins over whatever sanitized_litellm_params spread
in, without duplicating the kwarg.

Updates the regression test to seed request_data with a sentinel
custom_llm_provider so it actually exercises the overwrite path (the
previous test mocked transform_request with a minimal dict and never hit
the conflict).

* chore: trigger shin-agent re-eval on retargeted staging base

* chore: trigger shin-agent re-eval against updated Greptile state

* Add 1-hour cache write pricing for EU/AU/JP Bedrock Anthropic models

The 1-hour prompt-cache write tier
(`cache_creation_input_token_cost_above_1hr`) was added to the
us./global. variants of the Claude 4.5/4.6/4.7 family on Bedrock, but
the eu./au./jp. cross-region inference profiles were left without it.
AWS Bedrock pricing applies the same +10% regional premium across all
geo profiles, so eu./au./jp. should carry the same 1-hour rates as
us. (1.6x the 5-minute regional rate).

Without these fields, cost tracking on EU/AU/JP Bedrock 1-hour-TTL
prompt caching falls back to the 5-minute write rate and undercounts
spend by ~60% for European, Australian, and Japanese tenants.

Adds the 1-hour tier (and Sonnet 4.5's long-context >200K tier where
AWS publishes one) to 14 regional Bedrock entries in both
`model_prices_and_context_window.json` and the bundled
`model_prices_and_context_window_backup.json`:

  - eu./au.   Opus 4.6     ($11.00 / MTok)
  - eu./au.   Opus 4.7     ($11.00 / MTok)
  - eu./au./jp. Sonnet 4.6 ($6.60 / MTok)
  - eu./au./jp. Sonnet 4.5 ($6.60 / MTok regular, $13.20 / MTok LC)
  - eu./au./jp. Haiku 4.5  ($2.20 / MTok)

Also extends `tests/test_litellm/test_bedrock_anthropic_1hr_cache_pricing.py`
with a `REGIONAL_EXPECTED` parametrized block covering all 13 new
entries plus the existing 1.6x ratio invariant.

Note: `eu.anthropic.claude-opus-4-5-20251101-v1:0` carries the
wrong 5m rate today (base 6.25e-06 instead of regional 6.875e-06),
which would break the 1.6x ratio check. It is intentionally left out
of this PR so the scope stays "1-hour cache tier addition" — a
separate follow-up should correct the EU 5m rates for Opus 4.5.

---------

Co-authored-by: Terrajlz <info@jouleselectrictech.com>
Co-authored-by: Bruno Devaux <devaux.br@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>

* Add 1-hour cache write pricing tier for Vertex AI Anthropic models (#28569)

* fix(thinking): handle None thinking param in is_thinking_enabled (#28598)

Squash-merged by litellm-agent from Terrajlz's PR.

* feat(helm): support tpl rendering in podAnnotations (#28609)

Squash-merged by litellm-agent from devauxbr's PR.

* Forward custom_llm_provider through the Responses API bridge (Fixes #28505) (#28575)

* Forward custom_llm_provider through the Responses API bridge (Fixes #28505)

When a Chat Completions request to a GPT-5.4+ model contains both
`tools` and `reasoning_effort`, `completion()` auto-routes through
`responses_api_bridge`. The bridge handler called
`litellm.responses()` / `litellm.aresponses()` without forwarding the
already-resolved `custom_llm_provider`, so the downstream call
re-invoked `get_llm_provider()` with `custom_llm_provider=None` and
stripped a second provider prefix from a `provider/provider/model`
deployment string.

For a deployment configured as `openai/openai/openai/gpt-5.5`,
the bridge flow sent `openai/gpt-5.5` to the upstream API instead of
the correct `openai/openai/gpt-5.5`. Upstream APIs that enforce
model-name allow-lists rejected this as `key_model_access_denied`.

Fix: pass the locally-resolved `custom_llm_provider` into both the
sync `responses()` and async `aresponses()` calls so the downstream
`_resolve_model_provider_for_responses` sees an explicit provider
and skips the second prefix-strip.

New regression test
`tests/test_litellm/completion_extras/test_responses_bridge_provider_propagation.py`
pins both call sites: each must forward `custom_llm_provider`.

* fix(28505): set custom_llm_provider on request_data instead of as duplicate kwarg

Greptile flagged that the previous patch passed custom_llm_provider as an
explicit kwarg to responses()/aresponses() while request_data already
carried it via the spread of sanitized_litellm_params, which would raise
TypeError: got multiple values for keyword argument on every real bridge
call.

Switches to assigning request_data['custom_llm_provider'] before the call
so the resolved provider wins over whatever sanitized_litellm_params spread
in, without duplicating the kwarg.

Updates the regression test to seed request_data with a sentinel
custom_llm_provider so it actually exercises the overwrite path (the
previous test mocked transform_request with a minimal dict and never hit
the conflict).

* chore: trigger shin-agent re-eval on retargeted staging base

* chore: trigger shin-agent re-eval against updated Greptile state

* Add 1-hour cache write pricing tier for Vertex AI Anthropic models

GCP Vertex AI publishes a separate 1-hour cache write column for the
Claude family (1.6x the 5-minute write rate, matching the documented
Bedrock ratio). LiteLLM's Vertex AI Anthropic entries only carry the
5-minute tier, so any request that uses `cache_control: {"ttl": "1h"}`
on Vertex AI Claude is undercounted in cost tracking by ~60%.

The runtime side already supports the 1-hour tier — `VertexAIAnthropicConfig`
extends `AnthropicConfig`, populating `ephemeral_1h_input_tokens`, and
`_calculate_cache_creation_cost` reads `cache_creation_input_token_cost_above_1hr`.
Only the price registry was missing data.

Adds the field to 19 vertex_ai/claude-* entries across both
`model_prices_and_context_window.json` and the bundled
`model_prices_and_context_window_backup.json`:

  - Haiku 4.5 ($1.25 -> $2.00 / MTok)
  - Sonnet 3.7 / 4 / 4.5 / 4.6 ($3.75 -> $6.00 / MTok)
  - Opus 4.5 / 4.6 / 4.7 ($6.25 -> $10.00 / MTok)
  - Opus 4 / 4.1 ($18.75 -> $30.00 / MTok)

Adds `tests/test_litellm/test_vertex_anthropic_1hr_cache_pricing.py`
mirroring the Bedrock equivalent — pins each (5m, 1h) pair per model
and asserts the 1.6x ratio across the family.

Fixes #27781.

---------

Co-authored-by: Terrajlz <info@jouleselectrictech.com>
Co-authored-by: Bruno Devaux <devaux.br@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>

* Fix Gemini multimodal function responses (#29325)

Co-authored-by: shin-berri <shin-laptop@berri.ai>
Co-authored-by: yuneng-jiang <yuneng@berri.ai>

* address greptile review: add _transform_image_usage method and model-map supports_image_size flag

- Add _transform_image_usage instance method to GoogleImageGenConfig that
  delegates to transform_gemini_image_usage, fixing the regression test
- Replace hardcoded "2.5-flash" string check in supports_gemini_image_size
  with a get_model_info lookup on supports_image_size (default true)
- Add supports_image_size: false to all gemini-2.5-flash model entries in
  model_prices_and_context_window.json so capability is controlled via the
  model map rather than embedded in code

* fix test failures: schema validation, mypy type, model info plumbing, pricing test

- Add supports_image_size to ModelInfoBase TypedDict so get_model_info surfaces it
- Pass supports_image_size through _get_model_info_helper constructor call
- Fix supports_gemini_image_size to use value is not False (None means unset, defaults to True)
- Add supports_image_size to JSON schema in test_aaamodel_prices_and_context_window_json_is_valid
- Correct gemini-3.1-flash-lite pricing assertions in test to match JSON values

* Add Azure AI Kimi K2.6 metadata (#27052)

* Add Azure AI Kimi K2.6 metadata

* Scope Kimi metadata test cost map setup

* fall back to substring check for models not in model_prices_and_context_window.json

Models like gemini-2.5-flash-image-preview are not in the pricing JSON,
so get_model_info raises. Fall back to "2.5-flash" not in model when the
JSON has no explicit supports_image_size entry for the model.

* fix(inception): don't forward global litellm.api_key to Inception FIM

Match the Inception chat config: resolve only an Inception-specific key
(param, litellm.inception_key, or INCEPTION_API_KEY) for the text-completion
FIM path. The global litellm.api_key (often an OpenAI key) was both leaking
to api.inceptionlabs.ai and taking precedence over the configured Inception
key when set.

* fix(auth): enforce end-user budget on custom-auth path that skips common_checks

get_end_user_object() no longer raises BudgetExceededError, so custom-auth
deployments with custom_auth_run_common_checks unset (which skip the
centralized common_checks gate) stopped enforcing the end-user budget,
letting an over-budget end user keep making requests. Re-enforce the
budget in _run_post_custom_auth_checks on that path.

---------

Signed-off-by: José Luis Di Biase <josx@interorganic.com.ar>
Co-authored-by: Isha <72744901+IshaMeera@users.noreply.github.com>
Co-authored-by: aneeshsangvikar <aneeshsangvikar@fiddler.ai>
Co-authored-by: shin-berri <shin-laptop@berri.ai>
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
Co-authored-by: Aneesh-Fiddler <aneeshfiddler@gmail.com>
Co-authored-by: Suleiman Elkhoury <108065141+suleimanelkhoury@users.noreply.github.com>
Co-authored-by: Dmitriy Alergant <93501479+DmitriyAlergant@users.noreply.github.com>
Co-authored-by: Yanis Miraoui <yanis.miraoui19@imperial.ac.uk>
Co-authored-by: Lovro Seder <vrovro@gmail.com>
Co-authored-by: Thomas Mildner <12685945+Thomas-Mildner@users.noreply.github.com>
Co-authored-by: José Luis Di Biase <josx@interorganic.com.ar>
Co-authored-by: Lai Quang Huy <64073540+1qh@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Filippo Menghi <113345637+Cyberfilo@users.noreply.github.com>
Co-authored-by: Terrajlz <info@jouleselectrictech.com>
Co-authored-by: Bruno Devaux <devaux.br@gmail.com>
Co-authored-by: ZHONG Ziwen <67355585+zzw-math@users.noreply.github.com>
Co-authored-by: Emerson Gomes <emerson.gomes@thalesgroup.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
This commit is contained in:
Sameer Kankute 2026-06-03 23:31:51 +05:30 committed by GitHub
parent f3e2167730
commit c7ab9adde5
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
59 changed files with 3534 additions and 760 deletions

View file

@ -278,6 +278,7 @@ ovhcloud_key: Optional[str] = None
lemonade_key: Optional[str] = None
sap_service_key: Optional[str] = None
amazon_nova_api_key: Optional[str] = None
inception_key: Optional[str] = None
common_cloud_provider_auth_params: dict = {
"params": ["project", "region_name", "token"],
"providers": ["vertex_ai", "bedrock", "watsonx", "azure", "vertex_ai_beta"],
@ -551,6 +552,7 @@ cohere_models: Set = set()
cohere_chat_models: Set = set()
mistral_chat_models: Set = set()
text_completion_codestral_models: Set = set()
text_completion_inception_models: Set = set()
anthropic_models: Set = set()
openrouter_models: Set = set()
datarobot_models: Set = set()
@ -628,6 +630,7 @@ publicai_models: Set = set()
v0_models: Set = set()
morph_models: Set = set()
lambda_ai_models: Set = set()
inception_models: Set = set()
hyperbolic_models: Set = set()
black_forest_labs_models: Set = set()
recraft_models: Set = set()
@ -792,6 +795,8 @@ def add_known_models(model_cost_map: Optional[Dict] = None):
fireworks_ai_embedding_models.add(key)
elif value.get("litellm_provider") == "text-completion-codestral":
text_completion_codestral_models.add(key)
elif value.get("litellm_provider") == "text-completion-inception":
text_completion_inception_models.add(key)
elif value.get("litellm_provider") == "xai":
xai_models.add(key)
elif value.get("litellm_provider") == "zai":
@ -878,6 +883,8 @@ def add_known_models(model_cost_map: Optional[Dict] = None):
morph_models.add(key)
elif value.get("litellm_provider") == "lambda_ai":
lambda_ai_models.add(key)
elif value.get("litellm_provider") == "inception":
inception_models.add(key)
elif value.get("litellm_provider") == "hyperbolic":
hyperbolic_models.add(key)
elif value.get("litellm_provider") == "black_forest_labs":
@ -980,6 +987,7 @@ model_list = list(
| watsonx_models
| gemini_models
| text_completion_codestral_models
| text_completion_inception_models
| xai_models
| zai_models
| fal_ai_models
@ -1018,6 +1026,7 @@ model_list = list(
| v0_models
| morph_models
| lambda_ai_models
| inception_models
| black_forest_labs_models
| recraft_models
| cometapi_models
@ -1074,6 +1083,7 @@ models_by_provider: dict = {
"fireworks_ai": fireworks_ai_models | fireworks_ai_embedding_models,
"aleph_alpha": aleph_alpha_models,
"text-completion-codestral": text_completion_codestral_models,
"text-completion-inception": text_completion_inception_models,
"xai": xai_models,
"zai": zai_models,
"fal_ai": fal_ai_models,
@ -1118,6 +1128,7 @@ models_by_provider: dict = {
"v0": v0_models,
"morph": morph_models,
"lambda_ai": lambda_ai_models,
"inception": inception_models,
"hyperbolic": hyperbolic_models,
"black_forest_labs": black_forest_labs_models,
"recraft": recraft_models,
@ -1869,6 +1880,9 @@ if TYPE_CHECKING:
from .llms.codestral.completion.transformation import (
CodestralTextCompletionConfig as CodestralTextCompletionConfig,
)
from .llms.inception.completion.transformation import (
InceptionTextCompletionConfig as InceptionTextCompletionConfig,
)
from .llms.azure.azure import (
AzureOpenAIAssistantsAPIConfig as AzureOpenAIAssistantsAPIConfig,
)
@ -1937,6 +1951,9 @@ if TYPE_CHECKING:
from .llms.lambda_ai.chat.transformation import (
LambdaAIChatConfig as LambdaAIChatConfig,
)
from .llms.inception.chat.transformation import (
InceptionChatConfig as InceptionChatConfig,
)
from .llms.hyperbolic.chat.transformation import (
HyperbolicChatConfig as HyperbolicChatConfig,
)

View file

@ -267,6 +267,7 @@ LLM_CONFIG_NAMES = (
"AIMLChatConfig",
"VolcEngineChatConfig",
"CodestralTextCompletionConfig",
"InceptionTextCompletionConfig",
"AzureOpenAIAssistantsAPIConfig",
"HerokuChatConfig",
"CometAPIConfig",
@ -310,6 +311,7 @@ LLM_CONFIG_NAMES = (
"MorphChatConfig",
"RAGFlowConfig",
"LambdaAIChatConfig",
"InceptionChatConfig",
"HyperbolicChatConfig",
"VercelAIGatewayConfig",
"OVHCloudChatConfig",
@ -1040,6 +1042,10 @@ _LLM_CONFIGS_IMPORT_MAP = {
".llms.codestral.completion.transformation",
"CodestralTextCompletionConfig",
),
"InceptionTextCompletionConfig": (
".llms.inception.completion.transformation",
"InceptionTextCompletionConfig",
),
"AzureOpenAIAssistantsAPIConfig": (
".llms.azure.azure",
"AzureOpenAIAssistantsAPIConfig",
@ -1154,6 +1160,10 @@ _LLM_CONFIGS_IMPORT_MAP = {
"MorphChatConfig": (".llms.morph.chat.transformation", "MorphChatConfig"),
"RAGFlowConfig": (".llms.ragflow.chat.transformation", "RAGFlowConfig"),
"LambdaAIChatConfig": (".llms.lambda_ai.chat.transformation", "LambdaAIChatConfig"),
"InceptionChatConfig": (
".llms.inception.chat.transformation",
"InceptionChatConfig",
),
"HyperbolicChatConfig": (
".llms.hyperbolic.chat.transformation",
"HyperbolicChatConfig",

View file

@ -182,6 +182,14 @@ class ResponsesToCompletionBridgeHandler:
client=kwargs.get("client"),
)
# Pin the resolved provider so `responses()` doesn't re-run
# `get_llm_provider()` on the model string and strip a second
# provider prefix (see GitHub issue #28505). request_data already
# carries `custom_llm_provider` via the spread of
# `sanitized_litellm_params`; overwriting it on the dict (rather
# than adding an explicit kwarg) avoids the duplicate-keyword
# TypeError that would otherwise fire on the real bridge path.
request_data["custom_llm_provider"] = custom_llm_provider
result = responses(
**request_data,
)
@ -268,6 +276,13 @@ class ResponsesToCompletionBridgeHandler:
except Exception as e:
raise e
# Pin the resolved provider so `aresponses()` doesn't re-run
# `get_llm_provider()` on the model string and strip a second
# provider prefix (see GitHub issue #28505). Set on request_data
# rather than passed as a separate kwarg to avoid the duplicate-
# keyword TypeError when `sanitized_litellm_params` already
# carries `custom_llm_provider`.
request_data["custom_llm_provider"] = custom_llm_provider
result = await aresponses(
**request_data,
aresponses=True,

View file

@ -585,6 +585,7 @@ LITELLM_CHAT_PROVIDERS = [
"volcengine",
"codestral",
"text-completion-codestral",
"text-completion-inception",
"deepseek",
"sambanova",
"maritalk",
@ -620,6 +621,7 @@ LITELLM_CHAT_PROVIDERS = [
"oci",
"morph",
"lambda_ai",
"inception",
"vercel_ai_gateway",
"wandb",
"ovhcloud",
@ -779,6 +781,7 @@ openai_compatible_endpoints: List = [
"https://api.v0.dev/v1",
"https://api.morphllm.com/v1",
"https://api.lambda.ai/v1",
"https://api.inceptionlabs.ai/v1",
"https://api.hyperbolic.xyz/v1",
"https://ai-gateway.helicone.ai/",
"https://ai-gateway.vercel.sh/v1",
@ -835,6 +838,7 @@ openai_compatible_providers: List = [
"helicone",
"morph",
"lambda_ai",
"inception",
"hyperbolic",
"vercel_ai_gateway",
"aiml",

View file

@ -1012,6 +1012,13 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
litellm_params = kwargs.get("litellm_params", {}) or {}
_metadata = litellm_params.get("metadata", {}) or {}
proxy_span = _metadata.get("litellm_parent_otel_span", None)
# Fallback: check litellm_metadata (used by /v1/messages and other
# LITELLM_METADATA_ROUTES).
if proxy_span is None:
_litellm_metadata = litellm_params.get("litellm_metadata", {}) or {}
proxy_span = _litellm_metadata.get("litellm_parent_otel_span", None)
if (
proxy_span is not None
and getattr(proxy_span, "name", None) == LITELLM_PROXY_REQUEST_SPAN_NAME
@ -2718,6 +2725,13 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
_metadata = litellm_params.get("metadata", {}) or {}
parent_otel_span = _metadata.get("litellm_parent_otel_span", None)
# Fallback: check litellm_metadata (used by /v1/messages and other
# LITELLM_METADATA_ROUTES that store proxy-internal metadata
# separately from the provider's native "metadata" field).
if parent_otel_span is None:
_litellm_metadata = litellm_params.get("litellm_metadata", {}) or {}
parent_otel_span = _litellm_metadata.get("litellm_parent_otel_span", None)
# Priority 1: Explicit parent span from metadata
if parent_otel_span is not None:
verbose_logger.debug(

View file

@ -39,20 +39,32 @@ def extract_opik_metadata(
standard_logging_metadata: Dict[str, Any],
) -> Dict[str, Any]:
"""
Extract and merge Opik metadata from request and requester.
Merge Opik metadata from three sources in increasing priority order:
1. user_api_key_auth_metadata lowest priority (operator-level defaults)
2. litellm_metadata (request) overrides auth-key defaults
3. requester_metadata highest priority (e.g. proxy header overrides)
Args:
litellm_metadata: Metadata from litellm_params
standard_logging_metadata: Metadata from standard_logging_object
litellm_metadata: Metadata from litellm_params.mak
standard_logging_metadata: Metadata from standard_logging_object.
Returns:
Merged Opik metadata dictionary
Merged Opik metadata dictionary.
"""
opik_meta = litellm_metadata.get("opik", {}).copy()
# Start with auth-key defaults (lowest priority).
auth_meta = standard_logging_metadata.get("user_api_key_auth_metadata") or {}
opik_meta = (auth_meta.get("opik") or {}).copy()
# Request-level values override auth-key defaults.
request_opik = litellm_metadata.get("opik") or {}
opik_meta.update(request_opik)
# Requester-level values win over everything else.
requester_metadata = standard_logging_metadata.get("requester_metadata", {}) or {}
requester_opik = requester_metadata.get("opik", {}) or {}
opik_meta.update(requester_opik)
if requester_opik:
opik_meta.update(requester_opik)
_logging.verbose_logger.debug(
f"litellm_opik_metadata - {json.dumps(opik_meta, default=str)}"

View file

@ -373,6 +373,9 @@ def get_llm_provider( # noqa: PLR0915
elif endpoint == "https://api.lambda.ai/v1":
custom_llm_provider = "lambda_ai"
dynamic_api_key = get_secret_str("LAMBDA_API_KEY")
elif endpoint == "https://api.inceptionlabs.ai/v1":
custom_llm_provider = "inception"
dynamic_api_key = get_secret_str("INCEPTION_API_KEY")
elif endpoint == "https://api.hyperbolic.xyz/v1":
custom_llm_provider = "hyperbolic"
dynamic_api_key = get_secret_str("HYPERBOLIC_API_KEY")
@ -954,6 +957,13 @@ def _get_openai_compatible_provider_info( # noqa: PLR0915
) = litellm.LambdaAIChatConfig()._get_openai_compatible_provider_info(
api_base, api_key
)
elif custom_llm_provider == "inception":
(
api_base,
dynamic_api_key,
) = litellm.InceptionChatConfig()._get_openai_compatible_provider_info(
api_base, api_key
)
elif custom_llm_provider == "hyperbolic":
(
api_base,

View file

@ -34,6 +34,14 @@ _IMAGE_RESPONSE_CALL_TYPES = frozenset(
_VALID_DATA_RESIDENCIES = frozenset(r.value for r in DataResidency)
def _get_token_detail_value(details: object, key: str) -> Optional[int]:
if isinstance(details, dict):
value = details.get(key)
else:
value = getattr(details, key, None)
return value if isinstance(value, int) else None
def _is_above_128k(tokens: float) -> bool:
if tokens > 128000:
return True
@ -870,17 +878,47 @@ def calculate_image_response_cost_from_usage(
cached_tokens=0,
)
output_tokens_details = getattr(usage, "completion_tokens_details", None)
if output_tokens_details is None:
output_tokens_details = getattr(usage, "output_tokens_details", None)
if output_tokens_details is None:
completion_tokens_details = CompletionTokensDetailsWrapper(
text_tokens=0,
image_tokens=completion_tokens,
reasoning_tokens=0,
audio_tokens=0,
)
else:
text_tokens = _get_token_detail_value(output_tokens_details, "text_tokens") or 0
image_tokens = (
_get_token_detail_value(output_tokens_details, "image_tokens") or 0
)
audio_tokens = (
_get_token_detail_value(output_tokens_details, "audio_tokens") or 0
)
reasoning_tokens = (
_get_token_detail_value(output_tokens_details, "reasoning_tokens") or 0
)
known_output_tokens = (
text_tokens + image_tokens + audio_tokens + reasoning_tokens
)
if completion_tokens > known_output_tokens:
text_tokens += completion_tokens - known_output_tokens
completion_tokens_details = CompletionTokensDetailsWrapper(
text_tokens=text_tokens,
image_tokens=image_tokens,
reasoning_tokens=reasoning_tokens,
audio_tokens=audio_tokens,
)
normalized_usage = Usage(
prompt_tokens=prompt_tokens,
completion_tokens=completion_tokens,
total_tokens=total_tokens,
prompt_tokens_details=prompt_tokens_details,
completion_tokens_details=CompletionTokensDetailsWrapper(
text_tokens=0,
image_tokens=completion_tokens,
reasoning_tokens=0,
audio_tokens=0,
),
completion_tokens_details=completion_tokens_details,
)
prompt_cost, completion_cost = generic_cost_per_token(

View file

@ -1670,15 +1670,15 @@ def convert_to_gemini_tool_call_result( # noqa: PLR0915
if gemini_call_id:
_function_response["id"] = gemini_call_id
# Create part with function_response, and optionally inline_data for images (Computer Use)
_part: VertexPartType = {"function_response": _function_response}
# For Computer Use, if we have images/files, we need separate parts:
# - One part with function_response
# - One part per inline_data item
# Gemini's PartType is a oneof, so we can't have both in the same part
# For multimodal function responses, Gemini expects media parts nested
# inside functionResponse.parts instead of sibling content parts.
if inline_data_list:
return [_part] + [{"inline_data": d} for d in inline_data_list]
_function_response["parts"] = [
{"inline_data": inline_data} for inline_data in inline_data_list
]
return [_part]
return _part

View file

@ -1,6 +1,8 @@
import base64
import datetime
from typing import Any, Dict, List, Optional, Union
import json
import math
from typing import Any, Dict, List, Optional, Sequence, Union
import httpx
@ -12,6 +14,245 @@ from litellm.secret_managers.main import get_secret_str
from litellm.types.llms.openai import AllMessageValues
from litellm.types.utils import TokenCountResponse
GEMINI_IMAGE_ASPECT_RATIOS: Dict[str, float] = {
"1:1": 1 / 1,
"1:4": 1 / 4,
"1:8": 1 / 8,
"2:3": 2 / 3,
"3:2": 3 / 2,
"3:4": 3 / 4,
"4:1": 4 / 1,
"4:3": 4 / 3,
"4:5": 4 / 5,
"5:4": 5 / 4,
"8:1": 8 / 1,
"9:16": 9 / 16,
"16:9": 16 / 9,
"21:9": 21 / 9,
}
# Supported aspect ratio dimensions from Google Gemini image generation docs:
# https://ai.google.dev/gemini-api/docs/image-generation#aspect_ratios_and_image_size
GEMINI_IMAGE_SIZE_TO_ASPECT_RATIO: Dict[tuple[int, int], str] = {
(512, 512): "1:1",
(1024, 1024): "1:1",
(2048, 2048): "1:1",
(4096, 4096): "1:1",
(256, 1024): "1:4",
(512, 2048): "1:4",
(1024, 4096): "1:4",
(2048, 8192): "1:4",
(192, 1536): "1:8",
(384, 3072): "1:8",
(768, 6144): "1:8",
(1536, 12288): "1:8",
(424, 632): "2:3",
(848, 1264): "2:3",
(1696, 2528): "2:3",
(3392, 5056): "2:3",
(632, 424): "3:2",
(1264, 848): "3:2",
(2528, 1696): "3:2",
(5056, 3392): "3:2",
(448, 600): "3:4",
(896, 1200): "3:4",
(1792, 2400): "3:4",
(3584, 4800): "3:4",
(1024, 256): "4:1",
(2048, 512): "4:1",
(4096, 1024): "4:1",
(8192, 2048): "4:1",
(600, 448): "4:3",
(1200, 896): "4:3",
(2400, 1792): "4:3",
(4800, 3584): "4:3",
(464, 576): "4:5",
(928, 1152): "4:5",
(1856, 2304): "4:5",
(3712, 4608): "4:5",
(576, 464): "5:4",
(1152, 928): "5:4",
(2304, 1856): "5:4",
(4608, 3712): "5:4",
(1536, 192): "8:1",
(3072, 384): "8:1",
(6144, 768): "8:1",
(12288, 1536): "8:1",
(384, 688): "9:16",
(768, 1376): "9:16",
(1536, 2752): "9:16",
(3072, 5504): "9:16",
(688, 384): "16:9",
(1376, 768): "16:9",
(2752, 1536): "16:9",
(5504, 3072): "16:9",
(792, 336): "21:9",
(1584, 672): "21:9",
(3168, 1344): "21:9",
(6336, 2688): "21:9",
(1280, 896): "4:3",
(896, 1280): "3:4",
}
def map_openai_size_to_gemini_image_config(
size: str, model: str
) -> Optional[Dict[str, str]]:
dimensions = _parse_openai_image_size(size)
if dimensions is None:
return None
width, height = dimensions
image_config = {
"aspectRatio": _map_dimensions_to_gemini_aspect_ratio(width, height)
}
image_size = _map_dimensions_to_gemini_image_size(width, height)
if is_gemini_image_model(model):
if supports_gemini_image_size(model):
image_config["imageSize"] = image_size
else:
image_config["imageSize"] = image_size
return image_config
def supports_gemini_image_size(model: str) -> bool:
try:
model_info = litellm.get_model_info(model=model)
value = model_info.get("supports_image_size")
if value is not None:
return bool(value)
except Exception:
pass
return "2.5-flash" not in model
def is_gemini_image_model(model: str) -> bool:
base_model = model.split("/", 1)[-1]
return "gemini" in base_model
def map_openai_image_params_to_gemini(
params: Dict[str, Any],
model: str,
supported_params: Sequence[str],
optional_params: Optional[Dict[str, Any]] = None,
parse_image_config_string: bool = False,
) -> Dict[str, Any]:
optional_params = optional_params or {}
filtered_params = {
key: value for key, value in params.items() if key in supported_params
}
mapped_params: Dict[str, Any] = {}
if "n" in filtered_params and "n" not in optional_params:
mapped_params["sampleCount"] = filtered_params["n"]
if "size" in filtered_params and "size" not in optional_params:
image_config = map_openai_size_to_gemini_image_config(
filtered_params["size"],
model,
)
if image_config is not None:
if is_gemini_image_model(model):
mapped_params["imageConfig"] = image_config
else:
mapped_params["aspectRatio"] = image_config["aspectRatio"]
if "imageSize" in image_config:
mapped_params["imageSize"] = image_config["imageSize"]
image_config_param = filtered_params.get("imageConfig")
if isinstance(image_config_param, str) and parse_image_config_string:
try:
image_config_param = json.loads(image_config_param)
except json.JSONDecodeError as exc:
raise litellm.UnsupportedParamsError(
model=model,
message="`imageConfig` must be valid JSON when provided as a string.",
) from exc
if isinstance(image_config_param, dict):
mapped_params["imageConfig"] = image_config_param
for key, value in filtered_params.items():
if key not in ("n", "size", "imageConfig") and key not in optional_params:
mapped_params[key] = value
return mapped_params
def get_gemini_image_generation_config(
model: str,
optional_params: Dict[str, Any],
) -> Dict[str, Any]:
generation_config: Dict[str, Any] = {"response_modalities": ["IMAGE", "TEXT"]}
image_config: Dict[str, Any] = {}
if isinstance(optional_params.get("imageConfig"), dict):
image_config.update(optional_params["imageConfig"])
if not supports_gemini_image_size(model):
image_config.pop("imageSize", None)
if image_config:
generation_config["imageConfig"] = image_config
candidate_count = next(
(
optional_params[key]
for key in ("candidateCount", "candidate_count", "sampleCount", "n")
if optional_params.get(key) is not None
),
None,
)
if candidate_count is not None:
generation_config["candidateCount"] = candidate_count
return generation_config
def _parse_openai_image_size(size: str) -> Optional[tuple[int, int]]:
if size == "auto":
return None
width_str, separator, height_str = size.lower().partition("x")
if not separator:
return None
try:
width = int(width_str)
height = int(height_str)
except ValueError:
return None
if width <= 0 or height <= 0:
return None
return width, height
def _map_dimensions_to_gemini_aspect_ratio(width: int, height: int) -> str:
if (width, height) in GEMINI_IMAGE_SIZE_TO_ASPECT_RATIO:
return GEMINI_IMAGE_SIZE_TO_ASPECT_RATIO[(width, height)]
requested_ratio = width / height
return min(
GEMINI_IMAGE_ASPECT_RATIOS,
key=lambda aspect_ratio: abs(
math.log(GEMINI_IMAGE_ASPECT_RATIOS[aspect_ratio] / requested_ratio)
),
)
def _map_dimensions_to_gemini_image_size(width: int, height: int) -> str:
effective_square_side = math.sqrt(width * height)
if effective_square_side < 768:
return "512"
if effective_square_side < 1536:
return "1K"
if effective_square_side < 3072:
return "2K"
return "4K"
class GeminiError(BaseLLMException):
pass

View file

@ -4,8 +4,9 @@ Gemini Image Edit Cost Calculator
from typing import Any
import litellm
from litellm.types.utils import ImageResponse
from litellm.llms.gemini.image_generation.cost_calculator import (
cost_calculator as image_generation_cost_calculator,
)
def cost_calculator(
@ -15,20 +16,10 @@ def cost_calculator(
"""
Gemini image edit cost calculator.
Mirrors image generation pricing: charge per returned image based on
model metadata (`output_cost_per_image`).
Gemini image edits and generations share image response billing behavior:
use provider token usage when present, otherwise fall back to per-image pricing.
"""
model_info = litellm.get_model_info(
return image_generation_cost_calculator(
model=model,
custom_llm_provider="gemini",
image_response=image_response,
)
output_cost_per_image: float = model_info.get("output_cost_per_image") or 0.0
if not isinstance(image_response, ImageResponse):
raise ValueError(
f"image_response must be of type ImageResponse got type={type(image_response)}"
)
num_images = len(image_response.data or [])
return output_cost_per_image * num_images

View file

@ -7,10 +7,22 @@ from httpx._types import RequestFiles
from litellm.images.utils import ImageEditRequestUtils
from litellm.llms.base_llm.image_edit.transformation import BaseImageEditConfig
from litellm.llms.gemini.common_utils import (
get_gemini_image_generation_config,
map_openai_image_params_to_gemini,
)
from litellm.llms.gemini.image_usage_transformation import (
transform_gemini_image_usage,
)
from litellm.secret_managers.main import get_secret_str
from litellm.types.images.main import ImageEditOptionalRequestParams
from litellm.types.router import GenericLiteLLMParams
from litellm.types.utils import FileTypes, ImageObject, ImageResponse, OpenAIImage
from litellm.types.utils import (
FileTypes,
ImageObject,
ImageResponse,
OpenAIImage,
)
if TYPE_CHECKING:
from litellm.litellm_core_utils.litellm_logging import Logging as _LiteLLMLoggingObj
@ -22,7 +34,7 @@ else:
class GeminiImageEditConfig(BaseImageEditConfig):
DEFAULT_BASE_URL: str = "https://generativelanguage.googleapis.com/v1beta"
SUPPORTED_PARAMS: List[str] = ["size"]
SUPPORTED_PARAMS: List[str] = ["n", "size", "imageConfig"]
def get_supported_openai_params(self, model: str) -> List[str]:
return list(self.SUPPORTED_PARAMS)
@ -33,21 +45,12 @@ class GeminiImageEditConfig(BaseImageEditConfig):
model: str,
drop_params: bool,
) -> Dict[str, Any]:
supported_params = self.get_supported_openai_params(model)
filtered_params = {
key: value
for key, value in image_edit_optional_params.items()
if key in supported_params
}
mapped_params: Dict[str, Any] = {}
if "size" in filtered_params:
mapped_params["aspectRatio"] = self._map_size_to_aspect_ratio(
filtered_params["size"] # type: ignore[arg-type]
)
return mapped_params
return map_openai_image_params_to_gemini(
params=image_edit_optional_params, # type: ignore[arg-type]
model=model,
supported_params=self.get_supported_openai_params(model),
parse_image_config_string=True,
)
def validate_environment(
self,
@ -107,18 +110,10 @@ class GeminiImageEditConfig(BaseImageEditConfig):
request_body: Dict[str, Any] = {"contents": contents}
generation_config: Dict[str, Any] = {}
if "aspectRatio" in image_edit_optional_request_params:
# Move aspectRatio into imageConfig inside generationConfig
if "imageConfig" not in generation_config:
generation_config["imageConfig"] = {}
generation_config["imageConfig"]["aspectRatio"] = (
image_edit_optional_request_params["aspectRatio"]
)
if generation_config:
request_body["generationConfig"] = generation_config
request_body["generationConfig"] = get_gemini_image_generation_config(
model=model,
optional_params=image_edit_optional_request_params,
)
empty_files = cast(RequestFiles, [])
return request_body, empty_files
@ -156,18 +151,12 @@ class GeminiImageEditConfig(BaseImageEditConfig):
)
model_response.data = cast(List[OpenAIImage], data_list)
if "usageMetadata" in response_json:
model_response.usage = transform_gemini_image_usage(
response_json["usageMetadata"]
)
return model_response
def _map_size_to_aspect_ratio(self, size: str) -> str:
aspect_ratio_map = {
"1024x1024": "1:1",
"1792x1024": "16:9",
"1024x1792": "9:16",
"1280x896": "4:3",
"896x1280": "3:4",
}
return aspect_ratio_map.get(size, "1:1")
def _prepare_inline_image_parts(
self, image: Union[FileTypes, List[FileTypes]]
) -> List[Dict[str, Any]]:

View file

@ -5,18 +5,21 @@ import httpx
from litellm.llms.base_llm.image_generation.transformation import (
BaseImageGenerationConfig,
)
from litellm.llms.gemini.common_utils import (
get_gemini_image_generation_config,
is_gemini_image_model,
map_openai_image_params_to_gemini,
)
from litellm.llms.gemini.image_usage_transformation import (
transform_gemini_image_usage,
)
from litellm.secret_managers.main import get_secret_str
from litellm.types.llms.gemini import GeminiImageGenerationRequest
from litellm.types.llms.openai import (
AllMessageValues,
OpenAIImageGenerationOptionalParams,
)
from litellm.types.utils import (
ImageObject,
ImageResponse,
ImageUsage,
ImageUsageInputTokensDetails,
)
from litellm.types.utils import ImageObject, ImageResponse
if TYPE_CHECKING:
from litellm.litellm_core_utils.litellm_logging import Logging as _LiteLLMLoggingObj
@ -36,7 +39,10 @@ class GoogleImageGenConfig(BaseImageGenerationConfig):
Google AI Imagen API supported parameters
https://ai.google.dev/gemini-api/docs/imagen
"""
return ["n", "size"]
supported_params = ["n", "size"]
if is_gemini_image_model(model):
supported_params.append("imageConfig")
return supported_params # type: ignore[return-value]
def map_openai_params(
self,
@ -45,64 +51,11 @@ class GoogleImageGenConfig(BaseImageGenerationConfig):
model: str,
drop_params: bool,
) -> dict:
supported_params = self.get_supported_openai_params(model)
mapped_params = {}
for k, v in non_default_params.items():
if k not in optional_params.keys():
if k in supported_params:
# Map OpenAI parameters to Google format
if k == "n":
mapped_params["sampleCount"] = v
elif k == "size":
# Map OpenAI size format to Google aspectRatio
mapped_params["aspectRatio"] = self._map_size_to_aspect_ratio(v)
else:
mapped_params[k] = v
return mapped_params
def _map_size_to_aspect_ratio(self, size: str) -> str:
"""
https://ai.google.dev/gemini-api/docs/image-generation
"""
aspect_ratio_map = {
"1024x1024": "1:1",
"1792x1024": "16:9",
"1024x1792": "9:16",
"1280x896": "4:3",
"896x1280": "3:4",
}
return aspect_ratio_map.get(size, "1:1")
def _transform_image_usage(self, usage_metadata: dict) -> ImageUsage:
"""
Transform Gemini usageMetadata to ImageUsage format
"""
input_tokens_details = ImageUsageInputTokensDetails(
image_tokens=0,
text_tokens=0,
)
# Extract detailed token counts from promptTokensDetails
tokens_details = usage_metadata.get("promptTokensDetails", [])
for details in tokens_details:
if isinstance(details, dict):
modality = str(details.get("modality", "")).upper()
raw_token_count = details.get(
"tokenCount", details.get("token_count", 0)
)
token_count = raw_token_count if isinstance(raw_token_count, int) else 0
if modality == "TEXT":
input_tokens_details.text_tokens += token_count
elif modality == "IMAGE":
input_tokens_details.image_tokens += token_count
return ImageUsage(
input_tokens=usage_metadata.get("promptTokenCount", 0),
input_tokens_details=input_tokens_details,
output_tokens=usage_metadata.get("candidatesTokenCount", 0),
total_tokens=usage_metadata.get("totalTokenCount", 0),
return map_openai_image_params_to_gemini(
params=non_default_params,
model=model,
supported_params=self.get_supported_openai_params(model),
optional_params=optional_params,
)
def get_complete_url(
@ -127,7 +80,7 @@ class GoogleImageGenConfig(BaseImageGenerationConfig):
complete_url = complete_url.rstrip("/")
# Gemini Flash Image Preview models use generateContent endpoint
if "gemini" in model:
if is_gemini_image_model(model):
complete_url = f"{complete_url}/models/{model}:generateContent"
else:
# All other Imagen models use predict endpoint
@ -179,10 +132,13 @@ class GoogleImageGenConfig(BaseImageGenerationConfig):
}
"""
# For Gemini Flash Image Preview models, use standard Gemini format
if "gemini" in model:
if is_gemini_image_model(model):
request_body: dict = {
"contents": [{"parts": [{"text": prompt}]}],
"generationConfig": {"response_modalities": ["IMAGE", "TEXT"]},
"generationConfig": get_gemini_image_generation_config(
model=model,
optional_params=optional_params,
),
}
return request_body
else:
@ -200,6 +156,9 @@ class GoogleImageGenConfig(BaseImageGenerationConfig):
)
return request_body_obj.model_dump(exclude_none=True)
def _transform_image_usage(self, usage_metadata: dict):
return transform_gemini_image_usage(usage_metadata)
def transform_image_generation_response(
self,
model: str,
@ -229,7 +188,7 @@ class GoogleImageGenConfig(BaseImageGenerationConfig):
model_response.data = []
# Handle different response formats based on model
if "gemini" in model:
if is_gemini_image_model(model):
# Gemini Flash Image Preview models return in candidates format
candidates = response_data.get("candidates", [])
for candidate in candidates:
@ -255,7 +214,7 @@ class GoogleImageGenConfig(BaseImageGenerationConfig):
# Extract usage metadata for Gemini models
if "usageMetadata" in response_data:
model_response.usage = self._transform_image_usage(
model_response.usage = transform_gemini_image_usage(
response_data["usageMetadata"]
)
else:

View file

@ -0,0 +1,73 @@
from typing import Any
from litellm.types.utils import ImageUsage, ImageUsageInputTokensDetails
def _get_token_count(details: dict) -> int:
raw_token_count = details.get("tokenCount", details.get("token_count", 0))
return raw_token_count if isinstance(raw_token_count, int) else 0
def _get_modality_token_details(usage_metadata: dict, *details_keys: str) -> list:
for details_key in details_keys:
details = usage_metadata.get(details_key)
if isinstance(details, list):
return details
return []
def _sum_modality_token_details(
usage_metadata: dict, *details_keys: str
) -> ImageUsageInputTokensDetails:
tokens_details = ImageUsageInputTokensDetails(
image_tokens=0,
text_tokens=0,
)
for details in _get_modality_token_details(usage_metadata, *details_keys):
if isinstance(details, dict):
modality = str(details.get("modality", "")).upper()
token_count = _get_token_count(details)
if modality == "TEXT":
tokens_details.text_tokens += token_count
elif modality == "IMAGE":
tokens_details.image_tokens += token_count
return tokens_details
def transform_gemini_image_usage(usage_metadata: dict) -> ImageUsage:
"""
Transform Gemini usageMetadata to ImageUsage format.
"""
input_tokens_details = _sum_modality_token_details(
usage_metadata, "promptTokensDetails", "prompt_tokens_details"
)
output_tokens = usage_metadata.get("candidatesTokenCount", 0)
output_tokens_details = _sum_modality_token_details(
usage_metadata, "candidatesTokensDetails", "candidates_tokens_details"
)
if not _get_modality_token_details(
usage_metadata, "candidatesTokensDetails", "candidates_tokens_details"
):
output_tokens_details.image_tokens = output_tokens
else:
known_output_tokens = (
output_tokens_details.text_tokens + output_tokens_details.image_tokens
)
if output_tokens > known_output_tokens:
output_tokens_details.text_tokens += output_tokens - known_output_tokens
usage_payload: dict[str, Any] = {
"input_tokens": usage_metadata.get("promptTokenCount", 0),
"input_tokens_details": input_tokens_details,
"output_tokens": output_tokens,
"total_tokens": usage_metadata.get("totalTokenCount", 0),
"prompt_tokens": usage_metadata.get("promptTokenCount", 0),
"prompt_tokens_details": input_tokens_details.model_dump(),
"completion_tokens": output_tokens,
"completion_tokens_details": output_tokens_details.model_dump(),
"output_tokens_details": output_tokens_details.model_dump(),
}
return ImageUsage(**usage_payload)

View file

View file

View file

@ -0,0 +1,54 @@
"""
Translate from OpenAI's `/v1/chat/completions` to Inception's `/v1/chat/completions`
Inception Labs (https://www.inceptionlabs.ai) serves the Mercury family of
diffusion LLMs through an OpenAI-compatible API, so we only need to point the
OpenAI-like handler at the Inception API base and pick up the Inception API key.
"""
from typing import List, Optional, Tuple
import litellm
from litellm.secret_managers.main import get_secret_str
from ...openai_like.chat.transformation import OpenAILikeChatConfig
class InceptionChatConfig(OpenAILikeChatConfig):
"""
Inception is OpenAI-compatible with standard endpoints
"""
@property
def custom_llm_provider(self) -> Optional[str]:
return "inception"
def get_supported_openai_params(self, model: str) -> List:
return [
"max_tokens",
"max_completion_tokens",
"temperature",
"stop",
"tools",
"tool_choice",
"stream",
"stream_options",
"response_format",
"reasoning_effort",
"reasoning_summary",
"reasoning_summary_wait",
"diffusing",
"realtime",
]
def _get_openai_compatible_provider_info(
self, api_base: Optional[str], api_key: Optional[str]
) -> Tuple[Optional[str], Optional[str]]:
passed_api_base = api_base
api_base = api_base or get_secret_str("INCEPTION_API_BASE") or "https://api.inceptionlabs.ai/v1" # type: ignore
dynamic_api_key = api_key
if passed_api_base is None or api_key:
dynamic_api_key = (
api_key or litellm.inception_key or get_secret_str("INCEPTION_API_KEY")
)
return api_base, dynamic_api_key

View file

@ -0,0 +1,43 @@
"""
Inception fill-in-the-middle (FIM) completions.
Inception's FIM endpoint is OpenAI text-completion compatible: it takes a
`prompt` (prefix) plus an optional `suffix` and returns standard
`choices[].text`. It is served at `/v1/fim/completions` rather than
`/v1/completions`, so routing points the OpenAI client at the `/v1/fim` base
(see the `text-completion-inception` branch in `main.py`).
"""
from typing import List
from litellm.llms.openai.completion.transformation import OpenAITextCompletionConfig
class InceptionTextCompletionConfig(OpenAITextCompletionConfig):
def get_supported_openai_params(self, model: str) -> List:
return [
"suffix",
"max_tokens",
"max_completion_tokens",
"top_p",
"frequency_penalty",
"presence_penalty",
"stop",
"stream",
"stream_options",
]
def map_openai_params(
self,
non_default_params: dict,
optional_params: dict,
model: str,
drop_params: bool,
) -> dict:
supported_params = self.get_supported_openai_params(model)
for param, value in non_default_params.items():
if param == "max_completion_tokens":
optional_params["max_tokens"] = value
elif param in supported_params:
optional_params[param] = value
return optional_params

View file

@ -996,7 +996,19 @@ def _gemini_convert_messages_with_history( # noqa: PLR0915
excluded_keys=["thoughtSignature"],
):
assistant_content.append(gemini_tool_call_part)
last_message_with_tool_calls = assistant_msg
# Only record this as the active tool-call message when it actually
# carries tool calls. The `if` guard above is also entered for a
# text-only assistant message (`assistant_msg.get("tool_calls", [])
# is not None` is True for an empty list), so without this check a
# later assistant message with no tool calls would clobber the
# reference. The following tool result would then be matched against
# an assistant message that has no tool_calls, raising "Missing
# corresponding tool call for tool response message".
if (
assistant_msg.get("tool_calls")
or assistant_msg.get("function_call") is not None
):
last_message_with_tool_calls = assistant_msg
## HANDLE SERVER-SIDE TOOL INVOCATIONS (context circulation)
_psf = assistant_msg.get("provider_specific_fields")

View file

@ -641,6 +641,7 @@ async def acompletion( # noqa: PLR0915
if (
custom_llm_provider == "text-completion-openai"
or custom_llm_provider == "text-completion-codestral"
or custom_llm_provider == "text-completion-inception"
) and isinstance(response, TextCompletionResponse):
response = litellm.OpenAITextCompletionConfig().convert_to_chat_model_response_object(
response_object=response,
@ -3803,6 +3804,67 @@ def completion( # type: ignore # noqa: PLR0915
):
return _model_response
response = _model_response
elif custom_llm_provider == "text-completion-inception":
passed_api_base = (
api_base
or optional_params.pop("api_base", None)
or optional_params.pop("base_url", None)
)
api_base = (
passed_api_base
or get_secret_str("INCEPTION_API_BASE")
or "https://api.inceptionlabs.ai/v1"
)
# FIM is served at `/v1/fim/completions`; the OpenAI client appends
# `/completions`, so point it at the `/v1/fim` base.
api_base = api_base.rstrip("/")
if not api_base.endswith("/fim"):
api_base += "/fim"
# Don't forward the server-managed Inception key to a caller-supplied
# api_base; only resolve it for the default/server base, or when the
# caller passes their own key.
if passed_api_base is None or api_key:
api_key = (
api_key
or litellm.inception_key
or get_secret_str("INCEPTION_API_KEY")
)
_response = openai_text_completions.completion(
model=model,
messages=messages,
model_response=model_response,
print_verbose=print_verbose,
api_key=api_key, # type: ignore[arg-type]
custom_llm_provider="text-completion-inception",
api_base=api_base,
acompletion=acompletion,
client=client,
logging_obj=logging,
optional_params=optional_params,
litellm_params=litellm_params,
logger_fn=logger_fn,
timeout=timeout, # type: ignore
)
if (
optional_params.get("stream", False) is False
and acompletion is False
and text_completion is False
):
_response = litellm.OpenAITextCompletionConfig().convert_to_chat_model_response_object(
response_object=_response, model_response_object=model_response
)
if optional_params.get("stream", False) or acompletion is True:
logging.post_call(
input=messages,
api_key=api_key,
original_response=_response,
additional_args={"headers": headers},
)
response = _response
elif custom_llm_provider in ("sagemaker_chat", "sagemaker_nova"):
# boto3 reads keys from .env
# sagemaker_chat: HF Messages API endpoints

File diff suppressed because it is too large Load diff

View file

@ -319,36 +319,34 @@ async def create_agent(
Example Request:
```bash
curl -X POST "http://localhost:4000/agents" \\
curl -X POST "http://localhost:4000/v1/agents" \\
-H "Authorization: Bearer <your_api_key>" \\
-H "Content-Type: application/json" \\
-d '{
"agent": {
"agent_name": "my-custom-agent",
"agent_card_params": {
"protocolVersion": "1.0",
"name": "Hello World Agent",
"description": "Just a hello world agent",
"url": "http://localhost:9999/",
"version": "1.0.0",
"defaultInputModes": ["text"],
"defaultOutputModes": ["text"],
"capabilities": {
"streaming": true
},
"skills": [
{
"id": "hello_world",
"name": "Returns hello world",
"description": "just returns hello world",
"tags": ["hello world"],
"examples": ["hi", "hello world"]
}
]
"agent_name": "my-custom-agent",
"agent_card_params": {
"protocolVersion": "1.0",
"name": "Hello World Agent",
"description": "Just a hello world agent",
"url": "http://localhost:9999/",
"version": "1.0.0",
"defaultInputModes": ["text"],
"defaultOutputModes": ["text"],
"capabilities": {
"streaming": true
},
"litellm_params": {
"make_public": true
}
"skills": [
{
"id": "hello_world",
"name": "Returns hello world",
"description": "just returns hello world",
"tags": ["hello world"],
"examples": ["hi", "hello world"]
}
]
},
"litellm_params": {
"make_public": true
}
}'
```
@ -441,7 +439,7 @@ async def get_agent_by_id(
Example Request:
```bash
curl -X GET "http://localhost:4000/agents/123e4567-e89b-12d3-a456-426614174000" \\
curl -X GET "http://localhost:4000/v1/agents/123e4567-e89b-12d3-a456-426614174000" \\
-H "Authorization: Bearer <your_api_key>"
```
"""
@ -535,28 +533,26 @@ async def update_agent(
Example Request:
```bash
curl -X PUT "http://localhost:4000/agents/123e4567-e89b-12d3-a456-426614174000" \\
curl -X PUT "http://localhost:4000/v1/agents/123e4567-e89b-12d3-a456-426614174000" \\
-H "Authorization: Bearer <your_api_key>" \\
-H "Content-Type: application/json" \\
-d '{
"agent": {
"agent_name": "updated-agent",
"agent_card_params": {
"protocolVersion": "1.0",
"name": "Updated Agent",
"description": "Updated description",
"url": "http://localhost:9999/",
"version": "1.1.0",
"defaultInputModes": ["text"],
"defaultOutputModes": ["text"],
"capabilities": {
"streaming": true
},
"skills": []
"agent_name": "updated-agent",
"agent_card_params": {
"protocolVersion": "1.0",
"name": "Updated Agent",
"description": "Updated description",
"url": "http://localhost:9999/",
"version": "1.1.0",
"defaultInputModes": ["text"],
"defaultOutputModes": ["text"],
"capabilities": {
"streaming": true
},
"litellm_params": {
"make_public": false
}
"skills": []
},
"litellm_params": {
"make_public": false
}
}'
```
@ -645,28 +641,26 @@ async def patch_agent(
Example Request:
```bash
curl -X PUT "http://localhost:4000/agents/123e4567-e89b-12d3-a456-426614174000" \\
curl -X PATCH "http://localhost:4000/v1/agents/123e4567-e89b-12d3-a456-426614174000" \\
-H "Authorization: Bearer <your_api_key>" \\
-H "Content-Type: application/json" \\
-d '{
"agent": {
"agent_name": "updated-agent",
"agent_card_params": {
"protocolVersion": "1.0",
"name": "Updated Agent",
"description": "Updated description",
"url": "http://localhost:9999/",
"version": "1.1.0",
"defaultInputModes": ["text"],
"defaultOutputModes": ["text"],
"capabilities": {
"streaming": true
},
"skills": []
"agent_name": "updated-agent",
"agent_card_params": {
"protocolVersion": "1.0",
"name": "Updated Agent",
"description": "Updated description",
"url": "http://localhost:9999/",
"version": "1.1.0",
"defaultInputModes": ["text"],
"defaultOutputModes": ["text"],
"capabilities": {
"streaming": true
},
"litellm_params": {
"make_public": false
}
"skills": []
},
"litellm_params": {
"make_public": false
}
}'
```
@ -753,7 +747,7 @@ async def delete_agent(
Example Request:
```bash
curl -X DELETE "http://localhost:4000/agents/123e4567-e89b-12d3-a456-426614174000" \\
curl -X DELETE "http://localhost:4000/v1/agents/123e4567-e89b-12d3-a456-426614174000" \\
-H "Authorization: Bearer <your_api_key>"
```

View file

@ -1127,7 +1127,7 @@ async def get_end_user_object(
end_user_id: Optional[str],
prisma_client: Optional[PrismaClient],
user_api_key_cache: UserApiKeyCache,
route: str,
route: Optional[str] = "",
parent_otel_span: Optional[Span] = None,
proxy_logging_obj: Optional[ProxyLogging] = None,
) -> Optional[LiteLLM_EndUserTable]:
@ -1171,9 +1171,6 @@ async def get_end_user_object(
parent_otel_span=parent_otel_span,
)
# Check budget limits
await _check_end_user_budget(end_user_obj=return_obj, route=route)
return return_obj
# Fetch from database
@ -1204,14 +1201,9 @@ async def get_end_user_object(
model_type=LiteLLM_EndUserTable,
)
# Check budget limits
await _check_end_user_budget(end_user_obj=_response, route=route)
return _response
except Exception as e:
if isinstance(e, litellm.BudgetExceededError):
raise e
except Exception:
return None
@ -1308,8 +1300,6 @@ async def _end_user_id_exists_in_db(
)
if end_user_obj is not None:
return True
except litellm.BudgetExceededError:
raise
except Exception as e:
verbose_proxy_logger.debug(
f"end_user validation: get_end_user_object lookup failed: {e}"

View file

@ -30,6 +30,7 @@ from litellm.proxy._types import *
from litellm.proxy.auth.auth_checks import (
ExperimentalUIJWTToken,
_cache_key_object,
_check_end_user_budget,
_delete_cache_key_object,
_get_user_role,
_is_model_cost_zero,
@ -1762,8 +1763,7 @@ async def _user_api_key_auth_builder( # noqa: PLR0915
async def _safe_fetch(label: str, awaitable):
"""Run an awaitable and return its result. Re-raises authentication /
authorization failures (HTTPException, ProxyException,
BudgetExceededError which ``get_end_user_object`` raises for
end-user budget violations) so they propagate to the caller.
BudgetExceededError) so they propagate to the caller.
Other exceptions (e.g. transient DB errors fetching context) are
swallowed with a debug log and ``None`` is returned so
``common_checks`` can still run against whatever limits are recorded
@ -2584,6 +2584,14 @@ async def _run_post_custom_auth_checks(
user_api_key_cache=user_api_key_cache,
proxy_logging_obj=proxy_logging_obj,
)
# common_checks() enforces the end-user budget, but the centralized
# gate skips it for custom-auth deployments unless
# custom_auth_run_common_checks is set. Enforce it here on that path
# so an over-budget end user can't keep making requests.
if end_user_object is not None and not general_settings.get(
"custom_auth_run_common_checks", False
):
await _check_end_user_budget(end_user_obj=end_user_object, route=route)
# 2. Check token expiry
if valid_token.expires is not None:

View file

@ -9100,7 +9100,10 @@ class Router:
except Exception:
pass
# Three mutually exclusive scenarios for the model's metadata:
if custom_model_info is not None and litellm_model_name_model_info is not None:
# (1) It has both custom model_info set and exists in the built-in map
# merge with custom overriding built-in
model_info = cast(
ModelInfo,
_update_dictionary(
@ -9109,7 +9112,12 @@ class Router:
),
)
elif litellm_model_name_model_info is not None:
# (2) Built-in only — no custom pricing to merge
model_info = litellm_model_name_model_info
elif custom_model_info is not None:
# (3) Custom only — model not in built-in cost map yet
# custom_model_info already includes base_model defaults at this point, if applicable
model_info = cast(ModelInfo, custom_model_info)
return model_info

View file

@ -20,6 +20,7 @@ class ImageEditOptionalRequestParams(TypedDict, total=False):
response_format: Optional[Literal["url", "b64_json"]]
size: Optional[str]
user: Optional[str]
imageConfig: Optional[Dict[str, Any]]
class ImageEditRequestParams(ImageEditOptionalRequestParams, total=False):

View file

@ -1,7 +1,7 @@
from enum import Enum
from typing import Any, Dict, Iterable, List, Literal, Optional, Union
from typing import Any, Dict, List, Literal, Optional
from typing_extensions import Required, TypedDict
from typing_extensions import TypedDict
from .vertex_ai import (
GenerationConfig,
@ -171,6 +171,9 @@ class GeminiImageGenerationParameters(BaseModel):
aspectRatio: Optional[str] = None
"""Aspect ratio for generated images (e.g., '1:1', '16:9', '9:16', '4:3', '3:4')"""
imageSize: Optional[str] = None
"""Image size for generated images (e.g., '1K', '2K')"""
personGeneration: Optional[str] = None
"""Controls person generation in images"""

View file

@ -1084,6 +1084,7 @@ OpenAIImageGenerationOptionalParams = Literal[
"image_url",
"image_prompt_strength",
"aspect_ratio",
"imageConfig",
]
OpenAIImageEditOptionalParams = Literal[

View file

@ -20,6 +20,7 @@ class FunctionResponse(TypedDict, total=False):
id: str
name: Required[str]
response: Optional[dict]
parts: List["FunctionResponsePartType"]
class FunctionCall(TypedDict, total=False):
@ -40,6 +41,11 @@ class BlobType(TypedDict, total=False):
data: Required[str]
class FunctionResponsePartType(TypedDict, total=False):
inline_data: BlobType
file_data: FileDataType
class PartType(TypedDict, total=False):
text: str
inline_data: BlobType

View file

@ -148,6 +148,7 @@ class ProviderSpecificModelInfo(TypedDict, total=False):
supports_xhigh_reasoning_effort: Optional[bool]
supports_max_reasoning_effort: Optional[bool]
supports_output_config: Optional[bool]
supports_image_size: Optional[bool]
bedrock_output_config_effort_ceiling: Optional[
Literal["low", "medium", "high", "max", "xhigh"]
]
@ -3300,6 +3301,8 @@ class LlmProviders(str, Enum):
V0 = "v0"
MORPH = "morph"
LAMBDA_AI = "lambda_ai"
INCEPTION = "inception"
TEXT_COMPLETION_INCEPTION = "text-completion-inception"
DEEPSEEK = "deepseek"
SAMBANOVA = "sambanova"
MARITALK = "maritalk"

View file

@ -3147,6 +3147,7 @@ def get_optional_params_image_gen(
size: Optional[str] = None,
style: Optional[str] = None,
user: Optional[str] = None,
imageConfig: Optional[dict] = None,
custom_llm_provider: Optional[str] = None,
additional_drop_params: Optional[list] = None,
provider_config: Optional[BaseImageGenerationConfig] = None,
@ -3183,6 +3184,7 @@ def get_optional_params_image_gen(
"size": None,
"style": None,
"user": None,
"imageConfig": None,
}
non_default_params = _get_non_default_params(
@ -4547,6 +4549,18 @@ def get_optional_params( # noqa: PLR0915
),
)
elif custom_llm_provider == "text-completion-inception":
optional_params = litellm.InceptionTextCompletionConfig().map_openai_params(
non_default_params=non_default_params,
optional_params=optional_params,
model=model,
drop_params=(
drop_params
if drop_params is not None and isinstance(drop_params, bool)
else False
),
)
elif custom_llm_provider == "databricks":
optional_params = litellm.DatabricksConfig().map_openai_params(
non_default_params=non_default_params,
@ -6083,6 +6097,7 @@ def _get_model_info_helper( # noqa: PLR0915
"provider_specific_entry", None
),
uses_embed_content=_model_info.get("uses_embed_content", None),
supports_image_size=_model_info.get("supports_image_size", None),
)
except Exception as e:
verbose_logger.debug(f"Error getting model info: {e}")
@ -6637,6 +6652,14 @@ def validate_environment( # noqa: PLR0915
keys_in_environment = True
else:
missing_keys.append("CODESTRAL_API_KEY")
elif (
custom_llm_provider == "inception"
or custom_llm_provider == "text-completion-inception"
):
if "INCEPTION_API_KEY" in os.environ:
keys_in_environment = True
else:
missing_keys.append("INCEPTION_API_KEY")
elif custom_llm_provider == "deepseek":
if "DEEPSEEK_API_KEY" in os.environ:
keys_in_environment = True
@ -8291,6 +8314,7 @@ class ProviderConfigManager:
LlmProviders.XAI: (lambda: litellm.XAIChatConfig(), False),
LlmProviders.ZAI: (lambda: litellm.ZAIChatConfig(), False),
LlmProviders.LAMBDA_AI: (lambda: litellm.LambdaAIChatConfig(), False),
LlmProviders.INCEPTION: (lambda: litellm.InceptionChatConfig(), False),
LlmProviders.LLAMA: (lambda: litellm.LlamaAPIConfig(), False),
LlmProviders.TEXT_COMPLETION_OPENAI: (
lambda: litellm.OpenAITextCompletionConfig(),
@ -8356,6 +8380,10 @@ class ProviderConfigManager:
lambda: litellm.CodestralTextCompletionConfig(),
False,
),
LlmProviders.TEXT_COMPLETION_INCEPTION: (
lambda: litellm.InceptionTextCompletionConfig(),
False,
),
LlmProviders.SAMBANOVA: (lambda: litellm.SambanovaConfig(), False),
LlmProviders.MARITALK: (lambda: litellm.MaritalkConfig(), False),
LlmProviders.VLLM: (lambda: litellm.VLLMConfig(), False),
@ -8928,6 +8956,8 @@ class ProviderConfigManager:
return litellm.FireworksAITextCompletionConfig()
elif LlmProviders.TOGETHER_AI == provider:
return litellm.TogetherAITextCompletionConfig()
elif LlmProviders.TEXT_COMPLETION_INCEPTION == provider:
return litellm.InceptionTextCompletionConfig()
return litellm.OpenAITextCompletionConfig()
@staticmethod

View file

@ -1075,6 +1075,7 @@
},
"eu.anthropic.claude-opus-4-6-v1": {
"cache_creation_input_token_cost": 6.875e-06,
"cache_creation_input_token_cost_above_1hr": 1.1e-05,
"cache_read_input_token_cost": 5.5e-07,
"input_cost_per_token": 5.5e-06,
"litellm_provider": "bedrock_converse",
@ -1104,6 +1105,7 @@
},
"au.anthropic.claude-opus-4-6-v1": {
"cache_creation_input_token_cost": 6.875e-06,
"cache_creation_input_token_cost_above_1hr": 1.1e-05,
"cache_read_input_token_cost": 5.5e-07,
"input_cost_per_token": 5.5e-06,
"litellm_provider": "bedrock_converse",
@ -1241,6 +1243,7 @@
},
"eu.anthropic.claude-opus-4-7": {
"cache_creation_input_token_cost": 6.875e-06,
"cache_creation_input_token_cost_above_1hr": 1.1e-05,
"cache_read_input_token_cost": 5.5e-07,
"input_cost_per_token": 5.5e-06,
"litellm_provider": "bedrock_converse",
@ -1271,6 +1274,7 @@
},
"au.anthropic.claude-opus-4-7": {
"cache_creation_input_token_cost": 6.875e-06,
"cache_creation_input_token_cost_above_1hr": 1.1e-05,
"cache_read_input_token_cost": 5.5e-07,
"input_cost_per_token": 5.5e-06,
"litellm_provider": "bedrock_converse",
@ -1543,6 +1547,7 @@
},
"eu.anthropic.claude-sonnet-4-6": {
"cache_creation_input_token_cost": 4.125e-06,
"cache_creation_input_token_cost_above_1hr": 6.6e-06,
"cache_read_input_token_cost": 3.3e-07,
"input_cost_per_token": 3.3e-06,
"litellm_provider": "bedrock_converse",
@ -1571,6 +1576,7 @@
},
"au.anthropic.claude-sonnet-4-6": {
"cache_creation_input_token_cost": 4.125e-06,
"cache_creation_input_token_cost_above_1hr": 6.6e-06,
"cache_read_input_token_cost": 3.3e-07,
"input_cost_per_token": 3.3e-06,
"litellm_provider": "bedrock_converse",
@ -1599,6 +1605,7 @@
},
"jp.anthropic.claude-sonnet-4-6": {
"cache_creation_input_token_cost": 4.125e-06,
"cache_creation_input_token_cost_above_1hr": 6.6e-06,
"cache_read_input_token_cost": 3.3e-07,
"input_cost_per_token": 3.3e-06,
"litellm_provider": "bedrock_converse",
@ -1995,11 +2002,13 @@
},
"au.anthropic.claude-sonnet-4-5-20250929-v1:0": {
"cache_creation_input_token_cost": 4.125e-06,
"cache_creation_input_token_cost_above_1hr": 6.6e-06,
"cache_read_input_token_cost": 3.3e-07,
"input_cost_per_token": 3.3e-06,
"input_cost_per_token_above_200k_tokens": 6.6e-06,
"output_cost_per_token_above_200k_tokens": 2.475e-05,
"cache_creation_input_token_cost_above_200k_tokens": 8.25e-06,
"cache_creation_input_token_cost_above_1hr_above_200k_tokens": 1.32e-05,
"cache_read_input_token_cost_above_200k_tokens": 6.6e-07,
"litellm_provider": "bedrock_converse",
"max_input_tokens": 200000,
@ -7494,6 +7503,27 @@
"supports_video_input": true,
"supports_vision": true
},
"azure_ai/kimi-k2.6": {
"input_cost_per_token": 9.5e-07,
"litellm_provider": "azure_ai",
"max_input_tokens": 262144,
"max_output_tokens": 262144,
"max_tokens": 262144,
"mode": "chat",
"output_cost_per_token": 4e-06,
"source": "https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/introducing-kimi-k2-6-in-microsoft-foundry/4513125",
"supported_modalities": [
"text",
"image"
],
"supported_output_modalities": [
"text"
],
"supports_function_calling": true,
"supports_reasoning": true,
"supports_tool_choice": true,
"supports_vision": true
},
"azure_ai/ministral-3b": {
"input_cost_per_token": 4e-08,
"litellm_provider": "azure_ai",
@ -12682,7 +12712,8 @@
"litellm_provider": "deepinfra",
"mode": "chat",
"supports_tool_choice": true,
"supports_function_calling": true
"supports_function_calling": true,
"supports_image_size": false
},
"deepinfra/google/gemini-2.5-pro": {
"max_tokens": 1000000,
@ -13583,6 +13614,7 @@
},
"eu.anthropic.claude-haiku-4-5-20251001-v1:0": {
"cache_creation_input_token_cost": 1.375e-06,
"cache_creation_input_token_cost_above_1hr": 2.2e-06,
"cache_read_input_token_cost": 1.1e-07,
"input_cost_per_token": 1.1e-06,
"deprecation_date": "2026-10-15",
@ -13787,11 +13819,13 @@
},
"eu.anthropic.claude-sonnet-4-5-20250929-v1:0": {
"cache_creation_input_token_cost": 4.125e-06,
"cache_creation_input_token_cost_above_1hr": 6.6e-06,
"cache_read_input_token_cost": 3.3e-07,
"input_cost_per_token": 3.3e-06,
"input_cost_per_token_above_200k_tokens": 6.6e-06,
"output_cost_per_token_above_200k_tokens": 2.475e-05,
"cache_creation_input_token_cost_above_200k_tokens": 8.25e-06,
"cache_creation_input_token_cost_above_1hr_above_200k_tokens": 1.32e-05,
"cache_read_input_token_cost_above_200k_tokens": 6.6e-07,
"litellm_provider": "bedrock_converse",
"max_input_tokens": 200000,
@ -15006,7 +15040,8 @@
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
},
"supports_service_tier": true
"supports_service_tier": true,
"supports_image_size": false
},
"gemini-2.5-flash-image": {
"cache_read_input_token_cost": 3e-08,
@ -15056,7 +15091,8 @@
"supports_vision": true,
"supports_web_search": false,
"tpm": 8000000,
"supports_service_tier": true
"supports_service_tier": true,
"supports_image_size": false
},
"gemini-3-pro-image-preview": {
"input_cost_per_image": 0.0011,
@ -15345,7 +15381,8 @@
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
},
"supports_service_tier": true
"supports_service_tier": true,
"supports_image_size": false
},
"gemini-2.5-flash-lite-preview-09-2025": {
"cache_read_input_token_cost": 1e-08,
@ -15395,7 +15432,8 @@
"search_context_size_low": 0.035,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
}
},
"supports_image_size": false
},
"gemini-2.5-flash-preview-09-2025": {
"cache_read_input_token_cost": 7.5e-08,
@ -15445,7 +15483,8 @@
"search_context_size_low": 0.035,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
}
},
"supports_image_size": false
},
"gemini-live-2.5-flash-preview-native-audio-09-2025": {
"cache_read_input_token_cost": 7.5e-08,
@ -15596,7 +15635,8 @@
"search_context_size_low": 0.035,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
}
},
"supports_image_size": false
},
"gemini-2.5-pro": {
"cache_read_input_token_cost": 1.25e-07,
@ -16606,7 +16646,8 @@
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
},
"supports_service_tier": true
"supports_service_tier": true,
"supports_image_size": false
},
"gemini/gemini-2.5-flash-image": {
"cache_read_input_token_cost": 3e-08,
@ -16662,7 +16703,8 @@
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
},
"supports_service_tier": true
"supports_service_tier": true,
"supports_image_size": false
},
"gemini/gemini-3-pro-image-preview": {
"input_cost_per_image": 0.0011,
@ -16841,7 +16883,8 @@
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
},
"supports_service_tier": true
"supports_service_tier": true,
"supports_image_size": false
},
"gemini/gemini-2.5-flash-lite-preview-09-2025": {
"cache_read_input_token_cost": 1e-08,
@ -16893,7 +16936,8 @@
"search_context_size_low": 0.035,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
}
},
"supports_image_size": false
},
"gemini/gemini-2.5-flash-preview-09-2025": {
"cache_read_input_token_cost": 7.5e-08,
@ -16945,7 +16989,8 @@
"search_context_size_low": 0.035,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
}
},
"supports_image_size": false
},
"gemini/gemini-flash-latest": {
"cache_read_input_token_cost": 7.5e-08,
@ -17102,7 +17147,8 @@
"search_context_size_low": 0.035,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.035
}
},
"supports_image_size": false
},
"gemini/gemini-2.5-flash-preview-tts": {
"input_cost_per_token": 3e-07,
@ -23086,11 +23132,13 @@
},
"jp.anthropic.claude-sonnet-4-5-20250929-v1:0": {
"cache_creation_input_token_cost": 4.125e-06,
"cache_creation_input_token_cost_above_1hr": 6.6e-06,
"cache_read_input_token_cost": 3.3e-07,
"input_cost_per_token": 3.3e-06,
"input_cost_per_token_above_200k_tokens": 6.6e-06,
"output_cost_per_token_above_200k_tokens": 2.475e-05,
"cache_creation_input_token_cost_above_200k_tokens": 8.25e-06,
"cache_creation_input_token_cost_above_1hr_above_200k_tokens": 1.32e-05,
"cache_read_input_token_cost_above_200k_tokens": 6.6e-07,
"litellm_provider": "bedrock_converse",
"max_input_tokens": 200000,
@ -23116,6 +23164,7 @@
},
"jp.anthropic.claude-haiku-4-5-20251001-v1:0": {
"cache_creation_input_token_cost": 1.375e-06,
"cache_creation_input_token_cost_above_1hr": 2.2e-06,
"cache_read_input_token_cost": 1.1e-07,
"input_cost_per_token": 1.1e-06,
"litellm_provider": "bedrock_converse",
@ -23228,6 +23277,31 @@
"supports_system_messages": true,
"supports_tool_choice": true
},
"inception/mercury-2": {
"cache_read_input_token_cost": 2.5e-08,
"input_cost_per_token": 2.5e-07,
"litellm_provider": "inception",
"max_input_tokens": 128000,
"max_output_tokens": 50000,
"max_tokens": 50000,
"mode": "chat",
"output_cost_per_token": 7.5e-07,
"supports_function_calling": true,
"supports_prompt_caching": true,
"supports_response_schema": true,
"supports_system_messages": true,
"supports_tool_choice": true
},
"text-completion-inception/mercury-edit-2": {
"cache_read_input_token_cost": 2.5e-08,
"input_cost_per_token": 2.5e-07,
"litellm_provider": "text-completion-inception",
"max_input_tokens": 32000,
"max_output_tokens": 8192,
"max_tokens": 8192,
"mode": "completion",
"output_cost_per_token": 7.5e-07
},
"lambda_ai/deepseek-llama3.3-70b": {
"input_cost_per_token": 2e-07,
"litellm_provider": "lambda_ai",
@ -26449,7 +26523,8 @@
"supports_function_calling": true,
"supports_response_schema": true,
"supports_vision": true,
"supports_native_streaming": true
"supports_native_streaming": true,
"supports_image_size": false
},
"oci/google.gemini-2.5-pro": {
"input_cost_per_token": 1.25e-06,
@ -26477,7 +26552,8 @@
"supports_function_calling": true,
"supports_response_schema": true,
"supports_vision": true,
"supports_native_streaming": true
"supports_native_streaming": true,
"supports_image_size": false
},
"oci/cohere.command-a-vision": {
"input_cost_per_token": 1.56e-06,
@ -27513,7 +27589,8 @@
"supports_response_schema": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_vision": true
"supports_vision": true,
"supports_image_size": false
},
"openrouter/google/gemini-2.5-pro": {
"input_cost_per_audio_token": 7e-07,
@ -29383,7 +29460,8 @@
"mode": "responses",
"supports_web_search": true,
"supports_reasoning": false,
"supports_function_calling": true
"supports_function_calling": true,
"supports_image_size": false
},
"perplexity/xai/grok-4-1-fast-non-reasoning": {
"litellm_provider": "perplexity",
@ -29965,7 +30043,8 @@
"supports_vision": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_response_schema": true
"supports_response_schema": true,
"supports_image_size": false
},
"replicate/openai/gpt-oss-120b": {
"input_cost_per_token": 1.8e-07,
@ -31784,6 +31863,7 @@
},
"au.anthropic.claude-haiku-4-5-20251001-v1:0": {
"cache_creation_input_token_cost": 1.375e-06,
"cache_creation_input_token_cost_above_1hr": 2.2e-06,
"cache_read_input_token_cost": 1.1e-07,
"input_cost_per_token": 1.1e-06,
"litellm_provider": "bedrock_converse",
@ -32663,7 +32743,8 @@
"supports_vision": true,
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_response_schema": true
"supports_response_schema": true,
"supports_image_size": false
},
"vercel_ai_gateway/google/gemini-2.5-pro": {
"input_cost_per_token": 2.5e-06,
@ -33430,6 +33511,7 @@
},
"vertex_ai/claude-haiku-4-5": {
"cache_creation_input_token_cost": 1.25e-06,
"cache_creation_input_token_cost_above_1hr": 2e-06,
"cache_read_input_token_cost": 1e-07,
"input_cost_per_token": 1e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33451,6 +33533,7 @@
},
"vertex_ai/claude-haiku-4-5@20251001": {
"cache_creation_input_token_cost": 1.25e-06,
"cache_creation_input_token_cost_above_1hr": 2e-06,
"cache_read_input_token_cost": 1e-07,
"input_cost_per_token": 1e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33501,6 +33584,7 @@
},
"vertex_ai/claude-3-7-sonnet@20250219": {
"cache_creation_input_token_cost": 3.75e-06,
"cache_creation_input_token_cost_above_1hr": 6e-06,
"cache_read_input_token_cost": 3e-07,
"deprecation_date": "2026-05-11",
"input_cost_per_token": 3e-06,
@ -33600,6 +33684,7 @@
},
"vertex_ai/claude-opus-4": {
"cache_creation_input_token_cost": 1.875e-05,
"cache_creation_input_token_cost_above_1hr": 3e-05,
"cache_read_input_token_cost": 1.5e-06,
"input_cost_per_token": 1.5e-05,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33625,6 +33710,7 @@
},
"vertex_ai/claude-opus-4-1": {
"cache_creation_input_token_cost": 1.875e-05,
"cache_creation_input_token_cost_above_1hr": 3e-05,
"cache_read_input_token_cost": 1.5e-06,
"input_cost_per_token": 1.5e-05,
"input_cost_per_token_batches": 7.5e-06,
@ -33642,6 +33728,7 @@
},
"vertex_ai/claude-opus-4-1@20250805": {
"cache_creation_input_token_cost": 1.875e-05,
"cache_creation_input_token_cost_above_1hr": 3e-05,
"cache_read_input_token_cost": 1.5e-06,
"input_cost_per_token": 1.5e-05,
"input_cost_per_token_batches": 7.5e-06,
@ -33659,6 +33746,7 @@
},
"vertex_ai/claude-opus-4-5": {
"cache_creation_input_token_cost": 6.25e-06,
"cache_creation_input_token_cost_above_1hr": 1e-05,
"cache_read_input_token_cost": 5e-07,
"input_cost_per_token": 5e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33685,6 +33773,7 @@
},
"vertex_ai/claude-opus-4-5@20251101": {
"cache_creation_input_token_cost": 6.25e-06,
"cache_creation_input_token_cost_above_1hr": 1e-05,
"cache_read_input_token_cost": 5e-07,
"input_cost_per_token": 5e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33712,6 +33801,7 @@
},
"vertex_ai/claude-opus-4-6": {
"cache_creation_input_token_cost": 6.25e-06,
"cache_creation_input_token_cost_above_1hr": 1e-05,
"cache_read_input_token_cost": 5e-07,
"input_cost_per_token": 5e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33739,6 +33829,7 @@
},
"vertex_ai/claude-opus-4-6@default": {
"cache_creation_input_token_cost": 6.25e-06,
"cache_creation_input_token_cost_above_1hr": 1e-05,
"cache_read_input_token_cost": 5e-07,
"input_cost_per_token": 5e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33766,6 +33857,7 @@
},
"vertex_ai/claude-opus-4-7": {
"cache_creation_input_token_cost": 6.25e-06,
"cache_creation_input_token_cost_above_1hr": 1e-05,
"cache_read_input_token_cost": 5e-07,
"input_cost_per_token": 5e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33793,6 +33885,7 @@
},
"vertex_ai/claude-opus-4-7@default": {
"cache_creation_input_token_cost": 6.25e-06,
"cache_creation_input_token_cost_above_1hr": 1e-05,
"cache_read_input_token_cost": 5e-07,
"input_cost_per_token": 5e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33876,6 +33969,7 @@
},
"vertex_ai/claude-sonnet-4-5": {
"cache_creation_input_token_cost": 3.75e-06,
"cache_creation_input_token_cost_above_1hr": 6e-06,
"cache_read_input_token_cost": 3e-07,
"input_cost_per_token": 3e-06,
"input_cost_per_token_above_200k_tokens": 6e-06,
@ -33902,6 +33996,7 @@
},
"vertex_ai/claude-sonnet-4-6": {
"cache_creation_input_token_cost": 3.75e-06,
"cache_creation_input_token_cost_above_1hr": 6e-06,
"cache_read_input_token_cost": 3e-07,
"input_cost_per_token": 3e-06,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33929,6 +34024,7 @@
},
"vertex_ai/claude-sonnet-4-5@20250929": {
"cache_creation_input_token_cost": 3.75e-06,
"cache_creation_input_token_cost_above_1hr": 6e-06,
"cache_read_input_token_cost": 3e-07,
"input_cost_per_token": 3e-06,
"input_cost_per_token_above_200k_tokens": 6e-06,
@ -33956,6 +34052,7 @@
},
"vertex_ai/claude-opus-4@20250514": {
"cache_creation_input_token_cost": 1.875e-05,
"cache_creation_input_token_cost_above_1hr": 3e-05,
"cache_read_input_token_cost": 1.5e-06,
"input_cost_per_token": 1.5e-05,
"litellm_provider": "vertex_ai-anthropic_models",
@ -33981,6 +34078,7 @@
},
"vertex_ai/claude-sonnet-4": {
"cache_creation_input_token_cost": 3.75e-06,
"cache_creation_input_token_cost_above_1hr": 6e-06,
"cache_read_input_token_cost": 3e-07,
"input_cost_per_token": 3e-06,
"input_cost_per_token_above_200k_tokens": 6e-06,
@ -34010,6 +34108,7 @@
},
"vertex_ai/claude-sonnet-4@20250514": {
"cache_creation_input_token_cost": 3.75e-06,
"cache_creation_input_token_cost_above_1hr": 6e-06,
"cache_read_input_token_cost": 3e-07,
"input_cost_per_token": 3e-06,
"input_cost_per_token_above_200k_tokens": 6e-06,
@ -34217,7 +34316,8 @@
"supports_url_context": true,
"supports_vision": true,
"supports_web_search": false,
"tpm": 8000000
"tpm": 8000000,
"supports_image_size": false
},
"vertex_ai/gemini-3-pro-image-preview": {
"input_cost_per_image": 0.0011,
@ -41035,6 +41135,7 @@
},
"vertex_ai/claude-sonnet-4-6@default": {
"cache_creation_input_token_cost": 3.75e-06,
"cache_creation_input_token_cost_above_1hr": 6e-06,
"cache_read_input_token_cost": 3e-07,
"input_cost_per_token": 3e-06,
"litellm_provider": "vertex_ai-anthropic_models",

View file

@ -1273,6 +1273,24 @@
"interactions": true
}
},
"inception": {
"display_name": "Inception (`inception`)",
"url": "https://docs.litellm.ai/docs/providers/inception",
"endpoints": {
"chat_completions": true,
"messages": true,
"responses": true,
"embeddings": false,
"image_generations": false,
"audio_transcriptions": false,
"audio_speech": false,
"moderations": false,
"batches": false,
"rerank": false,
"a2a": true,
"interactions": true
}
},
"infinity": {
"display_name": "Infinity (`infinity`)",
"url": "https://docs.litellm.ai/docs/providers/infinity",

View file

@ -17,6 +17,66 @@ from litellm import completion
import json
GEMINI_3_IMAGE_SIZE_MAPPINGS = [
("512x512", "1:1", "512"),
("1024x1024", "1:1", "1K"),
("2048x2048", "1:1", "2K"),
("4096x4096", "1:1", "4K"),
("256x1024", "1:4", "512"),
("512x2048", "1:4", "1K"),
("1024x4096", "1:4", "2K"),
("2048x8192", "1:4", "4K"),
("192x1536", "1:8", "512"),
("384x3072", "1:8", "1K"),
("768x6144", "1:8", "2K"),
("1536x12288", "1:8", "4K"),
("424x632", "2:3", "512"),
("848x1264", "2:3", "1K"),
("1696x2528", "2:3", "2K"),
("3392x5056", "2:3", "4K"),
("632x424", "3:2", "512"),
("1264x848", "3:2", "1K"),
("2528x1696", "3:2", "2K"),
("5056x3392", "3:2", "4K"),
("448x600", "3:4", "512"),
("896x1200", "3:4", "1K"),
("1792x2400", "3:4", "2K"),
("3584x4800", "3:4", "4K"),
("1024x256", "4:1", "512"),
("2048x512", "4:1", "1K"),
("4096x1024", "4:1", "2K"),
("8192x2048", "4:1", "4K"),
("600x448", "4:3", "512"),
("1200x896", "4:3", "1K"),
("2400x1792", "4:3", "2K"),
("4800x3584", "4:3", "4K"),
("464x576", "4:5", "512"),
("928x1152", "4:5", "1K"),
("1856x2304", "4:5", "2K"),
("3712x4608", "4:5", "4K"),
("576x464", "5:4", "512"),
("1152x928", "5:4", "1K"),
("2304x1856", "5:4", "2K"),
("4608x3712", "5:4", "4K"),
("1536x192", "8:1", "512"),
("3072x384", "8:1", "1K"),
("6144x768", "8:1", "2K"),
("12288x1536", "8:1", "4K"),
("384x688", "9:16", "512"),
("768x1376", "9:16", "1K"),
("1536x2752", "9:16", "2K"),
("3072x5504", "9:16", "4K"),
("688x384", "16:9", "512"),
("1376x768", "16:9", "1K"),
("2752x1536", "16:9", "2K"),
("5504x3072", "16:9", "4K"),
("792x336", "21:9", "512"),
("1584x672", "21:9", "1K"),
("3168x1344", "21:9", "2K"),
("6336x2688", "21:9", "4K"),
]
class TestGoogleAIStudioGemini(BaseLLMChatTest):
def get_base_completion_call_args(self) -> dict:
return {"model": "gemini/gemini-2.5-flash"}
@ -365,6 +425,143 @@ def test_gemini_flash_image_preview_models(model_name: str):
]
@pytest.mark.parametrize(
"model, kwargs, expected_image_config",
[
(
"gemini/gemini-3-pro-image-preview",
{"imageConfig": {"aspectRatio": "16:9", "imageSize": "512px"}},
{"aspectRatio": "16:9", "imageSize": "512px"},
),
(
"gemini/gemini-2.5-flash-image",
{"size": "2048x2048"},
{"aspectRatio": "1:1"},
),
],
)
def test_gemini_image_generation_forwards_image_config(
model: str, kwargs: dict, expected_image_config: dict
):
from unittest.mock import patch, MagicMock
with patch(
"litellm.llms.custom_httpx.llm_http_handler.HTTPHandler.post"
) as mock_post:
mock_http_response = MagicMock()
mock_http_response.json.return_value = {
"candidates": [
{
"content": {
"parts": [{"inlineData": {"data": "test_base64_image_data"}}]
}
}
]
}
mock_http_response.status_code = 200
mock_post.return_value = mock_http_response
litellm.image_generation(
model=model,
prompt="Generate a simple test image",
api_key="test_api_key",
**kwargs,
)
request_data = mock_post.call_args.kwargs.get("json", {})
assert request_data["generationConfig"]["imageConfig"] == expected_image_config
def test_gemini_image_generation_image_config_takes_precedence_over_size():
from litellm.llms.gemini.image_generation.transformation import GoogleImageGenConfig
explicit_image_config = {"aspectRatio": "16:9", "imageSize": "2K"}
mapped_params = GoogleImageGenConfig().map_openai_params(
non_default_params={
"imageConfig": explicit_image_config,
"size": "768x1376",
},
optional_params={},
model="gemini-3-pro-image-preview",
drop_params=False,
)
assert mapped_params["imageConfig"] == explicit_image_config
def test_gemini_image_generation_ignores_non_dict_image_config():
from litellm.llms.gemini.image_generation.transformation import GoogleImageGenConfig
mapped_params = GoogleImageGenConfig().map_openai_params(
non_default_params={
"size": "768x1376",
"imageConfig": "not-a-dict",
},
optional_params={},
model="gemini-3-pro-image-preview",
drop_params=False,
)
assert mapped_params["imageConfig"] == {"aspectRatio": "9:16", "imageSize": "1K"}
@pytest.mark.parametrize(
"size, expected_aspect_ratio, expected_image_size",
GEMINI_3_IMAGE_SIZE_MAPPINGS,
)
def test_gemini_image_generation_openai_size_maps_to_google_table(
size: str, expected_aspect_ratio: str, expected_image_size: str
):
from litellm.llms.gemini.common_utils import (
map_openai_size_to_gemini_image_config,
)
assert map_openai_size_to_gemini_image_config(
size, "gemini-3-pro-image-preview"
) == {
"aspectRatio": expected_aspect_ratio,
"imageSize": expected_image_size,
}
@pytest.mark.parametrize(
"size, expected_aspect_ratio, expected_image_size",
[
("1000x1800", "9:16", "1K"),
("1800x1000", "16:9", "1K"),
("3000x3000", "1:1", "2K"),
("500x500", "1:1", "512"),
("1280x896", "4:3", "1K"),
("896x1280", "3:4", "1K"),
],
)
def test_gemini_image_generation_openai_size_snaps_to_nearest_option(
size: str, expected_aspect_ratio: str, expected_image_size: str
):
from litellm.llms.gemini.common_utils import (
map_openai_size_to_gemini_image_config,
)
assert map_openai_size_to_gemini_image_config(
size, "gemini-3-pro-image-preview"
) == {
"aspectRatio": expected_aspect_ratio,
"imageSize": expected_image_size,
}
@pytest.mark.parametrize("size", ["auto", "invalid", "0x1024", "1024x0"])
def test_gemini_image_generation_openai_size_auto_uses_google_defaults(size: str):
from litellm.llms.gemini.common_utils import (
map_openai_size_to_gemini_image_config,
)
assert map_openai_size_to_gemini_image_config(
size, "gemini-3-pro-image-preview"
) is None
def test_gemini_imagen_models_use_predict_endpoint():
"""
Test that Imagen models still use :predict endpoint (not broken by gemini-2.5-flash-image-preview fix)
@ -387,6 +584,7 @@ def test_gemini_imagen_models_use_predict_endpoint():
response = litellm.image_generation(
model="gemini/imagen-3.0-generate-001",
prompt="Generate a simple test image",
size="1280x896",
api_key="test_api_key",
)
@ -410,6 +608,9 @@ def test_gemini_imagen_models_use_predict_endpoint():
request_data = call_args.kwargs.get("json", {})
assert "instances" in request_data
assert "parameters" in request_data
assert request_data["parameters"]["aspectRatio"] == "4:3"
assert request_data["parameters"]["imageSize"] == "1K"
assert "imageConfig" not in request_data["parameters"]
def test_gemini_thinking():

View file

@ -38,8 +38,12 @@ from litellm.proxy.utils import CallInfo
@pytest.mark.asyncio
async def test_get_end_user_object(customer_spend, customer_budget):
"""
Scenario 1: normal
Scenario 2: user over budget
Scenario 1: normal - get_end_user_object returns the cached user
Scenario 2: user over budget - NOTE: budget enforcement now happens in
common_checks() via _check_end_user_budget(), not in get_end_user_object()
This test verifies that get_end_user_object correctly retrieves the end user
from cache. Budget enforcement is tested separately in test_check_end_user_budget().
"""
end_user_id = "my-test-customer"
_budget = LiteLLM_BudgetTable(max_budget=customer_budget)
@ -58,31 +62,62 @@ async def test_get_end_user_object(customer_spend, customer_budget):
value=end_user_obj,
model_type=LiteLLM_EndUserTable,
)
# get_end_user_object only fetches data - it no longer enforces budget
# Budget enforcement happens in common_checks() via _check_end_user_budget()
result = await get_end_user_object(
end_user_id=end_user_id,
prisma_client="RANDOM VALUE", # type: ignore
user_api_key_cache=_cache,
route="/v1/chat/completions",
)
assert result is not None
assert result.user_id == end_user_id
@pytest.mark.parametrize("customer_spend, customer_budget", [(0, 10), (10, 0)])
@pytest.mark.asyncio
async def test_check_end_user_budget(customer_spend, customer_budget):
"""
Test _check_end_user_budget enforcement:
- Scenario 1: customer_spend=0, customer_budget=10 - should pass (under budget)
- Scenario 2: customer_spend=10, customer_budget=0 - should fail (over budget)
Note: Budget enforcement for end users happens in common_checks() via
_check_end_user_budget(), not in get_end_user_object().
"""
from litellm.proxy.auth.auth_checks import _check_end_user_budget
_budget = LiteLLM_BudgetTable(max_budget=customer_budget)
end_user_obj = LiteLLM_EndUserTable(
user_id="my-test-customer",
spend=customer_spend,
litellm_budget_table=_budget,
blocked=False,
)
should_exceed = customer_spend > customer_budget
try:
await get_end_user_object(
end_user_id=end_user_id,
prisma_client="RANDOM VALUE", # type: ignore
user_api_key_cache=_cache,
await _check_end_user_budget(
end_user_obj=end_user_obj,
route="/v1/chat/completions",
)
if customer_spend > customer_budget:
if should_exceed:
pytest.fail(
"Expected call to fail. Customer Spend={}, Customer Budget={}".format(
"Expected BudgetExceededError. Customer Spend={}, Customer Budget={}".format(
customer_spend, customer_budget
)
)
except Exception as e:
if (
isinstance(e, litellm.BudgetExceededError)
and customer_spend > customer_budget
):
pass
else:
except litellm.BudgetExceededError as e:
if not should_exceed:
pytest.fail(
"Expected call to work. Customer Spend={}, Customer Budget={}, Error={}".format(
"Unexpected BudgetExceededError. Customer Spend={}, Customer Budget={}, Error={}".format(
customer_spend, customer_budget, str(e)
)
)
# Verify the error has correct info
assert e.current_cost == customer_spend
assert e.max_budget == customer_budget
@pytest.mark.parametrize(

View file

@ -134,9 +134,14 @@ async def test_explicit_budget_not_overridden_by_default():
@pytest.mark.asyncio
async def test_budget_enforcement_blocks_over_budget_users():
"""
Core scenario: Budget limits are actually enforced.
Core scenario: Budget limits are actually enforced via _check_end_user_budget.
Users who exceed their budget should be blocked.
Note: Budget enforcement happens in common_checks() via _check_end_user_budget(),
not in get_end_user_object(). get_end_user_object only fetches the user data.
"""
from litellm.proxy.auth.auth_checks import _check_end_user_budget
end_user_id = f"test_user_{uuid.uuid4().hex}"
default_budget_id = str(uuid.uuid4())
litellm.max_end_user_budget_id = default_budget_id
@ -170,12 +175,23 @@ async def test_budget_enforcement_blocks_over_budget_users():
mock_cache.async_get_cache = AsyncMock(return_value=None)
mock_cache.async_set_cache = AsyncMock()
# Should raise BudgetExceededError
# First, get the end user object (this just fetches data, doesn't enforce budget)
result = await get_end_user_object(
end_user_id=end_user_id,
prisma_client=mock_prisma_client,
user_api_key_cache=mock_cache,
route="/chat/completions",
)
# Verify user was fetched with default budget applied
assert result is not None
assert result.litellm_budget_table is not None
assert result.litellm_budget_table.max_budget == 10.0
# Now test budget enforcement separately via _check_end_user_budget
with pytest.raises(litellm.BudgetExceededError) as exc_info:
await get_end_user_object(
end_user_id=end_user_id,
prisma_client=mock_prisma_client,
user_api_key_cache=mock_cache,
await _check_end_user_budget(
end_user_obj=result,
route="/chat/completions",
)

View file

@ -804,6 +804,7 @@ def test_img_gen(mock_aimage_generation, client_no_auth):
"prompt": "A cute baby sea otter",
"n": 1,
"size": "1024x1024",
"imageConfig": {"aspectRatio": "9:16", "imageSize": "1K"},
}
response = client_no_auth.post("/v1/images/generations", json=test_data)
@ -813,6 +814,7 @@ def test_img_gen(mock_aimage_generation, client_no_auth):
prompt="A cute baby sea otter",
n=1,
size="1024x1024",
imageConfig={"aspectRatio": "9:16", "imageSize": "1K"},
metadata=mock.ANY,
proxy_server_request=mock.ANY,
secret_fields=mock.ANY,

View file

@ -0,0 +1,116 @@
"""
Regression test for https://github.com/BerriAI/litellm/issues/28505 -
the Responses API bridge double-strips the provider prefix from the
model name when a Chat Completions request has both `tools` and
`reasoning_effort`.
Root cause: the bridge handler called `litellm.responses()` /
`litellm.aresponses()` without passing the already-resolved
`custom_llm_provider`. The downstream call then re-invoked
`get_llm_provider()` with `custom_llm_provider=None`, which stripped
a second provider prefix from a `provider/provider/model` deployment
string.
This test pins both the sync and async bridge handler call sites:
the resolved `custom_llm_provider` must be forwarded to the underlying
`responses` / `aresponses` call so the provider isn't re-detected.
"""
from unittest.mock import MagicMock, patch
import pytest
from litellm.completion_extras.litellm_responses_transformation.handler import (
ResponsesToCompletionBridgeHandler,
)
def _validated_kwargs():
return {
"model": "openai/openai/openai/gpt-5.5",
"messages": [{"role": "user", "content": "hi"}],
"optional_params": {},
"litellm_params": {},
"headers": {},
"model_response": MagicMock(),
"logging_obj": MagicMock(),
"custom_llm_provider": "openai",
}
def test_sync_completion_forwards_custom_llm_provider():
handler = ResponsesToCompletionBridgeHandler()
handler.transformation_handler = MagicMock()
handler.transformation_handler.transform_request.return_value = {
"model": "openai/openai/openai/gpt-5.5",
"input": [],
# `_build_sanitized_litellm_params` spreads `custom_llm_provider` from
# `litellm_params` into request_data on the real bridge path. Seed
# it here so the test exercises the overwrite (not an explicit kwarg
# that would TypeError against an already-present key).
"custom_llm_provider": "should-be-overwritten",
}
handler.transformation_handler.transform_response.return_value = (
_validated_kwargs()["model_response"]
)
with (
patch.object(
handler, "validate_input_kwargs", return_value=_validated_kwargs()
),
patch(
"litellm.responses",
return_value=MagicMock(spec=[]),
) as mock_responses,
):
# The handler routes ResponsesAPIResponse through transform_response.
# We just want to verify the kwargs going INTO responses().
try:
handler.completion(acompletion=False)
except Exception:
# Downstream handling (transform_response, type checks) is not
# the subject of this test.
pass
assert mock_responses.called
kwargs = mock_responses.call_args.kwargs
assert kwargs.get("custom_llm_provider") == "openai", (
"sync bridge must forward custom_llm_provider to litellm.responses() "
"so the downstream get_llm_provider() call does not re-strip the "
"provider prefix on a provider/provider/model deployment string"
)
@pytest.mark.asyncio
async def test_async_completion_forwards_custom_llm_provider():
handler = ResponsesToCompletionBridgeHandler()
handler.transformation_handler = MagicMock()
handler.transformation_handler.transform_request.return_value = {
"model": "openai/openai/openai/gpt-5.5",
"input": [],
# `_build_sanitized_litellm_params` spreads `custom_llm_provider` from
# `litellm_params` into request_data on the real bridge path. Seed
# it here so the test exercises the overwrite (not an explicit kwarg
# that would TypeError against an already-present key).
"custom_llm_provider": "should-be-overwritten",
}
async def _fake_aresponses(**kwargs):
_fake_aresponses.kwargs = kwargs
return MagicMock(spec=[])
_fake_aresponses.kwargs = {}
with (
patch.object(
handler, "validate_input_kwargs", return_value=_validated_kwargs()
),
patch("litellm.aresponses", _fake_aresponses),
):
try:
await handler.acompletion()
except Exception:
pass
assert _fake_aresponses.kwargs.get("custom_llm_provider") == "openai", (
"async bridge must forward custom_llm_provider to litellm.aresponses() "
"so the downstream get_llm_provider() call does not re-strip the "
"provider prefix on a provider/provider/model deployment string"
)

View file

@ -0,0 +1,84 @@
from litellm.integrations.opik.opik_payload_builder.extractors import (
extract_opik_metadata,
)
def test_extract_opik_metadata_fills_missing_keys_from_auth_metadata():
litellm_metadata = {"opik": {"project_name": "my-proj"}}
standard_logging_metadata = {
"user_api_key_auth_metadata": {
"opik": {
"workspace": "auth-workspace",
"project_name": "auth-project",
}
}
}
result = extract_opik_metadata(
litellm_metadata=litellm_metadata,
standard_logging_metadata=standard_logging_metadata,
)
assert result == {
"project_name": "my-proj",
"workspace": "auth-workspace",
}
def test_extract_opik_metadata_request_metadata_overrides_auth_metadata():
litellm_metadata = {
"opik": {
"workspace": "request-workspace",
"thread_id": "request-thread",
}
}
standard_logging_metadata = {
"user_api_key_auth_metadata": {
"opik": {
"workspace": "auth-workspace",
"thread_id": "auth-thread",
"project_name": "auth-project",
}
}
}
result = extract_opik_metadata(
litellm_metadata=litellm_metadata,
standard_logging_metadata=standard_logging_metadata,
)
assert result == {
"workspace": "request-workspace",
"thread_id": "request-thread",
"project_name": "auth-project",
}
def test_extract_opik_metadata_requester_metadata_overrides_all_other_sources():
litellm_metadata = {"opik": {"project_name": "request-project"}}
standard_logging_metadata = {
"user_api_key_auth_metadata": {
"opik": {
"workspace": "auth-workspace",
"project_name": "auth-project",
}
},
"requester_metadata": {
"opik": {
"workspace": "requester-workspace",
"thread_id": "requester-thread",
"project_name": "requester-project",
}
},
}
result = extract_opik_metadata(
litellm_metadata=litellm_metadata,
standard_logging_metadata=standard_logging_metadata,
)
assert result == {
"project_name": "requester-project",
"workspace": "requester-workspace",
"thread_id": "requester-thread",
}

View file

@ -1263,7 +1263,6 @@ class TestOpenTelemetry(unittest.TestCase):
) as mock_get_headers,
patch.object(otel, "_get_tracer_with_dynamic_headers") as mock_get_tracer,
):
# Test case 1: With dynamic headers
mock_get_headers.return_value = {
"arize-space-id": "test-space",
@ -2668,7 +2667,7 @@ class TestOpenTelemetryExternalSpan(unittest.TestCase):
# Verify parent span is still recording after each call
self.assertTrue(
parent_span.is_recording(),
f"External span should still be recording after completion #{i+1}",
f"External span should still be recording after completion #{i + 1}",
)
# Verify all spans have the same trace_id
@ -5170,6 +5169,138 @@ class TestOpenTelemetryPreprocessingDuration(unittest.TestCase):
assert "litellm.preprocessing.duration_ms" not in self._attr(span, exp)
class TestGetSpanContextLitellmMetadataFallback(unittest.TestCase):
"""
Tests for _get_span_context() falling back to litellm_metadata.
On /v1/messages (Anthropic Messages API) and other LITELLM_METADATA_ROUTES,
litellm_parent_otel_span is stored in litellm_params["litellm_metadata"]
instead of litellm_params["metadata"]. _get_span_context() must check
both locations.
Fixes: https://github.com/BerriAI/litellm/issues/27934
"""
def test_span_context_from_metadata(self):
"""Parent span is found when stored in litellm_params['metadata'] (OpenAI path)."""
otel = OpenTelemetry()
mock_span = MagicMock()
mock_span.get_span_context.return_value = MagicMock(is_valid=True)
kwargs = {
"litellm_params": {
"metadata": {"litellm_parent_otel_span": mock_span},
}
}
ctx, detected_span = otel._get_span_context(kwargs)
self.assertIsNotNone(ctx)
# Should NOT fall through to "no parent context" path
self.assertIsNone(detected_span)
def test_span_context_from_litellm_metadata_fallback(self):
"""Parent span is found when stored in litellm_params['litellm_metadata'] (Anthropic path)."""
otel = OpenTelemetry()
mock_span = MagicMock()
mock_span.get_span_context.return_value = MagicMock(is_valid=True)
kwargs = {
"litellm_params": {
"metadata": {
"user_id": "test-user"
}, # Anthropic native metadata, no span
"litellm_metadata": {"litellm_parent_otel_span": mock_span},
}
}
ctx, detected_span = otel._get_span_context(kwargs)
self.assertIsNotNone(ctx)
self.assertIsNone(detected_span)
def test_span_context_metadata_takes_priority(self):
"""When both metadata and litellm_metadata have the span, metadata wins."""
otel = OpenTelemetry()
span_from_metadata = MagicMock(name="span_from_metadata")
span_from_metadata.get_span_context.return_value = MagicMock(is_valid=True)
span_from_litellm_metadata = MagicMock(name="span_from_litellm_metadata")
span_from_litellm_metadata.get_span_context.return_value = MagicMock(
is_valid=True
)
kwargs = {
"litellm_params": {
"metadata": {"litellm_parent_otel_span": span_from_metadata},
"litellm_metadata": {
"litellm_parent_otel_span": span_from_litellm_metadata
},
}
}
ctx, detected_span = otel._get_span_context(kwargs)
self.assertIsNotNone(ctx)
self.assertIsNone(detected_span)
# metadata span is found first, so get_span_context on the
# litellm_metadata span should never be called — proving
# metadata takes priority over litellm_metadata.
span_from_litellm_metadata.get_span_context.assert_not_called()
def test_span_context_no_parent_when_neither_has_span(self):
"""When neither metadata nor litellm_metadata has a span, returns (None, None)."""
otel = OpenTelemetry()
kwargs = {
"litellm_params": {
"metadata": {"user_id": "test-user"},
"litellm_metadata": {"some_key": "some_value"},
}
}
ctx, detected_span = otel._get_span_context(kwargs)
# No parent span in either metadata dict and no active span in test
# context, so both should be None.
self.assertIsNone(ctx)
self.assertIsNone(detected_span)
class TestEndProxySpanLitellmMetadataFallback(unittest.TestCase):
"""
Tests for _end_proxy_span_from_kwargs() falling back to litellm_metadata.
Fixes: https://github.com/BerriAI/litellm/issues/27934
"""
def test_end_proxy_span_from_metadata(self):
"""Proxy span is found and ended from litellm_params['metadata']."""
otel = OpenTelemetry()
mock_span = MagicMock()
mock_span.name = "Received Proxy Server Request"
mock_span.is_recording.return_value = True
kwargs = {
"litellm_params": {
"metadata": {"litellm_parent_otel_span": mock_span},
}
}
otel._end_proxy_span_from_kwargs(kwargs, end_time=datetime.now())
mock_span.end.assert_called_once()
def test_end_proxy_span_from_litellm_metadata(self):
"""Proxy span is found and ended from litellm_params['litellm_metadata'] (fallback)."""
otel = OpenTelemetry()
mock_span = MagicMock()
mock_span.name = "Received Proxy Server Request"
mock_span.is_recording.return_value = True
kwargs = {
"litellm_params": {
"metadata": {"user_id": "test-user"}, # No span here
"litellm_metadata": {"litellm_parent_otel_span": mock_span},
}
}
otel._end_proxy_span_from_kwargs(kwargs, end_time=datetime.now())
mock_span.end.assert_called_once()
class TestOpenTelemetryInferenceIdentityAttributes(unittest.TestCase):
"""team_metadata, http.route, and both model names (the user-facing
model_group alias and the dispatched provider model) must land on the

View file

@ -22,6 +22,18 @@ from litellm.litellm_core_utils.prompt_templates.factory import (
from litellm.types.llms.openai import ChatCompletionToolMessage
def _get_gemini_function_response_inline_data_parts(result):
assert isinstance(result, list), "expected Gemini parts list"
assert len(result) == 1, "multimodal function responses should stay in one part"
function_response_part = result[0]
assert (
"inline_data" not in function_response_part
), "inline_data should be nested under function_response.parts"
function_response = function_response_part["function_response"]
nested_parts = function_response["parts"]
return [part["inline_data"] for part in nested_parts if "inline_data" in part]
def test_ollama_pt_simple_messages():
"""Test basic functionality with simple text messages"""
messages = [
@ -615,8 +627,8 @@ def test_convert_gemini_tool_call_result_with_image_url():
message=message_str_format,
last_message_with_tool_calls=last_message_with_tool_calls,
)
# Should have inline_data for the image
assert isinstance(result, list) and any("inline_data" in p for p in result)
inline_parts = _get_gemini_function_response_inline_data_parts(result)
assert len(inline_parts) == 1
# Test with dict image_url format (OpenAI standard)
message_dict_format = ChatCompletionToolMessage(
@ -635,7 +647,8 @@ def test_convert_gemini_tool_call_result_with_image_url():
message=message_dict_format,
last_message_with_tool_calls=last_message_with_tool_calls,
)
assert isinstance(result2, list) and any("inline_data" in p for p in result2)
inline_parts = _get_gemini_function_response_inline_data_parts(result2)
assert len(inline_parts) == 1
def test_convert_gemini_tool_call_result_with_anthropic_image_block():
@ -677,11 +690,10 @@ def test_convert_gemini_tool_call_result_with_anthropic_image_block():
message=message,
last_message_with_tool_calls=last_message_with_tool_calls,
)
assert isinstance(result, list), "expected a list of parts"
inline_parts = [p for p in result if "inline_data" in p]
inline_parts = _get_gemini_function_response_inline_data_parts(result)
assert len(inline_parts) == 1, "expected exactly one inline_data part"
assert inline_parts[0]["inline_data"]["mime_type"] == "image/png"
assert inline_parts[0]["inline_data"]["data"] == tiny_png_b64
assert inline_parts[0]["mime_type"] == "image/png"
assert inline_parts[0]["data"] == tiny_png_b64
def test_convert_gemini_tool_call_result_with_multiple_anthropic_image_blocks():
@ -734,12 +746,11 @@ def test_convert_gemini_tool_call_result_with_multiple_anthropic_image_blocks():
message=message,
last_message_with_tool_calls=last_message_with_tool_calls,
)
assert isinstance(result, list), "expected a list of parts"
inline_parts = [p for p in result if "inline_data" in p]
inline_parts = _get_gemini_function_response_inline_data_parts(result)
assert (
len(inline_parts) == 2
), f"expected 2 inline_data parts, got {len(inline_parts)}"
mime_types = {p["inline_data"]["mime_type"] for p in inline_parts}
mime_types = {p["mime_type"] for p in inline_parts}
assert mime_types == {"image/png", "image/jpeg"}
@ -773,13 +784,12 @@ def test_convert_gemini_tool_call_result_with_data_url_string():
message=message,
last_message_with_tool_calls=last_message_with_tool_calls,
)
assert isinstance(result, list), "expected a list of parts"
inline_parts = [p for p in result if "inline_data" in p]
inline_parts = _get_gemini_function_response_inline_data_parts(result)
assert (
len(inline_parts) == 1
), "data-URL image string was not converted to inline_data"
assert inline_parts[0]["inline_data"]["mime_type"] == "image/png"
assert inline_parts[0]["inline_data"]["data"] == tiny_png_b64
assert inline_parts[0]["mime_type"] == "image/png"
assert inline_parts[0]["data"] == tiny_png_b64
def test_convert_gemini_tool_call_result_with_data_url_extra_params():
@ -811,12 +821,11 @@ def test_convert_gemini_tool_call_result_with_data_url_extra_params():
message=message,
last_message_with_tool_calls=last_message_with_tool_calls,
)
assert isinstance(result, list), "expected a list of parts"
inline_parts = [p for p in result if "inline_data" in p]
inline_parts = _get_gemini_function_response_inline_data_parts(result)
assert len(inline_parts) == 1
assert (
inline_parts[0]["inline_data"]["mime_type"] == "image/png"
), f"expected clean 'image/png', got '{inline_parts[0]['inline_data']['mime_type']}'"
inline_parts[0]["mime_type"] == "image/png"
), f"expected clean 'image/png', got '{inline_parts[0]['mime_type']}'"
def test_bedrock_tools_unpack_defs():

View file

@ -0,0 +1,76 @@
"""
Test Azure AI Kimi K2.6 model metadata.
"""
import json
from importlib.resources import files
import pytest
@pytest.fixture(scope="module")
def use_local_model_cost_map():
monkeypatch = pytest.MonkeyPatch()
monkeypatch.setenv("LITELLM_LOCAL_MODEL_COST_MAP", "True")
import litellm
from litellm.utils import _invalidate_model_cost_lowercase_map
original_model_cost = litellm.model_cost
litellm.model_cost = json.loads(
files("litellm")
.joinpath("model_prices_and_context_window_backup.json")
.read_text(encoding="utf-8")
)
litellm.get_model_info.cache_clear()
_invalidate_model_cost_lowercase_map()
try:
yield litellm
finally:
litellm.model_cost = original_model_cost
litellm.get_model_info.cache_clear()
_invalidate_model_cost_lowercase_map()
monkeypatch.undo()
def test_azure_ai_kimi_k26_model_info(use_local_model_cost_map):
model_info = use_local_model_cost_map.get_model_info(model="azure_ai/kimi-k2.6")
assert model_info["litellm_provider"] == "azure_ai"
assert model_info["mode"] == "chat"
assert model_info["max_input_tokens"] == 262144
assert model_info["max_output_tokens"] == 262144
assert model_info["max_tokens"] == 262144
assert model_info["input_cost_per_token"] == pytest.approx(9.5e-07)
assert model_info["output_cost_per_token"] == pytest.approx(4e-06)
assert model_info["supports_function_calling"] is True
assert model_info["supports_reasoning"] is True
assert model_info["supports_tool_choice"] is True
assert model_info["supports_vision"] is True
def test_azure_ai_kimi_k26_raw_model_cost_entry(use_local_model_cost_map):
model_info = use_local_model_cost_map.model_cost["azure_ai/kimi-k2.6"]
assert model_info["supported_modalities"] == ["text", "image"]
assert model_info["supported_output_modalities"] == ["text"]
assert model_info["supports_function_calling"] is True
assert model_info["supports_reasoning"] is True
assert model_info["supports_tool_choice"] is True
assert model_info["supports_vision"] is True
def test_azure_ai_kimi_k26_cost_per_token(use_local_model_cost_map):
from litellm.llms.azure_ai.cost_calculator import cost_per_token
from litellm.types.utils import Usage
usage = Usage(
prompt_tokens=1_000_000,
completion_tokens=1_000_000,
total_tokens=2_000_000,
)
prompt_cost, completion_cost = cost_per_token(model="kimi-k2.6", usage=usage)
assert prompt_cost == pytest.approx(0.95)
assert completion_cost == pytest.approx(4.0)

View file

@ -7,6 +7,8 @@ from unittest.mock import MagicMock
import httpx
import pytest
import litellm
from litellm.litellm_core_utils.litellm_logging import StandardLoggingPayloadSetup
from litellm.llms.gemini.image_edit.transformation import GeminiImageEditConfig
@ -19,6 +21,7 @@ class TestGeminiImageEditTransformation:
def test_map_openai_params(self) -> None:
optional_params: Dict[str, object] = {
"n": 2,
"size": "1792x1024",
"response_format": "b64_json",
"quality": "high",
@ -30,20 +33,77 @@ class TestGeminiImageEditTransformation:
drop_params=False,
)
assert mapped["aspectRatio"] == "16:9"
assert mapped["imageConfig"] == {"aspectRatio": "16:9"}
assert mapped["sampleCount"] == 2
assert "response_format" not in mapped
assert "quality" not in mapped
def test_map_openai_params_with_image_size_for_gemini_3(self) -> None:
optional_params: Dict[str, object] = {
"size": "768x1376",
}
mapped = self.config.map_openai_params(
image_edit_optional_params=optional_params, # type: ignore[arg-type]
model="gemini-3-pro-image-preview",
drop_params=False,
)
assert mapped["imageConfig"] == {"aspectRatio": "9:16", "imageSize": "1K"}
def test_map_openai_params_forwards_image_config_as_is(self) -> None:
optional_params: Dict[str, object] = {
"size": "1024x1024",
"imageConfig": {"aspectRatio": "16:9", "imageSize": "512px"},
}
mapped = self.config.map_openai_params(
image_edit_optional_params=optional_params, # type: ignore[arg-type]
model="gemini-3-pro-image-preview",
drop_params=False,
)
assert mapped["imageConfig"] == {"aspectRatio": "16:9", "imageSize": "512px"}
def test_map_openai_params_parses_form_image_config_json(self) -> None:
optional_params: Dict[str, object] = {
"imageConfig": '{"aspectRatio":"16:9","imageSize":"1K"}',
}
mapped = self.config.map_openai_params(
image_edit_optional_params=optional_params, # type: ignore[arg-type]
model="gemini-3-pro-image-preview",
drop_params=False,
)
assert mapped["imageConfig"] == {"aspectRatio": "16:9", "imageSize": "1K"}
def test_map_openai_params_rejects_malformed_form_image_config_json(
self,
) -> None:
optional_params: Dict[str, object] = {
"imageConfig": "{bad",
}
with pytest.raises(litellm.UnsupportedParamsError) as exc_info:
self.config.map_openai_params(
image_edit_optional_params=optional_params, # type: ignore[arg-type]
model="gemini-3-pro-image-preview",
drop_params=False,
)
assert "`imageConfig` must be valid JSON" in str(exc_info.value)
def test_transform_image_edit_request(self) -> None:
image_bytes = b"fake_image_data"
image = BytesIO(image_bytes)
optional_params = {
"sampleCount": 2,
"aspectRatio": "16:9",
"imageConfig": {"aspectRatio": "16:9", "imageSize": "2K"},
}
request_body, files = self.config.transform_image_edit_request(
model=self.model,
model="gemini-3-pro-image-preview",
prompt=self.prompt,
image=[image], # Gemini pipeline passes list of images
image_edit_optional_request_params=optional_params,
@ -61,7 +121,28 @@ class TestGeminiImageEditTransformation:
assert base64.b64decode(inline_data["data"]) == image_bytes
generation_config = request_body["generationConfig"]
assert generation_config["candidateCount"] == 2
assert generation_config["imageConfig"]["aspectRatio"] == "16:9"
assert generation_config["imageConfig"]["imageSize"] == "2K"
def test_transform_image_edit_request_omits_image_size_for_gemini_25(self) -> None:
image = BytesIO(b"fake_image_data")
optional_params = {
"imageConfig": {"aspectRatio": "16:9", "imageSize": "2K"},
}
request_body, _ = self.config.transform_image_edit_request(
model=self.model,
prompt=self.prompt,
image=[image],
image_edit_optional_request_params=optional_params,
litellm_params=MagicMock(),
headers={},
)
assert request_body["generationConfig"]["imageConfig"] == {
"aspectRatio": "16:9"
}
def test_transform_image_edit_request_multiple_images(self) -> None:
image_one = BytesIO(b"image_one")
@ -115,7 +196,16 @@ class TestGeminiImageEditTransformation:
]
}
},
]
],
"usageMetadata": {
"promptTokenCount": 35,
"candidatesTokenCount": 1716,
"totalTokenCount": 1751,
"promptTokensDetails": [
{"modality": "TEXT", "tokenCount": 30},
{"modality": "IMAGE", "tokenCount": 5},
],
},
}
mock_response = MagicMock(spec=httpx.Response)
@ -138,6 +228,19 @@ class TestGeminiImageEditTransformation:
"utf-8"
)
usage = image_response.model_dump()["usage"]
assert usage["input_tokens"] == 35
assert usage["output_tokens"] == 1716
assert usage["prompt_tokens"] == 35
assert usage["completion_tokens"] == 1716
assert usage["prompt_tokens_details"]["image_tokens"] == 5
assert usage["completion_tokens_details"]["image_tokens"] == 1716
logging_usage = StandardLoggingPayloadSetup.get_usage_as_dict(
response_obj=image_response.model_dump()
)
assert logging_usage["completion_tokens_details"]["image_tokens"] == 1716
def test_transform_image_edit_request_without_image_raises(self) -> None:
optional_params = {}

View file

@ -1,7 +1,23 @@
import os
import pytest
import litellm
from litellm.llms.gemini.cost_calculator import cost_per_web_search_request
from litellm.types.utils import PromptTokensDetailsWrapper, Usage
from litellm.llms.gemini.image_edit.cost_calculator import (
cost_calculator as gemini_image_edit_cost_calculator,
)
from litellm.llms.gemini.image_generation.cost_calculator import (
cost_calculator as gemini_image_generation_cost_calculator,
)
from litellm.types.utils import (
ImageObject,
ImageResponse,
ImageUsage,
ImageUsageInputTokensDetails,
PromptTokensDetailsWrapper,
Usage,
)
def _make_usage(web_search_requests: int) -> Usage:
@ -63,3 +79,171 @@ def test_no_usage_details():
usage = Usage(prompt_tokens=100, completion_tokens=50, total_tokens=150)
cost = cost_per_web_search_request(usage=usage, model_info=model_info)
assert cost == 0.0
def test_gemini_image_edit_cost_prefers_token_usage_metadata():
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
model = "gemini/gemini-3-pro-image-preview"
model_info = litellm.get_model_info(model=model, custom_llm_provider="gemini")
input_text_tokens = 20
input_image_tokens = 1120
output_image_tokens = 1120
prompt_tokens = input_text_tokens + input_image_tokens
image_response = ImageResponse(
data=[ImageObject(b64_json="img1"), ImageObject(b64_json="img2")],
usage=ImageUsage(
input_tokens=prompt_tokens,
input_tokens_details=ImageUsageInputTokensDetails(
text_tokens=input_text_tokens,
image_tokens=input_image_tokens,
),
output_tokens=output_image_tokens,
total_tokens=prompt_tokens + output_image_tokens,
),
)
cost = gemini_image_edit_cost_calculator(
model=model,
image_response=image_response,
)
expected_cost = (
prompt_tokens * model_info["input_cost_per_token"]
+ output_image_tokens * model_info["output_cost_per_image_token"]
)
flat_image_cost = (
len(image_response.data or []) * model_info["output_cost_per_image"]
)
assert round(cost, 10) == round(expected_cost, 10)
assert cost != flat_image_cost
def test_gemini_image_edit_cost_uses_output_token_details():
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
model = "gemini/gemini-3-pro-image-preview"
model_info = litellm.get_model_info(model=model, custom_llm_provider="gemini")
input_text_tokens = 20
output_text_tokens = 213
output_image_tokens = 1120
output_tokens = output_text_tokens + output_image_tokens
image_response = ImageResponse(
data=[ImageObject(b64_json="img1")],
usage=ImageUsage(
input_tokens=input_text_tokens,
input_tokens_details=ImageUsageInputTokensDetails(
text_tokens=input_text_tokens,
image_tokens=0,
),
output_tokens=output_tokens,
total_tokens=input_text_tokens + output_tokens,
prompt_tokens=input_text_tokens,
completion_tokens=output_tokens,
prompt_tokens_details={
"text_tokens": input_text_tokens,
"image_tokens": 0,
},
completion_tokens_details={
"text_tokens": output_text_tokens,
"image_tokens": output_image_tokens,
},
output_tokens_details={
"text_tokens": output_text_tokens,
"image_tokens": output_image_tokens,
},
),
)
cost = gemini_image_edit_cost_calculator(
model=model,
image_response=image_response,
)
expected_cost = (
input_text_tokens * model_info["input_cost_per_token"]
+ output_text_tokens * model_info["output_cost_per_token"]
+ output_image_tokens * model_info["output_cost_per_image_token"]
)
all_output_as_image_cost = (
input_text_tokens * model_info["input_cost_per_token"]
+ (output_text_tokens + output_image_tokens)
* model_info["output_cost_per_image_token"]
)
assert round(cost, 10) == round(expected_cost, 10)
assert cost != all_output_as_image_cost
def test_gemini_image_generation_cost_uses_output_token_details():
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
model = "gemini/gemini-3-pro-image-preview"
model_info = litellm.get_model_info(model=model, custom_llm_provider="gemini")
input_text_tokens = 20
output_text_tokens = 213
output_image_tokens = 1120
output_tokens = output_text_tokens + output_image_tokens
image_response = ImageResponse(
data=[ImageObject(b64_json="img1")],
usage=ImageUsage(
input_tokens=input_text_tokens,
input_tokens_details=ImageUsageInputTokensDetails(
text_tokens=input_text_tokens,
image_tokens=0,
),
output_tokens=output_tokens,
total_tokens=input_text_tokens + output_tokens,
prompt_tokens=input_text_tokens,
completion_tokens=output_tokens,
prompt_tokens_details={
"text_tokens": input_text_tokens,
"image_tokens": 0,
},
completion_tokens_details={
"text_tokens": output_text_tokens,
"image_tokens": output_image_tokens,
},
output_tokens_details={
"text_tokens": output_text_tokens,
"image_tokens": output_image_tokens,
},
),
)
cost = gemini_image_generation_cost_calculator(
model=model,
image_response=image_response,
)
expected_cost = (
input_text_tokens * model_info["input_cost_per_token"]
+ output_text_tokens * model_info["output_cost_per_token"]
+ output_image_tokens * model_info["output_cost_per_image_token"]
)
all_output_as_image_cost = (
input_text_tokens * model_info["input_cost_per_token"]
+ (output_text_tokens + output_image_tokens)
* model_info["output_cost_per_image_token"]
)
assert round(cost, 10) == round(expected_cost, 10)
assert cost != all_output_as_image_cost
def test_gemini_image_edit_cost_falls_back_to_flat_image_pricing():
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
model = "gemini/gemini-3-pro-image-preview"
model_info = litellm.get_model_info(model=model, custom_llm_provider="gemini")
image_response = ImageResponse(
data=[ImageObject(b64_json="img1"), ImageObject(b64_json="img2")]
)
cost = gemini_image_edit_cost_calculator(
model=model,
image_response=image_response,
)
assert cost == len(image_response.data or []) * model_info["output_cost_per_image"]

View file

@ -0,0 +1,240 @@
import httpx
from litellm.litellm_core_utils.litellm_logging import StandardLoggingPayloadSetup
from litellm.llms.gemini.image_generation.transformation import GoogleImageGenConfig
from litellm.types.utils import ImageResponse
def test_gemini_image_generation_request_uses_shared_generation_config():
config = GoogleImageGenConfig()
request = config.transform_image_generation_request(
model="gemini-3.1-flash-image-preview",
prompt="Generate a simple app icon",
optional_params={
"sampleCount": 2,
"imageConfig": {"aspectRatio": "16:9", "imageSize": "2K"},
},
litellm_params={},
headers={},
)
assert request["contents"][0]["parts"] == [{"text": "Generate a simple app icon"}]
assert request["generationConfig"] == {
"response_modalities": ["IMAGE", "TEXT"],
"imageConfig": {"aspectRatio": "16:9", "imageSize": "2K"},
"candidateCount": 2,
}
def test_gemini_image_generation_map_openai_params_maps_n_size_and_image_config():
config = GoogleImageGenConfig()
mapped = config.map_openai_params(
non_default_params={
"n": 2,
"size": "768x1376",
"imageConfig": {"aspectRatio": "1:1", "imageSize": "512"},
},
optional_params={},
model="gemini-3.1-flash-image-preview",
drop_params=False,
)
assert mapped == {
"sampleCount": 2,
"imageConfig": {"aspectRatio": "1:1", "imageSize": "512"},
}
def test_imagen_generation_with_provider_prefix_uses_imagen_params_and_response():
config = GoogleImageGenConfig()
mapped = config.map_openai_params(
non_default_params={
"n": 1,
"size": "1024x1024",
},
optional_params={},
model="gemini/imagen-4.0-generate-001",
drop_params=False,
)
assert mapped == {
"sampleCount": 1,
"aspectRatio": "1:1",
"imageSize": "1K",
}
request = config.transform_image_generation_request(
model="gemini/imagen-4.0-generate-001",
prompt="Generate a simple app icon",
optional_params=mapped,
litellm_params={},
headers={},
)
assert request == {
"instances": [{"prompt": "Generate a simple app icon"}],
"parameters": {
"sampleCount": 1,
"aspectRatio": "1:1",
"imageSize": "1K",
},
}
result = config.transform_image_generation_response(
model="gemini/imagen-4.0-generate-001",
raw_response=httpx.Response(
status_code=200,
json={
"predictions": [
{
"bytesBase64Encoded": "fake-imagen-image",
}
]
},
),
model_response=ImageResponse(data=[]),
logging_obj=None,
request_data={},
optional_params={},
litellm_params={},
encoding=None,
)
assert result.data is not None
assert result.data[0].b64_json == "fake-imagen-image"
def test_imagen_generation_forwards_mapped_openai_size_image_size():
config = GoogleImageGenConfig()
mapped = config.map_openai_params(
non_default_params={
"size": "512x512",
},
optional_params={},
model="gemini/imagen-4.0-generate-001",
drop_params=False,
)
assert mapped == {"aspectRatio": "1:1", "imageSize": "512"}
request = config.transform_image_generation_request(
model="gemini/imagen-4.0-generate-001",
prompt="Generate a simple app icon",
optional_params=mapped,
litellm_params={},
headers={},
)
assert request == {
"instances": [{"prompt": "Generate a simple app icon"}],
"parameters": {"aspectRatio": "1:1", "imageSize": "512"},
}
def test_gemini_image_generation_usage_includes_chat_token_details():
config = GoogleImageGenConfig()
raw_response = httpx.Response(
status_code=200,
json={
"candidates": [
{
"content": {
"parts": [
{
"inlineData": {
"mimeType": "image/png",
"data": "fake-image",
}
}
]
}
}
],
"usageMetadata": {
"promptTokenCount": 35,
"candidatesTokenCount": 1716,
"totalTokenCount": 1751,
"promptTokensDetails": [
{"modality": "TEXT", "tokenCount": 30},
{"modality": "IMAGE", "tokenCount": 5},
],
"candidatesTokensDetails": [
{"modality": "TEXT", "tokenCount": 213},
{"modality": "IMAGE", "tokenCount": 1120},
],
},
},
)
result = config.transform_image_generation_response(
model="gemini-3.1-flash-image-preview",
raw_response=raw_response,
model_response=ImageResponse(data=[]),
logging_obj=None,
request_data={},
optional_params={},
litellm_params={},
encoding=None,
)
usage = result.model_dump()["usage"]
assert usage["input_tokens"] == 35
assert usage["output_tokens"] == 1716
assert usage["prompt_tokens"] == 35
assert usage["completion_tokens"] == 1716
assert usage["prompt_tokens_details"]["image_tokens"] == 5
assert usage["completion_tokens_details"]["text_tokens"] == 596
assert usage["completion_tokens_details"]["image_tokens"] == 1120
assert usage["output_tokens_details"]["text_tokens"] == 596
assert usage["output_tokens_details"]["image_tokens"] == 1120
logging_usage = StandardLoggingPayloadSetup.get_usage_as_dict(
response_obj=result.model_dump()
)
assert logging_usage["completion_tokens_details"]["text_tokens"] == 596
assert logging_usage["completion_tokens_details"]["image_tokens"] == 1120
def test_gemini_image_generation_usage_without_output_details_treats_output_as_image():
config = GoogleImageGenConfig()
raw_response = httpx.Response(
status_code=200,
json={
"candidates": [
{
"content": {
"parts": [
{
"inlineData": {
"mimeType": "image/png",
"data": "fake-image",
}
}
]
}
}
],
"usageMetadata": {
"promptTokenCount": 35,
"candidatesTokenCount": 1716,
"totalTokenCount": 1751,
"promptTokensDetails": [{"modality": "TEXT", "tokenCount": 35}],
},
},
)
result = config.transform_image_generation_response(
model="gemini-3.1-flash-image-preview",
raw_response=raw_response,
model_response=ImageResponse(data=[]),
logging_obj=None,
request_data={},
optional_params={},
litellm_params={},
encoding=None,
)
usage = result.model_dump()["usage"]
assert usage["completion_tokens_details"]["text_tokens"] == 0
assert usage["completion_tokens_details"]["image_tokens"] == 1716

View file

@ -0,0 +1,326 @@
"""
Tests for Inception (Mercury) chat provider integration
"""
import json
import os
from unittest import mock
import httpx
import litellm
from litellm.llms.inception.chat.transformation import InceptionChatConfig
def test_inception_config_initialization():
config = InceptionChatConfig()
assert config.custom_llm_provider == "inception"
def test_inception_chat_supports_diffusion_params():
"""The chat config must expose Inception's diffusion-LLM request controls"""
params = InceptionChatConfig().get_supported_openai_params("mercury-2")
for p in (
"reasoning_effort",
"reasoning_summary",
"reasoning_summary_wait",
"diffusing",
"realtime",
"tools",
"tool_choice",
"response_format",
):
assert p in params, f"{p} should be a supported chat param"
def test_inception_chat_sends_diffusion_params_in_body():
"""reasoning_effort (incl. `instant`) and the diffusion flags reach the request body"""
captured = {}
def fake_send(self, request, **kwargs):
captured["body"] = json.loads(request.content.decode())
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=json.dumps(
{
"id": "c-1",
"object": "chat.completion",
"created": 1,
"model": "mercury-2",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "hi"},
"finish_reason": "stop",
}
],
"usage": {
"prompt_tokens": 5,
"completion_tokens": 1,
"total_tokens": 6,
},
}
).encode(),
)
with mock.patch("httpx.Client.send", new=fake_send):
litellm.completion(
model="inception/mercury-2",
messages=[{"role": "user", "content": "hi"}],
api_key="sk-x",
reasoning_effort="instant",
reasoning_summary=True,
reasoning_summary_wait=True,
diffusing=True,
realtime=True,
max_completion_tokens=128,
)
body = captured["body"]
assert body["reasoning_effort"] == "instant"
assert body["reasoning_summary"] is True
assert body["reasoning_summary_wait"] is True
assert body["diffusing"] is True
assert body["realtime"] is True
assert body["max_tokens"] == 128 # max_completion_tokens mapped to max_tokens
def test_inception_chat_response_surfaces_reasoning_and_usage():
"""reasoning_summary / warning survive, and reasoning_tokens maps to usage details"""
def fake_send(self, request, **kwargs):
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=json.dumps(
{
"id": "c-1",
"object": "chat.completion",
"created": 1,
"model": "mercury-2",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "answer"},
"finish_reason": "stop",
}
],
"usage": {
"prompt_tokens": 5,
"completion_tokens": 2,
"total_tokens": 7,
"reasoning_tokens": 4,
"cached_input_tokens": 3,
},
"reasoning_summary": {
"content": "step by step",
"status": "complete",
},
"warning": "heads up",
}
).encode(),
)
with mock.patch("httpx.Client.send", new=fake_send):
r = litellm.completion(
model="inception/mercury-2",
messages=[{"role": "user", "content": "hi"}],
api_key="sk-x",
)
assert r.reasoning_summary == {"content": "step by step", "status": "complete"}
assert r.warning == "heads up"
assert r.usage.completion_tokens_details.reasoning_tokens == 4
assert r.usage.model_extra.get("cached_input_tokens") == 3
def test_inception_get_openai_compatible_provider_info():
config = InceptionChatConfig()
with mock.patch.dict(os.environ, {}, clear=True):
with mock.patch.object(litellm, "inception_key", None):
api_base, api_key = config._get_openai_compatible_provider_info(None, None)
assert api_base == "https://api.inceptionlabs.ai/v1"
assert api_key is None
with mock.patch.dict(
os.environ,
{
"INCEPTION_API_KEY": "test-key",
"INCEPTION_API_BASE": "https://custom.inceptionlabs.ai/v1",
},
):
api_base, api_key = config._get_openai_compatible_provider_info(None, None)
assert api_base == "https://custom.inceptionlabs.ai/v1"
assert api_key == "test-key"
with mock.patch.dict(
os.environ,
{
"INCEPTION_API_KEY": "env-key",
"INCEPTION_API_BASE": "https://env.inceptionlabs.ai/v1",
},
):
api_base, api_key = config._get_openai_compatible_provider_info(
"https://param.inceptionlabs.ai/v1", "param-key"
)
assert api_base == "https://param.inceptionlabs.ai/v1"
assert api_key == "param-key"
def test_inception_key_module_attr_fallback():
"""litellm.inception_key is used when no param/env key is provided"""
config = InceptionChatConfig()
with mock.patch.dict(os.environ, {}, clear=True):
with mock.patch.object(litellm, "inception_key", "module-attr-key"):
_, api_key = config._get_openai_compatible_provider_info(None, None)
assert api_key == "module-attr-key"
def test_inception_does_not_leak_key_to_caller_api_base():
"""
The server-managed Inception key must not be forwarded to a caller-supplied
api_base. It is only resolved for the default/server base, or when the
caller also supplies their own key.
"""
config = InceptionChatConfig()
with mock.patch.dict(
os.environ, {"INCEPTION_API_KEY": "server-secret"}, clear=True
):
with mock.patch.object(litellm, "inception_key", "module-secret"):
# caller overrides api_base without a key -> server key withheld
api_base, api_key = config._get_openai_compatible_provider_info(
"https://attacker.example/v1", None
)
assert api_base == "https://attacker.example/v1"
assert api_key is None
# caller overrides api_base AND supplies their own key -> used as-is
_, api_key = config._get_openai_compatible_provider_info(
"https://attacker.example/v1", "caller-key"
)
assert api_key == "caller-key"
# default/server base -> server-managed key resolved
_, api_key = config._get_openai_compatible_provider_info(None, None)
assert api_key == "module-secret"
def test_get_llm_provider_inception():
from litellm.litellm_core_utils.get_llm_provider_logic import get_llm_provider
model, provider, _, _ = get_llm_provider("inception/mercury-2")
assert model == "mercury-2"
assert provider == "inception"
model, provider, _, api_base = get_llm_provider(
"mercury-2", api_base="https://api.inceptionlabs.ai/v1"
)
assert model == "mercury-2"
assert provider == "inception"
assert api_base == "https://api.inceptionlabs.ai/v1"
def test_inception_in_provider_lists():
assert "inception" in litellm.openai_compatible_providers
assert "inception" in litellm.provider_list
assert "https://api.inceptionlabs.ai/v1" in litellm.openai_compatible_endpoints
def test_inception_model_configuration():
from litellm import get_model_info
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
litellm.inception_models = set()
litellm.add_known_models()
info = get_model_info("inception/mercury-2")
assert info.get("litellm_provider") == "inception"
assert info.get("mode") == "chat"
assert info.get("max_input_tokens") == 128000
assert info.get("input_cost_per_token") == 2.5e-07
assert info.get("output_cost_per_token") == 7.5e-07
assert info.get("cache_read_input_token_cost") == 2.5e-08
assert info.get("supports_function_calling") is True
assert info.get("supports_tool_choice") is True
assert info.get("supports_response_schema") is True
def test_inception_model_list_populated():
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
litellm.inception_models = set()
litellm.add_known_models()
assert "inception/mercury-2" in litellm.inception_models
for model in litellm.inception_models:
assert model.startswith("inception/")
def test_inception_completion_targets_inception_endpoint():
"""
End-to-end: a completion routed through the inception provider must hit
Inception's base URL and path, send a Bearer token, strip the
`inception/` prefix from the model name, and forward tool_choice.
"""
captured = {}
def fake_send(self, request, **kwargs):
captured["url"] = str(request.url)
captured["auth"] = request.headers.get("authorization")
captured["body"] = json.loads(request.content.decode())
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=json.dumps(
{
"id": "cmpl-1",
"object": "chat.completion",
"created": 1,
"model": "mercury-2",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "hi"},
"finish_reason": "stop",
}
],
"usage": {
"prompt_tokens": 5,
"completion_tokens": 1,
"total_tokens": 6,
},
}
).encode(),
)
tools = [
{
"type": "function",
"function": {
"name": "f",
"parameters": {"type": "object", "properties": {}},
},
}
]
with mock.patch("httpx.Client.send", new=fake_send):
response = litellm.completion(
model="inception/mercury-2",
messages=[{"role": "user", "content": "hello"}],
api_key="sk-test-fake-123",
tools=tools,
tool_choice="auto",
)
assert captured["url"] == "https://api.inceptionlabs.ai/v1/chat/completions"
assert captured["auth"] == "Bearer sk-test-fake-123"
assert captured["body"]["model"] == "mercury-2"
assert captured["body"]["tool_choice"] == "auto"
assert response.choices[0].message.content == "hi"

View file

@ -0,0 +1,300 @@
"""
Tests for Inception (Mercury) fill-in-the-middle (FIM) provider integration
"""
import json
import os
from unittest import mock
import httpx
import pytest
import litellm
from litellm.llms.inception.completion.transformation import (
InceptionTextCompletionConfig,
)
def _fim_response_bytes():
return json.dumps(
{
"id": "fim-1",
"object": "text_completion",
"created": 1,
"model": "mercury-edit-2",
"choices": [
{"text": "a + b", "index": 0, "finish_reason": "stop", "logprobs": None}
],
"usage": {"prompt_tokens": 5, "completion_tokens": 3, "total_tokens": 8},
}
).encode()
def test_inception_fim_supports_suffix_param():
"""The FIM config must keep `suffix` (otherwise FIM requests lose context)"""
config = InceptionTextCompletionConfig()
assert "suffix" in config.get_supported_openai_params("mercury-edit-2")
mapped = config.map_openai_params(
non_default_params={"suffix": "\n return x", "max_completion_tokens": 50},
optional_params={},
model="mercury-edit-2",
drop_params=False,
)
assert mapped["suffix"] == "\n return x"
assert mapped["max_tokens"] == 50
def test_inception_fim_supported_params_match_schema():
"""FIM exposes the OpenAI subset of Inception's FIMCompletionRequest only"""
params = InceptionTextCompletionConfig().get_supported_openai_params(
"mercury-edit-2"
)
for p in ("suffix", "top_p", "frequency_penalty", "presence_penalty", "stop"):
assert p in params
# Chat-only sampling controls are not part of Inception's FIM schema
for p in ("temperature", "seed", "logprobs", "n", "user"):
assert p not in params
def test_text_completion_inception_in_provider_lists():
from litellm.types.utils import LlmProviders
assert LlmProviders.TEXT_COMPLETION_INCEPTION == "text-completion-inception"
assert "text-completion-inception" in litellm.provider_list
def test_inception_get_supported_openai_params_dispatch():
"""litellm.get_supported_openai_params routes the FIM provider to our config"""
params = litellm.get_supported_openai_params(
model="mercury-edit-2", custom_llm_provider="text-completion-inception"
)
assert "suffix" in params
assert "temperature" not in params
@pytest.mark.parametrize("provider", ["inception", "text-completion-inception"])
def test_inception_validate_environment(provider):
model = (
"inception/mercury-2"
if provider == "inception"
else "text-completion-inception/mercury-edit-2"
)
with mock.patch.dict(os.environ, {}, clear=True):
result = litellm.validate_environment(model)
assert result["keys_in_environment"] is False
assert "INCEPTION_API_KEY" in result["missing_keys"]
with mock.patch.dict(os.environ, {"INCEPTION_API_KEY": "sk-x"}, clear=True):
result = litellm.validate_environment(model)
assert result["keys_in_environment"] is True
def test_inception_completion_endpoint_returns_chat_object():
"""
Calling chat `completion()` with the FIM provider converts the text
completion result into a chat-shaped ModelResponse.
"""
def fake_send(self, request, **kwargs):
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=_fim_response_bytes(),
)
with mock.patch("httpx.Client.send", new=fake_send):
r = litellm.completion(
model="text-completion-inception/mercury-edit-2",
messages=[{"role": "user", "content": "def add(a, b): return "}],
api_key="sk-x",
)
assert r.choices[0].message.content == "a + b"
@pytest.mark.asyncio
async def test_inception_fim_async():
"""async FIM path (acompletion) hits Inception's /v1/fim/completions"""
captured = {}
async def fake_asend(self, request, **kwargs):
captured["url"] = str(request.url)
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=_fim_response_bytes(),
)
with mock.patch("httpx.AsyncClient.send", new=fake_asend):
r = await litellm.atext_completion(
model="text-completion-inception/mercury-edit-2",
prompt="def add(a, b): return ",
suffix="\n",
api_key="sk-x",
max_tokens=10,
)
assert captured["url"] == "https://api.inceptionlabs.ai/v1/fim/completions"
assert r.choices[0].text == "a + b"
def test_inception_fim_model_configuration():
from litellm import get_model_info
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
litellm.model_cost = litellm.get_model_cost_map(url="")
litellm.text_completion_inception_models = set()
litellm.add_known_models()
assert (
"text-completion-inception/mercury-edit-2"
in litellm.text_completion_inception_models
)
info = get_model_info("text-completion-inception/mercury-edit-2")
assert info.get("litellm_provider") == "text-completion-inception"
assert info.get("mode") == "completion"
assert info.get("max_input_tokens") == 32000
def test_inception_fim_targets_fim_endpoint():
"""
End-to-end: a FIM request must hit `/v1/fim/completions` (NOT
`/v1/completions`), carry the `suffix`, and parse the standard `text` field.
"""
captured = {}
def fake_send(self, request, **kwargs):
captured["url"] = str(request.url)
captured["auth"] = request.headers.get("authorization")
captured["body"] = json.loads(request.content.decode())
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=json.dumps(
{
"id": "fim-1",
"object": "text_completion",
"created": 1,
"model": "mercury-edit-2",
"choices": [
{
"text": "a + b",
"index": 0,
"finish_reason": "stop",
"logprobs": None,
}
],
"usage": {
"prompt_tokens": 5,
"completion_tokens": 3,
"total_tokens": 8,
},
}
).encode(),
)
with mock.patch("httpx.Client.send", new=fake_send):
response = litellm.text_completion(
model="text-completion-inception/mercury-edit-2",
prompt="def add(a, b):\n return ",
suffix="\n",
api_key="sk-fim-fake",
max_tokens=20,
)
assert captured["url"] == "https://api.inceptionlabs.ai/v1/fim/completions"
assert captured["auth"] == "Bearer sk-fim-fake"
assert captured["body"]["model"] == "mercury-edit-2"
assert captured["body"]["suffix"] == "\n"
assert "prompt" in captured["body"]
assert response.choices[0].text == "a + b"
def test_inception_fim_does_not_leak_global_api_key():
"""
Regression: the global litellm.api_key (commonly an OpenAI key) must not be
forwarded to Inception. Only an Inception-specific key (param,
litellm.inception_key, or INCEPTION_API_KEY) may be sent to the Inception base.
"""
captured = {}
def fake_send(self, request, **kwargs):
captured["auth"] = request.headers.get("authorization")
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=_fim_response_bytes(),
)
with mock.patch.dict(
os.environ, {"INCEPTION_API_KEY": "sk-inception-correct"}, clear=True
):
with mock.patch.object(litellm, "inception_key", None):
with mock.patch.object(litellm, "api_key", "sk-global-should-not-leak"):
with mock.patch("httpx.Client.send", new=fake_send):
litellm.text_completion(
model="text-completion-inception/mercury-edit-2",
prompt="def add(a, b): return ",
max_tokens=10,
)
assert captured["auth"] == "Bearer sk-inception-correct"
def test_inception_fim_extra_body_forwards_vllm_params():
"""top_k / repetition_penalty are reachable via extra_body (not OpenAI params)"""
captured = {}
def fake_send(self, request, **kwargs):
captured["body"] = json.loads(request.content.decode())
return httpx.Response(
status_code=200,
request=request,
headers={"content-type": "application/json"},
content=json.dumps(
{
"id": "f-1",
"object": "text_completion",
"created": 1,
"model": "mercury-edit-2",
"choices": [
{
"text": "x",
"index": 0,
"finish_reason": "stop",
"logprobs": None,
}
],
"usage": {
"prompt_tokens": 2,
"completion_tokens": 1,
"total_tokens": 3,
},
}
).encode(),
)
with mock.patch("httpx.Client.send", new=fake_send):
litellm.text_completion(
model="text-completion-inception/mercury-edit-2",
prompt="def f(",
suffix=")",
api_key="sk-x",
top_p=0.9,
extra_body={"top_k": 40, "repetition_penalty": 1.1},
)
body = captured["body"]
assert body["top_p"] == 0.9
assert body["top_k"] == 40
assert body["repetition_penalty"] == 1.1

View file

@ -0,0 +1,57 @@
"""
Regression test for tool-call / tool-result matching in the Gemini message converter.
When an assistant message that contains tool_calls is followed by a *second* assistant
message that has no tool_calls (e.g. the model emits a short narration turn after the
tool call but before the tool result), the converter used to overwrite its
`last_message_with_tool_calls` reference with the text-only assistant message. The
subsequent tool result could then no longer be matched to its tool call, and conversion
failed with:
Exception: Missing corresponding tool call for tool response message.
This happens for any OpenAI-style history with that shape, independent of provider/model.
"""
import pytest
from litellm.llms.vertex_ai.gemini.transformation import (
_gemini_convert_messages_with_history,
)
def _messages_with_text_assistant_between_tool_call_and_result():
return [
{"role": "user", "content": "list the files"},
{
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {"name": "shell", "arguments": '{"command": ["ls"]}'},
}
],
},
# text-only assistant message in between (no tool_calls)
{"role": "assistant", "content": "Running the command now."},
{"role": "tool", "tool_call_id": "call_abc123", "content": "math.py"},
]
def test_tool_result_matches_tool_call_with_text_assistant_in_between():
messages = _messages_with_text_assistant_between_tool_call_and_result()
# Should not raise "Missing corresponding tool call for tool response message".
contents = _gemini_convert_messages_with_history(messages=messages)
# The function response must be present and carry the correct tool name.
function_responses = [
part["function_response"]
for content in contents
for part in content["parts"]
if isinstance(part, dict) and part.get("function_response")
]
assert function_responses, f"expected a functionResponse part, got: {contents}"
assert function_responses[0]["name"] == "shell"

View file

@ -1154,44 +1154,82 @@ def test_convert_tool_response_with_base64_image():
]
}
# Convert tool response (returns list when image is present)
# Convert tool response with nested multimodal functionResponse.parts.
result = convert_to_gemini_tool_call_result(
tool_message, last_message_with_tool_calls
)
# Verify results - should be a list with 2 parts (function_response + inline_data)
assert isinstance(
result, list
), f"Expected list when image present, got {type(result)}"
assert len(result) == 2, f"Expected 2 parts, got {len(result)}"
# Find function_response part and inline_data part
function_response_part = None
inline_data_part = None
for part in result:
if "function_response" in part:
function_response_part = part
elif "inline_data" in part:
inline_data_part = part
# Check function_response exists
assert function_response_part is not None, "Missing function_response part"
function_response = function_response_part["function_response"]
assert isinstance(result, list), "Should return a parts list when media is present"
assert len(result) == 1, "Should return one function_response part"
result_part = result[0]
assert "function_response" in result_part
assert "inline_data" not in result_part
function_response = result_part["function_response"]
assert function_response["name"] == "click_at"
assert "response" in function_response
# Verify JSON response is parsed correctly
assert "url" in function_response["response"]
assert function_response["response"]["url"] == "https://example.com"
# Check inline_data exists
assert inline_data_part is not None, "Missing inline_data part"
inline_data: BlobType = inline_data_part["inline_data"]
# Check inline_data is nested under functionResponse.parts.
assert "parts" in function_response
assert len(function_response["parts"]) == 1
inline_data: BlobType = function_response["parts"][0]["inline_data"]
assert "data" in inline_data
assert "mime_type" in inline_data
assert inline_data["mime_type"] == "image/png"
assert inline_data["data"] == test_image_base64
def test_gemini_history_nests_multimodal_tool_response_parts():
"""Full history conversion should not emit sibling inline_data tool result parts."""
test_image_base64 = "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="
messages = [
{"role": "user", "content": "Get me an image"},
{
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": "call_get_image",
"type": "function",
"function": {"name": "get_image", "arguments": "{}"},
}
],
},
{
"role": "tool",
"tool_call_id": "call_get_image",
"content": [
{"type": "text", "text": '{"image_ref": "inline"}'},
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": test_image_base64,
},
},
],
},
]
contents = _gemini_convert_messages_with_history(messages=messages)
tool_response_parts = contents[-1]["parts"]
assert len(tool_response_parts) == 1
assert "inline_data" not in tool_response_parts[0]
function_response = tool_response_parts[0]["function_response"]
assert function_response["parts"] == [
{
"inline_data": {
"data": test_image_base64,
"mime_type": "image/png",
}
}
]
def test_convert_tool_response_with_url_image():
"""Test tool response with HTTP URL image (will download and convert)."""
import pytest
@ -1225,24 +1263,20 @@ def test_convert_tool_response_with_url_image():
tool_message, last_message_with_tool_calls
)
# Should be a list with 2 parts when image is present
assert isinstance(
result, list
), f"Expected list when image present, got {type(result)}"
assert len(result) == 2, f"Expected 2 parts, got {len(result)}"
# Find parts
function_response_part = next(p for p in result if "function_response" in p)
inline_data_part = next(p for p in result if "inline_data" in p)
# Check function_response exists
assert function_response_part is not None, "Missing function_response part"
function_response = function_response_part["function_response"]
), "Should return a parts list when media is present"
assert len(result) == 1, "Should return one function_response part"
result_part = result[0]
assert "function_response" in result_part
assert "inline_data" not in result_part
function_response = result_part["function_response"]
assert function_response["name"] == "type_text_at"
# Check inline_data exists (URL should be downloaded and converted)
assert inline_data_part is not None, "Missing inline_data part"
inline_data: BlobType = inline_data_part["inline_data"]
# Check inline_data is nested under functionResponse.parts.
assert "parts" in function_response
assert len(function_response["parts"]) == 1
inline_data: BlobType = function_response["parts"][0]["inline_data"]
assert "data" in inline_data
assert "mime_type" in inline_data
except Exception as e:
@ -1558,38 +1592,27 @@ def test_convert_tool_response_with_pdf_file():
]
}
# Convert tool response (returns list when file is present)
# Convert tool response with nested multimodal functionResponse.parts.
result = convert_to_gemini_tool_call_result(
tool_message, last_message_with_tool_calls
)
# Verify results - should be a list with 2 parts (function_response + inline_data)
assert isinstance(
result, list
), f"Expected list when file present, got {type(result)}"
assert len(result) == 2, f"Expected 2 parts, got {len(result)}"
# Find function_response part and inline_data part
function_response_part = None
inline_data_part = None
for part in result:
if "function_response" in part:
function_response_part = part
elif "inline_data" in part:
inline_data_part = part
# Check function_response exists
assert function_response_part is not None, "Missing function_response part"
function_response = function_response_part["function_response"]
assert isinstance(result, list), "Should return a parts list when media is present"
assert len(result) == 1, "Should return one function_response part"
result_part = result[0]
assert "function_response" in result_part
assert "inline_data" not in result_part
function_response = result_part["function_response"]
assert function_response["name"] == "analyze_document"
assert "response" in function_response
# Verify JSON response is parsed correctly
assert "status" in function_response["response"]
assert function_response["response"]["status"] == "success"
# Check inline_data exists
assert inline_data_part is not None, "Missing inline_data part"
inline_data: BlobType = inline_data_part["inline_data"]
# Check inline_data is nested under functionResponse.parts.
assert "parts" in function_response
assert len(function_response["parts"]) == 1
inline_data: BlobType = function_response["parts"][0]["inline_data"]
assert "data" in inline_data
assert "mime_type" in inline_data
assert inline_data["mime_type"] == "application/pdf"
@ -1624,21 +1647,13 @@ def test_convert_tool_response_with_input_file_type():
tool_message, last_message_with_tool_calls
)
# Verify results
assert isinstance(
result, list
), f"Expected list when file present, got {type(result)}"
assert len(result) == 2, f"Expected 2 parts, got {len(result)}"
# Find inline_data part
inline_data_part = None
for part in result:
if "inline_data" in part:
inline_data_part = part
# Check inline_data exists
assert inline_data_part is not None, "Missing inline_data part"
assert inline_data_part["inline_data"]["mime_type"] == "application/pdf"
# Check inline_data is nested under functionResponse.parts.
assert isinstance(result, list), "Should return a parts list when media is present"
assert len(result) == 1, "Should return one function_response part"
function_response = result[0]["function_response"]
assert (
function_response["parts"][0]["inline_data"]["mime_type"] == "application/pdf"
)
def test_convert_tool_response_with_nested_file_object():
@ -1669,21 +1684,11 @@ def test_convert_tool_response_with_nested_file_object():
tool_message, last_message_with_tool_calls
)
# Verify results - should be a list with 2 parts
assert isinstance(
result, list
), f"Expected list when file present, got {type(result)}"
assert len(result) == 2, f"Expected 2 parts, got {len(result)}"
# Find inline_data part
inline_data_part = None
for part in result:
if "inline_data" in part:
inline_data_part = part
# Check inline_data exists
assert inline_data_part is not None, "Missing inline_data part"
inline_data: BlobType = inline_data_part["inline_data"]
# Check inline_data is nested under functionResponse.parts.
assert isinstance(result, list), "Should return a parts list when media is present"
assert len(result) == 1, "Should return one function_response part"
function_response = result[0]["function_response"]
inline_data: BlobType = function_response["parts"][0]["inline_data"]
assert "data" in inline_data
assert "mime_type" in inline_data
assert inline_data["mime_type"] == "application/pdf"

View file

@ -3391,30 +3391,44 @@ async def test_resolve_end_user_swallows_db_errors_and_returns_none(
@pytest.mark.asyncio
async def test_resolve_end_user_reraises_budget_exceeded(
async def test_resolve_end_user(
_validate_flag_on, monkeypatch
):
"""BudgetExceededError from get_end_user_object must bubble up so the
auth path enforces spend limits instead of silently dropping the id."""
import litellm
"""Verify that resolve_and_validate_end_user_id does NOT raise BudgetExceededError.
Note: As of the refactor that moved _check_end_user_budget out of
get_end_user_object, budget enforcement now happens in common_checks().
The end-user validation path should return the user ID regardless of budget status.
Budget enforcement for end users happens later in common_checks() via
_check_end_user_budget(), which respects skip_budget_checks for zero-cost models.
This test verifies that even when get_end_user_object returns a user with a budget,
resolve_and_validate_end_user_id does not block the request - budget enforcement
is deferred to common_checks() where skip_budget_checks logic can be applied.
"""
from litellm.proxy.auth import auth_checks
from litellm.proxy.auth.auth_checks import resolve_and_validate_end_user_id
# Mock get_end_user_object to return a user with budget info
# (simulating a user who may have exceeded their budget)
mock_end_user = MagicMock()
mock_end_user.user_id = "customer-over-budget"
monkeypatch.setattr(
auth_checks,
"get_end_user_object",
AsyncMock(
side_effect=litellm.BudgetExceededError(current_cost=10.0, max_budget=5.0)
),
AsyncMock(return_value=mock_end_user),
)
cache = _validation_cache()
with pytest.raises(litellm.BudgetExceededError):
await resolve_and_validate_end_user_id(
raw_end_user_id="customer-over-budget",
prisma_client=MagicMock(),
user_api_key_cache=cache,
)
# resolve_and_validate_end_user_id should return the user ID without raising
# BudgetExceededError - budget enforcement happens in common_checks()
result = await resolve_and_validate_end_user_id(
raw_end_user_id="customer-over-budget",
prisma_client=MagicMock(),
user_api_key_cache=cache,
)
assert result == "customer-over-budget"
@pytest.mark.asyncio

View file

@ -5,7 +5,11 @@ from litellm.proxy.auth.user_api_key_auth import (
_run_post_custom_auth_checks,
update_valid_token_with_end_user_params,
)
from litellm.proxy._types import UserAPIKeyAuth
from litellm.proxy._types import (
LiteLLM_BudgetTable,
LiteLLM_EndUserTable,
UserAPIKeyAuth,
)
@pytest.mark.asyncio
@ -88,6 +92,85 @@ async def test_custom_auth_run_post_custom_auth_checks_with_end_user_budget_exce
mock_budget_check.assert_awaited_once()
@pytest.mark.asyncio
async def test_custom_auth_enforces_end_user_budget_when_common_checks_skipped():
# custom-auth deployments with custom_auth_run_common_checks unset skip
# common_checks() (and its end-user budget enforcement) in the centralized
# gate, so the helper must enforce the end-user budget itself. Regression:
# an over-budget end user must be rejected on this path.
valid_token = UserAPIKeyAuth(token="test_token", end_user_id="customer-1")
over_budget_end_user = LiteLLM_EndUserTable(
user_id="customer-1",
blocked=False,
spend=0.0,
litellm_budget_table=LiteLLM_BudgetTable(max_budget=1.0),
)
async def mock_get_current_spend(counter_key, fallback_spend):
if counter_key == "spend:end_user:customer-1":
return 5.0
return fallback_spend
with (
patch(
"litellm.proxy.auth.user_api_key_auth.get_end_user_object",
new_callable=AsyncMock,
return_value=over_budget_end_user,
),
patch("litellm.proxy.proxy_server.get_current_spend", mock_get_current_spend),
patch("litellm.proxy.proxy_server.general_settings", {}),
):
with pytest.raises(litellm.BudgetExceededError):
await _run_post_custom_auth_checks(
valid_token=valid_token,
request=None,
request_data={"model": "gpt-4"},
route="/v1/chat/completions",
parent_otel_span=None,
)
@pytest.mark.asyncio
async def test_custom_auth_defers_end_user_budget_to_common_checks_when_enabled():
# With custom_auth_run_common_checks set, the wrapper's common_checks()
# enforces the end-user budget, so the helper must not double-enforce it.
valid_token = UserAPIKeyAuth(token="test_token", end_user_id="customer-1")
end_user_obj = LiteLLM_EndUserTable(
user_id="customer-1",
blocked=False,
spend=0.0,
litellm_budget_table=LiteLLM_BudgetTable(max_budget=1.0),
)
with (
patch(
"litellm.proxy.auth.user_api_key_auth.get_end_user_object",
new_callable=AsyncMock,
return_value=end_user_obj,
),
patch(
"litellm.proxy.auth.user_api_key_auth._check_end_user_budget",
new_callable=AsyncMock,
) as mock_check,
patch(
"litellm.proxy.auth.user_api_key_auth._enforce_key_and_fallback_model_access",
new_callable=AsyncMock,
),
patch(
"litellm.proxy.proxy_server.general_settings",
{"custom_auth_run_common_checks": True},
),
):
await _run_post_custom_auth_checks(
valid_token=valid_token,
request=None,
request_data={"model": "gpt-4"},
route="/v1/chat/completions",
parent_otel_span=None,
)
mock_check.assert_not_awaited()
def test_update_valid_token_does_not_override_custom_auth_values_with_none():
"""
Greptile feedback: if custom auth sets end_user_model_max_budget on the token,

View file

@ -72,9 +72,40 @@ US_EXPECTED = [
("us.anthropic.claude-haiku-4-5-20251001-v1:0", 2.2e-06, None),
]
# EU/AU/JP cross-region inference profiles carry the same +10% regional
# premium as US (per AWS Bedrock pricing). Coverage list filters to entries
# that actually exist in the pricing JSON - e.g. Opus 4.6 has no JP profile.
REGIONAL_EXPECTED = [
# Opus 4.6 - $11.00 / MTok (eu/au only; no jp profile)
("eu.anthropic.claude-opus-4-6-v1", 1.1e-05, None),
("au.anthropic.claude-opus-4-6-v1", 1.1e-05, None),
# Opus 4.7 - $11.00 / MTok (eu/au; jp is added in #28567)
("eu.anthropic.claude-opus-4-7", 1.1e-05, None),
("au.anthropic.claude-opus-4-7", 1.1e-05, None),
# Sonnet 4.6 - $6.60 / MTok
("eu.anthropic.claude-sonnet-4-6", 6.6e-06, None),
("au.anthropic.claude-sonnet-4-6", 6.6e-06, None),
("jp.anthropic.claude-sonnet-4-6", 6.6e-06, None),
# Sonnet 4.5 - $6.60 / MTok with $13.20 / MTok long-context tier
("eu.anthropic.claude-sonnet-4-5-20250929-v1:0", 6.6e-06, 1.32e-05),
("au.anthropic.claude-sonnet-4-5-20250929-v1:0", 6.6e-06, 1.32e-05),
("jp.anthropic.claude-sonnet-4-5-20250929-v1:0", 6.6e-06, 1.32e-05),
# Haiku 4.5 - $2.20 / MTok
("eu.anthropic.claude-haiku-4-5-20251001-v1:0", 2.2e-06, None),
("au.anthropic.claude-haiku-4-5-20251001-v1:0", 2.2e-06, None),
("jp.anthropic.claude-haiku-4-5-20251001-v1:0", 2.2e-06, None),
# Note: eu.anthropic.claude-opus-4-5-20251101-v1:0 is intentionally NOT
# in this list. The existing entry carries base/global 5m rates
# (5e-06 / 6.25e-06) instead of the +10% regional premium (5.5e-06 /
# 6.875e-06), which would make the 1.6x 5m-to-1h invariant fail.
# Fixing the EU 5m rates first is left to a follow-up so this PR
# stays scoped to the 1-hour cache tier addition.
]
@pytest.mark.parametrize(
"model_key, expected_1hr, expected_1hr_lc", GLOBAL_EXPECTED + US_EXPECTED
"model_key, expected_1hr, expected_1hr_lc",
GLOBAL_EXPECTED + US_EXPECTED + REGIONAL_EXPECTED,
)
def test_bedrock_anthropic_1hr_cache_write_pricing(
model_data, model_key, expected_1hr, expected_1hr_lc

View file

@ -2120,11 +2120,11 @@ def test_gemini_3_1_flash_lite_pricing():
):
model_info = litellm.model_cost.get(model_name)
assert model_info is not None, f"Missing model pricing entry: {model_name}"
assert model_info["input_cost_per_token"] == 4.5e-07
assert model_info["input_cost_per_audio_token"] == 9e-07
assert model_info["output_cost_per_token"] == 2.7e-06
assert model_info["output_cost_per_reasoning_token"] == 2.7e-06
assert model_info["cache_read_input_token_cost"] == 4.5e-08
assert model_info["input_cost_per_token"] == 2.5e-07
assert model_info["input_cost_per_audio_token"] == 5e-07
assert model_info["output_cost_per_token"] == 1.5e-06
assert model_info["output_cost_per_reasoning_token"] == 1.5e-06
assert model_info["cache_read_input_token_cost"] == 2.5e-08
assert model_info["max_input_tokens"] == 1048576

View file

@ -2376,6 +2376,74 @@ def test_get_deployment_model_info_base_model_flow():
# Should return None when no model info is found
assert result is None
# Test Case 6: custom_model_info present but litellm_model_name_model_info is None
# (model has custom pricing in config but is not in built-in model_prices_and_context_window.json)
mock_custom_pricing_only = {
"input_cost_per_token": 1.74e-06,
"output_cost_per_token": 3.48e-06,
"cache_read_input_token_cost": 1.45e-08,
"mode": "chat",
}
with patch.object(
litellm,
"model_cost",
{"custom-model-id": mock_custom_pricing_only},
):
with patch.object(litellm, "get_model_info") as mock_get_model_info:
# Model NOT in built-in cost map — raise exception
mock_get_model_info.side_effect = Exception("Model not in cost map")
result = router.get_deployment_model_info(
model_id="custom-model-id", model_name="unknown-model"
)
# Should return custom_model_info even when litellm_model_name_model_info is None
assert result is not None
assert result["input_cost_per_token"] == 1.74e-06
assert result["output_cost_per_token"] == 3.48e-06
assert result["cache_read_input_token_cost"] == 1.45e-08
assert result["mode"] == "chat"
# Test Case 7: custom_model_info with base_model but litellm_model_name_model_info None
mock_custom_with_base = {
"base_model": "some-base-model",
"input_cost_per_token": 0.01,
"output_cost_per_token": 0.02,
}
mock_base_info = {
"key": "some-base-model",
"max_tokens": 8192,
"mode": "chat",
"litellm_provider": "openai",
}
with patch.object(
litellm,
"model_cost",
{"custom-with-base": mock_custom_with_base},
):
with patch.object(litellm, "get_model_info") as mock_get_model_info:
def get_info_side_effect(model):
if model == "some-base-model":
return mock_base_info
raise Exception("Model not in cost map")
mock_get_model_info.side_effect = get_info_side_effect
result = router.get_deployment_model_info(
model_id="custom-with-base", model_name="unknown-model"
)
# Should return custom_model_info merged with base model info
assert result is not None
assert (
result["input_cost_per_token"] == 0.01
) # From custom (overrides base)
assert result["max_tokens"] == 8192 # From base model
assert result["litellm_provider"] == "openai" # From base model
print("✓ All base model flow test cases passed!")

View file

@ -928,6 +928,7 @@ def test_aaamodel_prices_and_context_window_json_is_valid():
},
},
"supports_native_streaming": {"type": "boolean"},
"supports_image_size": {"type": "boolean"},
"supports_native_structured_output": {"type": "boolean"},
"tiered_pricing": {
"type": "array",

View file

@ -11,7 +11,7 @@ const inter = Inter({ subsets: ["latin"] });
export const metadata: Metadata = {
title: "LiteLLM Dashboard",
description: "LiteLLM Proxy Admin UI",
icons: { icon: "./favicon.ico" },
icons: { icon: "/get_favicon" },
};
export default function RootLayout({