feat(bedrock-batch): route /v1/embeddings JSONL to Titan v2 modelInput (#28875)

* fix(thinking): handle None thinking param in is_thinking_enabled (#28598)

Squash-merged by litellm-agent from Terrajlz's PR.

* feat(helm): support tpl rendering in podAnnotations (#28609)

Squash-merged by litellm-agent from devauxbr's PR.

* feat(bedrock-batch): route /v1/embeddings JSONL to Titan v2 modelInput

Changes vs main:
- BedrockFilesConfig now detects OpenAI batch JSONL lines whose `url`
  is /v1/embeddings (with body-shape fallback) and routes them through
  a new `_map_openai_embedding_to_bedrock_params` helper instead of the
  chat-completion transformer that silently produces an invalid body.
- The embedding helper currently supports Amazon Titan Text Embeddings V2
  only. Other embed models (Titan G1, Titan Multimodal, Cohere Embed,
  Nova Multimodal Embeddings) raise NotImplementedError with a clear
  message; each will get a dedicated branch + tests in follow-up PRs to
  keep schema-specific risks isolated.
- Validation refuses pre-tokenized inputs (List[int], List[List[int]])
  and multi-element string lists with explicit errors so callers emit
  one JSONL line per embedding instead of relying on us to fan out.
- Titan v2 model id match tolerates "bedrock/" prefix, cross-region
  inference profile prefix ("us.", "eu.", etc.), and ARN forms; the
  marker boundary check rejects lookalikes like "titan-embed-text-v20".
- Tests cover happy path (fixtures), dimensions/encoding_format mapping,
  body-shape fallback, single-element list unwrap, error paths
  (missing input, multi-element list, unsupported model, pre-tokenized),
  mixed chat+embedding batch, and the model-id boundary check.

* fix(bedrock-batch): trust explicit url over body shape in record routing

Greptile-flagged gap in `_is_embedding_record`: when an OpenAI batch
JSONL line carries an explicit `url` pointing to a non-embedding
endpoint (e.g. `/v1/chat/completions`) AND its body happens to have
`input` without `messages`, the body-shape fallback would mis-route
that record to the embedding transformer and corrupt the modelInput.

Changes vs previous commit:
- `_is_embedding_record` now short-circuits to NOT-embedding whenever
  `url` is non-empty and not equal to `/v1/embeddings`. The body-shape
  fallback only runs when `url` is missing or empty. Docstring updated
  to spell out the precedence rules.
- Two new tests cover the case: direct helper assertion that an
  explicit chat url plus an input-bearing body returns False, plus an
  end-to-end check that the resulting modelInput contains no `inputText`
  key. A second test asserts the same short-circuit for arbitrary
  non-embeddings urls (`/v1/completions`, `/v1/responses`).

28/28 tests pass (was 26/26 before this commit + 2 new).

* refactor(bedrock-batch): extract embedding input normalization helper

Splits the input-shape validation out of
`_map_openai_embedding_to_bedrock_params` into a new static helper
`_coerce_embedding_input_to_string`. Same semantics; the goal is to make
the validation testable in isolation and to give future
embedding-provider branches (Titan G1, Cohere) a reusable shaping
function instead of duplicating type checks.

- Helper accepts `str`, single-element `list[str]`, and raises
  `ValueError` / `NotImplementedError` with actionable messages for
  None, multi-element lists, pre-tokenized inputs (`list[int]` /
  `list[list[int]]`), and other unsupported types.
- New unit test exercises the helper directly across happy paths,
  None / missing input, multi-element string list, multi-element int
  list (caught as 'one input per JSONL record' since we can't
  disambiguate from 'multiple strings' without more context),
  pre-tokenized single-element list-of-list, single-element list of
  bare int, and dict input.

29/29 tests in the file still pass.

* fix(bedrock-batch): layer registry mode check + drop unused provider param

Greptile-flagged issues on Titan v2 model detection:

1. Hardcoded substring without registry consultation conflicts with the
   project convention of treating `model_prices_and_context_window.json`
   as the source of truth for model capability flags.

   Fix: `_is_titan_v2_embed_model` now layers a registry mode check on
   top of the existing marker boundary. When `get_model_info` resolves
   the id, we additionally require `mode == "embedding"` so a malformed
   id whose path-component matches the marker but whose registered mode
   is "chat" doesn't slip through. Registry silence (cross-region
   inference profile prefixes like `us.amazon.titan-embed-text-v2:0`,
   ARN forms) keeps the substring-only behavior because the registry
   genuinely can't normalize those ids today. The substring is still
   needed because `mode == "embedding"` alone doesn't distinguish
   Titan v2's InvokeModel schema from Cohere Embed, Nova Multimodal,
   or Titan G1 - all also embedding mode, all with incompatible bodies.

2. `provider` parameter on `_map_openai_embedding_to_bedrock_params`
   was accepted but never used.

   Fix: dropped from the signature and the call site.

Also adds `_lookup_registry_mode` static helper (mirrors the one in the
sibling batches transformer) so the registry try/except shape lives in
one place instead of being inlined into the detector.

5 new tests pin the layered behavior:
- registry mode=chat overrides the marker match (rejected)
- registry mode=embedding + marker match (accepted)
- registry silent + marker match for cross-region and ARN ids (accepted)
- direct `_lookup_registry_mode` coverage across all return paths

33/33 tests in the file pass.

* fix(bedrock-batch): drive Titan v2 detection from registry provider_specific_entry

Addresses Greptile's remaining policy concern on PR #28875: the hardcoded
`_TITAN_V2_EMBED_MODEL_MARKER` substring required a code change to
register new Titan v2 variants. Now the registry is the source of truth.

Changes:
- model_prices_and_context_window.json + model_prices_and_context_window_backup.json
  Adds `provider_specific_entry.bedrock_invocation_schema = "titan_v2"`
  to the `amazon.titan-embed-text-v2:0` entry. Uses the existing
  `provider_specific_entry` escape-hatch field (already surfaced by
  `get_model_info` via `ModelInfo.provider_specific_entry`) rather than
  introducing a new top-level field that would need a corresponding
  change in `utils.py::get_model_info` to flow through.

- BedrockFilesConfig._is_titan_v2_embed_model now reads
  `get_model_info(model).provider_specific_entry.bedrock_invocation_schema`
  first. When the registry resolves the id we trust the field;
  registered ids without (or with a different) schema value are rejected
  outright - no substring second-chance for registered ids. The
  substring fallback only runs when `get_model_info` raises, which
  covers cross-region inference profile prefixes
  (`us.amazon.titan-embed-text-v2:0`) and Bedrock ARN forms - neither of
  which the registry normalizes today.

- _lookup_registry_mode helper renamed to _lookup_provider_specific_field
  and generalized: takes a field name, reads
  `provider_specific_entry[field]` defensively. Future embed-schema
  branches (Titan G1, Cohere, Nova Multimodal) will share this helper.

- New module-level constants: _BEDROCK_INVOCATION_SCHEMA_FIELD names the
  registry key; _TITAN_V2_INVOCATION_SCHEMA names the value.

- _map_openai_embedding_to_bedrock_params: dropped the unused `provider`
  parameter (separate Greptile flag).

Tests updated for the new nested-field semantics. 34/34 tests pass; one
end-to-end integration check confirms that with
LITELLM_LOCAL_MODEL_COST_MAP=true, the registry returns the new field
and all 7 representative model ids classify correctly.

* chore(lint): apply Black formatting to is_thinking_enabled

Pre-existing Black violation on the shin_agent_oss_staging_05_22_2026
base branch - the lint CI check on this PR fails on
`litellm/llms/base_llm/chat/transformation.py` even with zero changes
from our side. Applying the formatter's preferred line-break inside
`is_thinking_enabled` unblocks the lint check without touching the
method's logic.

---------

Co-authored-by: Terrajlz <info@jouleselectrictech.com>
Co-authored-by: Bruno Devaux <devaux.br@gmail.com>
This commit is contained in:
OS-joaocastilho 2026-06-01 11:44:01 +01:00 • committed by GitHub
parent 8ed909236b
commit fd554501ef
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
6 changed files with 908 additions and 7 deletions

View file

@ -233,6 +233,259 @@ class BedrockFilesConfig(BaseAWSLLM, BaseFilesConfig):
# example; add others here as they adopt the same schema.
CONVERSE_INVOKE_PROVIDERS = ("nova",)
# OpenAI batch URL that signals an embedding request. Per OpenAI Batch API
# spec, every JSONL record carries a `url` field; we use it as the
# authoritative signal to route the line to the embedding code path
# instead of inferring from the presence of `input` vs `messages`.
OPENAI_EMBEDDINGS_URL = "/v1/embeddings"
@staticmethod
def _is_embedding_record(openai_jsonl_record: Dict[str, Any]) -> bool:
"""
Decide whether an OpenAI batch JSONL line is an embedding request.
Precedence (strict - any explicit `url` short-circuits):
1. `url == "/v1/embeddings"` -> embedding. Authoritative per the
OpenAI Batch API spec.
2. Any other non-empty `url` (e.g. `/v1/chat/completions`) -> NOT
embedding. We trust the caller's explicit signal even if the
body would otherwise suggest embedding; misrouting a chat
record into the embedding transformer would corrupt the
modelInput, while a chat-shaped body sent to the chat path
either succeeds or fails cleanly inside that transformer.
3. `url` missing/empty -> fall back to body shape. Requires
`input` present AND `messages` absent so a malformed record
carrying both keys routes to the chat path (safer default:
Anthropic transforms ignore unknown top-level keys, whereas
the embedding transformer would silently drop the messages).
"""
url = openai_jsonl_record.get("url")
if url == BedrockFilesConfig.OPENAI_EMBEDDINGS_URL:
return True
if url:
return False
body = openai_jsonl_record.get("body", {})
if not isinstance(body, dict):
return False
return "input" in body and "messages" not in body
# Identifier for the Bedrock Titan v2 InvokeModel body schema as stored
# in `model_prices_and_context_window.json`. Centralized so future
# embedding-schema variants can add their own value
# (e.g. `cohere_v3`, `titan_g1`, `titan_multimodal`) without touching
# the detection logic.
_TITAN_V2_INVOCATION_SCHEMA = "titan_v2"
# Substring marker used as a fallback when the registry can't resolve
# the model id - notably cross-region inference profile prefixes
# (`us.amazon.titan-embed-text-v2:0`) and Bedrock ARN forms, which
# `get_model_info` doesn't normalize today.
_TITAN_V2_EMBED_MODEL_MARKER = "titan-embed-text-v2"
# Nested field name under `provider_specific_entry` that identifies the
# Bedrock InvokeModel body schema for batch inference.
# `provider_specific_entry` is the registry's escape hatch for fields
# `get_model_info` doesn't promote to top-level - exactly what we need
# here. Documented in the `sample_spec` entry of
# `model_prices_and_context_window.json` and surfaced by
# `get_model_info` (see `ModelInfo.provider_specific_entry`).
_BEDROCK_INVOCATION_SCHEMA_FIELD = "bedrock_invocation_schema"
@staticmethod
def _is_titan_v2_embed_model(model: str) -> bool:
"""
True iff `model` refers to Amazon Titan Text Embeddings V2.
Resolution order:
1. `model_prices_and_context_window.json` via `get_model_info`.
The Titan v2 registry entry carries an explicit
`provider_specific_entry.bedrock_invocation_schema` discriminator
(`"titan_v2"`). When the registry resolves the id we trust that
field as the source of truth - no hardcoded model-id comparison
needed.
2. Substring fallback (`titan-embed-text-v2` followed by `:`, `/`,
or end-of-string) for ids the registry can't normalize. This
catches cross-region inference profile prefixes
(`us.amazon.titan-embed-text-v2:0`) and Bedrock ARN forms; the
marker boundary check rejects lookalikes like
`titan-embed-text-v20` or `titan-embed-text-v2-experimental`.
Tolerant of common id shapes:
- "amazon.titan-embed-text-v2:0"
- "bedrock/amazon.titan-embed-text-v2:0"
- "us.amazon.titan-embed-text-v2:0" (cross-region inference profile)
- ARN forms ending in ".../amazon.titan-embed-text-v2:0"
"""
# Registry-driven path: when get_model_info resolves the id we trust
# the registry's discriminator. A resolved id with a different (or
# absent) schema value here is intentionally not given a substring
# second-chance - the registry is authoritative for ids it knows.
registry_schema = BedrockFilesConfig._lookup_provider_specific_field(
model, BedrockFilesConfig._BEDROCK_INVOCATION_SCHEMA_FIELD
)
if registry_schema is not None:
return registry_schema == BedrockFilesConfig._TITAN_V2_INVOCATION_SCHEMA
# Registry silence -> substring fallback for unmapped ids only.
normalized = model.lower()
if normalized.startswith("bedrock/"):
normalized = normalized[len("bedrock/") :]
marker = BedrockFilesConfig._TITAN_V2_EMBED_MODEL_MARKER
idx = normalized.find(marker)
if idx < 0:
return False
end = idx + len(marker)
return end == len(normalized) or normalized[end] in (":", "/")
@staticmethod
def _lookup_provider_specific_field(model_id: str, field: str) -> Optional[str]:
"""
Read a nested string field from the registry entry's
`provider_specific_entry` dict via `litellm.get_model_info`.
Returns the field's string value when:
- the registry resolves `model_id`,
- the entry exposes `provider_specific_entry` as a dict, and
- that dict has `field` mapped to a non-empty string.
Otherwise returns `None`.
Isolating this means feature detectors (Titan v2 today, future
Cohere Embed / Nova Multimodal branches) share one defensive
try/except shape instead of duplicating it. The `None` return
covers every realistic failure mode: `get_model_info` raises
(cross-region profile prefixes, Bedrock ARN forms, unreleased
models), returns a non-dict, has no `provider_specific_entry`, or
the requested field is missing / non-string / empty.
"""
try:
from litellm import get_model_info
info = get_model_info(model_id)
except Exception:
return None
if not isinstance(info, dict):
return None
provider_specific = info.get("provider_specific_entry")
if not isinstance(provider_specific, dict):
return None
value = provider_specific.get(field)
return value if isinstance(value, str) and value else None
@staticmethod
def _coerce_embedding_input_to_string(raw_input: Any, model: str = "") -> str:
"""
Normalize an OpenAI /v1/embeddings `input` field into the single
string that Bedrock Titan v2 InvokeModel expects in `inputText`.
Accepts: a string, or a single-element list containing one string.
Rejects (with actionable messages):
- None / missing -> ValueError
- Multi-element string lists -> ValueError, prompts caller to
emit one JSONL line per input
- Pre-tokenized inputs (List[int], List[List[int]]) -> NotImplementedError
- Any other type -> ValueError
Extracted so the validation can be exercised in isolation and so
future embedding-provider branches (Titan G1, Cohere) can reuse it
without duplicating the type-shaping logic.
"""
if raw_input is None:
raise ValueError(
"Embedding batch record is missing required `input` field: "
f"model={model}"
)
# Bedrock InvokeModel for Titan v2 takes exactly one string `inputText`
# per call. Pre-tokenized inputs and multi-element string lists are
# explicitly unsupported so callers emit one JSONL line per embedding
# instead of relying on us to silently fan out or concatenate.
if isinstance(raw_input, list):
if len(raw_input) == 1:
candidate = raw_input[0]
else:
raise ValueError(
"Bedrock batch embedding requires one input per JSONL "
"record. Got a list with "
f"{len(raw_input)} items for model={model}; emit one "
"JSONL line per input string instead."
)
else:
candidate = raw_input
# Catches pre-tokenized inputs (List[int] from OpenAI spec, or a
# single int slipping past the list-unwrap above).
# NOTE: bool is a subclass of int but treating True/False as a token
# is meaningless either way, so the broad check is fine.
if isinstance(candidate, (list, int)):
raise NotImplementedError(
"Bedrock Titan v2 batch embedding does not support "
"pre-tokenized integer inputs. Pass `input` as a string "
f"(model={model})."
)
if not isinstance(candidate, str):
raise ValueError(
"Bedrock batch embedding `input` must be a string (or a "
"single-element list of strings). Got type "
f"{type(candidate).__name__} for model={model}."
)
return candidate
def _map_openai_embedding_to_bedrock_params(
self,
openai_request_body: Dict[str, Any],
) -> Dict[str, Any]:
"""
Transform an OpenAI /v1/embeddings request body into the
Bedrock InvokeModel `modelInput` for embedding models that AWS
supports via batch inference (CreateModelInvocationJob).
Currently routes Amazon Titan Text Embeddings V2 only; other
embedding providers (Titan G1, Titan Multimodal, Cohere Embed,
Nova Multimodal Embeddings) raise NotImplementedError until they
get a dedicated branch. Splitting them keeps PR scope tight and
lets each model's request schema be exercised by its own tests.
AWS docs (Titan v2 InvokeModel body):
https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-titan-embed-text.html
"""
from litellm.llms.bedrock.embed.amazon_titan_v2_transformation import (
AmazonTitanV2Config,
)
_model = openai_request_body.get("model", "")
if not self._is_titan_v2_embed_model(_model):
# Refuse early instead of silently shaping the body for the wrong
# provider. The synchronous /v1/embeddings path supports more
# models, but each has a different InvokeModel schema; mapping
# them here without dedicated tests would risk corrupt batches.
raise NotImplementedError(
"Bedrock batch embedding currently supports only Amazon "
"Titan Text Embeddings V2 (model id contains "
f"'titan-embed-text-v2'). Got model={_model!r}. Track other "
"embedding models in https://github.com/BerriAI/litellm/issues."
)
input_text = self._coerce_embedding_input_to_string(
openai_request_body.get("input"), model=_model
)
# Map OpenAI-style params (dimensions, encoding_format) onto the
# Titan v2 schema (dimensions, embeddingTypes) via the embed config
# so this stays in sync with the synchronous /v1/embeddings path.
non_default_params = {
k: v for k, v in openai_request_body.items() if k not in ("model", "input")
}
titan_config = AmazonTitanV2Config()
inference_params = titan_config.map_openai_params(
non_default_params=non_default_params,
optional_params={},
)
return dict(
titan_config._transform_request(
input=input_text, inference_params=inference_params
)
)
def _map_openai_to_bedrock_params(
self,
openai_request_body: Dict[str, Any],
@ -349,10 +602,19 @@ class BedrockFilesConfig(BaseAWSLLM, BaseFilesConfig):
# Determine provider from model name
provider = self.get_bedrock_invoke_provider(model)
# Transform to Bedrock modelInput format
model_input = self._map_openai_to_bedrock_params(
openai_request_body=openai_body, provider=provider
)
# Route to the embedding transformer when the OpenAI batch line
# targets /v1/embeddings; otherwise fall back to the existing
# chat-completion path. We branch here (rather than inside
# `_map_openai_to_bedrock_params`) so the chat helper keeps its
# narrow contract and the embedding helper can evolve independently.
if self._is_embedding_record(_openai_jsonl_content):
model_input = self._map_openai_embedding_to_bedrock_params(
openai_request_body=openai_body
)
else:
model_input = self._map_openai_to_bedrock_params(
openai_request_body=openai_body, provider=provider
)
# Create Bedrock batch record
record_id = _openai_jsonl_content.get(

View file

@ -577,7 +577,10 @@
"max_tokens": 8192,
"mode": "embedding",
"output_cost_per_token": 0.0,
"output_vector_size": 1024
"output_vector_size": 1024,
"provider_specific_entry": {
"bedrock_invocation_schema": "titan_v2"
}
},
"amazon.titan-image-generator-v1": {
"input_cost_per_image": 0.0,

View file

@ -577,7 +577,10 @@
"max_tokens": 8192,
"mode": "embedding",
"output_cost_per_token": 0.0,
"output_vector_size": 1024
"output_vector_size": 1024,
"provider_specific_entry": {
"bedrock_invocation_schema": "titan_v2"
}
},
"amazon.titan-image-generator-v1": {
"input_cost_per_image": 0.0,

View file

@ -0,0 +1,3 @@
{"recordId": "embed-1", "modelInput": {"inputText": "Hello world"}}
{"recordId": "embed-2", "modelInput": {"inputText": "Another document to embed", "dimensions": 512}}
{"recordId": "embed-3", "modelInput": {"inputText": "Single element list", "embeddingTypes": ["binary"]}}

View file

@ -0,0 +1,3 @@
{"custom_id": "embed-1", "method": "POST", "url": "/v1/embeddings", "body": {"model": "bedrock/amazon.titan-embed-text-v2:0", "input": "Hello world"}}
{"custom_id": "embed-2", "method": "POST", "url": "/v1/embeddings", "body": {"model": "bedrock/amazon.titan-embed-text-v2:0", "input": "Another document to embed", "dimensions": 512}}
{"custom_id": "embed-3", "method": "POST", "url": "/v1/embeddings", "body": {"model": "bedrock/amazon.titan-embed-text-v2:0", "input": ["Single element list"], "encoding_format": "base64"}}

View file

@ -426,7 +426,7 @@ class TestBedrockFilesTransformation:
"s3_bucket_name": "litellm-batch-352026",
"s3_region_name": "us-gov-west-1",
}
# aws_region_name set to something different — s3_region_name must still win
# aws_region_name set to something different - s3_region_name must still win
optional_params = {"aws_region_name": "us-east-1"}
captured_optional_params: dict = {}
@ -482,3 +482,630 @@ class TestBedrockFilesTransformation:
assert "messages" in model_input
assert "max_tokens" in model_input
assert model_input["max_tokens"] == 10
class TestBedrockFilesEmbeddingTransformation:
"""
Tests for routing OpenAI /v1/embeddings batch JSONL records through the
Titan v2 transformer so AWS Bedrock's CreateModelInvocationJob receives
a valid modelInput body.
Scope is intentionally Titan v2 only - other embedding models will get
their own follow-up PRs/tests so each schema is exercised in isolation.
"""
def test_titan_v2_embedding_jsonl_matches_fixture(self):
"""Round-trip the input fixture against the expected Bedrock output."""
import json
import os
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
here = os.path.dirname(__file__)
with open(os.path.join(here, "input_batch_embeddings.jsonl")) as f:
openai_jsonl = [json.loads(line) for line in f if line.strip()]
with open(os.path.join(here, "expected_bedrock_batch_embeddings.jsonl")) as f:
expected = [json.loads(line) for line in f if line.strip()]
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
openai_jsonl
)
assert result == expected
def test_titan_v2_simple_string_input(self):
"""Single string `input` maps to `{"inputText": <str>}` with no extras."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": "Hello",
},
}
]
)
assert result == [{"recordId": "e1", "modelInput": {"inputText": "Hello"}}]
def test_titan_v2_dimensions_and_encoding_format(self):
"""OpenAI `dimensions` / `encoding_format` map to Titan v2 schema."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": "Hi",
"dimensions": 256,
"encoding_format": "float",
},
}
]
)
model_input = result[0]["modelInput"]
assert model_input["inputText"] == "Hi"
assert model_input["dimensions"] == 256
assert model_input["embeddingTypes"] == ["float"]
def test_embedding_routing_falls_back_to_body_shape(self):
"""Records without `url` still route via `input` presence."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": "Hello",
},
}
]
)
assert result[0]["modelInput"] == {"inputText": "Hello"}
def test_embedding_single_element_list_input_is_accepted(self):
"""A single-element list maps to the same shape as a bare string."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": ["only one"],
},
}
]
)
assert result[0]["modelInput"]["inputText"] == "only one"
def test_embedding_multi_input_list_raises(self):
"""Multi-element `input` lists are rejected with a clear message."""
import pytest
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
with pytest.raises(ValueError, match="one input per JSONL record"):
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": ["a", "b"],
},
}
]
)
def test_embedding_missing_input_raises(self):
"""A record routed to /v1/embeddings without `input` is an error."""
import pytest
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
with pytest.raises(ValueError, match="missing required `input`"):
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {"model": "bedrock/amazon.titan-embed-text-v2:0"},
}
]
)
def test_mixed_chat_and_embedding_in_same_batch(self):
"""Chat and embedding records in the same JSONL each take their path."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "chat-1",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
"messages": [{"role": "user", "content": "Hi"}],
"max_tokens": 5,
},
},
{
"custom_id": "embed-1",
"method": "POST",
"url": "/v1/embeddings",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": "Hi",
},
},
]
)
assert result[0]["recordId"] == "chat-1"
assert "messages" in result[0]["modelInput"]
assert result[0]["modelInput"]["anthropic_version"] == "bedrock-2023-05-31"
assert result[1]["recordId"] == "embed-1"
assert result[1]["modelInput"] == {"inputText": "Hi"}
def test_unsupported_embedding_model_raises_not_implemented(self):
"""Cohere/Nova/Titan-G1 embed get a clear NotImplementedError, not a corrupt body."""
import pytest
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
for unsupported_model in (
"bedrock/cohere.embed-english-v3",
"bedrock/amazon.titan-embed-text-v1",
"bedrock/amazon.titan-embed-image-v1",
"bedrock/amazon.nova-2-multimodal-embeddings-v1:0",
):
with pytest.raises(NotImplementedError, match="titan-embed-text-v2"):
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {"model": unsupported_model, "input": "Hi"},
}
]
)
def test_titan_v2_model_name_variants_route_correctly(self):
"""All common Titan v2 model id shapes route through the embedding path."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
for model_id in (
"amazon.titan-embed-text-v2:0",
"bedrock/amazon.titan-embed-text-v2:0",
"us.amazon.titan-embed-text-v2:0",
"bedrock/us.amazon.titan-embed-text-v2:0",
):
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {"model": model_id, "input": "Hi"},
}
]
)
assert result[0]["modelInput"] == {
"inputText": "Hi"
}, f"model id {model_id} did not route to Titan v2 embedding path"
def test_pretokenized_input_list_of_ints_raises(self):
"""`input: List[int]` (pre-tokenized) is rejected, not silently mis-shaped."""
import pytest
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
with pytest.raises(
(NotImplementedError, ValueError), match=r"pre-tokenized|one input per"
):
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": [1, 2, 3],
},
}
]
)
def test_pretokenized_single_wrapped_list_raises(self):
"""`input: List[List[int]]` with one element is rejected as pre-tokenized."""
import pytest
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
with pytest.raises(NotImplementedError, match="pre-tokenized"):
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {
"model": "bedrock/amazon.titan-embed-text-v2:0",
"input": [[1, 2, 3]],
},
}
]
)
def test_record_with_both_input_and_messages_routes_to_chat(self):
"""If a record has both fields, chat wins (safer default - see helper docstring)."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "ambiguous-1",
"body": {
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
"messages": [{"role": "user", "content": "Hi"}],
"input": "this should be ignored by chat path",
"max_tokens": 5,
},
}
]
)
assert "messages" in result[0]["modelInput"]
assert "inputText" not in result[0]["modelInput"]
def test_url_embeddings_with_missing_input_raises_not_chat_error(self):
"""url says embed, body lacks input → embedding-path error, not chat-path crash."""
import pytest
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
config = BedrockFilesConfig()
with pytest.raises(ValueError, match="missing required `input`"):
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "e1",
"method": "POST",
"url": "/v1/embeddings",
"body": {"model": "bedrock/amazon.titan-embed-text-v2:0"},
}
]
)
def test_titan_v2_marker_boundary_rejects_lookalikes(self):
"""The marker must end at `:`, `/`, or end-of-string to avoid false positives."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
# Look-alikes that must NOT route through the Titan v2 path
for model in (
"bedrock/amazon.titan-embed-text-v20:0",
"bedrock/amazon.titan-embed-text-v2-experimental:0",
"bedrock/amazon.titan-embed-text-v2foo",
):
assert not BedrockFilesConfig._is_titan_v2_embed_model(
model
), f"{model} unexpectedly matched the Titan v2 marker"
# Real Titan v2 ids that MUST match
for model in (
"amazon.titan-embed-text-v2:0",
"bedrock/amazon.titan-embed-text-v2:0",
"us.amazon.titan-embed-text-v2:0",
"arn:aws:bedrock:us-east-1:123:foundation-model/amazon.titan-embed-text-v2:0",
):
assert BedrockFilesConfig._is_titan_v2_embed_model(
model
), f"{model} unexpectedly missed the Titan v2 marker"
def test_titan_v2_accepted_when_registry_schema_field_matches(self, mocker):
"""Registry-driven happy path: nested
`provider_specific_entry.bedrock_invocation_schema == "titan_v2"`
is the authoritative signal."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
mocker.patch(
"litellm.get_model_info",
return_value={
"provider_specific_entry": {"bedrock_invocation_schema": "titan_v2"}
},
)
assert BedrockFilesConfig._is_titan_v2_embed_model(
"amazon.titan-embed-text-v2:0"
)
def test_titan_v2_rejected_when_registry_schema_field_differs(self, mocker):
"""Registry resolves with a different schema value (e.g. a hypothetical
Cohere Embed entry) -> reject. Registry is authoritative; no substring
second-chance for ids the registry knows."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
mocker.patch(
"litellm.get_model_info",
return_value={
"provider_specific_entry": {"bedrock_invocation_schema": "cohere_v3"}
},
)
# Even though the model id looks like Titan v2, the registry says
# otherwise and we trust it.
assert not BedrockFilesConfig._is_titan_v2_embed_model(
"amazon.titan-embed-text-v2:0"
)
def test_titan_v2_falls_back_to_marker_when_registry_lacks_schema_field(
self, mocker
):
"""Registry resolves but the entry has no
`provider_specific_entry.bedrock_invocation_schema` field yet (e.g.
a stale local registry) -> fall through to substring."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
# No provider_specific_entry at all
mocker.patch(
"litellm.get_model_info",
return_value={"mode": "embedding"},
)
assert BedrockFilesConfig._is_titan_v2_embed_model(
"amazon.titan-embed-text-v2:0"
)
# provider_specific_entry present but missing the schema key
mocker.patch(
"litellm.get_model_info",
return_value={
"mode": "embedding",
"provider_specific_entry": {"unrelated": "value"},
},
)
assert BedrockFilesConfig._is_titan_v2_embed_model(
"amazon.titan-embed-text-v2:0"
)
def test_titan_v2_accepted_when_registry_silent(self, mocker):
"""Marker-only match is fine for ids the registry can't resolve
(cross-region profile prefixes, ARN forms)."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
mocker.patch("litellm.get_model_info", side_effect=Exception("not mapped"))
assert BedrockFilesConfig._is_titan_v2_embed_model(
"us.amazon.titan-embed-text-v2:0"
)
assert BedrockFilesConfig._is_titan_v2_embed_model(
"arn:aws:bedrock:us-east-1:123:foundation-model/amazon.titan-embed-text-v2:0"
)
def test_lookup_provider_specific_field_helper(self, mocker):
"""Direct coverage of the nested registry field helper."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
# Happy path: returns the nested field's string value
mocker.patch(
"litellm.get_model_info",
return_value={
"provider_specific_entry": {"bedrock_invocation_schema": "titan_v2"}
},
)
assert (
BedrockFilesConfig._lookup_provider_specific_field(
"anything", "bedrock_invocation_schema"
)
== "titan_v2"
)
# Registry raises -> None
mocker.patch("litellm.get_model_info", side_effect=Exception("not mapped"))
assert (
BedrockFilesConfig._lookup_provider_specific_field("anything", "any")
is None
)
# Registry returns non-dict -> None
mocker.patch("litellm.get_model_info", return_value="not a dict")
assert (
BedrockFilesConfig._lookup_provider_specific_field("anything", "any")
is None
)
# Registry returns dict without provider_specific_entry -> None
mocker.patch("litellm.get_model_info", return_value={"mode": "embedding"})
assert (
BedrockFilesConfig._lookup_provider_specific_field(
"anything", "bedrock_invocation_schema"
)
is None
)
# provider_specific_entry exists but isn't a dict -> None
mocker.patch(
"litellm.get_model_info",
return_value={"provider_specific_entry": "not a dict"},
)
assert (
BedrockFilesConfig._lookup_provider_specific_field(
"anything", "bedrock_invocation_schema"
)
is None
)
# provider_specific_entry dict missing the requested field -> None
mocker.patch(
"litellm.get_model_info",
return_value={"provider_specific_entry": {"unrelated": "x"}},
)
assert (
BedrockFilesConfig._lookup_provider_specific_field(
"anything", "bedrock_invocation_schema"
)
is None
)
# Non-string nested value -> None
mocker.patch(
"litellm.get_model_info",
return_value={"provider_specific_entry": {"bedrock_invocation_schema": 42}},
)
assert (
BedrockFilesConfig._lookup_provider_specific_field(
"anything", "bedrock_invocation_schema"
)
is None
)
# Empty-string nested value -> None
mocker.patch(
"litellm.get_model_info",
return_value={"provider_specific_entry": {"bedrock_invocation_schema": ""}},
)
assert (
BedrockFilesConfig._lookup_provider_specific_field(
"anything", "bedrock_invocation_schema"
)
is None
)
def test_is_embedding_record_helper(self):
"""Helper detects embeddings via `url` first, then by body shape."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
assert BedrockFilesConfig._is_embedding_record(
{"url": "/v1/embeddings", "body": {"input": "x"}}
)
# body-only fallback
assert BedrockFilesConfig._is_embedding_record({"body": {"input": "x"}})
# chat shape
assert not BedrockFilesConfig._is_embedding_record(
{"url": "/v1/chat/completions", "body": {"messages": []}}
)
# ambiguous body without `input` is treated as not-embedding
assert not BedrockFilesConfig._is_embedding_record({"body": {}})
def test_explicit_chat_url_with_input_body_short_circuits_to_chat(self):
"""Explicit url=/v1/chat/completions wins even if body looks like embedding.
Without this short-circuit, a chat record whose body happens to carry
`input` (and no `messages`) would be mis-routed to the embedding
transformer, corrupting the modelInput.
"""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
# Direct helper assertion
assert not BedrockFilesConfig._is_embedding_record(
{
"url": "/v1/chat/completions",
"body": {
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
"input": "this would mis-route under the old precedence",
},
}
)
# End-to-end: a record like this routes through the chat path. We
# just need to make sure we DON'T silently produce an inputText
# body and call it a chat completion.
config = BedrockFilesConfig()
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
[
{
"custom_id": "explicit-chat-with-input",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
"messages": [{"role": "user", "content": "Hi"}],
"input": "should not become inputText",
"max_tokens": 5,
},
}
]
)
model_input = result[0]["modelInput"]
assert (
"inputText" not in model_input
), "explicit chat URL must not produce an embedding-shaped modelInput"
def test_coerce_embedding_input_helper_isolated(self):
"""Direct coverage of the extracted input-normalization helper."""
import pytest
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
# Happy paths
assert BedrockFilesConfig._coerce_embedding_input_to_string("hello") == "hello"
assert (
BedrockFilesConfig._coerce_embedding_input_to_string(["hello"]) == "hello"
)
# Error paths
with pytest.raises(ValueError, match="missing required `input`"):
BedrockFilesConfig._coerce_embedding_input_to_string(None, model="m")
with pytest.raises(ValueError, match="one input per JSONL record"):
BedrockFilesConfig._coerce_embedding_input_to_string(["a", "b"])
# A multi-element list of ints is rejected as "one input per JSONL
# record" too - we can't tell if it's pre-tokenized or "3 strings"
# without more context, so the most-actionable error wins.
with pytest.raises(ValueError, match="one input per JSONL record"):
BedrockFilesConfig._coerce_embedding_input_to_string([1, 2, 3])
# Single-element list wrapping a token list -> pre-tokenized error.
with pytest.raises(NotImplementedError, match="pre-tokenized"):
BedrockFilesConfig._coerce_embedding_input_to_string([[1, 2, 3]])
# Single-element list wrapping a bare int -> pre-tokenized error.
with pytest.raises(NotImplementedError, match="pre-tokenized"):
BedrockFilesConfig._coerce_embedding_input_to_string([42])
with pytest.raises(ValueError, match="must be a string"):
BedrockFilesConfig._coerce_embedding_input_to_string({"unsupported": True})
def test_other_non_embedding_urls_route_to_chat(self):
"""Any non-/v1/embeddings url short-circuits to chat path."""
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
# /v1/completions (legacy completions endpoint)
assert not BedrockFilesConfig._is_embedding_record(
{"url": "/v1/completions", "body": {"input": "x"}}
)
# Arbitrary unknown url - caller's explicit signal still wins
assert not BedrockFilesConfig._is_embedding_record(
{"url": "/v1/responses", "body": {"input": "x"}}
)