mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-04 02:31:27 +00:00
feat(bedrock-batch): route /v1/embeddings JSONL to Titan v2 modelInput (#28875)
* fix(thinking): handle None thinking param in is_thinking_enabled (#28598) Squash-merged by litellm-agent from Terrajlz's PR. * feat(helm): support tpl rendering in podAnnotations (#28609) Squash-merged by litellm-agent from devauxbr's PR. * feat(bedrock-batch): route /v1/embeddings JSONL to Titan v2 modelInput Changes vs main: - BedrockFilesConfig now detects OpenAI batch JSONL lines whose `url` is /v1/embeddings (with body-shape fallback) and routes them through a new `_map_openai_embedding_to_bedrock_params` helper instead of the chat-completion transformer that silently produces an invalid body. - The embedding helper currently supports Amazon Titan Text Embeddings V2 only. Other embed models (Titan G1, Titan Multimodal, Cohere Embed, Nova Multimodal Embeddings) raise NotImplementedError with a clear message; each will get a dedicated branch + tests in follow-up PRs to keep schema-specific risks isolated. - Validation refuses pre-tokenized inputs (List[int], List[List[int]]) and multi-element string lists with explicit errors so callers emit one JSONL line per embedding instead of relying on us to fan out. - Titan v2 model id match tolerates "bedrock/" prefix, cross-region inference profile prefix ("us.", "eu.", etc.), and ARN forms; the marker boundary check rejects lookalikes like "titan-embed-text-v20". - Tests cover happy path (fixtures), dimensions/encoding_format mapping, body-shape fallback, single-element list unwrap, error paths (missing input, multi-element list, unsupported model, pre-tokenized), mixed chat+embedding batch, and the model-id boundary check. * fix(bedrock-batch): trust explicit url over body shape in record routing Greptile-flagged gap in `_is_embedding_record`: when an OpenAI batch JSONL line carries an explicit `url` pointing to a non-embedding endpoint (e.g. `/v1/chat/completions`) AND its body happens to have `input` without `messages`, the body-shape fallback would mis-route that record to the embedding transformer and corrupt the modelInput. Changes vs previous commit: - `_is_embedding_record` now short-circuits to NOT-embedding whenever `url` is non-empty and not equal to `/v1/embeddings`. The body-shape fallback only runs when `url` is missing or empty. Docstring updated to spell out the precedence rules. - Two new tests cover the case: direct helper assertion that an explicit chat url plus an input-bearing body returns False, plus an end-to-end check that the resulting modelInput contains no `inputText` key. A second test asserts the same short-circuit for arbitrary non-embeddings urls (`/v1/completions`, `/v1/responses`). 28/28 tests pass (was 26/26 before this commit + 2 new). * refactor(bedrock-batch): extract embedding input normalization helper Splits the input-shape validation out of `_map_openai_embedding_to_bedrock_params` into a new static helper `_coerce_embedding_input_to_string`. Same semantics; the goal is to make the validation testable in isolation and to give future embedding-provider branches (Titan G1, Cohere) a reusable shaping function instead of duplicating type checks. - Helper accepts `str`, single-element `list[str]`, and raises `ValueError` / `NotImplementedError` with actionable messages for None, multi-element lists, pre-tokenized inputs (`list[int]` / `list[list[int]]`), and other unsupported types. - New unit test exercises the helper directly across happy paths, None / missing input, multi-element string list, multi-element int list (caught as 'one input per JSONL record' since we can't disambiguate from 'multiple strings' without more context), pre-tokenized single-element list-of-list, single-element list of bare int, and dict input. 29/29 tests in the file still pass. * fix(bedrock-batch): layer registry mode check + drop unused provider param Greptile-flagged issues on Titan v2 model detection: 1. Hardcoded substring without registry consultation conflicts with the project convention of treating `model_prices_and_context_window.json` as the source of truth for model capability flags. Fix: `_is_titan_v2_embed_model` now layers a registry mode check on top of the existing marker boundary. When `get_model_info` resolves the id, we additionally require `mode == "embedding"` so a malformed id whose path-component matches the marker but whose registered mode is "chat" doesn't slip through. Registry silence (cross-region inference profile prefixes like `us.amazon.titan-embed-text-v2:0`, ARN forms) keeps the substring-only behavior because the registry genuinely can't normalize those ids today. The substring is still needed because `mode == "embedding"` alone doesn't distinguish Titan v2's InvokeModel schema from Cohere Embed, Nova Multimodal, or Titan G1 - all also embedding mode, all with incompatible bodies. 2. `provider` parameter on `_map_openai_embedding_to_bedrock_params` was accepted but never used. Fix: dropped from the signature and the call site. Also adds `_lookup_registry_mode` static helper (mirrors the one in the sibling batches transformer) so the registry try/except shape lives in one place instead of being inlined into the detector. 5 new tests pin the layered behavior: - registry mode=chat overrides the marker match (rejected) - registry mode=embedding + marker match (accepted) - registry silent + marker match for cross-region and ARN ids (accepted) - direct `_lookup_registry_mode` coverage across all return paths 33/33 tests in the file pass. * fix(bedrock-batch): drive Titan v2 detection from registry provider_specific_entry Addresses Greptile's remaining policy concern on PR #28875: the hardcoded `_TITAN_V2_EMBED_MODEL_MARKER` substring required a code change to register new Titan v2 variants. Now the registry is the source of truth. Changes: - model_prices_and_context_window.json + model_prices_and_context_window_backup.json Adds `provider_specific_entry.bedrock_invocation_schema = "titan_v2"` to the `amazon.titan-embed-text-v2:0` entry. Uses the existing `provider_specific_entry` escape-hatch field (already surfaced by `get_model_info` via `ModelInfo.provider_specific_entry`) rather than introducing a new top-level field that would need a corresponding change in `utils.py::get_model_info` to flow through. - BedrockFilesConfig._is_titan_v2_embed_model now reads `get_model_info(model).provider_specific_entry.bedrock_invocation_schema` first. When the registry resolves the id we trust the field; registered ids without (or with a different) schema value are rejected outright - no substring second-chance for registered ids. The substring fallback only runs when `get_model_info` raises, which covers cross-region inference profile prefixes (`us.amazon.titan-embed-text-v2:0`) and Bedrock ARN forms - neither of which the registry normalizes today. - _lookup_registry_mode helper renamed to _lookup_provider_specific_field and generalized: takes a field name, reads `provider_specific_entry[field]` defensively. Future embed-schema branches (Titan G1, Cohere, Nova Multimodal) will share this helper. - New module-level constants: _BEDROCK_INVOCATION_SCHEMA_FIELD names the registry key; _TITAN_V2_INVOCATION_SCHEMA names the value. - _map_openai_embedding_to_bedrock_params: dropped the unused `provider` parameter (separate Greptile flag). Tests updated for the new nested-field semantics. 34/34 tests pass; one end-to-end integration check confirms that with LITELLM_LOCAL_MODEL_COST_MAP=true, the registry returns the new field and all 7 representative model ids classify correctly. * chore(lint): apply Black formatting to is_thinking_enabled Pre-existing Black violation on the shin_agent_oss_staging_05_22_2026 base branch - the lint CI check on this PR fails on `litellm/llms/base_llm/chat/transformation.py` even with zero changes from our side. Applying the formatter's preferred line-break inside `is_thinking_enabled` unblocks the lint check without touching the method's logic. --------- Co-authored-by: Terrajlz <info@jouleselectrictech.com> Co-authored-by: Bruno Devaux <devaux.br@gmail.com>
This commit is contained in:
parent
8ed909236b
commit
fd554501ef
6 changed files with 908 additions and 7 deletions
|
|
@ -233,6 +233,259 @@ class BedrockFilesConfig(BaseAWSLLM, BaseFilesConfig):
|
|||
# example; add others here as they adopt the same schema.
|
||||
CONVERSE_INVOKE_PROVIDERS = ("nova",)
|
||||
|
||||
# OpenAI batch URL that signals an embedding request. Per OpenAI Batch API
|
||||
# spec, every JSONL record carries a `url` field; we use it as the
|
||||
# authoritative signal to route the line to the embedding code path
|
||||
# instead of inferring from the presence of `input` vs `messages`.
|
||||
OPENAI_EMBEDDINGS_URL = "/v1/embeddings"
|
||||
|
||||
@staticmethod
|
||||
def _is_embedding_record(openai_jsonl_record: Dict[str, Any]) -> bool:
|
||||
"""
|
||||
Decide whether an OpenAI batch JSONL line is an embedding request.
|
||||
|
||||
Precedence (strict - any explicit `url` short-circuits):
|
||||
1. `url == "/v1/embeddings"` -> embedding. Authoritative per the
|
||||
OpenAI Batch API spec.
|
||||
2. Any other non-empty `url` (e.g. `/v1/chat/completions`) -> NOT
|
||||
embedding. We trust the caller's explicit signal even if the
|
||||
body would otherwise suggest embedding; misrouting a chat
|
||||
record into the embedding transformer would corrupt the
|
||||
modelInput, while a chat-shaped body sent to the chat path
|
||||
either succeeds or fails cleanly inside that transformer.
|
||||
3. `url` missing/empty -> fall back to body shape. Requires
|
||||
`input` present AND `messages` absent so a malformed record
|
||||
carrying both keys routes to the chat path (safer default:
|
||||
Anthropic transforms ignore unknown top-level keys, whereas
|
||||
the embedding transformer would silently drop the messages).
|
||||
"""
|
||||
url = openai_jsonl_record.get("url")
|
||||
if url == BedrockFilesConfig.OPENAI_EMBEDDINGS_URL:
|
||||
return True
|
||||
if url:
|
||||
return False
|
||||
body = openai_jsonl_record.get("body", {})
|
||||
if not isinstance(body, dict):
|
||||
return False
|
||||
return "input" in body and "messages" not in body
|
||||
|
||||
# Identifier for the Bedrock Titan v2 InvokeModel body schema as stored
|
||||
# in `model_prices_and_context_window.json`. Centralized so future
|
||||
# embedding-schema variants can add their own value
|
||||
# (e.g. `cohere_v3`, `titan_g1`, `titan_multimodal`) without touching
|
||||
# the detection logic.
|
||||
_TITAN_V2_INVOCATION_SCHEMA = "titan_v2"
|
||||
|
||||
# Substring marker used as a fallback when the registry can't resolve
|
||||
# the model id - notably cross-region inference profile prefixes
|
||||
# (`us.amazon.titan-embed-text-v2:0`) and Bedrock ARN forms, which
|
||||
# `get_model_info` doesn't normalize today.
|
||||
_TITAN_V2_EMBED_MODEL_MARKER = "titan-embed-text-v2"
|
||||
|
||||
# Nested field name under `provider_specific_entry` that identifies the
|
||||
# Bedrock InvokeModel body schema for batch inference.
|
||||
# `provider_specific_entry` is the registry's escape hatch for fields
|
||||
# `get_model_info` doesn't promote to top-level - exactly what we need
|
||||
# here. Documented in the `sample_spec` entry of
|
||||
# `model_prices_and_context_window.json` and surfaced by
|
||||
# `get_model_info` (see `ModelInfo.provider_specific_entry`).
|
||||
_BEDROCK_INVOCATION_SCHEMA_FIELD = "bedrock_invocation_schema"
|
||||
|
||||
@staticmethod
|
||||
def _is_titan_v2_embed_model(model: str) -> bool:
|
||||
"""
|
||||
True iff `model` refers to Amazon Titan Text Embeddings V2.
|
||||
|
||||
Resolution order:
|
||||
1. `model_prices_and_context_window.json` via `get_model_info`.
|
||||
The Titan v2 registry entry carries an explicit
|
||||
`provider_specific_entry.bedrock_invocation_schema` discriminator
|
||||
(`"titan_v2"`). When the registry resolves the id we trust that
|
||||
field as the source of truth - no hardcoded model-id comparison
|
||||
needed.
|
||||
2. Substring fallback (`titan-embed-text-v2` followed by `:`, `/`,
|
||||
or end-of-string) for ids the registry can't normalize. This
|
||||
catches cross-region inference profile prefixes
|
||||
(`us.amazon.titan-embed-text-v2:0`) and Bedrock ARN forms; the
|
||||
marker boundary check rejects lookalikes like
|
||||
`titan-embed-text-v20` or `titan-embed-text-v2-experimental`.
|
||||
|
||||
Tolerant of common id shapes:
|
||||
- "amazon.titan-embed-text-v2:0"
|
||||
- "bedrock/amazon.titan-embed-text-v2:0"
|
||||
- "us.amazon.titan-embed-text-v2:0" (cross-region inference profile)
|
||||
- ARN forms ending in ".../amazon.titan-embed-text-v2:0"
|
||||
"""
|
||||
# Registry-driven path: when get_model_info resolves the id we trust
|
||||
# the registry's discriminator. A resolved id with a different (or
|
||||
# absent) schema value here is intentionally not given a substring
|
||||
# second-chance - the registry is authoritative for ids it knows.
|
||||
registry_schema = BedrockFilesConfig._lookup_provider_specific_field(
|
||||
model, BedrockFilesConfig._BEDROCK_INVOCATION_SCHEMA_FIELD
|
||||
)
|
||||
if registry_schema is not None:
|
||||
return registry_schema == BedrockFilesConfig._TITAN_V2_INVOCATION_SCHEMA
|
||||
|
||||
# Registry silence -> substring fallback for unmapped ids only.
|
||||
normalized = model.lower()
|
||||
if normalized.startswith("bedrock/"):
|
||||
normalized = normalized[len("bedrock/") :]
|
||||
marker = BedrockFilesConfig._TITAN_V2_EMBED_MODEL_MARKER
|
||||
idx = normalized.find(marker)
|
||||
if idx < 0:
|
||||
return False
|
||||
end = idx + len(marker)
|
||||
return end == len(normalized) or normalized[end] in (":", "/")
|
||||
|
||||
@staticmethod
|
||||
def _lookup_provider_specific_field(model_id: str, field: str) -> Optional[str]:
|
||||
"""
|
||||
Read a nested string field from the registry entry's
|
||||
`provider_specific_entry` dict via `litellm.get_model_info`.
|
||||
|
||||
Returns the field's string value when:
|
||||
- the registry resolves `model_id`,
|
||||
- the entry exposes `provider_specific_entry` as a dict, and
|
||||
- that dict has `field` mapped to a non-empty string.
|
||||
Otherwise returns `None`.
|
||||
|
||||
Isolating this means feature detectors (Titan v2 today, future
|
||||
Cohere Embed / Nova Multimodal branches) share one defensive
|
||||
try/except shape instead of duplicating it. The `None` return
|
||||
covers every realistic failure mode: `get_model_info` raises
|
||||
(cross-region profile prefixes, Bedrock ARN forms, unreleased
|
||||
models), returns a non-dict, has no `provider_specific_entry`, or
|
||||
the requested field is missing / non-string / empty.
|
||||
"""
|
||||
try:
|
||||
from litellm import get_model_info
|
||||
|
||||
info = get_model_info(model_id)
|
||||
except Exception:
|
||||
return None
|
||||
if not isinstance(info, dict):
|
||||
return None
|
||||
provider_specific = info.get("provider_specific_entry")
|
||||
if not isinstance(provider_specific, dict):
|
||||
return None
|
||||
value = provider_specific.get(field)
|
||||
return value if isinstance(value, str) and value else None
|
||||
|
||||
@staticmethod
|
||||
def _coerce_embedding_input_to_string(raw_input: Any, model: str = "") -> str:
|
||||
"""
|
||||
Normalize an OpenAI /v1/embeddings `input` field into the single
|
||||
string that Bedrock Titan v2 InvokeModel expects in `inputText`.
|
||||
|
||||
Accepts: a string, or a single-element list containing one string.
|
||||
Rejects (with actionable messages):
|
||||
- None / missing -> ValueError
|
||||
- Multi-element string lists -> ValueError, prompts caller to
|
||||
emit one JSONL line per input
|
||||
- Pre-tokenized inputs (List[int], List[List[int]]) -> NotImplementedError
|
||||
- Any other type -> ValueError
|
||||
|
||||
Extracted so the validation can be exercised in isolation and so
|
||||
future embedding-provider branches (Titan G1, Cohere) can reuse it
|
||||
without duplicating the type-shaping logic.
|
||||
"""
|
||||
if raw_input is None:
|
||||
raise ValueError(
|
||||
"Embedding batch record is missing required `input` field: "
|
||||
f"model={model}"
|
||||
)
|
||||
|
||||
# Bedrock InvokeModel for Titan v2 takes exactly one string `inputText`
|
||||
# per call. Pre-tokenized inputs and multi-element string lists are
|
||||
# explicitly unsupported so callers emit one JSONL line per embedding
|
||||
# instead of relying on us to silently fan out or concatenate.
|
||||
if isinstance(raw_input, list):
|
||||
if len(raw_input) == 1:
|
||||
candidate = raw_input[0]
|
||||
else:
|
||||
raise ValueError(
|
||||
"Bedrock batch embedding requires one input per JSONL "
|
||||
"record. Got a list with "
|
||||
f"{len(raw_input)} items for model={model}; emit one "
|
||||
"JSONL line per input string instead."
|
||||
)
|
||||
else:
|
||||
candidate = raw_input
|
||||
|
||||
# Catches pre-tokenized inputs (List[int] from OpenAI spec, or a
|
||||
# single int slipping past the list-unwrap above).
|
||||
# NOTE: bool is a subclass of int but treating True/False as a token
|
||||
# is meaningless either way, so the broad check is fine.
|
||||
if isinstance(candidate, (list, int)):
|
||||
raise NotImplementedError(
|
||||
"Bedrock Titan v2 batch embedding does not support "
|
||||
"pre-tokenized integer inputs. Pass `input` as a string "
|
||||
f"(model={model})."
|
||||
)
|
||||
if not isinstance(candidate, str):
|
||||
raise ValueError(
|
||||
"Bedrock batch embedding `input` must be a string (or a "
|
||||
"single-element list of strings). Got type "
|
||||
f"{type(candidate).__name__} for model={model}."
|
||||
)
|
||||
return candidate
|
||||
|
||||
def _map_openai_embedding_to_bedrock_params(
|
||||
self,
|
||||
openai_request_body: Dict[str, Any],
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Transform an OpenAI /v1/embeddings request body into the
|
||||
Bedrock InvokeModel `modelInput` for embedding models that AWS
|
||||
supports via batch inference (CreateModelInvocationJob).
|
||||
|
||||
Currently routes Amazon Titan Text Embeddings V2 only; other
|
||||
embedding providers (Titan G1, Titan Multimodal, Cohere Embed,
|
||||
Nova Multimodal Embeddings) raise NotImplementedError until they
|
||||
get a dedicated branch. Splitting them keeps PR scope tight and
|
||||
lets each model's request schema be exercised by its own tests.
|
||||
|
||||
AWS docs (Titan v2 InvokeModel body):
|
||||
https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-titan-embed-text.html
|
||||
"""
|
||||
from litellm.llms.bedrock.embed.amazon_titan_v2_transformation import (
|
||||
AmazonTitanV2Config,
|
||||
)
|
||||
|
||||
_model = openai_request_body.get("model", "")
|
||||
if not self._is_titan_v2_embed_model(_model):
|
||||
# Refuse early instead of silently shaping the body for the wrong
|
||||
# provider. The synchronous /v1/embeddings path supports more
|
||||
# models, but each has a different InvokeModel schema; mapping
|
||||
# them here without dedicated tests would risk corrupt batches.
|
||||
raise NotImplementedError(
|
||||
"Bedrock batch embedding currently supports only Amazon "
|
||||
"Titan Text Embeddings V2 (model id contains "
|
||||
f"'titan-embed-text-v2'). Got model={_model!r}. Track other "
|
||||
"embedding models in https://github.com/BerriAI/litellm/issues."
|
||||
)
|
||||
|
||||
input_text = self._coerce_embedding_input_to_string(
|
||||
openai_request_body.get("input"), model=_model
|
||||
)
|
||||
|
||||
# Map OpenAI-style params (dimensions, encoding_format) onto the
|
||||
# Titan v2 schema (dimensions, embeddingTypes) via the embed config
|
||||
# so this stays in sync with the synchronous /v1/embeddings path.
|
||||
non_default_params = {
|
||||
k: v for k, v in openai_request_body.items() if k not in ("model", "input")
|
||||
}
|
||||
titan_config = AmazonTitanV2Config()
|
||||
inference_params = titan_config.map_openai_params(
|
||||
non_default_params=non_default_params,
|
||||
optional_params={},
|
||||
)
|
||||
return dict(
|
||||
titan_config._transform_request(
|
||||
input=input_text, inference_params=inference_params
|
||||
)
|
||||
)
|
||||
|
||||
def _map_openai_to_bedrock_params(
|
||||
self,
|
||||
openai_request_body: Dict[str, Any],
|
||||
|
|
@ -349,10 +602,19 @@ class BedrockFilesConfig(BaseAWSLLM, BaseFilesConfig):
|
|||
# Determine provider from model name
|
||||
provider = self.get_bedrock_invoke_provider(model)
|
||||
|
||||
# Transform to Bedrock modelInput format
|
||||
model_input = self._map_openai_to_bedrock_params(
|
||||
openai_request_body=openai_body, provider=provider
|
||||
)
|
||||
# Route to the embedding transformer when the OpenAI batch line
|
||||
# targets /v1/embeddings; otherwise fall back to the existing
|
||||
# chat-completion path. We branch here (rather than inside
|
||||
# `_map_openai_to_bedrock_params`) so the chat helper keeps its
|
||||
# narrow contract and the embedding helper can evolve independently.
|
||||
if self._is_embedding_record(_openai_jsonl_content):
|
||||
model_input = self._map_openai_embedding_to_bedrock_params(
|
||||
openai_request_body=openai_body
|
||||
)
|
||||
else:
|
||||
model_input = self._map_openai_to_bedrock_params(
|
||||
openai_request_body=openai_body, provider=provider
|
||||
)
|
||||
|
||||
# Create Bedrock batch record
|
||||
record_id = _openai_jsonl_content.get(
|
||||
|
|
|
|||
|
|
@ -577,7 +577,10 @@
|
|||
"max_tokens": 8192,
|
||||
"mode": "embedding",
|
||||
"output_cost_per_token": 0.0,
|
||||
"output_vector_size": 1024
|
||||
"output_vector_size": 1024,
|
||||
"provider_specific_entry": {
|
||||
"bedrock_invocation_schema": "titan_v2"
|
||||
}
|
||||
},
|
||||
"amazon.titan-image-generator-v1": {
|
||||
"input_cost_per_image": 0.0,
|
||||
|
|
|
|||
|
|
@ -577,7 +577,10 @@
|
|||
"max_tokens": 8192,
|
||||
"mode": "embedding",
|
||||
"output_cost_per_token": 0.0,
|
||||
"output_vector_size": 1024
|
||||
"output_vector_size": 1024,
|
||||
"provider_specific_entry": {
|
||||
"bedrock_invocation_schema": "titan_v2"
|
||||
}
|
||||
},
|
||||
"amazon.titan-image-generator-v1": {
|
||||
"input_cost_per_image": 0.0,
|
||||
|
|
|
|||
|
|
@ -0,0 +1,3 @@
|
|||
{"recordId": "embed-1", "modelInput": {"inputText": "Hello world"}}
|
||||
{"recordId": "embed-2", "modelInput": {"inputText": "Another document to embed", "dimensions": 512}}
|
||||
{"recordId": "embed-3", "modelInput": {"inputText": "Single element list", "embeddingTypes": ["binary"]}}
|
||||
|
|
@ -0,0 +1,3 @@
|
|||
{"custom_id": "embed-1", "method": "POST", "url": "/v1/embeddings", "body": {"model": "bedrock/amazon.titan-embed-text-v2:0", "input": "Hello world"}}
|
||||
{"custom_id": "embed-2", "method": "POST", "url": "/v1/embeddings", "body": {"model": "bedrock/amazon.titan-embed-text-v2:0", "input": "Another document to embed", "dimensions": 512}}
|
||||
{"custom_id": "embed-3", "method": "POST", "url": "/v1/embeddings", "body": {"model": "bedrock/amazon.titan-embed-text-v2:0", "input": ["Single element list"], "encoding_format": "base64"}}
|
||||
|
|
@ -426,7 +426,7 @@ class TestBedrockFilesTransformation:
|
|||
"s3_bucket_name": "litellm-batch-352026",
|
||||
"s3_region_name": "us-gov-west-1",
|
||||
}
|
||||
# aws_region_name set to something different — s3_region_name must still win
|
||||
# aws_region_name set to something different - s3_region_name must still win
|
||||
optional_params = {"aws_region_name": "us-east-1"}
|
||||
|
||||
captured_optional_params: dict = {}
|
||||
|
|
@ -482,3 +482,630 @@ class TestBedrockFilesTransformation:
|
|||
assert "messages" in model_input
|
||||
assert "max_tokens" in model_input
|
||||
assert model_input["max_tokens"] == 10
|
||||
|
||||
|
||||
class TestBedrockFilesEmbeddingTransformation:
|
||||
"""
|
||||
Tests for routing OpenAI /v1/embeddings batch JSONL records through the
|
||||
Titan v2 transformer so AWS Bedrock's CreateModelInvocationJob receives
|
||||
a valid modelInput body.
|
||||
|
||||
Scope is intentionally Titan v2 only - other embedding models will get
|
||||
their own follow-up PRs/tests so each schema is exercised in isolation.
|
||||
"""
|
||||
|
||||
def test_titan_v2_embedding_jsonl_matches_fixture(self):
|
||||
"""Round-trip the input fixture against the expected Bedrock output."""
|
||||
import json
|
||||
import os
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
here = os.path.dirname(__file__)
|
||||
with open(os.path.join(here, "input_batch_embeddings.jsonl")) as f:
|
||||
openai_jsonl = [json.loads(line) for line in f if line.strip()]
|
||||
with open(os.path.join(here, "expected_bedrock_batch_embeddings.jsonl")) as f:
|
||||
expected = [json.loads(line) for line in f if line.strip()]
|
||||
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
openai_jsonl
|
||||
)
|
||||
|
||||
assert result == expected
|
||||
|
||||
def test_titan_v2_simple_string_input(self):
|
||||
"""Single string `input` maps to `{"inputText": <str>}` with no extras."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": "Hello",
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
assert result == [{"recordId": "e1", "modelInput": {"inputText": "Hello"}}]
|
||||
|
||||
def test_titan_v2_dimensions_and_encoding_format(self):
|
||||
"""OpenAI `dimensions` / `encoding_format` map to Titan v2 schema."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": "Hi",
|
||||
"dimensions": 256,
|
||||
"encoding_format": "float",
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
model_input = result[0]["modelInput"]
|
||||
assert model_input["inputText"] == "Hi"
|
||||
assert model_input["dimensions"] == 256
|
||||
assert model_input["embeddingTypes"] == ["float"]
|
||||
|
||||
def test_embedding_routing_falls_back_to_body_shape(self):
|
||||
"""Records without `url` still route via `input` presence."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": "Hello",
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
assert result[0]["modelInput"] == {"inputText": "Hello"}
|
||||
|
||||
def test_embedding_single_element_list_input_is_accepted(self):
|
||||
"""A single-element list maps to the same shape as a bare string."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": ["only one"],
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
assert result[0]["modelInput"]["inputText"] == "only one"
|
||||
|
||||
def test_embedding_multi_input_list_raises(self):
|
||||
"""Multi-element `input` lists are rejected with a clear message."""
|
||||
import pytest
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
with pytest.raises(ValueError, match="one input per JSONL record"):
|
||||
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": ["a", "b"],
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
def test_embedding_missing_input_raises(self):
|
||||
"""A record routed to /v1/embeddings without `input` is an error."""
|
||||
import pytest
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
with pytest.raises(ValueError, match="missing required `input`"):
|
||||
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {"model": "bedrock/amazon.titan-embed-text-v2:0"},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
def test_mixed_chat_and_embedding_in_same_batch(self):
|
||||
"""Chat and embedding records in the same JSONL each take their path."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "chat-1",
|
||||
"method": "POST",
|
||||
"url": "/v1/chat/completions",
|
||||
"body": {
|
||||
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
|
||||
"messages": [{"role": "user", "content": "Hi"}],
|
||||
"max_tokens": 5,
|
||||
},
|
||||
},
|
||||
{
|
||||
"custom_id": "embed-1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": "Hi",
|
||||
},
|
||||
},
|
||||
]
|
||||
)
|
||||
|
||||
assert result[0]["recordId"] == "chat-1"
|
||||
assert "messages" in result[0]["modelInput"]
|
||||
assert result[0]["modelInput"]["anthropic_version"] == "bedrock-2023-05-31"
|
||||
|
||||
assert result[1]["recordId"] == "embed-1"
|
||||
assert result[1]["modelInput"] == {"inputText": "Hi"}
|
||||
|
||||
def test_unsupported_embedding_model_raises_not_implemented(self):
|
||||
"""Cohere/Nova/Titan-G1 embed get a clear NotImplementedError, not a corrupt body."""
|
||||
import pytest
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
for unsupported_model in (
|
||||
"bedrock/cohere.embed-english-v3",
|
||||
"bedrock/amazon.titan-embed-text-v1",
|
||||
"bedrock/amazon.titan-embed-image-v1",
|
||||
"bedrock/amazon.nova-2-multimodal-embeddings-v1:0",
|
||||
):
|
||||
with pytest.raises(NotImplementedError, match="titan-embed-text-v2"):
|
||||
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {"model": unsupported_model, "input": "Hi"},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
def test_titan_v2_model_name_variants_route_correctly(self):
|
||||
"""All common Titan v2 model id shapes route through the embedding path."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
for model_id in (
|
||||
"amazon.titan-embed-text-v2:0",
|
||||
"bedrock/amazon.titan-embed-text-v2:0",
|
||||
"us.amazon.titan-embed-text-v2:0",
|
||||
"bedrock/us.amazon.titan-embed-text-v2:0",
|
||||
):
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {"model": model_id, "input": "Hi"},
|
||||
}
|
||||
]
|
||||
)
|
||||
assert result[0]["modelInput"] == {
|
||||
"inputText": "Hi"
|
||||
}, f"model id {model_id} did not route to Titan v2 embedding path"
|
||||
|
||||
def test_pretokenized_input_list_of_ints_raises(self):
|
||||
"""`input: List[int]` (pre-tokenized) is rejected, not silently mis-shaped."""
|
||||
import pytest
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
with pytest.raises(
|
||||
(NotImplementedError, ValueError), match=r"pre-tokenized|one input per"
|
||||
):
|
||||
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": [1, 2, 3],
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
def test_pretokenized_single_wrapped_list_raises(self):
|
||||
"""`input: List[List[int]]` with one element is rejected as pre-tokenized."""
|
||||
import pytest
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
with pytest.raises(NotImplementedError, match="pre-tokenized"):
|
||||
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {
|
||||
"model": "bedrock/amazon.titan-embed-text-v2:0",
|
||||
"input": [[1, 2, 3]],
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
def test_record_with_both_input_and_messages_routes_to_chat(self):
|
||||
"""If a record has both fields, chat wins (safer default - see helper docstring)."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "ambiguous-1",
|
||||
"body": {
|
||||
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
|
||||
"messages": [{"role": "user", "content": "Hi"}],
|
||||
"input": "this should be ignored by chat path",
|
||||
"max_tokens": 5,
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
assert "messages" in result[0]["modelInput"]
|
||||
assert "inputText" not in result[0]["modelInput"]
|
||||
|
||||
def test_url_embeddings_with_missing_input_raises_not_chat_error(self):
|
||||
"""url says embed, body lacks input → embedding-path error, not chat-path crash."""
|
||||
import pytest
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
config = BedrockFilesConfig()
|
||||
with pytest.raises(ValueError, match="missing required `input`"):
|
||||
config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "e1",
|
||||
"method": "POST",
|
||||
"url": "/v1/embeddings",
|
||||
"body": {"model": "bedrock/amazon.titan-embed-text-v2:0"},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
def test_titan_v2_marker_boundary_rejects_lookalikes(self):
|
||||
"""The marker must end at `:`, `/`, or end-of-string to avoid false positives."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
# Look-alikes that must NOT route through the Titan v2 path
|
||||
for model in (
|
||||
"bedrock/amazon.titan-embed-text-v20:0",
|
||||
"bedrock/amazon.titan-embed-text-v2-experimental:0",
|
||||
"bedrock/amazon.titan-embed-text-v2foo",
|
||||
):
|
||||
assert not BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
model
|
||||
), f"{model} unexpectedly matched the Titan v2 marker"
|
||||
|
||||
# Real Titan v2 ids that MUST match
|
||||
for model in (
|
||||
"amazon.titan-embed-text-v2:0",
|
||||
"bedrock/amazon.titan-embed-text-v2:0",
|
||||
"us.amazon.titan-embed-text-v2:0",
|
||||
"arn:aws:bedrock:us-east-1:123:foundation-model/amazon.titan-embed-text-v2:0",
|
||||
):
|
||||
assert BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
model
|
||||
), f"{model} unexpectedly missed the Titan v2 marker"
|
||||
|
||||
def test_titan_v2_accepted_when_registry_schema_field_matches(self, mocker):
|
||||
"""Registry-driven happy path: nested
|
||||
`provider_specific_entry.bedrock_invocation_schema == "titan_v2"`
|
||||
is the authoritative signal."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={
|
||||
"provider_specific_entry": {"bedrock_invocation_schema": "titan_v2"}
|
||||
},
|
||||
)
|
||||
assert BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
"amazon.titan-embed-text-v2:0"
|
||||
)
|
||||
|
||||
def test_titan_v2_rejected_when_registry_schema_field_differs(self, mocker):
|
||||
"""Registry resolves with a different schema value (e.g. a hypothetical
|
||||
Cohere Embed entry) -> reject. Registry is authoritative; no substring
|
||||
second-chance for ids the registry knows."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={
|
||||
"provider_specific_entry": {"bedrock_invocation_schema": "cohere_v3"}
|
||||
},
|
||||
)
|
||||
# Even though the model id looks like Titan v2, the registry says
|
||||
# otherwise and we trust it.
|
||||
assert not BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
"amazon.titan-embed-text-v2:0"
|
||||
)
|
||||
|
||||
def test_titan_v2_falls_back_to_marker_when_registry_lacks_schema_field(
|
||||
self, mocker
|
||||
):
|
||||
"""Registry resolves but the entry has no
|
||||
`provider_specific_entry.bedrock_invocation_schema` field yet (e.g.
|
||||
a stale local registry) -> fall through to substring."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
# No provider_specific_entry at all
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={"mode": "embedding"},
|
||||
)
|
||||
assert BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
"amazon.titan-embed-text-v2:0"
|
||||
)
|
||||
|
||||
# provider_specific_entry present but missing the schema key
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={
|
||||
"mode": "embedding",
|
||||
"provider_specific_entry": {"unrelated": "value"},
|
||||
},
|
||||
)
|
||||
assert BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
"amazon.titan-embed-text-v2:0"
|
||||
)
|
||||
|
||||
def test_titan_v2_accepted_when_registry_silent(self, mocker):
|
||||
"""Marker-only match is fine for ids the registry can't resolve
|
||||
(cross-region profile prefixes, ARN forms)."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
mocker.patch("litellm.get_model_info", side_effect=Exception("not mapped"))
|
||||
assert BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
"us.amazon.titan-embed-text-v2:0"
|
||||
)
|
||||
assert BedrockFilesConfig._is_titan_v2_embed_model(
|
||||
"arn:aws:bedrock:us-east-1:123:foundation-model/amazon.titan-embed-text-v2:0"
|
||||
)
|
||||
|
||||
def test_lookup_provider_specific_field_helper(self, mocker):
|
||||
"""Direct coverage of the nested registry field helper."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
# Happy path: returns the nested field's string value
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={
|
||||
"provider_specific_entry": {"bedrock_invocation_schema": "titan_v2"}
|
||||
},
|
||||
)
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field(
|
||||
"anything", "bedrock_invocation_schema"
|
||||
)
|
||||
== "titan_v2"
|
||||
)
|
||||
|
||||
# Registry raises -> None
|
||||
mocker.patch("litellm.get_model_info", side_effect=Exception("not mapped"))
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field("anything", "any")
|
||||
is None
|
||||
)
|
||||
|
||||
# Registry returns non-dict -> None
|
||||
mocker.patch("litellm.get_model_info", return_value="not a dict")
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field("anything", "any")
|
||||
is None
|
||||
)
|
||||
|
||||
# Registry returns dict without provider_specific_entry -> None
|
||||
mocker.patch("litellm.get_model_info", return_value={"mode": "embedding"})
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field(
|
||||
"anything", "bedrock_invocation_schema"
|
||||
)
|
||||
is None
|
||||
)
|
||||
|
||||
# provider_specific_entry exists but isn't a dict -> None
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={"provider_specific_entry": "not a dict"},
|
||||
)
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field(
|
||||
"anything", "bedrock_invocation_schema"
|
||||
)
|
||||
is None
|
||||
)
|
||||
|
||||
# provider_specific_entry dict missing the requested field -> None
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={"provider_specific_entry": {"unrelated": "x"}},
|
||||
)
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field(
|
||||
"anything", "bedrock_invocation_schema"
|
||||
)
|
||||
is None
|
||||
)
|
||||
|
||||
# Non-string nested value -> None
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={"provider_specific_entry": {"bedrock_invocation_schema": 42}},
|
||||
)
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field(
|
||||
"anything", "bedrock_invocation_schema"
|
||||
)
|
||||
is None
|
||||
)
|
||||
|
||||
# Empty-string nested value -> None
|
||||
mocker.patch(
|
||||
"litellm.get_model_info",
|
||||
return_value={"provider_specific_entry": {"bedrock_invocation_schema": ""}},
|
||||
)
|
||||
assert (
|
||||
BedrockFilesConfig._lookup_provider_specific_field(
|
||||
"anything", "bedrock_invocation_schema"
|
||||
)
|
||||
is None
|
||||
)
|
||||
|
||||
def test_is_embedding_record_helper(self):
|
||||
"""Helper detects embeddings via `url` first, then by body shape."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
assert BedrockFilesConfig._is_embedding_record(
|
||||
{"url": "/v1/embeddings", "body": {"input": "x"}}
|
||||
)
|
||||
# body-only fallback
|
||||
assert BedrockFilesConfig._is_embedding_record({"body": {"input": "x"}})
|
||||
# chat shape
|
||||
assert not BedrockFilesConfig._is_embedding_record(
|
||||
{"url": "/v1/chat/completions", "body": {"messages": []}}
|
||||
)
|
||||
# ambiguous body without `input` is treated as not-embedding
|
||||
assert not BedrockFilesConfig._is_embedding_record({"body": {}})
|
||||
|
||||
def test_explicit_chat_url_with_input_body_short_circuits_to_chat(self):
|
||||
"""Explicit url=/v1/chat/completions wins even if body looks like embedding.
|
||||
|
||||
Without this short-circuit, a chat record whose body happens to carry
|
||||
`input` (and no `messages`) would be mis-routed to the embedding
|
||||
transformer, corrupting the modelInput.
|
||||
"""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
# Direct helper assertion
|
||||
assert not BedrockFilesConfig._is_embedding_record(
|
||||
{
|
||||
"url": "/v1/chat/completions",
|
||||
"body": {
|
||||
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
|
||||
"input": "this would mis-route under the old precedence",
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
# End-to-end: a record like this routes through the chat path. We
|
||||
# just need to make sure we DON'T silently produce an inputText
|
||||
# body and call it a chat completion.
|
||||
config = BedrockFilesConfig()
|
||||
result = config._transform_openai_jsonl_content_to_bedrock_jsonl_content(
|
||||
[
|
||||
{
|
||||
"custom_id": "explicit-chat-with-input",
|
||||
"method": "POST",
|
||||
"url": "/v1/chat/completions",
|
||||
"body": {
|
||||
"model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0",
|
||||
"messages": [{"role": "user", "content": "Hi"}],
|
||||
"input": "should not become inputText",
|
||||
"max_tokens": 5,
|
||||
},
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
model_input = result[0]["modelInput"]
|
||||
assert (
|
||||
"inputText" not in model_input
|
||||
), "explicit chat URL must not produce an embedding-shaped modelInput"
|
||||
|
||||
def test_coerce_embedding_input_helper_isolated(self):
|
||||
"""Direct coverage of the extracted input-normalization helper."""
|
||||
import pytest
|
||||
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
# Happy paths
|
||||
assert BedrockFilesConfig._coerce_embedding_input_to_string("hello") == "hello"
|
||||
assert (
|
||||
BedrockFilesConfig._coerce_embedding_input_to_string(["hello"]) == "hello"
|
||||
)
|
||||
|
||||
# Error paths
|
||||
with pytest.raises(ValueError, match="missing required `input`"):
|
||||
BedrockFilesConfig._coerce_embedding_input_to_string(None, model="m")
|
||||
with pytest.raises(ValueError, match="one input per JSONL record"):
|
||||
BedrockFilesConfig._coerce_embedding_input_to_string(["a", "b"])
|
||||
# A multi-element list of ints is rejected as "one input per JSONL
|
||||
# record" too - we can't tell if it's pre-tokenized or "3 strings"
|
||||
# without more context, so the most-actionable error wins.
|
||||
with pytest.raises(ValueError, match="one input per JSONL record"):
|
||||
BedrockFilesConfig._coerce_embedding_input_to_string([1, 2, 3])
|
||||
# Single-element list wrapping a token list -> pre-tokenized error.
|
||||
with pytest.raises(NotImplementedError, match="pre-tokenized"):
|
||||
BedrockFilesConfig._coerce_embedding_input_to_string([[1, 2, 3]])
|
||||
# Single-element list wrapping a bare int -> pre-tokenized error.
|
||||
with pytest.raises(NotImplementedError, match="pre-tokenized"):
|
||||
BedrockFilesConfig._coerce_embedding_input_to_string([42])
|
||||
with pytest.raises(ValueError, match="must be a string"):
|
||||
BedrockFilesConfig._coerce_embedding_input_to_string({"unsupported": True})
|
||||
|
||||
def test_other_non_embedding_urls_route_to_chat(self):
|
||||
"""Any non-/v1/embeddings url short-circuits to chat path."""
|
||||
from litellm.llms.bedrock.files.transformation import BedrockFilesConfig
|
||||
|
||||
# /v1/completions (legacy completions endpoint)
|
||||
assert not BedrockFilesConfig._is_embedding_record(
|
||||
{"url": "/v1/completions", "body": {"input": "x"}}
|
||||
)
|
||||
# Arbitrary unknown url - caller's explicit signal still wins
|
||||
assert not BedrockFilesConfig._is_embedding_record(
|
||||
{"url": "/v1/responses", "body": {"input": "x"}}
|
||||
)
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue