litellm/litellm/llms
yucheng-berri b2202cb1aa
feat(guardrails): streaming text transformation in generic_guardrail_api (#33110)
* feat(guardrails): support streaming text transformation in generic_guardrail_api

* chore(guardrails): address PR review feedback

* fix(guardrails): fail closed on tool-call and prefix-rewrite leaks in streaming transform

* fix(guardrails): address Bugbot review on streaming transform correctness

* fix(guardrails): coerce holdback in handler for in-process guardrails

* fix(guardrails): harden streaming transform (holdback coercion, tool-call passthrough, n>1 finish_reason)

* test(guardrails): targeted _mode_matches coverage for all guardrail_mode shapes

* fix(guardrails): inspect streamed tool calls and harden incremental_diff edge cases

* test: move ComplianceChecker mode tests to the compliance PR

* fix(guardrails): strip content from tool-call passthrough so streamed text can't bypass the transform

* fix(guardrails): four correctness fixes for incremental_diff streaming path

Four bug fixes on top of the OSS PR's incremental_diff streaming text
transformation, all inside the incremental_diff code paths only. No
existing block_only, non-streaming, or pre_call behavior is touched.

Fix #1 — Mixed content+tool_call finish_reason ordering
  _tool_call_passthrough_chunk now takes an optional finish_reason_per_choice
  map. For a choice carrying both delta.content and delta.tool_calls,
  finish_reason is stripped from the passthrough and recorded on the map so
  the final synthetic text chunk delivers it. Without this, SSE-compliant
  clients stopping at finish_reason drop the guardrailed text — defeating
  the redaction the whole feature exists for. (Greptile P1 twice, Veria.)

Fix #2 — Choice index sort in _process_streaming_transform
  indices/texts_to_check were derived from dict insertion order. For n>1
  streams where choice 1 emits before choice 0, guardrail-returned texts
  aligned to the input order mapped back to the wrong choice indices on
  write-back — wrong text goes to wrong choice. Sort raw_by_index.keys()
  up front so realignment is deterministic. (Bugbot Medium.)

Fix #3 — Cross-chunk pre-tool-call text flush
  With default streaming_sampling_rate=5, text chunks followed by a pure
  tool-call chunk carrying finish_reason='tool_calls' would emit the
  passthrough with finish_reason before any transformed text delta had
  fired. Same failure mode as fix #1 but cross-chunk. Now we flush any
  accumulated text via _round(is_final=False) BEFORE yielding the
  tool-call passthrough. (Greptile P1.)

Fix #4 — Terminator chunk for deferred finish_reason on empty mutated_text
  _build_transform_chunk returned None early when mutated_text_per_choice
  was empty. If a mixed content+tool_call chunk had deferred its
  finish_reason (via fix #1) and the guardrail then suppressed the text
  (empty return), the deferred finish_reason was never delivered. Now on
  is_final=True with empty mutated_text_per_choice, we emit a terminator
  carrying finish_reason per choice from finish_reason_per_choice.
  (Bugbot High.)

Also normalized Optional[X] → X | None across the OSS PR's added surface
via ruff UP045 autofix to keep the strict-rule gate within budget. Pure
mechanical typing style change, no semantic effect.

Regression tests for all four fixes:
- test_mixed_chunk_finish_reason_arrives_after_transformed_text (#1)
- test_text_flush_precedes_tool_call_passthrough (#3)
- test_final_finish_reason_flushed_when_guardrail_suppresses_text (#4)
- test_transform_sends_texts_sorted_by_choice_index (#2)

All fixes reachable only when streaming_transform_mode == 'incremental_diff'
is configured (via _run_incremental_transform_stream) or when a
StreamTransformSink is present (via _process_streaming_transform). Verified
scope-clean: no changes to block_only, non-streaming, pre_call, moderation,
or sibling guardrails.

---------

Co-authored-by: Marton Schneider <marton@schneider.co.nl>
2026-07-14 17:38:11 -07:00
..
a2a fix(a2a): populate response usage in a2a chat transformation (#31980) 2026-07-03 09:28:36 +05:30
ai21/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
aiml style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
aiohttp_openai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
amazon_nova style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
anthropic fix(anthropic): use native output capability (#33235) 2026-07-14 14:23:49 -07:00
apiserpent style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
aws_polly style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
azure fix(azure): build responses input_items url with path before query string (#32270) 2026-07-06 14:02:44 -07:00
azure_ai fix(anthropic): override custom_llm_provider in provider config subclasses so capability probes use the right namespace 2026-07-11 12:15:59 -07:00
base_llm feat(guardrails): streaming text transformation in generic_guardrail_api (#33110) 2026-07-14 17:38:11 -07:00
baseten style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
bedrock Merge pull request #32956 from BerriAI/litellm_fix_lit3859_wif_bridge 2026-07-13 11:50:45 -07:00
bedrock_mantle style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
black_forest_labs style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
brave/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
bytez style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cerebras fix: add reasoning param support for GPT OSS cerebras 2026-02-02 17:20:04 +05:30
chatgpt style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
clarifai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cloudflare/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
codestral/completion style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cohere style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
cometapi style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
compactifai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
custom_httpx fix(responses): surface upstream error status on get instead of 500 (#32287) 2026-07-06 22:45:25 -07:00
dashscope fix(cost): coerce string tiered-pricing costs and share tier helper 2026-07-11 14:05:26 -07:00
databricks fix(anthropic): override custom_llm_provider in provider config subclasses so capability probes use the right namespace 2026-07-11 12:15:59 -07:00
dataforseo/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
datarobot/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
deepgram style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
deepinfra style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
deepseek test: litellm fix failing tests (#32577) 2026-07-09 13:54:45 -07:00
deprecated_providers style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
docker_model_runner/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
duckduckgo/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
e2b style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
elevenlabs style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
empower/chat LiteLLM Common Base LLM Config (pt.3): Move all OAI compatible providers to base llm config (#7148) 2024-12-10 17:12:42 -08:00
exa_ai/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
fal_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
fastcrw style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
featherless_ai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
firecrawl style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
fireworks_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
friendliai/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
galadriel/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
gdc feat(gdc): implement Google Distributed Cloud (GDC) Gemini provider (#31895) 2026-07-01 17:31:07 -07:00
gemini fix(realtime): stop second Gemini Live setup, retry hung handshake, close guardrail bypass (#31519) 2026-06-28 08:52:20 +05:30
gigachat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
github/chat (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
github_copilot fix(anthropic): override custom_llm_provider in provider config subclasses so capability probes use the right namespace 2026-07-11 12:15:59 -07:00
google_pse/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
gradient_ai/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
groq style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
heroku/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
hosted_vllm style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
huggingface style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
hyperbolic Revert "Litellm dev 07 21 2025 p1 (#12848)" 2025-07-22 18:28:36 -07:00
inception style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
infinity style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
jina_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
lambda_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
langflow style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
langgraph style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
lemonade style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
linkup style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
litellm_proxy style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
llamafile/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
lm_studio style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
manus style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
meta_llama/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
milvus/vector_stores style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
minimax style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
mistral style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
modelscope style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
moonshot/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
morph style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
nebius Integration with Nebius AI Studio added (#11143) 2025-05-27 11:05:22 -07:00
nlp_cloud style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
novita/chat build(deps-dev): bump black to 26.3.1 and apply formatting (#28525) 2026-05-21 17:24:18 -07:00
nscale/chat style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
nvidia_nim style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
nvidia_riva style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
oci chore(lint): zero out crash-class pyright rules and ban new type: ignore comments (#32152) 2026-07-04 16:56:12 -07:00
ollama chore(lint): zero out crash-class pyright rules and ban new type: ignore comments (#32152) 2026-07-04 16:56:12 -07:00
oobabooga style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
openai feat(guardrails): streaming text transformation in generic_guardrail_api (#33110) 2026-07-14 17:38:11 -07:00
openai_like fix(anthropic): override custom_llm_provider in provider config subclasses so capability probes use the right namespace 2026-07-11 12:15:59 -07:00
openrouter style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
opensandbox style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
ovhcloud style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
parallel_ai/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
pass_through style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
perplexity style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
petals style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
pg_vector/vector_stores style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
predibase style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
ragflow style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
recraft style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
reducto style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
replicate style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
runwayml style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
s3_vectors style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
sagemaker style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
sambanova style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
sap chore(lint): zero out crash-class pyright rules and ban new type: ignore comments (#32152) 2026-07-04 16:56:12 -07:00
scaleway/audio_transcription style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
searchapi style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
searxng style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
serper/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
snowflake style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
soniox style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
stability style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
tavily/search style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
tencent feat(tencent): add Tencent TokenHub as a provider (#31903) 2026-07-02 18:31:59 -07:00
tinyfish/search feat(tinyfish): make search provider permissive, attribute errors (#31997) 2026-07-03 10:17:11 -07:00
together_ai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
topaz style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
triton style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
v0 style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
vercel_ai_gateway style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
vertex_ai fix(gemini): map video response modality instead of MODALITY_UNSPECIFIED 2026-07-14 14:11:52 -07:00
vllm style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
volcengine style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
voyage style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
wandb (feat): Add W&B Inference to LiteLLM 2025-09-11 00:07:30 +05:30
watsonx style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
xai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
xinference/image_generation style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
you_com style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
zai style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
__init__.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
base.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
custom_llm.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
maritalk.py style: unify ruff format width on 120 (#31518) 2026-06-27 12:39:29 -07:00
README.md LiteLLM Minor Fixes and Improvements (09/13/2024) (#5689) 2024-09-14 10:02:55 -07:00

File Structure

August 27th, 2024

To make it easy to see how calls are transformed for each model/provider:

we are working on moving all supported litellm providers to a folder structure, where folder name is the supported litellm provider name.

Each folder will contain a *_transformation.py file, which has all the request/response transformation logic, making it easy to see how calls are modified.

E.g. cohere/, bedrock/.