litellm/litellm
devin-ai-integration[bot] e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
..
_v2 feat(cache): select Rust caching through explicit cache objects (#43601) 2026-09-29 00:01:44 +00:00
a2a_protocol Merge remote-tracking branch 'origin/main' into litellm_decrease_anys_opus5_r5 2026-09-21 12:57:38 -07:00
anthropic_interface fix(proxy): send a real error event when a /v1/messages stream fails (#41826) 2026-09-25 15:36:35 -07:00
assistants refactor(types): replace Any with proven types in 5 files (#43304) 2026-09-27 01:28:02 -07:00
batch_completion
batches refactor: daily fresh tech debt cleanup, rolling PR (2026-09-25) (#43151) 2026-09-26 01:32:55 -07:00
caching feat(cache): select Rust caching through explicit cache objects (#43601) 2026-09-29 00:01:44 +00:00
chat_completions refactor(rust): remove delivery routing abstraction (#43514) 2026-09-28 03:31:22 +00:00
completion_extras fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870) 2026-09-29 12:54:17 -07:00
compression merge: main into litellm_headroom_protect_cached_prefix 2026-09-15 04:12:28 +00:00
containers test: finish the non-proxy half of tests/test_litellm (#43281) 2026-09-25 22:43:41 -07:00
embeddings feat(embeddings): add native dispatch foundation (#42799) 2026-09-23 21:11:09 +00:00
endpoints/speech/speech_to_completion_bridge
evals
experimental_mcp_client feat(mcp): configure protocol versions and capability discovery (#43169) 2026-09-25 13:13:52 -07:00
files feat(xai): add native xAI batches and files support (#42812) 2026-09-25 15:35:20 -07:00
fine_tuning
google_genai fix(google_genai): preserve proxy_server_request in completion adapter (#43536) 2026-09-28 21:29:55 -07:00
images fix(params): validate stream_chunk_size once, before any provider call (#43222) 2026-09-26 23:01:20 +00:00
integrations fix(otel): send cache and reasoning tokens in langfuse usage_details (#43553) 2026-09-28 21:30:56 -07:00
interactions feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
litellm_core_utils fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870) 2026-09-29 12:54:17 -07:00
llms fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870) 2026-09-29 12:54:17 -07:00
messages refactor(rust): remove delivery routing abstraction (#43514) 2026-09-28 03:31:22 +00:00
models fix(autorouter): compare historical and new savings consistently (#43348) 2026-09-29 12:40:19 -07:00
ocr refactor(ocr): remove the Python OCR execution path and require the Rust route (#43081) 2026-09-24 18:18:50 -07:00
passthrough fix(responses): fall back on pre-output stream drops, fail truncated streams, honor request_timeout (#43133) 2026-09-25 14:32:41 -07:00
proxy fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870) 2026-09-29 12:54:17 -07:00
proxy_auth
rag fix(types): read upstream headers through a typed helper 2026-09-21 13:16:58 -07:00
realtime_api refactor(types): replace Any with proven types in 13 files (#42937) 2026-09-24 09:02:02 -07:00
repositories feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
rerank_api feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
responses refactor(types): replace Any with proven types in 8 files (#43551) 2026-09-28 04:26:41 -07:00
router_strategy feat(router): opt in to prompt-cache cost routing (#43232) 2026-09-26 18:13:13 -07:00
router_utils fix(router): match provider-prefixed fallback keys for bare model groups served by wildcard deployments (#43062) 2026-09-24 18:08:56 -07:00
rust_bridge refactor: clean up fresh tech debt from 2026-09-28 (#43674) 2026-09-29 02:15:12 -07:00
sandbox
search
secret_managers fix(ci): stop stale CI reds, keep unit tests off the host env, retry CyberArk policy conflicts (#43294) 2026-09-26 09:25:13 -07:00
skills
types fix(autorouter): compare historical and new savings consistently (#43348) 2026-09-29 12:40:19 -07:00
vector_store_files
vector_stores fix(vector_stores): keep config-defined vector stores listed and read-only (#42574) 2026-09-23 04:02:16 +00:00
videos
__init__.py fix(proxy): log key owner identity on expired key auth failures (#43105) 2026-09-28 08:18:48 -07:00
_internal_context.py fix(otel): detach post-response service spans by request phase, name redis spans by operation (#43237) 2026-09-26 10:10:12 -07:00
_lazy_imports.py feat(tokenizer): preserve Python defaults with opt-in Rust dispatch (#42174) 2026-09-22 04:41:11 +00:00
_lazy_imports_registry.py refactor(anthropic): rename experimental_pass_through to pass_through (#43329) 2026-09-26 13:00:50 -07:00
_logging.py feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
_redis.py fix(redis): authenticate sync clusters with IAM credential providers (#40204) 2026-09-23 17:23:55 -05:00
_redis_credential_provider.py
_service_logger.py fix(otel): detach post-response service spans by request phase, name redis spans by operation (#43237) 2026-09-26 10:10:12 -07:00
_uuid.py
_version.py
anthropic_beta_headers_config.json fix(vertex_ai): forward the per-turn-control beta for per-message output_config (#43558) 2026-09-28 21:35:08 -07:00
anthropic_beta_headers_manager.py fix: drop a blank anthropic-beta header before it reaches the provider 2026-09-19 20:13:59 -07:00
blog_posts.json
budget_manager.py
constants.py feat: add model leaderboard page (#43649) 2026-09-28 19:40:04 -07:00
cost.json
cost_calculator.py fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870) 2026-09-29 12:54:17 -07:00
exceptions.py fix(spend): return 400 from /spend/calculate for a model with no pricing row (#42497) 2026-09-22 15:44:00 -07:00
main.py fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614) 2026-09-29 11:27:14 -07:00
model_prices_and_context_window_backup.json fix(model-prices): align Azure, Bedrock, Copilot, Gemini, Groq, OpenAI and OpenRouter entries with official docs (#43598) 2026-09-29 12:45:22 -07:00
policy_templates_backup.json
provider_endpoints_support_backup.json feat(providers): add Prism provider (internal copy of #40914) (#41961) 2026-09-28 16:07:09 -07:00
py.typed
router.py fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870) 2026-09-29 12:54:17 -07:00
scheduler.py
setup_wizard.py fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys (#43587) 2026-09-28 19:12:31 +00:00
timeout.py
utils.py fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614) 2026-09-29 11:27:14 -07:00