chunk_parser built ModelResponseStream without passing usage, so the
cache_read_input_tokens and cache_creation_input_tokens that Databricks
returns for Anthropic models never reached the cost calculator. Every
streamed request was billed at the full input rate even when served
from cache.
ModelResponseStream already coerces a usage dict into Usage, which maps
those keys into prompt_tokens_details, so passing the chunk's usage
through is sufficient.
* fix(databricks): split parallel tool calls so each tool message follows tool_calls
Databricks OpenAI-compatible serving (e.g. GPT models) 400s with "messages with
role 'tool' must be a response to a preceeding message with 'tool_calls'" when an
assistant turn makes parallel tool calls. LiteLLM faithfully sends one assistant
message holding all tool_calls followed by one 'tool' message per result, so every
result after the first is preceded by another 'tool' message rather than the
assistant tool_calls message, which Databricks rejects.
Re-emit each result immediately after an assistant message that carries only its
matching tool_call, turning assistant(tool_calls=[A, B]), tool(A), tool(B) into
assistant(tool_calls=[A]), tool(A), assistant(tool_calls=[B]), tool(B). The
rewrite is a no-op when the turn is already valid (single call), the group is
incomplete, or ids don't line up, so no tool call is ever dropped. Scoped to
non-Claude models, matching the existing OpenAI-shaped transformation path.
* style(databricks): use builtin list generics in parallel tool-call split
Switch the List[...] annotations introduced by _split_parallel_tool_calls
to lowercase list[...] so the UP006 strict-rule budget stays within its
ceiling.
* fix(vertex): stream Model Garden Gemma/Qwen responses correctly through /v1/messages
* test(vertex): cover _CombinedChunkSplitter defensive branches
* test(databricks): rename test file to avoid duplicate basename collision
* fix(databricks,anthropic): defensive token defaults; document single-mode splitter
Address greptile P2 concerns:
- databricks: default usage token fields to 0 when constructing
ChatCompletionUsageBlock from a partially populated usage block — matches
the defensive pattern used in ollama/vertex_ai/cohere/bedrock.
- _CombinedChunkSplitter: clarify in the docstring that an instance is
single-mode (sync or async, not both), since the two iteration paths hold
independent upstream iterator references.
Co-authored-by: Claude <claude@anthropic.com>
---------
Co-authored-by: Steven Kessler <9701252+stvnksslr@users.noreply.github.com>
Co-authored-by: Claude <claude@anthropic.com>
- Use custom_endpoint=False so Databricks SDK auth fallback works
(custom_endpoint=True was blocking it). The api_base returned by
databricks_validate_environment is discarded since get_complete_url
builds the URL separately.
- Remove unused verbose_logger import
- Remove unused json import in tests
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Databricks supports the Responses API natively for GPT models, but litellm
was falling back to the completion transformation handler which converts
responses requests to chat completion calls, losing response schema enforcement.
This adds DatabricksResponsesAPIConfig that passes responses API requests
directly to Databricks' /responses endpoint for GPT models, while non-GPT
models (Claude, Llama, etc.) continue using the completion transformation path.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add OAuth M2M (Machine-to-Machine) authentication via DATABRICKS_CLIENT_ID and DATABRICKS_CLIENT_SECRET
- Add Databricks SDK auto-auth with automatic credential discovery
- Add sensitive data redaction for secure logging (tokens, API keys, secrets)
- Add custom user_agent parameter for partner attribution in Databricks telemetry
- Support user_agent in LiteLLM Proxy via config.yaml litellm_params
- Add 49 mocked unit tests for all new functionality
- Add 13 E2E tests for real-world validation (skipped in CI)
- Update documentation with new features and examples
* update databricks pricing and add DBU<>USD test
* Refactor test_databricks_pricing.py
Removed unnecessary sys.path modification and cleaned up comments.