Add key-name-based regex patterns (master_key, database_url, auth_token,
etc.) to SecretRedactionFilter so secrets embedded in dict/config dumps
are redacted by key name, regardless of value format.
Fixes a leak where general_settings containing master_key and
database_url was logged in full because the secret values didn't match
any existing value-format regex pattern.
Address Greptile review feedback:
1. Replace opt-out `disable_default_reasoning_summary` with existing opt-in
`reasoning_auto_summary` flag — avoids backwards-incompatible change where
all users routing thinking-enabled requests would silently get a changed
reasoning_effort shape (string -> dict) on upgrade.
2. Add default summary injection to `_translate_thinking_to_openai` — this path
was the only one missing it, causing inconsistent behavior for
litellm.completion() callers using the Anthropic adapter.
3. Narrow `except Exception` to `except (ValueError, TypeError, AttributeError)`
in tests to avoid masking genuine failures.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove incorrect supports_prompt_caching from gpt-4-0314 (predates the feature)
- Make data-URL detection case-insensitive in Gemini tool call result conversion
- Mock show_banner/generate_feedback_box in max_budget tests to prevent real I/O
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Remove early return in get_llm_provider_logic.py that prevented
the 'openrouter/' prefix from being stripped. The early return was
intended for 'native OpenRouter models' like 'openrouter/free',
but no such models exist in the model registry — all OpenRouter
models are multi-segment (e.g. 'openrouter/anthropic/claude-3.5-sonnet')
and need the prefix stripped before being sent to the OpenRouter API.
This regression was introduced in v1.82.3 and caused 400 Bad Request
errors for all OpenRouter models.
- Update load_local_model_cost_map to use project root fallback for dev
- Keep main's validation, aliases, and source info tracking
- Remove backup JSON (purpose of this PR)
* fix(moonshot): preserve reasoning_content on Pydantic Message objects in multi-turn tool calls
The condition 'reasoning_content not in msg' doesn't work correctly for
Pydantic Message objects because they don't support the 'in' operator
like dicts do. This caused reasoning_content to be stripped from
assistant messages in multi-turn conversation history.
Changed the condition to use msg.get('reasoning_content') instead,
which works correctly for both dicts and Pydantic models.
Fixes#23765
* added newline eof
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Update tests/test_litellm/llms/moonshot/test_moonshot_chat_transformation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Simplify assertions in test_moonshot_chat_transformation
Removed redundant assertions for non-assistant messages.
---------
Co-authored-by: BillionClaw <267901332+BillionClaw@users.noreply.github.com>
Co-authored-by: Aarish Alam <arishalam121@gmail.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
- Set output_cost_per_second to 0.0 (was 0.0001) for whisper-1 and
azure/whisper-1: transcription is billed on input duration only,
not output duration
- Fix cost_per_second() in openai/cost_calculation.py: change elif to if
so input_cost_per_second is evaluated independently of output_cost_per_second,
and remove the erroneous completion_cost = 0.0 assignment that masked
any previously-set output cost
- Add TestCostPerSecondArithmetic unit tests covering both cost fields,
the None-guard, and zero-duration edge case
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
The Gemini REST API documents the embedding task type parameter as
camelCase `taskType`. The existing transformation functions convert
`dimensions` to `outputDimensionality` but miss the parallel
`task_type` to `taskType` conversion. This adds that conversion to
both `transform_openai_input_gemini_content` (batchEmbedContents path)
and `transform_openai_input_gemini_embed_content` (embedContent path).
Fixes#24190
Non-streaming paths call _process_hidden_params_and_response_cost; streaming
assembles the full response later and skipped that, so litellm_params.metadata
lacked hidden_params (e.g. response_cost for OTEL/OpenSearch).
- Add _merge_hidden_params_from_response_into_metadata and call it from
success_handler and async_success_handler after cost is set, before
_build_standard_logging_payload.
- Unit tests for merge helper.
Tests: pytest tests/test_litellm/litellm_core_utils/test_litellm_logging.py
Made-with: Cursor
The /v1/messages/count_tokens endpoint was hardcoding the Bedrock runtime
URL, ignoring api_base and aws_bedrock_runtime_endpoint settings. This
aligns it with invoke/converse handlers by using the existing
get_runtime_endpoint() method for consistent endpoint resolution.
Signed-off-by: stias <seokjun.yang@mycraft.kr>
Adds a control plane capability that enables a central admin instance
to manage multiple regional worker proxies from a single UI.
Backend:
- Worker registry loaded from YAML config (worker_id, name, url)
- /.well-known/litellm-ui-config exposes is_control_plane and workers list
- /v3/login + /v3/login/exchange: opaque code exchange for cross-origin
username/password auth (JWT never in URL/logs, single-use 60s TTL)
- SSO cookie handoff with return_to → opaque code → exchange
- _validate_return_to: full origin validation (scheme+hostname+port)
- Startup warning when control_plane_url set without Redis
- Both /v3 endpoints gated behind control_plane_url config
Frontend:
- Worker selector dropdown on login page (gated behind is_control_plane)
- Cross-origin SSO code exchange handling on callback
- switchToWorkerUrl: localStorage-persisted worker URL for API calls
- useWorker hook: shared worker state management
- WorkerDropdown in navbar for switching workers
- Logout/switch clears worker state from localStorage
Tests:
- 7 tests for /v3/login + /v3/login/exchange
- 10 tests for _validate_return_to
- 2 tests for control plane discovery endpoint