Some providers charge different per-token rates depending on the time of
day. DeepSeek, for example, has historically discounted its chat and
reasoner models during an off-peak window (16:30-00:30 UTC). LiteLLM's
cost map only modeled static per-token pricing, so cost tracking could
not stay accurate for these providers.
This adds optional off-peak pricing to a model entry: input_cost_per_token_off_peak,
output_cost_per_token_off_peak, cache_read_input_token_cost_off_peak, and an
off_peak_hours_utc window expressed as "HH:MM-HH:MM" in UTC (the window may
wrap past midnight). When the current UTC time falls inside the window, the
cost calculator uses the off-peak rates and otherwise falls back to the
standard rates, so existing models are unaffected. The fields are also
accepted as custom pricing on a deployment, so they can be set from the
proxy config or the SDK.
The window check is a pure function that takes the current time as an
argument, which keeps the regression tests deterministic without patching
the clock.
When a pre-call filter left no order-2 deployments, target_order matching
fell through to the remaining healthy list and reselected the failed
primary. Prompt-cache and deployment affinity also pinned that hop back
to order 1. Match the requested order strictly, skip those pins while
target_order is set, and keep target_order across retries of that hop.
Default remains true so lite claude turns tool search back on through
a proxy. An explicit false or auto in the env or settings is left alone
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
Claude Code turns tool search off when ANTHROPIC_BASE_URL is a proxy.
lite claude, lite up, login --config-claude, and autoroute now force
ENABLE_TOOL_SEARCH=true so MCP tools stay deferred through the proxy
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
_image_sources had no test asserting what it extracts. The existing image tests
live on the Bedrock side and all use base64 without a media_type, which is the one
path the fix left unchanged, so both behaviors it does change went unverified: the
url shape reaching the guardrail at all, and base64 arriving as a data URI.
Against the pre-fix extractor the url case sees [] and the media_type case sees
['AAAA'] instead of ['data:image/png;base64,AAAA'].
The remaining three assert behavior the fix deliberately preserves -- bare base64
passed through, a file source yielding nothing, a malformed source dropped rather
than handed on for a consumer to choke on.
Each message carries a text block because a message with no text never reaches the
guardrail, which would make every source shape look equally dropped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
_image_sources returned source["data"] only. An Anthropic image block has three
shapes (types/llms/anthropic.py:259) and only the base64 one carries "data", so
{"type": "url", "url": ...} yielded nothing and the image never reached any
guardrail at all.
This is not Bedrock-specific. Five guardrails consume
GenericGuardrailAPIInputs["images"] (vigil_guard, custom_code, deepkeep, straiker,
generic_guardrail_api) and every one of them was blind to url sources on
/v1/messages.
base64 now returns a data URI rather than the bare payload. A consumer otherwise
has no way to recover media_type, and an API like Bedrock's ApplyGuardrail needs
the format to build its request.
The file shape stays unresolvable here: the bytes live behind the Files API and
this extractor has no client to fetch them. Documented rather than silently
dropped, so a consumer treating a missing entry as "no image to scan" is a known
gap and not a surprise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The recursion detector in code-quality reads the staticmethod
_written_back_request_fields calling the module-level function of the
same name as a recursive call. Renaming the module-level helper to
_patch_or_convert_request_fields removes the shadowing and describes
what it does: patch changed rows in place, else fall back to full
conversion.
The staging merge tightened the ruff-strict and type-discipline budgets, so the
two `data: dict` parameters the write-back helpers introduced (LIT001) and the
17-branch `process_input_messages` (C901) no longer fit.
Fold the duplicated string/list guardrail flow into one path: a pure
`_extract_guardrail_inputs` builds the guardrail payload, the write-back
helpers become pure functions returning `_RequestFields` (patched input items
plus the resulting instructions value), and the request dict is only mutated
in `process_input_messages` itself. `_apply_guardrail_responses_to_input`
takes Sequence views since it only reads. A non-list `structured_messages`
payload now falls through to the plain texts write-back, matching the
pre-write-back behavior for guardrails that never touch structured messages.
Handler file deltas vs the merge base: LIT001 57 -> 53, LIT002 42 -> 40,
LIT010 28 -> 18, C901 3 -> 3.
Resolves budget-ratchet conflicts by taking staging's tighter limits and reworks the embedding raw-response helpers so the branch stays net-negative on the LIT001/LIT002 ceilings staging lowered: the request methods now return the LegacyAPIResponse and each caller keeps a single dict(headers) conversion.
The route-level RBAC in litellm/proxy/auth/route_checks.py 403s
/model/new, /model/update, and /model/delete for proxy_admin_viewer on
the session role alone, before ModelManagementAuthChecks' team-admin
carve-out can run. A view-only session therefore gets no model write
affordance, team admin or not.