Previously, unrecognized check keys (e.g. snake_case typos like
`content_filter` instead of `contentFilter`) were silently dropped,
causing the guardrail to fall back to ApplyGuardrail mode without any
indication. Now logs a WARNING listing the unknown keys and the valid
set, so operators can catch misconfigurations before they reach Bedrock.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The prior refactor attempt deleted ui_sso.py but could not push the
new content due to output-size limits (noted in the PARTIAL marker).
This left proxy_server.py and tests unable to import
get_disabled_non_admin_personal_key_creation and router, causing all
test suites to fail at collection time (exit code 4).
Restores ui_sso.py from main until the refactor can be completed
properly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
A chunk with empty choices and no usage carries no content, billing, or
finish reason, so translating it emitted a premature message_delta and
broke Anthropic SSE ordering. Skip such keepalives in both the sync and
async iterator loops; usage-only trailing chunks (choices=[] with usage)
still flow through.
- asqav: bundle start_time/end_time into a timing tuple to reduce
_build_and_append from 6 to 5 args
- milvus_ingestion: absorb unused content_type into **_ to bring store()
within max-args limit while preserving the base class interface
- test_guardrail_usage_config: remove flagged_count from _metric signature
(hardcoded to 0 in the returned object)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Extract highest_risk_percentage from response_payload in
_log_guardrail_result instead of passing as a param
- Introduce DataSourceConfig dataclass to bundle name/enabled/priority
for URLDataSource, VectorStoreDataSource, FactCheckDataSource
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fix strict-budget violations introduced by new files:
- Replace typing.Any with object/proper types in callback signatures,
guardrail hooks, usage endpoints, and common request processing
- Use datetime.now(timezone.utc) instead of datetime.now() for
timezone-aware datetimes in bedrock_guardrails and bias_hallucination_estimator
- Remove noqa suppressions in favour of actual fixes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
New files introduced in this staging batch (asqav, bedrock_guardrails,
bias_hallucination_estimator, milvus_ingestion, usage_endpoints) exceeded
the ruff strict-rule budget for ANN401/BLE001/DTZ005. Add targeted
# noqa comments to bring totals back within their configured ceilings.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Forward all Perplexity search params instead of a hardcoded subset
PerplexitySearchConfig.transform_search_request only copied four keys
(max_results, search_domain_filter, max_tokens_per_page, country) into
the outgoing request body and silently dropped everything else, so
documented Search API parameters like search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size and max_tokens never reached Perplexity even though
callers could set them.
Perplexity's native parameter names already match LiteLLM's unified search
spec, so there is nothing to remap; the transformation now passes every set
optional parameter through as-is, the same approach the Exa AI search
transformation already takes. None-valued params are still omitted.
* fix: update litellm/llms/perplexity/search/transformation.py
add key != "query"
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Add search transformation tests and extend PerplexitySearchRequest
Adds unit tests asserting the Perplexity Search request body forwards the
full documented parameter set (search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size, max_tokens and the original four), omits None/unset
params, passes through arbitrary params, and never lets an optional_params
"query" key override the query argument.
Extends the PerplexitySearchRequest TypedDict with those documented fields
so it no longer advertises only the original four.
* Use builtin list[str] for new search_language_filter field
The UP006 strict-budget gate is over its ceiling on the base branch, so
any net-new typing.List usage fails CI. Type the newly added
search_language_filter field with the builtin list[str] generic instead of
List[str] so the change adds no new UP006 violations.
---------
Co-authored-by: Mehmet Can Şakiroğlu <can.sakiroglu@getmidas.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix(anthropic): guard empty choices[] chunks in the messages streaming bridge
OpenAI/Azure-compatible backends emit a trailing usage-only chunk with
choices=[]. The anthropic /v1/messages streaming adapter assumed every chunk
has choices[0], so it crashed mid-stream with IndexError. Guard the choices[0]
accesses in the streaming path and route usage-only chunks into the message
delta. Fixes#30761.
Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>
* style: black-format empty-choices guard + test
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(otel): don't crash set_attributes on non-dict (MCP) response_obj
MCP tool calls pass a Pydantic CallToolResult, but set_attributes accesses
response_obj via .get() throughout. The AttributeError was caught but skipped
writing the span output. Normalize a non-dict response_obj to a dict (model_dump)
so the span is fully emitted. Fixes#30651.
Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>
* test(otel): cover model_dump-raises and non-serializable fallbacks for #30651
* style: black-format the non-dict response test
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
_calculate_tiered_cost resolved a tier's per-token rate with
`tier.get(cost_key) or tier.get(fallback_cost_key, 0)`. The `or`
short-circuits on a falsy 0.0, so a tier that legitimately prices cached
reads (or reasoning tokens) at 0.0 was silently billed at the full
fallback rate, at both the in-range and overflow sites.
Add _resolve_tier_cost_per_token, which only falls back when the primary
key is absent (is None), mirroring the flat-pricing path that already
guards correctly. Uses the X | None annotation style to stay within the
ruff strict-rule budget.
Re-submit of #30653, which was reverted from litellm_internal_staging
because the previous Optional[str] annotation pushed the UP045 count over
the ruff-strict-budget.json ceiling.
* feat: local-first tamper-evident audit log callback (asqav)
* fix(asqav): remove unused import, drop dead checkpoint path, update tests
- Remove unused `httpxSpecialProvider` import (F401 lint fix)
- Remove cloud checkpoint feature: no /v1/checkpoints or /api/v1/checkpoints
endpoint exists in the Asqav cloud API; the path 404s on prod
- Drop the `api_key`/`checkpoint_interval` constructor params and
`_schedule_checkpoint` method that backed the dead path
- Update tests: remove checkpoint-specific stubs and test cases,
rename tests that now have broader applicability
- seq restore on restart already present in `_load_chain_tail`; the
`test_seq_counter_restored_after_restart` test confirms the behaviour
* docs(asqav): remove stale cloud-checkpoint sentence from _build_and_append docstring
* fix(asqav): file perms 0600, proxy identity metadata, multi-worker doc
- _write_record: create audit log via os.open(O_CREAT, 0o600) and chmod
existing file to 0600 before append; prevents other local users reading
the log under a permissive umask (Veria ~line 296)
- _extract_loggable: merge proxy identity fields (user_api_key_user_id,
team_id, org_id, key_alias) from kwargs["litellm_params"]["metadata"],
filtering sensitive keys (user_api_key, Authorization) (Veria ~line 89)
- AsqavLogger docstring: document single-writer assumption and multi-worker
limitation; recommend single audit-writer process or fcntl-based wrapper
for multi-worker proxy deployments (Veria ~line 188)
- tests: add three anti-vacuous regression tests that fail against unfixed
code (file perms, proxy identity attribution, docstring guard)
Items already correct before this commit (no code change needed):
- seq counter restore: _load_chain_tail already sets _call_count from
last_record.get("seq", -1)+1 (Greptile ~line 223)
- write inside lock: _write_record called inside with self._lock: block
(Greptile P1 concurrency)
* style: apply black formatting to asqav integration
* feat(rag): add Milvus vector store ingestion support
Adds write/ingest support for self-hosted Milvus to complement the existing
Milvus search provider. /rag/ingest now accepts custom_llm_provider=milvus.
- MilvusRAGIngestion implements the store() step via the Milvus REST API v2
(entities/insert), reusing the base upload/ocr/chunk/embed pipeline
- Auto-creates the collection via quick setup (dynamic fields) when missing
- Embeddings generated through litellm embedding API (any provider)
- api_key optional for auth-less self-hosted Milvus; supports db_name/partition
- Registered in INGESTION_REGISTRY; MilvusVectorStoreOptions added to types
- 16 unit tests (mocked REST) + env-gated integration test
* fix(rag): authorize Milvus collection_name as vector_store_id on ingest
Milvus ingestion writes to collection_name (falling back to vector_store_id),
but /rag/ingest only authorized fields named vector_store_id. A request with
custom_llm_provider=milvus and collection_name set to another team's managed
collection bypassed assert_user_can_access_vector_store_id. Normalize
collection_name into vector_store_id before authorization.
* fix(rag): close Milvus collection_name authz bypass and address review
Resolves the Greptile review on the Milvus RAG ingestion path:
- P0 (security): vector-store-id normalization for authorization now always
mirrors collection_name onto vector_store_id for Milvus, not only when
vector_store_id is absent. A request pairing a collection_name the caller
cannot access with a vector_store_id they can no longer bypasses
assert_user_can_access_vector_store_id. Adds a test for the both-fields case.
- P1: removes the provider-specific `custom_llm_provider == "milvus"` branch
from proxy/rag_endpoints/endpoints.py. BaseRAGIngestion now exposes a
normalize_authorized_vector_store_id classmethod (no-op by default) that
MilvusRAGIngestion overrides; the proxy dispatches generically via
get_ingestion_class.
- P2: removes the embed() side-effect that mutated self.embedding_config on
first call. The default model is set once in MilvusRAGIngestion.__init__ and
the class inherits BaseRAGIngestion.embed. Drops the now-unused top-level
`import litellm` (also clears the CodeQL import/import-from warning).
* fix(rag): block view-only role from auto-creating Milvus collections
Require INTERNAL_USER_VIEW_ONLY ingest targets to resolve to an existing
managed vector store. Presence of vector_store_id was insufficient: Milvus
normalization mirrors collection_name onto vector_store_id and unknown ids
pass authorization as provider-native targets, letting a view-only caller
trigger Milvus auto_create_collection for a brand-new collection.
* fix(rag): authorize Milvus db_name via server env only
Milvus db_name selects the write target's database namespace but the proxy
only authorizes collection_name/vector_store_id. A caller with access to a
managed collection could set db_name to redirect writes/auto-create into
another Milvus database using the server's credentials, outside the
per-collection authorization boundary.
Resolve db_name from MILVUS_DB_NAME (server-side) only; never from the
request. Drop db_name from MilvusVectorStoreOptions and add a regression
test asserting a request-supplied db_name is ignored.
* fix(rag): authorize Milvus partition_name via server env only
* fix(rag): scope view-only ingest guard to auto-creating providers
The view-only ingest guard required every vector_store_id to resolve to a
litellm-managed store, which broke INTERNAL_USER_VIEW_ONLY callers writing to
provider-native ids (e.g. OpenAI vs_*) that are not in the managed registry
Only providers that can create a store on ingest (Milvus with
auto_create_collection) let a view-only caller bring a brand-new store into
existence, so the managed-store requirement now applies only to those. Each
ingestion class declares this via can_auto_create_vector_store and the proxy
dispatches to it instead of hardcoding provider logic. Providers that only
write to a pre-existing store keep accepting their provider-native ids
unchanged
Also drops the banned typing imports from the new milvus_ingestion module so
it stays within the strict-rule budget gate after the rebase onto
litellm_internal_staging
* fix(rag): bind Milvus api_key fallback to server-resolved api_base
A named credential can carry api_base while leaving api_key unset, which
slips a request-controlled endpoint past the proxy's api_base block. The
constructor then fell back to MILVUS_API_KEY independently, sending the
server token to that endpoint. Only fall back to the env token when
api_base also comes from MILVUS_API_BASE.
* fix(rag): require managed store for view-only Milvus ingest regardless of auto_create flag
can_auto_create_vector_store read the request-supplied auto_create_collection
flag, so a view-only key could set it to false, name any existing unmanaged
collection, and skip the managed-store resolution check in
_assert_view_only_role_cannot_create_vector_store. Report the provider's
capability instead: Milvus can always auto-create, so a view-only target must
always resolve to a managed vector store.
* style(rag): apply black formatting to Milvus ingest files
* style(rag): modernize typing to satisfy ruff strict-rule budget
Use PEP 585/604 builtins (dict, tuple, X | None) in the Milvus ingestion and
RAG endpoint helpers so the strict-rule budget delta (UP006/UP035/UP045) stays
under the lowered ceiling pulled in from staging.
* style(rag): drop redundant quoted annotations to satisfy UP037 budget
* chore(rag): retrigger CI after transient artifact-download 403
* fix(rag): block credential hydration from overriding authorized write target
* fix(guardrails): show config guardrails in usage details
* test(guardrails): cover config guardrail helper branches
* fix(guardrails): persist guardrail_info and resolve config guardrails in logs
Address review findings on the config-guardrail usage work:
- initialize_guardrail dropped guardrail_info when building the in-memory
Guardrail, so description and type were always empty for config guardrails
in production; persist the field.
- guardrails_usage_logs only resolved a logical name for DB-backed guardrails,
so logs for a config guardrail queried by UUID were always empty; fall back
to the in-memory list like the detail endpoint does.
- _get_config_loaded_guardrails now expresses an explicit allow (source ==
"config") instead of a double-negative skip, and _get_guardrail_dict_field
dispatches on type so a falsy-but-valid value (e.g. {}) is not dropped.
Tests exercise the real handler path (initialize_guardrail through
_get_config_loaded_guardrails) rather than mocking it, so they fail if
guardrail_info is dropped or the logs fallback is removed.
* Implement Bias and Hallucination Estimator with Grounding Checker, Risk Scorer, and Utility Functions
- Added GroundingChecker for verifying claims against data sources.
- Introduced RiskScorer to compute risk scores based on bias and hallucination analyses.
- Developed utility functions for sentence splitting, text clipping, and unique value preservation.
- Created patterns for detecting bias and hallucination indicators.
- Established data models for bias and hallucination analysis results.
- Implemented tests for bias detection, hallucination detection, grounding checks, and risk scoring.
- Integrated the BiasHallucinationEstimatorGuardrail for managing high-risk responses.
* Refactor Bias Hallucination Estimator: Enhance logging, remove unused parameters, and improve concurrency handling
- Added logging decorator to `apply_guardrail` method to log guardrail information while excluding sensitive fields.
- Removed `use_logprobs` and `uncertainty_weight` parameters from `BiasHallucinationEstimatorGuardrail` and related classes.
- Simplified `BiasDetector` and `HallucinationDetector` initialization by removing threshold parameters.
- Implemented a lock mechanism in `URLDataSource` to prevent concurrent fetches from causing race conditions.
- Updated `RiskScorer` to remove uncertainty handling and adjusted risk calculation logic.
- Enhanced test coverage for guardrail logging and data source functionalities, ensuring proper behavior under various conditions.
* Refactor bias hallucination estimator code for improved readability and consistency
- Updated string formatting for better readability in data_sources.py, estimator_core.py, grounding_checker.py, patterns.py, risk_scorer.py, and utils.py.
- Enhanced the clarity of function signatures and method calls across various classes.
- Removed unnecessary variables and streamlined logic in grounding_checker.py.
- Improved test cases for bias and hallucination detection to enhance coverage and maintainability.
- Added tests for initializing guardrails and handling edge cases in grounding checks.
* fix(bias-hallucination-estimator): 100% patch coverage, ruff strict gate passing
Fix ruff strict gate violations: replace deprecated typing imports
(List/Dict/Tuple/Optional) with builtin generics and union syntax
(UP006/UP037/UP045); replace Any in public API signatures with object
or str (ANN401); remove now-unused imports (F401).
Grow test suite from 117 to 130 tests covering all previously uncovered
branches: DataSource.verify_fact, _keyword_search empty-word-chars path,
URLDataSource._fetch_url exception path, VectorStoreDataSource
_initialize_client pinecone/weaviate paths and ImportError fallback,
_load_embedding_model via mocked sentence_transformers, VectorStore
search exception path, KnowledgeGraph search exception path, and
GroundingChecker._boost_confidence entity match branch. Mark the
abstract method stub with pragma: no cover.
All 8 files now at 100% patch coverage; strict gate clean.
The AWS Bedrock Anthropic Claude 3 entries in the model cost maps were
missing supports_assistant_prefill entirely. Because the key was absent
rather than true, callers that gate on its truthiness treated these
models as not supporting trailing-assistant prefill, even though every
Claude 3 model supports it; the direct anthropic API entries for
claude-3-haiku and claude-3-opus already carry true
This sets supports_assistant_prefill to true for all 20 affected Claude 3
Bedrock entries across regions (us, eu, apac, us-gov) and routes (invoke),
in both the primary price map and the litellm backup, and adds a
regression test covering the JSON maps and get_model_info
Fixes#30863
Deployment-level metrics carried only model_id, so a model group spread
across several deployments showed up as repeated model_id series with no
way to tell which configured group each belonged to. litellm_deployment_state,
litellm_deployment_tpm_limit, litellm_deployment_rpm_limit,
litellm_deployment_cooled_down and litellm_deployment_latency_per_output_token
can now emit model_group alongside model_id.
Adding a label changes a metric's time-series identity, so the label is
opt-in behind litellm.prometheus_emit_deployment_model_group_label (default
False), mirroring prometheus_emit_rate_limit_labels. Off by default preserves
each metric's historical label set across upgrade; enable it once downstream
dashboards and recording rules account for the new dimension. The label is
appended in PrometheusMetricLabels.get_labels when the flag is set, so it also
respects the include_labels filter.
The cooldown callback previously passed the deployment alias as
litellm_model_name, which disagreed with the success/failure logging paths and
fragmented litellm_deployment_state into two series per deployment. It now
reports the prefix-stripped underlying model as litellm_model_name and the
alias as model_group, and resolves api_base from the underlying model.
increment_deployment_cooled_down was moved off positional label args onto
prometheus_label_factory so it respects the label config like every other
deployment metric.
Fixes#30748
* fix(proxy): bill partial usage when a streaming request is cancelled
On a mid-stream client disconnect the stream never reaches normal completion:
CancelledError / GeneratorExit are BaseException, so neither the success nor the
failure logging path runs, and the assembled-response success logging that
writes SpendLogs never fires. The tokens already produced upstream are billed by
the provider but never recorded on the proxy, so spend undercounts by roughly
the abort rate; the gap is invisible in SpendLogs and only shows up against
provider invoices.
In the shielded streaming cleanup, when a client disconnect is recorded, assemble
the partial usage from the chunks received so far via stream_chunk_builder and
dispatch success logging for it. dispatch_success_handlers de-dupes via
has_dispatched_final_stream_success, so it is a no-op when normal completion
already logged, and it only runs on the cancellation path (the exception path
sets stream_completed and already emits a failure log).
Reported and root-caused in production by @jingyu-lin. Complements #30522, which
releases the budget reservation on the same cancellation path.
* fix(proxy): close upstream stream before billing partial usage on disconnect
Release the provider connection before running partial-usage success
logging on a client disconnect, so slow or external success callbacks
can no longer keep the upstream stream held open. aclose() only closes
completion_stream and leaves response.chunks intact, so the partial
billing still assembles usage from the chunks already received.
---------
Co-authored-by: Bytechoreographer <Bytechoreographer@users.noreply.github.com>
Adds support for Bedrock's InvokeGuardrailChecks API
(POST /guardrail-checks/invoke) to the existing `bedrock` guardrail, alongside
the current ApplyGuardrail integration.
Changes vs litellm_internal_staging:
- BedrockGuardrail calls InvokeGuardrailChecks when `checks` is configured
(inline contentFilter / promptAttack / sensitiveInformation safeguards; no
guardrailIdentifier or guardrail resource required); otherwise the
ApplyGuardrail path runs unchanged. make_bedrock_api_request is now a thin
dispatcher, and the shared signed-POST / transport-error handling is factored
into _sign_and_post so the two API paths cannot drift.
- Detect-only scores are mapped to block decisions via configurable per-check
thresholds (content_filter_threshold / prompt_attack_threshold /
pii_confidence_threshold, default 0.5, range [0,1]); set a threshold to null
to make that check detect-only (logged, never blocks). disable_exception_on_block
is honored (returns a normal-string error the proxy turns into a mock response).
- PII location offsets (beginOffset/endOffset/messageIndex/contentIndex) are
stripped before the response is logged; block details and tracing carry only
category/type labels and numeric scores, never raw input.
- Adds structured BedrockChecksConfigModel and the threshold fields to
BedrockGuardrailConfigModel, forwarded through initialize_bedrock; configuring
both `checks` and `guardrailIdentifier` raises a clear error.
- Adds request/response TypedDicts for the new API.
- Adds mocked unit tests covering block/allow/detect-only for all three checks,
INPUT and OUTPUT scanning, dispatcher routing, request shape/path, PII offset
stripping, empty-message short-circuit, error and config-validation paths.
* feat(proxy): add LITELLM_DISABLE_ACCESS_LOG_PATHS to drop noisy access logs
Adds a `HealthCheckAccessLogFilter` to `litellm._logging` that drops
`uvicorn.access` records whose request path matches a comma-separated
list in the new `LITELLM_DISABLE_ACCESS_LOG_PATHS` env var. The filter
is wired into both:
- the JSON `log_config` produced by `_get_uvicorn_json_log_config()`
(used when `JSON_LOGS=true`), via `dictConfig`'s native filter
binding on the access handler, and
- the plain `uvicorn.access` logger at module import (covers the
non-JSON code path).
This is useful when the proxy runs behind k8s liveness/readiness
probes, ALB health pings, or Prometheus `/metrics/` scrapes that
otherwise drown real request logs at a multi-line-per-second rate.
Example:
LITELLM_DISABLE_ACCESS_LOG_PATHS="/,/health/liveliness,/health/readiness,/metrics/"
Path matching is exact (after stripping any query string) and only
applies to the `uvicorn.access` logger -- application logs and
`uvicorn.error` are untouched.
Default behaviour is unchanged: when the env var is empty/unset the
filter short-circuits and all access lines are emitted.
Tests cover env-unset pass-through, configured-path drops, query
string stripping, and robustness to malformed log records.
* fix(proxy): only wire healthcheck filter in JSON log config when paths configured
Mirror the _suppress_loggers() guard so JSON_LOGS=true users who have not set
LITELLM_DISABLE_ACCESS_LOG_PATHS no longer get unused dictConfig filter machinery
---------
Co-authored-by: Huynh Duc Tran <ducth6@tcbs.com.vn>
* feat: declarative fallback generalizations for unknown models
Unknown or newly-released models previously degraded (missed cost lookups,
wrong supports_* flags, broken provider routing) and were patched with one-off
hardcoded regexes scattered across Python. This adds a single data-driven source
of truth: a fallback_generalizations block in model_prices_and_context_window.json
holding ordered, case-insensitive regex rules that map a model name to the
metadata to apply when it has no exact entry.
A new fallback_generalizations module owns the rules and a compiled-regex cache
that is built once and invalidated on reload, so the O(n) scan runs only on a
cache miss. get_llm_provider now routes an otherwise-unknown model via the first
matching rule's litellm_provider, replacing the hardcoded _CLAUDE_PATTERN and
_matches_claude_model_pattern. _get_model_info_helper falls back to a matching
rule's model_info after the exact lookups miss, so get_model_info and the
supports_* helpers resolve unknown models from the same rule. get_model_cost_map
extracts the block out of the returned map, and the integrity check now counts
real model entries (excluding reserved meta keys) so the new key cannot mask a
genuinely shrunk upstream file.
The top level of the file stays a flat map of models so existing litellm releases
that fetch the live file keep working and keep receiving updates; the block ships
in both the root file and the bundled backup. An anthropic-claude rule reproduces
the old future-claude routing and additionally supplies capability flags and a
context window
https://claude.ai/code/session_01G8Jro8dPLktwnaaSJwVDpo
* refactor(anthropic): derive adaptive-thinking from a version threshold; harden generalizations
Replace the per-minor-version _is_claude_4_6_model / _is_claude_4_7_model substring
matchers with a single _claude_version_at_least predicate that parses the Claude
family version from the model name and compares against 4.6. This covers 4.8/4.9/5.x
without a code change (the old matchers missed 4.8 entirely) while keeping an explicit
supports_adaptive_thinking flag authoritative when present, so there is one source of
truth. The two direct call sites in the chat transformation now route through
_is_adaptive_thinking_model instead of the deleted matchers.
Also address review feedback on the generalizations module: return a copy of the
matched model_info so a future caller cannot mutate the compiled-rule cache, document
that patterns are matched with re.search and must anchor with ^ and $, and reindent
the fallback_generalizations block to the file's 2-space style in both JSON files.
https://claude.ai/code/session_01G8Jro8dPLktwnaaSJwVDpo
* fix(anthropic): surface adaptive-thinking from the cost map; fix date misparse
supports_adaptive_thinking shipped in the model cost map but was never declared
on ModelInfo nor copied during construction, so get_model_info (and the supports_*
factory) silently dropped it for every provider-prefixed or generalized name; only
a bare base entry resolved. Wire it through ModelInfo like the other capability
flags and backfill the flag onto the genuine Claude 4.6/4.7/4.8 entries across
providers so the data, not code, declares the capability. The anthropic-claude
fallback rule also carries the flag (and now accepts a dotted minor, e.g. 4.6) so
an unmapped future Claude degrades to adaptive thinking without a code change.
Tighten the Claude version parser so an eight-digit date suffix
(claude-opus-4-20250514, the non-adaptive Opus 4.0) is no longer read as minor
4.20250514. The cost map stays authoritative; the version check is only a fallback
for provider-prefixed names (bedrock/invoke routes, -v1-less ids) that resolve to
no mapped entry and so cannot be reached by an exact lookup or the bare-name rule.
https://claude.ai/code/session_01G8Jro8dPLktwnaaSJwVDpo
* fix(anthropic): date-safe adaptive-thinking version fallback, conservative fallback pricing, ruff strict gate
Reconcile adaptive-thinking detection after merging litellm_internal_staging.
Keep the cost-map resolver (_supports_model_capability) as the source of truth and
add a date-safe opus/sonnet/haiku >= 4.6 name version as a fallback for
provider-prefixed ids the cost map cannot resolve (e.g.
bedrock/invoke/us.anthropic.claude-opus-4-6). A two-digit cap on the minor keeps an
eight-digit date suffix from being misread as a minor version, so the dated Claude
4.0 release stays non-adaptive
Price the shipped anthropic-claude fallback rule at the Opus tier so an unknown or
newly released Claude is over-costed rather than billed as free
Drop the module-level global state in fallback_generalizations (PLW0603) in favor of
a small registry object, and switch its annotations plus the new utils helper to
builtin generics (UP006), bringing the ruff strict-rule totals back under ceiling
* refactor(anthropic): drive adaptive-thinking version gate from a declarative rule
Replace the bespoke _claude_version_at_least heuristic with a version-gated fallback_generalizations rule. Unmapped Claude ids now resolve adaptive thinking purely from the cost map: an explicit entry, or the new self-contained anthropic-claude-adaptive-thinking rule that matches opus/sonnet/haiku >= 4.6 (covering 5.x, 6.x and beyond with no code change). New families ship via Price Data Reload instead of a code edit
The rule carries the same Opus-tier pricing as the broad anthropic-claude rule plus supports_adaptive_thinking, and is matched first; the broad rule stays version-neutral, so an unmapped >= 4.6 Claude resolves to full pricing and the adaptive flag from one rule, while a sub-4.6 alias such as claude-opus-4-0 is still priced yet stays non-adaptive. The regex caps the minor at two digits so a dated 4.0 id (...-4-20250514) is never read as a >= 4.6 minor
* refactor(anthropic): dedupe adaptive-thinking rule via declarative extends
The version-gated anthropic-claude-adaptive-thinking rule duplicated the
broad anthropic-claude rule's entire Opus-tier price block because rules do
not merge: first match wins and returns one rule's whole model_info, so the
adaptive rule had to be self-contained.
Add a declarative extends field to fallback_generalizations: a rule names a
parent and inherits its model_info, with its own keys overriding. Inheritance
is resolved once at install time against each rule's raw model_info, so the
adaptive rule now carries only its delta (supports_adaptive_thinking) and
inherits pricing from the broad rule. Runtime matching, provider routing and
gating are unchanged; the broad rule stays anchored and first-match-wins still
holds.
* docs(anthropic): add ignored description key documenting each generalization regex
* fix(anthropic): drop fabricated pricing from the anthropic-claude fallback rule
Per review feedback, the base rule no longer carries input/output/cache costs, and the
adaptive-thinking rule that extends it inherits that no-pricing model_info. Pricing an
unmapped model at a guessed tier reports a confidently-wrong cost without the caller
knowing; dropping it keeps the standard unpriced behavior (zero, not a fabricated
number) so a missing price stays visible. The rules still supply provider routing,
context window, and capability flags, so a brand-new Claude can still be called and its
capabilities (including adaptive thinking for >= 4.6) resolved. Description and tests
updated to match
* fix(realtime): stop sending a second Gemini Live setup on follow-up session.update
Gemini Live (BidiGenerateContent) accepts setup as the first-and-only client
message; a second setup closes the socket with 1007 Request contains an invalid
argument. The AI Studio Gemini path forwarded every client session.update after
the first as a follow-up setup, and GA clients (pipecat) send several while
configuring the session, so the second one tore the session down before the
first turn. Callers saw silence after the first response, exponential per-turn
latency from reconnect/retry churn, and intermittent 1011 errors.
Drop subsequent session.updates instead of resending setup, matching what the
Vertex subclass already does. Tools and instructions must ride on the first
session.update before any conversation content.
Adds regression tests covering the plain follow-up, a follow-up that adds tools
(the case the previous identical-only dedup still forwarded), and the guardrail
create_response=False warning path.
* fix(realtime): retry the backend open handshake instead of failing with 1011
The upstream Live API open handshake (e.g. Gemini Live) intermittently hangs;
waiting longer never recovers a hung attempt, but a fresh attempt almost always
connects in ~1s. The proxy opened the backend websocket once with the default
open_timeout and no retry, so a single slow handshake surfaced to the caller as
a fatal 1011 internal error and dropped the call.
Bound each open attempt with a short open_timeout and retry; a bounded attempt
that already timed out spaces out the next try, so no backoff is needed.
Deterministic handshake-status rejections (auth/4xx) are not retried, and the
retry only ever wraps the open, never a live session.
Adds tests for retry-then-succeed, raise-after-max-attempts, and
no-retry-on-auth-failure.
* fix(realtime): close guardrail bypass + surface handshake status; drop obsolete tests
Three review fixes on the Gemini Live realtime path.
Transcription-guardrail bypass: Gemini Live rejects a second setup (1007), so once
the initial setup is sent the guardrail's automaticActivityDetection.disabled=true
can no longer be delivered as a follow-up session.update. With that follow-up now
dropped, the model's auto-response stayed enabled and a realtime_input_transcription
guardrail was bypassed (the model answered before the proxy could gate the turn).
Fold the disable into the one-and-only setup instead: the handler injects it into
the auto-sent setup (gemini_live_defer_setup false) and _send_to_backend injects it
into the deferred first setup. OpenAI sessions accept follow-up updates and are left
untouched.
Backend handshake status: the open-retry treated only InvalidStatusCode as
deterministic; websockets>=15 raises InvalidStatus for a rejected client handshake,
so a 401/403 fell into the broad WebSocketException branch and was retried before
the caller closed the client with 1011 instead of the upstream status. Treat both as
non-retryable.
Obsolete tests: the four tests asserting a follow-up session.update is merged and
re-sent as a second setup asserted behavior that crashes Gemini Live with 1007
(verified directly against the API). Removed; the drop is covered by new regression
tests.
* style(realtime): reformat changed files to ruff line-length 120
Post-merge with litellm_internal_staging, which unified ruff format width to 120
(#31518). The realtime change set was formatted at 88, so the changed lines
tripped the whole-file ruff format check. Reformat with ruff 0.15.3 at the repo's
120 width; no logic changes.
tests/e2e/CLAUDE.md captures the harness code-style rules (suite-as-a-class, shared transport, typed pydantic models, Result/unwrap, markers, typing) and the coverage-registry naming grammar; CONTRIBUTING.md gets the Contributors Guide intro and a Setup section
* fix(router): persist global retry_policy via /config/update (LIT-3152)
The Admin UI Model Retry Settings tab POSTs
{router_settings: {retry_policy: {...}}} to /config/update, but the
field was dropped on two write-side layers so it never reached the
router. UpdateRouterConfig did not declare retry_policy, so
dict(exclude_none=True) stripped it before the DB upsert. And even when
fed directly, Router.update_settings had no "retry_policy" entry in
_allowed_settings, so the assignment was a silent no-op. The DB row
stayed at {"model_group_alias": {}}, llm_router.retry_policy stayed
None, and the UI fell back to defaultRetry = num_retries = 2 on refresh.
Declare retry_policy on UpdateRouterConfig as a plain dict, and add a
retry_policy branch to update_settings that coerces dict payloads to
RetryPolicy before setattr, mirroring Router.__init__. get_settings
already lists retry_policy, so reads work once writes land.
* fix(router): guard retry_policy type in update_settings
Mirror Router.__init__ semantics in update_settings: only assign
retry_policy when it is None or a RetryPolicy (after dict coercion).
Previously a non-dict, non-RetryPolicy value (e.g. a YAML typo like
retry_policy: 5 flowing through /config/update) was stored verbatim,
deferring the failure to request time in get_num_retries_from_retry_policy
instead of being dropped at write time.
* refactor(ui): harden Model Retry Settings flow and validate retry_policy at the boundary
Types UpdateRouterConfig.retry_policy as RetryPolicy and model_group_retry_policy as Dict[str, RetryPolicy] so /config/update validates the payload and rejects malformed counts instead of silently persisting them; the apply path in update_settings keeps coercing the stored dict back to RetryPolicy
Makes the Model Retry Settings tab the single owner of retry_policy and model_group_retry_policy so the generic Router Settings page no longer renders or writes them, replaces the fire-and-forget save with a react-query mutation that only shows the success toast after the write resolves, surfaces real errors, disables Save while in flight, and re-reads authoritative state on success, and sends both the global and per-group policies atomically so edits in the inactive scope are no longer dropped
Decouples the retry-scope selector from the All Models filter and defaults it to Global, seeds the displayed default from num_retries (falling back to 2), and gives per-group rows real inherit semantics so an empty input shows the global value as a placeholder with a Reset control, keeping 0 ("no retries") distinct from inheriting the global value
* fix(keys): align router_settings examples with typed RetryPolicy and resync UI artifacts
model_group_retry_policy is now Dict[str, RetryPolicy], so the {"max_retries": 5} sample in the key-generate test and the /key/generate and /key/update docstrings no longer validate; they now use a valid {"gpt-4": {"RateLimitErrorRetries": 5}} shape.
Regenerated eslint-metrics.json (no-explicit-any drifted 2027 -> 2026) and schema.d.ts (new RetryPolicy schema, retry_policy field, model_group_retry_policy value type) so the UI build and api-types-sync checks pass
* test(router): pin retry_policy persistence end to end (LIT-3152)
The existing retry_policy tests exercise UpdateRouterConfig and Router.update_settings in isolation, so they would all still pass if a regression flipped ConfigYAML.router_settings back to a loose dict or stopped add_deployment from applying the stored row. This drives the real handler chain an Admin UI save triggers: update_config writes the LiteLLM_Config row, the apply path forwards it to the live router, and get_config serializes it back, pinning retry_policy across persist, apply, and read-back.
* fix(teams): use valid model_group_retry_policy example in router_settings docstring
Same stale {"max_retries": 5} example the key endpoints carried; model_group_retry_policy maps a model group to a RetryPolicy, so the team /team/new and /team/update docs now show {"gpt-4": {"RateLimitErrorRetries": 5}}. Regenerated schema.d.ts to match.
* fix(ui): load retry settings via deferred fetch to satisfy set-state-in-effect
The Model Retry Settings effect called loadRetrySettings synchronously; eslint-plugin-react-hooks (react-hooks/set-state-in-effect) traces into it and flags the setState calls, failing frontend-lint. Split the loader into fetchRouterSettings + applyRouterSettings and run the fetch in an inline async IIFE with a cancellation flag, so state is applied in the post-await callback rather than on the effect's synchronous path. Behavior is unchanged and onSuccess still refreshes via loadRetrySettings.
* fix(ui): match CI rendering of RateLimitError 429 docstring in generated schema
gen:api run on a dev env (python 3.13 / newer fastapi) rendered the RateLimitError response description with 4-space indentation, but CI regenerates it with 8-space under its frozen python 3.12 toolchain, which is the canonical committed form. The Check UI API Types Sync job regenerates and diffs, so restore that block to the CI rendering; verified byte-identical to the pre-existing committed version.
* fix(ui): pin RateLimitError 429 docstring to CI's frozen schema rendering
Base #29619 regenerated schema.d.ts on a newer FastAPI that renders the RateLimitError response description at 4-space indent, but the Check UI API Types Sync job regenerates under the frozen python 3.12 toolchain, which renders 8-space. Merging base pulled in the 4-space form; restore the 8-space rendering so the generated types match what CI produces (verified byte-identical to the pre-#29619 committed form), which also corrects the base drift once this PR merges.