_resolve_cli_session_budget was returning None for users with no team_id
and no personal budget, issuing uncapped JWT sessions. Cap these with
litellm.max_ui_session_budget to match the behaviour of the team path.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- bias_hallucination_estimator: forward all session options to GuardrailSessionConfig
- gemini: guard null text in _extract_audio_response_from_parts
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- asqav: redact sensitive keys from top-level metadata before writing audit record
- bedrock: raise ValueError when checks block contains only unrecognized/empty keys
- milvus: always fall back to MILVUS_API_KEY env when config api_key is absent
- ui_sso: use UserRepository instead of raw prisma_client.db.litellm_usertable.count()
- openrouter: use setdefault("store", False) to preserve explicit caller-set store=True
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Resolve merge conflicts from rebasing onto litellm_internal_staging
- Add HEADROOM to SupportedGuardrailIntegrations enum (dropped during conflict resolution)
- Fix cli_poll_key to pass max_budget to get_cli_jwt_auth_token: look up user/team budgets and fall back to max_ui_session_budget when neither has one
- Run ruff format on all changed files to fix lint failures
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
upsert_sso_user had the same bare except-Exception swallowing pattern as
get_user_info_from_db — the 403 from _enforce_free_sso_user_limit was
caught and discarded. Also passes prisma_client through to insert_sso_user
so the limit check actually runs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Three fixes for tests added in this PR:
1. Extract _enforce_free_sso_user_limit() — inline free-tier user
limit check was only in google_login; now available as a standalone
function so insert_sso_user and login flows can share the logic.
2. insert_sso_user now accepts an optional prisma_client param and
calls _enforce_free_sso_user_limit (block_at_limit=True) before
inserting, so a non-premium proxy cannot onboard a 6th SSO user.
3. get_user_info_from_db re-raises ProxyException instead of swallowing
it — a 403 from the user limit check was being caught by the bare
except and silently returning None.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously, unrecognized check keys (e.g. snake_case typos like
`content_filter` instead of `contentFilter`) were silently dropped,
causing the guardrail to fall back to ApplyGuardrail mode without any
indication. Now logs a WARNING listing the unknown keys and the valid
set, so operators can catch misconfigurations before they reach Bedrock.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The prior refactor attempt deleted ui_sso.py but could not push the
new content due to output-size limits (noted in the PARTIAL marker).
This left proxy_server.py and tests unable to import
get_disabled_non_admin_personal_key_creation and router, causing all
test suites to fail at collection time (exit code 4).
Restores ui_sso.py from main until the refactor can be completed
properly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
A chunk with empty choices and no usage carries no content, billing, or
finish reason, so translating it emitted a premature message_delta and
broke Anthropic SSE ordering. Skip such keepalives in both the sync and
async iterator loops; usage-only trailing chunks (choices=[] with usage)
still flow through.
- asqav: bundle start_time/end_time into a timing tuple to reduce
_build_and_append from 6 to 5 args
- milvus_ingestion: absorb unused content_type into **_ to bring store()
within max-args limit while preserving the base class interface
- test_guardrail_usage_config: remove flagged_count from _metric signature
(hardcoded to 0 in the returned object)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Extract highest_risk_percentage from response_payload in
_log_guardrail_result instead of passing as a param
- Introduce DataSourceConfig dataclass to bundle name/enabled/priority
for URLDataSource, VectorStoreDataSource, FactCheckDataSource
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fix strict-budget violations introduced by new files:
- Replace typing.Any with object/proper types in callback signatures,
guardrail hooks, usage endpoints, and common request processing
- Use datetime.now(timezone.utc) instead of datetime.now() for
timezone-aware datetimes in bedrock_guardrails and bias_hallucination_estimator
- Remove noqa suppressions in favour of actual fixes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
New files introduced in this staging batch (asqav, bedrock_guardrails,
bias_hallucination_estimator, milvus_ingestion, usage_endpoints) exceeded
the ruff strict-rule budget for ANN401/BLE001/DTZ005. Add targeted
# noqa comments to bring totals back within their configured ceilings.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Forward all Perplexity search params instead of a hardcoded subset
PerplexitySearchConfig.transform_search_request only copied four keys
(max_results, search_domain_filter, max_tokens_per_page, country) into
the outgoing request body and silently dropped everything else, so
documented Search API parameters like search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size and max_tokens never reached Perplexity even though
callers could set them.
Perplexity's native parameter names already match LiteLLM's unified search
spec, so there is nothing to remap; the transformation now passes every set
optional parameter through as-is, the same approach the Exa AI search
transformation already takes. None-valued params are still omitted.
* fix: update litellm/llms/perplexity/search/transformation.py
add key != "query"
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Add search transformation tests and extend PerplexitySearchRequest
Adds unit tests asserting the Perplexity Search request body forwards the
full documented parameter set (search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size, max_tokens and the original four), omits None/unset
params, passes through arbitrary params, and never lets an optional_params
"query" key override the query argument.
Extends the PerplexitySearchRequest TypedDict with those documented fields
so it no longer advertises only the original four.
* Use builtin list[str] for new search_language_filter field
The UP006 strict-budget gate is over its ceiling on the base branch, so
any net-new typing.List usage fails CI. Type the newly added
search_language_filter field with the builtin list[str] generic instead of
List[str] so the change adds no new UP006 violations.
---------
Co-authored-by: Mehmet Can Şakiroğlu <can.sakiroglu@getmidas.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix(anthropic): guard empty choices[] chunks in the messages streaming bridge
OpenAI/Azure-compatible backends emit a trailing usage-only chunk with
choices=[]. The anthropic /v1/messages streaming adapter assumed every chunk
has choices[0], so it crashed mid-stream with IndexError. Guard the choices[0]
accesses in the streaming path and route usage-only chunks into the message
delta. Fixes#30761.
Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>
* style: black-format empty-choices guard + test
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(otel): don't crash set_attributes on non-dict (MCP) response_obj
MCP tool calls pass a Pydantic CallToolResult, but set_attributes accesses
response_obj via .get() throughout. The AttributeError was caught but skipped
writing the span output. Normalize a non-dict response_obj to a dict (model_dump)
so the span is fully emitted. Fixes#30651.
Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>
* test(otel): cover model_dump-raises and non-serializable fallbacks for #30651
* style: black-format the non-dict response test
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
_calculate_tiered_cost resolved a tier's per-token rate with
`tier.get(cost_key) or tier.get(fallback_cost_key, 0)`. The `or`
short-circuits on a falsy 0.0, so a tier that legitimately prices cached
reads (or reasoning tokens) at 0.0 was silently billed at the full
fallback rate, at both the in-range and overflow sites.
Add _resolve_tier_cost_per_token, which only falls back when the primary
key is absent (is None), mirroring the flat-pricing path that already
guards correctly. Uses the X | None annotation style to stay within the
ruff strict-rule budget.
Re-submit of #30653, which was reverted from litellm_internal_staging
because the previous Optional[str] annotation pushed the UP045 count over
the ruff-strict-budget.json ceiling.
* feat: local-first tamper-evident audit log callback (asqav)
* fix(asqav): remove unused import, drop dead checkpoint path, update tests
- Remove unused `httpxSpecialProvider` import (F401 lint fix)
- Remove cloud checkpoint feature: no /v1/checkpoints or /api/v1/checkpoints
endpoint exists in the Asqav cloud API; the path 404s on prod
- Drop the `api_key`/`checkpoint_interval` constructor params and
`_schedule_checkpoint` method that backed the dead path
- Update tests: remove checkpoint-specific stubs and test cases,
rename tests that now have broader applicability
- seq restore on restart already present in `_load_chain_tail`; the
`test_seq_counter_restored_after_restart` test confirms the behaviour
* docs(asqav): remove stale cloud-checkpoint sentence from _build_and_append docstring
* fix(asqav): file perms 0600, proxy identity metadata, multi-worker doc
- _write_record: create audit log via os.open(O_CREAT, 0o600) and chmod
existing file to 0600 before append; prevents other local users reading
the log under a permissive umask (Veria ~line 296)
- _extract_loggable: merge proxy identity fields (user_api_key_user_id,
team_id, org_id, key_alias) from kwargs["litellm_params"]["metadata"],
filtering sensitive keys (user_api_key, Authorization) (Veria ~line 89)
- AsqavLogger docstring: document single-writer assumption and multi-worker
limitation; recommend single audit-writer process or fcntl-based wrapper
for multi-worker proxy deployments (Veria ~line 188)
- tests: add three anti-vacuous regression tests that fail against unfixed
code (file perms, proxy identity attribution, docstring guard)
Items already correct before this commit (no code change needed):
- seq counter restore: _load_chain_tail already sets _call_count from
last_record.get("seq", -1)+1 (Greptile ~line 223)
- write inside lock: _write_record called inside with self._lock: block
(Greptile P1 concurrency)
* style: apply black formatting to asqav integration
* feat(rag): add Milvus vector store ingestion support
Adds write/ingest support for self-hosted Milvus to complement the existing
Milvus search provider. /rag/ingest now accepts custom_llm_provider=milvus.
- MilvusRAGIngestion implements the store() step via the Milvus REST API v2
(entities/insert), reusing the base upload/ocr/chunk/embed pipeline
- Auto-creates the collection via quick setup (dynamic fields) when missing
- Embeddings generated through litellm embedding API (any provider)
- api_key optional for auth-less self-hosted Milvus; supports db_name/partition
- Registered in INGESTION_REGISTRY; MilvusVectorStoreOptions added to types
- 16 unit tests (mocked REST) + env-gated integration test
* fix(rag): authorize Milvus collection_name as vector_store_id on ingest
Milvus ingestion writes to collection_name (falling back to vector_store_id),
but /rag/ingest only authorized fields named vector_store_id. A request with
custom_llm_provider=milvus and collection_name set to another team's managed
collection bypassed assert_user_can_access_vector_store_id. Normalize
collection_name into vector_store_id before authorization.
* fix(rag): close Milvus collection_name authz bypass and address review
Resolves the Greptile review on the Milvus RAG ingestion path:
- P0 (security): vector-store-id normalization for authorization now always
mirrors collection_name onto vector_store_id for Milvus, not only when
vector_store_id is absent. A request pairing a collection_name the caller
cannot access with a vector_store_id they can no longer bypasses
assert_user_can_access_vector_store_id. Adds a test for the both-fields case.
- P1: removes the provider-specific `custom_llm_provider == "milvus"` branch
from proxy/rag_endpoints/endpoints.py. BaseRAGIngestion now exposes a
normalize_authorized_vector_store_id classmethod (no-op by default) that
MilvusRAGIngestion overrides; the proxy dispatches generically via
get_ingestion_class.
- P2: removes the embed() side-effect that mutated self.embedding_config on
first call. The default model is set once in MilvusRAGIngestion.__init__ and
the class inherits BaseRAGIngestion.embed. Drops the now-unused top-level
`import litellm` (also clears the CodeQL import/import-from warning).
* fix(rag): block view-only role from auto-creating Milvus collections
Require INTERNAL_USER_VIEW_ONLY ingest targets to resolve to an existing
managed vector store. Presence of vector_store_id was insufficient: Milvus
normalization mirrors collection_name onto vector_store_id and unknown ids
pass authorization as provider-native targets, letting a view-only caller
trigger Milvus auto_create_collection for a brand-new collection.
* fix(rag): authorize Milvus db_name via server env only
Milvus db_name selects the write target's database namespace but the proxy
only authorizes collection_name/vector_store_id. A caller with access to a
managed collection could set db_name to redirect writes/auto-create into
another Milvus database using the server's credentials, outside the
per-collection authorization boundary.
Resolve db_name from MILVUS_DB_NAME (server-side) only; never from the
request. Drop db_name from MilvusVectorStoreOptions and add a regression
test asserting a request-supplied db_name is ignored.
* fix(rag): authorize Milvus partition_name via server env only
* fix(rag): scope view-only ingest guard to auto-creating providers
The view-only ingest guard required every vector_store_id to resolve to a
litellm-managed store, which broke INTERNAL_USER_VIEW_ONLY callers writing to
provider-native ids (e.g. OpenAI vs_*) that are not in the managed registry
Only providers that can create a store on ingest (Milvus with
auto_create_collection) let a view-only caller bring a brand-new store into
existence, so the managed-store requirement now applies only to those. Each
ingestion class declares this via can_auto_create_vector_store and the proxy
dispatches to it instead of hardcoding provider logic. Providers that only
write to a pre-existing store keep accepting their provider-native ids
unchanged
Also drops the banned typing imports from the new milvus_ingestion module so
it stays within the strict-rule budget gate after the rebase onto
litellm_internal_staging
* fix(rag): bind Milvus api_key fallback to server-resolved api_base
A named credential can carry api_base while leaving api_key unset, which
slips a request-controlled endpoint past the proxy's api_base block. The
constructor then fell back to MILVUS_API_KEY independently, sending the
server token to that endpoint. Only fall back to the env token when
api_base also comes from MILVUS_API_BASE.
* fix(rag): require managed store for view-only Milvus ingest regardless of auto_create flag
can_auto_create_vector_store read the request-supplied auto_create_collection
flag, so a view-only key could set it to false, name any existing unmanaged
collection, and skip the managed-store resolution check in
_assert_view_only_role_cannot_create_vector_store. Report the provider's
capability instead: Milvus can always auto-create, so a view-only target must
always resolve to a managed vector store.
* style(rag): apply black formatting to Milvus ingest files
* style(rag): modernize typing to satisfy ruff strict-rule budget
Use PEP 585/604 builtins (dict, tuple, X | None) in the Milvus ingestion and
RAG endpoint helpers so the strict-rule budget delta (UP006/UP035/UP045) stays
under the lowered ceiling pulled in from staging.
* style(rag): drop redundant quoted annotations to satisfy UP037 budget
* chore(rag): retrigger CI after transient artifact-download 403
* fix(rag): block credential hydration from overriding authorized write target
* fix(guardrails): show config guardrails in usage details
* test(guardrails): cover config guardrail helper branches
* fix(guardrails): persist guardrail_info and resolve config guardrails in logs
Address review findings on the config-guardrail usage work:
- initialize_guardrail dropped guardrail_info when building the in-memory
Guardrail, so description and type were always empty for config guardrails
in production; persist the field.
- guardrails_usage_logs only resolved a logical name for DB-backed guardrails,
so logs for a config guardrail queried by UUID were always empty; fall back
to the in-memory list like the detail endpoint does.
- _get_config_loaded_guardrails now expresses an explicit allow (source ==
"config") instead of a double-negative skip, and _get_guardrail_dict_field
dispatches on type so a falsy-but-valid value (e.g. {}) is not dropped.
Tests exercise the real handler path (initialize_guardrail through
_get_config_loaded_guardrails) rather than mocking it, so they fail if
guardrail_info is dropped or the logs fallback is removed.
* Implement Bias and Hallucination Estimator with Grounding Checker, Risk Scorer, and Utility Functions
- Added GroundingChecker for verifying claims against data sources.
- Introduced RiskScorer to compute risk scores based on bias and hallucination analyses.
- Developed utility functions for sentence splitting, text clipping, and unique value preservation.
- Created patterns for detecting bias and hallucination indicators.
- Established data models for bias and hallucination analysis results.
- Implemented tests for bias detection, hallucination detection, grounding checks, and risk scoring.
- Integrated the BiasHallucinationEstimatorGuardrail for managing high-risk responses.
* Refactor Bias Hallucination Estimator: Enhance logging, remove unused parameters, and improve concurrency handling
- Added logging decorator to `apply_guardrail` method to log guardrail information while excluding sensitive fields.
- Removed `use_logprobs` and `uncertainty_weight` parameters from `BiasHallucinationEstimatorGuardrail` and related classes.
- Simplified `BiasDetector` and `HallucinationDetector` initialization by removing threshold parameters.
- Implemented a lock mechanism in `URLDataSource` to prevent concurrent fetches from causing race conditions.
- Updated `RiskScorer` to remove uncertainty handling and adjusted risk calculation logic.
- Enhanced test coverage for guardrail logging and data source functionalities, ensuring proper behavior under various conditions.
* Refactor bias hallucination estimator code for improved readability and consistency
- Updated string formatting for better readability in data_sources.py, estimator_core.py, grounding_checker.py, patterns.py, risk_scorer.py, and utils.py.
- Enhanced the clarity of function signatures and method calls across various classes.
- Removed unnecessary variables and streamlined logic in grounding_checker.py.
- Improved test cases for bias and hallucination detection to enhance coverage and maintainability.
- Added tests for initializing guardrails and handling edge cases in grounding checks.
* fix(bias-hallucination-estimator): 100% patch coverage, ruff strict gate passing
Fix ruff strict gate violations: replace deprecated typing imports
(List/Dict/Tuple/Optional) with builtin generics and union syntax
(UP006/UP037/UP045); replace Any in public API signatures with object
or str (ANN401); remove now-unused imports (F401).
Grow test suite from 117 to 130 tests covering all previously uncovered
branches: DataSource.verify_fact, _keyword_search empty-word-chars path,
URLDataSource._fetch_url exception path, VectorStoreDataSource
_initialize_client pinecone/weaviate paths and ImportError fallback,
_load_embedding_model via mocked sentence_transformers, VectorStore
search exception path, KnowledgeGraph search exception path, and
GroundingChecker._boost_confidence entity match branch. Mark the
abstract method stub with pragma: no cover.
All 8 files now at 100% patch coverage; strict gate clean.
The AWS Bedrock Anthropic Claude 3 entries in the model cost maps were
missing supports_assistant_prefill entirely. Because the key was absent
rather than true, callers that gate on its truthiness treated these
models as not supporting trailing-assistant prefill, even though every
Claude 3 model supports it; the direct anthropic API entries for
claude-3-haiku and claude-3-opus already carry true
This sets supports_assistant_prefill to true for all 20 affected Claude 3
Bedrock entries across regions (us, eu, apac, us-gov) and routes (invoke),
in both the primary price map and the litellm backup, and adds a
regression test covering the JSON maps and get_model_info
Fixes#30863
Deployment-level metrics carried only model_id, so a model group spread
across several deployments showed up as repeated model_id series with no
way to tell which configured group each belonged to. litellm_deployment_state,
litellm_deployment_tpm_limit, litellm_deployment_rpm_limit,
litellm_deployment_cooled_down and litellm_deployment_latency_per_output_token
can now emit model_group alongside model_id.
Adding a label changes a metric's time-series identity, so the label is
opt-in behind litellm.prometheus_emit_deployment_model_group_label (default
False), mirroring prometheus_emit_rate_limit_labels. Off by default preserves
each metric's historical label set across upgrade; enable it once downstream
dashboards and recording rules account for the new dimension. The label is
appended in PrometheusMetricLabels.get_labels when the flag is set, so it also
respects the include_labels filter.
The cooldown callback previously passed the deployment alias as
litellm_model_name, which disagreed with the success/failure logging paths and
fragmented litellm_deployment_state into two series per deployment. It now
reports the prefix-stripped underlying model as litellm_model_name and the
alias as model_group, and resolves api_base from the underlying model.
increment_deployment_cooled_down was moved off positional label args onto
prometheus_label_factory so it respects the label config like every other
deployment metric.
Fixes#30748
* fix(proxy): bill partial usage when a streaming request is cancelled
On a mid-stream client disconnect the stream never reaches normal completion:
CancelledError / GeneratorExit are BaseException, so neither the success nor the
failure logging path runs, and the assembled-response success logging that
writes SpendLogs never fires. The tokens already produced upstream are billed by
the provider but never recorded on the proxy, so spend undercounts by roughly
the abort rate; the gap is invisible in SpendLogs and only shows up against
provider invoices.
In the shielded streaming cleanup, when a client disconnect is recorded, assemble
the partial usage from the chunks received so far via stream_chunk_builder and
dispatch success logging for it. dispatch_success_handlers de-dupes via
has_dispatched_final_stream_success, so it is a no-op when normal completion
already logged, and it only runs on the cancellation path (the exception path
sets stream_completed and already emits a failure log).
Reported and root-caused in production by @jingyu-lin. Complements #30522, which
releases the budget reservation on the same cancellation path.
* fix(proxy): close upstream stream before billing partial usage on disconnect
Release the provider connection before running partial-usage success
logging on a client disconnect, so slow or external success callbacks
can no longer keep the upstream stream held open. aclose() only closes
completion_stream and leaves response.chunks intact, so the partial
billing still assembles usage from the chunks already received.
---------
Co-authored-by: Bytechoreographer <Bytechoreographer@users.noreply.github.com>
Adds support for Bedrock's InvokeGuardrailChecks API
(POST /guardrail-checks/invoke) to the existing `bedrock` guardrail, alongside
the current ApplyGuardrail integration.
Changes vs litellm_internal_staging:
- BedrockGuardrail calls InvokeGuardrailChecks when `checks` is configured
(inline contentFilter / promptAttack / sensitiveInformation safeguards; no
guardrailIdentifier or guardrail resource required); otherwise the
ApplyGuardrail path runs unchanged. make_bedrock_api_request is now a thin
dispatcher, and the shared signed-POST / transport-error handling is factored
into _sign_and_post so the two API paths cannot drift.
- Detect-only scores are mapped to block decisions via configurable per-check
thresholds (content_filter_threshold / prompt_attack_threshold /
pii_confidence_threshold, default 0.5, range [0,1]); set a threshold to null
to make that check detect-only (logged, never blocks). disable_exception_on_block
is honored (returns a normal-string error the proxy turns into a mock response).
- PII location offsets (beginOffset/endOffset/messageIndex/contentIndex) are
stripped before the response is logged; block details and tracing carry only
category/type labels and numeric scores, never raw input.
- Adds structured BedrockChecksConfigModel and the threshold fields to
BedrockGuardrailConfigModel, forwarded through initialize_bedrock; configuring
both `checks` and `guardrailIdentifier` raises a clear error.
- Adds request/response TypedDicts for the new API.
- Adds mocked unit tests covering block/allow/detect-only for all three checks,
INPUT and OUTPUT scanning, dispatcher routing, request shape/path, PII offset
stripping, empty-message short-circuit, error and config-validation paths.