- bedrock: fix misleading warning message in _normalize_checks; partial matches
stay in InvokeGuardrailChecks mode, they do not fall back to ApplyGuardrail
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
_resolve_cli_session_budget was returning None for users with no team_id
and no personal budget, issuing uncapped JWT sessions. Cap these with
litellm.max_ui_session_budget to match the behaviour of the team path.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- bias_hallucination_estimator: forward all session options to GuardrailSessionConfig
- gemini: guard null text in _extract_audio_response_from_parts
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- asqav: redact sensitive keys from top-level metadata before writing audit record
- bedrock: raise ValueError when checks block contains only unrecognized/empty keys
- milvus: always fall back to MILVUS_API_KEY env when config api_key is absent
- ui_sso: use UserRepository instead of raw prisma_client.db.litellm_usertable.count()
- openrouter: use setdefault("store", False) to preserve explicit caller-set store=True
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Resolve merge conflicts from rebasing onto litellm_internal_staging
- Add HEADROOM to SupportedGuardrailIntegrations enum (dropped during conflict resolution)
- Fix cli_poll_key to pass max_budget to get_cli_jwt_auth_token: look up user/team budgets and fall back to max_ui_session_budget when neither has one
- Run ruff format on all changed files to fix lint failures
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
upsert_sso_user had the same bare except-Exception swallowing pattern as
get_user_info_from_db — the 403 from _enforce_free_sso_user_limit was
caught and discarded. Also passes prisma_client through to insert_sso_user
so the limit check actually runs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Three fixes for tests added in this PR:
1. Extract _enforce_free_sso_user_limit() — inline free-tier user
limit check was only in google_login; now available as a standalone
function so insert_sso_user and login flows can share the logic.
2. insert_sso_user now accepts an optional prisma_client param and
calls _enforce_free_sso_user_limit (block_at_limit=True) before
inserting, so a non-premium proxy cannot onboard a 6th SSO user.
3. get_user_info_from_db re-raises ProxyException instead of swallowing
it — a 403 from the user limit check was being caught by the bare
except and silently returning None.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously, unrecognized check keys (e.g. snake_case typos like
`content_filter` instead of `contentFilter`) were silently dropped,
causing the guardrail to fall back to ApplyGuardrail mode without any
indication. Now logs a WARNING listing the unknown keys and the valid
set, so operators can catch misconfigurations before they reach Bedrock.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The prior refactor attempt deleted ui_sso.py but could not push the
new content due to output-size limits (noted in the PARTIAL marker).
This left proxy_server.py and tests unable to import
get_disabled_non_admin_personal_key_creation and router, causing all
test suites to fail at collection time (exit code 4).
Restores ui_sso.py from main until the refactor can be completed
properly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
A chunk with empty choices and no usage carries no content, billing, or
finish reason, so translating it emitted a premature message_delta and
broke Anthropic SSE ordering. Skip such keepalives in both the sync and
async iterator loops; usage-only trailing chunks (choices=[] with usage)
still flow through.
- asqav: bundle start_time/end_time into a timing tuple to reduce
_build_and_append from 6 to 5 args
- milvus_ingestion: absorb unused content_type into **_ to bring store()
within max-args limit while preserving the base class interface
- test_guardrail_usage_config: remove flagged_count from _metric signature
(hardcoded to 0 in the returned object)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Extract highest_risk_percentage from response_payload in
_log_guardrail_result instead of passing as a param
- Introduce DataSourceConfig dataclass to bundle name/enabled/priority
for URLDataSource, VectorStoreDataSource, FactCheckDataSource
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fix strict-budget violations introduced by new files:
- Replace typing.Any with object/proper types in callback signatures,
guardrail hooks, usage endpoints, and common request processing
- Use datetime.now(timezone.utc) instead of datetime.now() for
timezone-aware datetimes in bedrock_guardrails and bias_hallucination_estimator
- Remove noqa suppressions in favour of actual fixes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
New files introduced in this staging batch (asqav, bedrock_guardrails,
bias_hallucination_estimator, milvus_ingestion, usage_endpoints) exceeded
the ruff strict-rule budget for ANN401/BLE001/DTZ005. Add targeted
# noqa comments to bring totals back within their configured ceilings.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Forward all Perplexity search params instead of a hardcoded subset
PerplexitySearchConfig.transform_search_request only copied four keys
(max_results, search_domain_filter, max_tokens_per_page, country) into
the outgoing request body and silently dropped everything else, so
documented Search API parameters like search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size and max_tokens never reached Perplexity even though
callers could set them.
Perplexity's native parameter names already match LiteLLM's unified search
spec, so there is nothing to remap; the transformation now passes every set
optional parameter through as-is, the same approach the Exa AI search
transformation already takes. None-valued params are still omitted.
* fix: update litellm/llms/perplexity/search/transformation.py
add key != "query"
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Add search transformation tests and extend PerplexitySearchRequest
Adds unit tests asserting the Perplexity Search request body forwards the
full documented parameter set (search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size, max_tokens and the original four), omits None/unset
params, passes through arbitrary params, and never lets an optional_params
"query" key override the query argument.
Extends the PerplexitySearchRequest TypedDict with those documented fields
so it no longer advertises only the original four.
* Use builtin list[str] for new search_language_filter field
The UP006 strict-budget gate is over its ceiling on the base branch, so
any net-new typing.List usage fails CI. Type the newly added
search_language_filter field with the builtin list[str] generic instead of
List[str] so the change adds no new UP006 violations.
---------
Co-authored-by: Mehmet Can Şakiroğlu <can.sakiroglu@getmidas.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix(anthropic): guard empty choices[] chunks in the messages streaming bridge
OpenAI/Azure-compatible backends emit a trailing usage-only chunk with
choices=[]. The anthropic /v1/messages streaming adapter assumed every chunk
has choices[0], so it crashed mid-stream with IndexError. Guard the choices[0]
accesses in the streaming path and route usage-only chunks into the message
delta. Fixes#30761.
Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>
* style: black-format empty-choices guard + test
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(otel): don't crash set_attributes on non-dict (MCP) response_obj
MCP tool calls pass a Pydantic CallToolResult, but set_attributes accesses
response_obj via .get() throughout. The AttributeError was caught but skipped
writing the span output. Normalize a non-dict response_obj to a dict (model_dump)
so the span is fully emitted. Fixes#30651.
Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>
* test(otel): cover model_dump-raises and non-serializable fallbacks for #30651
* style: black-format the non-dict response test
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
_calculate_tiered_cost resolved a tier's per-token rate with
`tier.get(cost_key) or tier.get(fallback_cost_key, 0)`. The `or`
short-circuits on a falsy 0.0, so a tier that legitimately prices cached
reads (or reasoning tokens) at 0.0 was silently billed at the full
fallback rate, at both the in-range and overflow sites.
Add _resolve_tier_cost_per_token, which only falls back when the primary
key is absent (is None), mirroring the flat-pricing path that already
guards correctly. Uses the X | None annotation style to stay within the
ruff strict-rule budget.
Re-submit of #30653, which was reverted from litellm_internal_staging
because the previous Optional[str] annotation pushed the UP045 count over
the ruff-strict-budget.json ceiling.
* feat: local-first tamper-evident audit log callback (asqav)
* fix(asqav): remove unused import, drop dead checkpoint path, update tests
- Remove unused `httpxSpecialProvider` import (F401 lint fix)
- Remove cloud checkpoint feature: no /v1/checkpoints or /api/v1/checkpoints
endpoint exists in the Asqav cloud API; the path 404s on prod
- Drop the `api_key`/`checkpoint_interval` constructor params and
`_schedule_checkpoint` method that backed the dead path
- Update tests: remove checkpoint-specific stubs and test cases,
rename tests that now have broader applicability
- seq restore on restart already present in `_load_chain_tail`; the
`test_seq_counter_restored_after_restart` test confirms the behaviour
* docs(asqav): remove stale cloud-checkpoint sentence from _build_and_append docstring
* fix(asqav): file perms 0600, proxy identity metadata, multi-worker doc
- _write_record: create audit log via os.open(O_CREAT, 0o600) and chmod
existing file to 0600 before append; prevents other local users reading
the log under a permissive umask (Veria ~line 296)
- _extract_loggable: merge proxy identity fields (user_api_key_user_id,
team_id, org_id, key_alias) from kwargs["litellm_params"]["metadata"],
filtering sensitive keys (user_api_key, Authorization) (Veria ~line 89)
- AsqavLogger docstring: document single-writer assumption and multi-worker
limitation; recommend single audit-writer process or fcntl-based wrapper
for multi-worker proxy deployments (Veria ~line 188)
- tests: add three anti-vacuous regression tests that fail against unfixed
code (file perms, proxy identity attribution, docstring guard)
Items already correct before this commit (no code change needed):
- seq counter restore: _load_chain_tail already sets _call_count from
last_record.get("seq", -1)+1 (Greptile ~line 223)
- write inside lock: _write_record called inside with self._lock: block
(Greptile P1 concurrency)
* style: apply black formatting to asqav integration
* feat(rag): add Milvus vector store ingestion support
Adds write/ingest support for self-hosted Milvus to complement the existing
Milvus search provider. /rag/ingest now accepts custom_llm_provider=milvus.
- MilvusRAGIngestion implements the store() step via the Milvus REST API v2
(entities/insert), reusing the base upload/ocr/chunk/embed pipeline
- Auto-creates the collection via quick setup (dynamic fields) when missing
- Embeddings generated through litellm embedding API (any provider)
- api_key optional for auth-less self-hosted Milvus; supports db_name/partition
- Registered in INGESTION_REGISTRY; MilvusVectorStoreOptions added to types
- 16 unit tests (mocked REST) + env-gated integration test
* fix(rag): authorize Milvus collection_name as vector_store_id on ingest
Milvus ingestion writes to collection_name (falling back to vector_store_id),
but /rag/ingest only authorized fields named vector_store_id. A request with
custom_llm_provider=milvus and collection_name set to another team's managed
collection bypassed assert_user_can_access_vector_store_id. Normalize
collection_name into vector_store_id before authorization.
* fix(rag): close Milvus collection_name authz bypass and address review
Resolves the Greptile review on the Milvus RAG ingestion path:
- P0 (security): vector-store-id normalization for authorization now always
mirrors collection_name onto vector_store_id for Milvus, not only when
vector_store_id is absent. A request pairing a collection_name the caller
cannot access with a vector_store_id they can no longer bypasses
assert_user_can_access_vector_store_id. Adds a test for the both-fields case.
- P1: removes the provider-specific `custom_llm_provider == "milvus"` branch
from proxy/rag_endpoints/endpoints.py. BaseRAGIngestion now exposes a
normalize_authorized_vector_store_id classmethod (no-op by default) that
MilvusRAGIngestion overrides; the proxy dispatches generically via
get_ingestion_class.
- P2: removes the embed() side-effect that mutated self.embedding_config on
first call. The default model is set once in MilvusRAGIngestion.__init__ and
the class inherits BaseRAGIngestion.embed. Drops the now-unused top-level
`import litellm` (also clears the CodeQL import/import-from warning).
* fix(rag): block view-only role from auto-creating Milvus collections
Require INTERNAL_USER_VIEW_ONLY ingest targets to resolve to an existing
managed vector store. Presence of vector_store_id was insufficient: Milvus
normalization mirrors collection_name onto vector_store_id and unknown ids
pass authorization as provider-native targets, letting a view-only caller
trigger Milvus auto_create_collection for a brand-new collection.
* fix(rag): authorize Milvus db_name via server env only
Milvus db_name selects the write target's database namespace but the proxy
only authorizes collection_name/vector_store_id. A caller with access to a
managed collection could set db_name to redirect writes/auto-create into
another Milvus database using the server's credentials, outside the
per-collection authorization boundary.
Resolve db_name from MILVUS_DB_NAME (server-side) only; never from the
request. Drop db_name from MilvusVectorStoreOptions and add a regression
test asserting a request-supplied db_name is ignored.
* fix(rag): authorize Milvus partition_name via server env only
* fix(rag): scope view-only ingest guard to auto-creating providers
The view-only ingest guard required every vector_store_id to resolve to a
litellm-managed store, which broke INTERNAL_USER_VIEW_ONLY callers writing to
provider-native ids (e.g. OpenAI vs_*) that are not in the managed registry
Only providers that can create a store on ingest (Milvus with
auto_create_collection) let a view-only caller bring a brand-new store into
existence, so the managed-store requirement now applies only to those. Each
ingestion class declares this via can_auto_create_vector_store and the proxy
dispatches to it instead of hardcoding provider logic. Providers that only
write to a pre-existing store keep accepting their provider-native ids
unchanged
Also drops the banned typing imports from the new milvus_ingestion module so
it stays within the strict-rule budget gate after the rebase onto
litellm_internal_staging
* fix(rag): bind Milvus api_key fallback to server-resolved api_base
A named credential can carry api_base while leaving api_key unset, which
slips a request-controlled endpoint past the proxy's api_base block. The
constructor then fell back to MILVUS_API_KEY independently, sending the
server token to that endpoint. Only fall back to the env token when
api_base also comes from MILVUS_API_BASE.
* fix(rag): require managed store for view-only Milvus ingest regardless of auto_create flag
can_auto_create_vector_store read the request-supplied auto_create_collection
flag, so a view-only key could set it to false, name any existing unmanaged
collection, and skip the managed-store resolution check in
_assert_view_only_role_cannot_create_vector_store. Report the provider's
capability instead: Milvus can always auto-create, so a view-only target must
always resolve to a managed vector store.
* style(rag): apply black formatting to Milvus ingest files
* style(rag): modernize typing to satisfy ruff strict-rule budget
Use PEP 585/604 builtins (dict, tuple, X | None) in the Milvus ingestion and
RAG endpoint helpers so the strict-rule budget delta (UP006/UP035/UP045) stays
under the lowered ceiling pulled in from staging.
* style(rag): drop redundant quoted annotations to satisfy UP037 budget
* chore(rag): retrigger CI after transient artifact-download 403
* fix(rag): block credential hydration from overriding authorized write target
* fix(guardrails): show config guardrails in usage details
* test(guardrails): cover config guardrail helper branches
* fix(guardrails): persist guardrail_info and resolve config guardrails in logs
Address review findings on the config-guardrail usage work:
- initialize_guardrail dropped guardrail_info when building the in-memory
Guardrail, so description and type were always empty for config guardrails
in production; persist the field.
- guardrails_usage_logs only resolved a logical name for DB-backed guardrails,
so logs for a config guardrail queried by UUID were always empty; fall back
to the in-memory list like the detail endpoint does.
- _get_config_loaded_guardrails now expresses an explicit allow (source ==
"config") instead of a double-negative skip, and _get_guardrail_dict_field
dispatches on type so a falsy-but-valid value (e.g. {}) is not dropped.
Tests exercise the real handler path (initialize_guardrail through
_get_config_loaded_guardrails) rather than mocking it, so they fail if
guardrail_info is dropped or the logs fallback is removed.
* Implement Bias and Hallucination Estimator with Grounding Checker, Risk Scorer, and Utility Functions
- Added GroundingChecker for verifying claims against data sources.
- Introduced RiskScorer to compute risk scores based on bias and hallucination analyses.
- Developed utility functions for sentence splitting, text clipping, and unique value preservation.
- Created patterns for detecting bias and hallucination indicators.
- Established data models for bias and hallucination analysis results.
- Implemented tests for bias detection, hallucination detection, grounding checks, and risk scoring.
- Integrated the BiasHallucinationEstimatorGuardrail for managing high-risk responses.
* Refactor Bias Hallucination Estimator: Enhance logging, remove unused parameters, and improve concurrency handling
- Added logging decorator to `apply_guardrail` method to log guardrail information while excluding sensitive fields.
- Removed `use_logprobs` and `uncertainty_weight` parameters from `BiasHallucinationEstimatorGuardrail` and related classes.
- Simplified `BiasDetector` and `HallucinationDetector` initialization by removing threshold parameters.
- Implemented a lock mechanism in `URLDataSource` to prevent concurrent fetches from causing race conditions.
- Updated `RiskScorer` to remove uncertainty handling and adjusted risk calculation logic.
- Enhanced test coverage for guardrail logging and data source functionalities, ensuring proper behavior under various conditions.
* Refactor bias hallucination estimator code for improved readability and consistency
- Updated string formatting for better readability in data_sources.py, estimator_core.py, grounding_checker.py, patterns.py, risk_scorer.py, and utils.py.
- Enhanced the clarity of function signatures and method calls across various classes.
- Removed unnecessary variables and streamlined logic in grounding_checker.py.
- Improved test cases for bias and hallucination detection to enhance coverage and maintainability.
- Added tests for initializing guardrails and handling edge cases in grounding checks.
* fix(bias-hallucination-estimator): 100% patch coverage, ruff strict gate passing
Fix ruff strict gate violations: replace deprecated typing imports
(List/Dict/Tuple/Optional) with builtin generics and union syntax
(UP006/UP037/UP045); replace Any in public API signatures with object
or str (ANN401); remove now-unused imports (F401).
Grow test suite from 117 to 130 tests covering all previously uncovered
branches: DataSource.verify_fact, _keyword_search empty-word-chars path,
URLDataSource._fetch_url exception path, VectorStoreDataSource
_initialize_client pinecone/weaviate paths and ImportError fallback,
_load_embedding_model via mocked sentence_transformers, VectorStore
search exception path, KnowledgeGraph search exception path, and
GroundingChecker._boost_confidence entity match branch. Mark the
abstract method stub with pragma: no cover.
All 8 files now at 100% patch coverage; strict gate clean.
The AWS Bedrock Anthropic Claude 3 entries in the model cost maps were
missing supports_assistant_prefill entirely. Because the key was absent
rather than true, callers that gate on its truthiness treated these
models as not supporting trailing-assistant prefill, even though every
Claude 3 model supports it; the direct anthropic API entries for
claude-3-haiku and claude-3-opus already carry true
This sets supports_assistant_prefill to true for all 20 affected Claude 3
Bedrock entries across regions (us, eu, apac, us-gov) and routes (invoke),
in both the primary price map and the litellm backup, and adds a
regression test covering the JSON maps and get_model_info
Fixes#30863
Deployment-level metrics carried only model_id, so a model group spread
across several deployments showed up as repeated model_id series with no
way to tell which configured group each belonged to. litellm_deployment_state,
litellm_deployment_tpm_limit, litellm_deployment_rpm_limit,
litellm_deployment_cooled_down and litellm_deployment_latency_per_output_token
can now emit model_group alongside model_id.
Adding a label changes a metric's time-series identity, so the label is
opt-in behind litellm.prometheus_emit_deployment_model_group_label (default
False), mirroring prometheus_emit_rate_limit_labels. Off by default preserves
each metric's historical label set across upgrade; enable it once downstream
dashboards and recording rules account for the new dimension. The label is
appended in PrometheusMetricLabels.get_labels when the flag is set, so it also
respects the include_labels filter.
The cooldown callback previously passed the deployment alias as
litellm_model_name, which disagreed with the success/failure logging paths and
fragmented litellm_deployment_state into two series per deployment. It now
reports the prefix-stripped underlying model as litellm_model_name and the
alias as model_group, and resolves api_base from the underlying model.
increment_deployment_cooled_down was moved off positional label args onto
prometheus_label_factory so it respects the label config like every other
deployment metric.
Fixes#30748