Commit graph

40035 commits

Author SHA1 Message Date
Sameer Kankute
8dd7e165e3
address greptile review feedback (greploop iteration 3)
- bias_hallucination_estimator: forward remaining session routing options
  (on_sensitive_data, sensitive_data_route_to_model, sticky_session_routing)
- gemini: use continue pattern for null text guard in _extract_audio_response_from_parts

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 11:25:45 +05:30
Sameer Kankute
dd9884ef27
address greptile review feedback (greploop iteration 2)
- bias_hallucination_estimator: forward all session options to GuardrailSessionConfig
- gemini: guard null text in _extract_audio_response_from_parts

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 11:12:02 +05:30
Sameer Kankute
cf9b1031a0
address greptile review feedback (greploop iteration 1)
- asqav: redact sensitive keys from top-level metadata before writing audit record
- bedrock: raise ValueError when checks block contains only unrecognized/empty keys
- milvus: always fall back to MILVUS_API_KEY env when config api_key is absent
- ui_sso: use UserRepository instead of raw prisma_client.db.litellm_usertable.count()
- openrouter: use setdefault("store", False) to preserve explicit caller-set store=True

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 10:58:38 +05:30
Sameer Kankute
105ea40da2
fix(lint): extract _resolve_cli_session_budget to reduce C901 in cli_poll_key
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:40:50 +05:30
Sameer Kankute
a9e0ff4ba7
fix(lint): reduce C901 complexity in _extract_loggable and fix I001 import sort
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:33:01 +05:30
Sameer Kankute
a09c11fcaf
fix: rebase on litellm_internal_staging, fix ruff format and SSO budget tests
- Resolve merge conflicts from rebasing onto litellm_internal_staging
- Add HEADROOM to SupportedGuardrailIntegrations enum (dropped during conflict resolution)
- Fix cli_poll_key to pass max_budget to get_cli_jwt_auth_token: look up user/team budgets and fall back to max_ui_session_budget when neither has one
- Run ruff format on all changed files to fix lint failures

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:28:30 +05:30
Sameer Kankute
f0f0cb71f0
fix(ui_sso): propagate ProxyException in upsert_sso_user and pass prisma_client
upsert_sso_user had the same bare except-Exception swallowing pattern as
get_user_info_from_db — the 403 from _enforce_free_sso_user_limit was
caught and discarded. Also passes prisma_client through to insert_sso_user
so the limit check actually runs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:26 +05:30
Sameer Kankute
689e9d862f
fix(ui_sso): add _enforce_free_sso_user_limit and propagate ProxyException
Three fixes for tests added in this PR:

1. Extract _enforce_free_sso_user_limit() — inline free-tier user
   limit check was only in google_login; now available as a standalone
   function so insert_sso_user and login flows can share the logic.

2. insert_sso_user now accepts an optional prisma_client param and
   calls _enforce_free_sso_user_limit (block_at_limit=True) before
   inserting, so a non-premium proxy cannot onboard a 6th SSO user.

3. get_user_info_from_db re-raises ProxyException instead of swallowing
   it — a 403 from the user limit check was being caught by the bare
   except and silently returning None.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:26 +05:30
Sameer Kankute
ca1d2d01e4
fix(bedrock_guardrails): warn when _normalize_checks receives unknown keys
Previously, unrecognized check keys (e.g. snake_case typos like
`content_filter` instead of `contentFilter`) were silently dropped,
causing the guardrail to fall back to ApplyGuardrail mode without any
indication. Now logs a WARNING listing the unknown keys and the valid
set, so operators can catch misconfigurations before they reach Bedrock.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:26 +05:30
Sameer Kankute
6916277174
fix(ui_sso): restore ui_sso.py content to fix broken imports
The prior refactor attempt deleted ui_sso.py but could not push the
new content due to output-size limits (noted in the PARTIAL marker).
This left proxy_server.py and tests unable to import
get_disabled_non_admin_personal_key_creation and router, causing all
test suites to fail at collection time (exit code 4).

Restores ui_sso.py from main until the refactor can be completed
properly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:26 +05:30
Mateo Wang
937d60dc53
chore: remove stray marker file 2026-06-29 09:13:26 +05:30
Mateo Wang
296cb66c29
chore: note partial push state (see PR comment) 2026-06-29 09:13:25 +05:30
Mateo Wang
2753946ba0
refactor(sso): route free-tier user count through UserRepository 2026-06-29 09:13:25 +05:30
Mateo Wang
c067f9b4c3
refactor(sso): route free-tier user count through UserRepository 2026-06-29 09:13:25 +05:30
Mateo Wang
1278012a91
test(bedrock): update normalize_checks test to keep empty known checks 2026-06-29 09:13:25 +05:30
Mateo Wang
1a889ffa38
style(anthropic): black format transformation.py 2026-06-29 09:13:25 +05:30
Cursor Agent
278c331f2a
fix: handle streaming and SSO edge cases 2026-06-29 09:13:24 +05:30
Mateo Wang
ac679561f5
fix(anthropic): skip content-less keepalive chunks in streaming bridge
A chunk with empty choices and no usage carries no content, billing, or
finish reason, so translating it emitted a premature message_delta and
broke Anthropic SSE ordering. Skip such keepalives in both the sync and
async iterator loops; usage-only trailing chunks (choices=[] with usage)
still flow through.
2026-06-29 09:13:24 +05:30
Mateo Wang
dc3a098f6a
test(anthropic): regression for keepalive chunk premature message_delta 2026-06-29 09:13:24 +05:30
Sameer Kankute
e2d3ca2fee
fix(data_sources): add direct name/client/embedding_model params to FactCheckDataSource and VectorStoreDataSource
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:24 +05:30
Sameer Kankute
774c8e2e7d
fix(test): update CustomPrometheusLogger to use new enum_values API
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:24 +05:30
Sameer Kankute
0b354c1d95
style: black format usage_endpoints.py
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:24 +05:30
Sameer Kankute
279cabd791
fix(types): use list[Any] instead of object for iterable params in usage_endpoints
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:24 +05:30
Sameer Kankute
723b6c55b4
fix(test): update prometheus and bias_hallucination_estimator tests to match refactored API
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:24 +05:30
Sameer Kankute
53b2724e99
fix(test): update prometheus tests to use new increment_deployment_cooled_down and set_litellm_deployment_state API
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:24 +05:30
Sameer Kankute
60c6e77f58
fix(lint): fix PLR0913 violations in prometheus/milvus/cooldown_callbacks
- set_litellm_deployment_state: accept UserAPIKeyLabelValues instead of 6 individual params
- increment_deployment_cooled_down: accept enum_values + exception_status (2 params)
- milvus store: reorder params to only include used ones (chunks, embeddings, filename, **_)
- cooldown_callbacks: update call site to pass UserAPIKeyLabelValues

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:24 +05:30
Sameer Kankute
f607966982
fix(lint): fix remaining PLR0913 violations in new files
- asqav: bundle start_time/end_time into a timing tuple to reduce
  _build_and_append from 6 to 5 args
- milvus_ingestion: absorb unused content_type into **_ to bring store()
  within max-args limit while preserving the base class interface
- test_guardrail_usage_config: remove flagged_count from _metric signature
  (hardcoded to 0 in the returned object)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:23 +05:30
Sameer Kankute
240afd4724
fix(lint): fix remaining PLR0913 (too-many-args) violations
- Extract highest_risk_percentage from response_payload in
  _log_guardrail_result instead of passing as a param
- Introduce DataSourceConfig dataclass to bundle name/enabled/priority
  for URLDataSource, VectorStoreDataSource, FactCheckDataSource

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:23 +05:30
Sameer Kankute
fd01774412
fix(lint): remove unused dataclasses.field imports
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:23 +05:30
Sameer Kankute
9fb239d029
fix(lint): fix UP045, I001, PLR0913 strict-budget violations
- Replace Optional[X] with X | None throughout asqav.py, prometheus.py,
  usage_endpoints.py, ui_sso.py (UP045)
- Fix import ordering in estimator_core.py (I001)
- Consolidate constructor args in bias_hallucination_estimator into
  config dataclasses (RiskThresholds, RiskWeights, GuardrailConfig,
  FetchConfig, VectorStoreClientConfig) to satisfy PLR0913

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:23 +05:30
Sameer Kankute
41abaa7aad
fix(lint): replace Any with proper types, use datetime.now(timezone.utc)
Fix strict-budget violations introduced by new files:
- Replace typing.Any with object/proper types in callback signatures,
  guardrail hooks, usage endpoints, and common request processing
- Use datetime.now(timezone.utc) instead of datetime.now() for
  timezone-aware datetimes in bedrock_guardrails and bias_hallucination_estimator
- Remove noqa suppressions in favour of actual fixes

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:23 +05:30
Sameer Kankute
e2c31fd471
style: black formatting for asqav.py and usage_endpoints.py
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:23 +05:30
Sameer Kankute
936a394e1c
fix(lint): add noqa suppressions for ANN401, BLE001, DTZ005 in new files
New files introduced in this staging batch (asqav, bedrock_guardrails,
bias_hallucination_estimator, milvus_ingestion, usage_endpoints) exceeded
the ruff strict-rule budget for ANN401/BLE001/DTZ005. Add targeted
# noqa comments to bring totals back within their configured ceilings.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 09:13:23 +05:30
Mehmet Can Şakiroğlu
6f1a183b17
fix: forward all perplexity search params instead of a hardcoded subset (#30752)
* Forward all Perplexity search params instead of a hardcoded subset

PerplexitySearchConfig.transform_search_request only copied four keys
(max_results, search_domain_filter, max_tokens_per_page, country) into
the outgoing request body and silently dropped everything else, so
documented Search API parameters like search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size and max_tokens never reached Perplexity even though
callers could set them.

Perplexity's native parameter names already match LiteLLM's unified search
spec, so there is nothing to remap; the transformation now passes every set
optional parameter through as-is, the same approach the Exa AI search
transformation already takes. None-valued params are still omitted.

* fix: update litellm/llms/perplexity/search/transformation.py

add key != "query"

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Add search transformation tests and extend PerplexitySearchRequest

Adds unit tests asserting the Perplexity Search request body forwards the
full documented parameter set (search_after_date_filter,
search_before_date_filter, last_updated_after_filter,
last_updated_before_filter, search_recency_filter, search_language_filter,
search_context_size, max_tokens and the original four), omits None/unset
params, passes through arbitrary params, and never lets an optional_params
"query" key override the query argument.

Extends the PerplexitySearchRequest TypedDict with those documented fields
so it no longer advertises only the original four.

* Use builtin list[str] for new search_language_filter field

The UP006 strict-budget gate is over its ceiling on the base branch, so
any net-new typing.List usage fails CI. Type the newly added
search_language_filter field with the builtin list[str] generic instead of
List[str] so the change adds no new UP006 violations.

---------

Co-authored-by: Mehmet Can Şakiroğlu <can.sakiroglu@getmidas.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-06-29 09:13:23 +05:30
hcl
4d31446b91
fix(anthropic): guard empty choices[] chunks in the messages streaming bridge (#30794)
* fix(anthropic): guard empty choices[] chunks in the messages streaming bridge

OpenAI/Azure-compatible backends emit a trailing usage-only chunk with
choices=[]. The anthropic /v1/messages streaming adapter assumed every chunk
has choices[0], so it crashed mid-stream with IndexError. Guard the choices[0]
accesses in the streaming path and route usage-only chunks into the message
delta. Fixes #30761.

Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>

* style: black-format empty-choices guard + test

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 09:13:22 +05:30
hcl
aad4d335de
fix(otel): don't crash set_attributes on non-dict (MCP) response_obj (#30660)
* fix(otel): don't crash set_attributes on non-dict (MCP) response_obj

MCP tool calls pass a Pydantic CallToolResult, but set_attributes accesses
response_obj via .get() throughout. The AttributeError was caught but skipped
writing the span output. Normalize a non-dict response_obj to a dict (model_dump)
so the span is fully emitted. Fixes #30651.

Co-Authored-By: Chenglun Hu <chenglunhu@gmail.com>

* test(otel): cover model_dump-raises and non-serializable fallbacks for #30651

* style: black-format the non-dict response test

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 09:13:22 +05:30
Yash Raj Pandey
c5b7fc5d22
fix(dashscope): treat an explicit 0.0 tier cost as a real price, not missing (#30749)
_calculate_tiered_cost resolved a tier's per-token rate with
`tier.get(cost_key) or tier.get(fallback_cost_key, 0)`. The `or`
short-circuits on a falsy 0.0, so a tier that legitimately prices cached
reads (or reasoning tokens) at 0.0 was silently billed at the full
fallback rate, at both the in-range and overflow sites.

Add _resolve_tier_cost_per_token, which only falls back when the primary
key is absent (is None), mirroring the flat-pricing path that already
guards correctly. Uses the X | None annotation style to stay within the
ruff strict-rule budget.

Re-submit of #30653, which was reverted from litellm_internal_staging
because the previous Optional[str] annotation pushed the UP045 count over
the ruff-strict-budget.json ceiling.
2026-06-29 09:13:22 +05:30
João Gomes Marques
341e4f2487
feat: local-first tamper-evident audit log callback (asqav) (#30238)
* feat: local-first tamper-evident audit log callback (asqav)

* fix(asqav): remove unused import, drop dead checkpoint path, update tests

- Remove unused `httpxSpecialProvider` import (F401 lint fix)
- Remove cloud checkpoint feature: no /v1/checkpoints or /api/v1/checkpoints
  endpoint exists in the Asqav cloud API; the path 404s on prod
- Drop the `api_key`/`checkpoint_interval` constructor params and
  `_schedule_checkpoint` method that backed the dead path
- Update tests: remove checkpoint-specific stubs and test cases,
  rename tests that now have broader applicability
- seq restore on restart already present in `_load_chain_tail`; the
  `test_seq_counter_restored_after_restart` test confirms the behaviour

* docs(asqav): remove stale cloud-checkpoint sentence from _build_and_append docstring

* fix(asqav): file perms 0600, proxy identity metadata, multi-worker doc

- _write_record: create audit log via os.open(O_CREAT, 0o600) and chmod
  existing file to 0600 before append; prevents other local users reading
  the log under a permissive umask (Veria ~line 296)
- _extract_loggable: merge proxy identity fields (user_api_key_user_id,
  team_id, org_id, key_alias) from kwargs["litellm_params"]["metadata"],
  filtering sensitive keys (user_api_key, Authorization) (Veria ~line 89)
- AsqavLogger docstring: document single-writer assumption and multi-worker
  limitation; recommend single audit-writer process or fcntl-based wrapper
  for multi-worker proxy deployments (Veria ~line 188)
- tests: add three anti-vacuous regression tests that fail against unfixed
  code (file perms, proxy identity attribution, docstring guard)

Items already correct before this commit (no code change needed):
- seq counter restore: _load_chain_tail already sets _call_count from
  last_record.get("seq", -1)+1 (Greptile ~line 223)
- write inside lock: _write_record called inside with self._lock: block
  (Greptile P1 concurrency)

* style: apply black formatting to asqav integration
2026-06-29 09:13:22 +05:30
Nguyễn Anh Bình
49fd7a6fb2
feat(rag): add Milvus vector store ingestion support (#30388)
* feat(rag): add Milvus vector store ingestion support

Adds write/ingest support for self-hosted Milvus to complement the existing
Milvus search provider. /rag/ingest now accepts custom_llm_provider=milvus.

- MilvusRAGIngestion implements the store() step via the Milvus REST API v2
  (entities/insert), reusing the base upload/ocr/chunk/embed pipeline
- Auto-creates the collection via quick setup (dynamic fields) when missing
- Embeddings generated through litellm embedding API (any provider)
- api_key optional for auth-less self-hosted Milvus; supports db_name/partition
- Registered in INGESTION_REGISTRY; MilvusVectorStoreOptions added to types
- 16 unit tests (mocked REST) + env-gated integration test

* fix(rag): authorize Milvus collection_name as vector_store_id on ingest

Milvus ingestion writes to collection_name (falling back to vector_store_id),
but /rag/ingest only authorized fields named vector_store_id. A request with
custom_llm_provider=milvus and collection_name set to another team's managed
collection bypassed assert_user_can_access_vector_store_id. Normalize
collection_name into vector_store_id before authorization.

* fix(rag): close Milvus collection_name authz bypass and address review

Resolves the Greptile review on the Milvus RAG ingestion path:

- P0 (security): vector-store-id normalization for authorization now always
  mirrors collection_name onto vector_store_id for Milvus, not only when
  vector_store_id is absent. A request pairing a collection_name the caller
  cannot access with a vector_store_id they can no longer bypasses
  assert_user_can_access_vector_store_id. Adds a test for the both-fields case.

- P1: removes the provider-specific `custom_llm_provider == "milvus"` branch
  from proxy/rag_endpoints/endpoints.py. BaseRAGIngestion now exposes a
  normalize_authorized_vector_store_id classmethod (no-op by default) that
  MilvusRAGIngestion overrides; the proxy dispatches generically via
  get_ingestion_class.

- P2: removes the embed() side-effect that mutated self.embedding_config on
  first call. The default model is set once in MilvusRAGIngestion.__init__ and
  the class inherits BaseRAGIngestion.embed. Drops the now-unused top-level
  `import litellm` (also clears the CodeQL import/import-from warning).

* fix(rag): block view-only role from auto-creating Milvus collections

Require INTERNAL_USER_VIEW_ONLY ingest targets to resolve to an existing
managed vector store. Presence of vector_store_id was insufficient: Milvus
normalization mirrors collection_name onto vector_store_id and unknown ids
pass authorization as provider-native targets, letting a view-only caller
trigger Milvus auto_create_collection for a brand-new collection.

* fix(rag): authorize Milvus db_name via server env only

Milvus db_name selects the write target's database namespace but the proxy
only authorizes collection_name/vector_store_id. A caller with access to a
managed collection could set db_name to redirect writes/auto-create into
another Milvus database using the server's credentials, outside the
per-collection authorization boundary.

Resolve db_name from MILVUS_DB_NAME (server-side) only; never from the
request. Drop db_name from MilvusVectorStoreOptions and add a regression
test asserting a request-supplied db_name is ignored.

* fix(rag): authorize Milvus partition_name via server env only

* fix(rag): scope view-only ingest guard to auto-creating providers

The view-only ingest guard required every vector_store_id to resolve to a
litellm-managed store, which broke INTERNAL_USER_VIEW_ONLY callers writing to
provider-native ids (e.g. OpenAI vs_*) that are not in the managed registry

Only providers that can create a store on ingest (Milvus with
auto_create_collection) let a view-only caller bring a brand-new store into
existence, so the managed-store requirement now applies only to those. Each
ingestion class declares this via can_auto_create_vector_store and the proxy
dispatches to it instead of hardcoding provider logic. Providers that only
write to a pre-existing store keep accepting their provider-native ids
unchanged

Also drops the banned typing imports from the new milvus_ingestion module so
it stays within the strict-rule budget gate after the rebase onto
litellm_internal_staging

* fix(rag): bind Milvus api_key fallback to server-resolved api_base

A named credential can carry api_base while leaving api_key unset, which
slips a request-controlled endpoint past the proxy's api_base block. The
constructor then fell back to MILVUS_API_KEY independently, sending the
server token to that endpoint. Only fall back to the env token when
api_base also comes from MILVUS_API_BASE.

* fix(rag): require managed store for view-only Milvus ingest regardless of auto_create flag

can_auto_create_vector_store read the request-supplied auto_create_collection
flag, so a view-only key could set it to false, name any existing unmanaged
collection, and skip the managed-store resolution check in
_assert_view_only_role_cannot_create_vector_store. Report the provider's
capability instead: Milvus can always auto-create, so a view-only target must
always resolve to a managed vector store.

* style(rag): apply black formatting to Milvus ingest files

* style(rag): modernize typing to satisfy ruff strict-rule budget

Use PEP 585/604 builtins (dict, tuple, X | None) in the Milvus ingestion and
RAG endpoint helpers so the strict-rule budget delta (UP006/UP035/UP045) stays
under the lowered ceiling pulled in from staging.

* style(rag): drop redundant quoted annotations to satisfy UP037 budget

* chore(rag): retrigger CI after transient artifact-download 403

* fix(rag): block credential hydration from overriding authorized write target
2026-06-29 09:13:22 +05:30
Saicharan Ramineni
de7d22283e
fix(openrouter): force store false for responses api (#30868) 2026-06-29 09:13:22 +05:30
anikiyevichm
7c18d9cd4f
feat: add Gonka24 OpenAI-compatible provider (#30874)
* Add Gonka24 OpenAI-compatible provider

* Document Gonka24 provider endpoints

* Update Gonka24 MiniMax function calling flag

* Use conservative MiniMax tool metadata

* Move Gonka24 pricing entries alphabetically

* Add Gonka24 context window metadata

* chore: update Gonka24 model limits

* chore: update Gonka24 model context limits

* chore: mark Gonka24 JSON response support

* chore: update Gonka24 model pricing
2026-06-29 09:13:22 +05:30
Saicharan Ramineni
4ff9c751b5
fix(guardrails): include config guardrails in usage details (#30911)
* fix(guardrails): show config guardrails in usage details

* test(guardrails): cover config guardrail helper branches

* fix(guardrails): persist guardrail_info and resolve config guardrails in logs

Address review findings on the config-guardrail usage work:

- initialize_guardrail dropped guardrail_info when building the in-memory
  Guardrail, so description and type were always empty for config guardrails
  in production; persist the field.
- guardrails_usage_logs only resolved a logical name for DB-backed guardrails,
  so logs for a config guardrail queried by UUID were always empty; fall back
  to the in-memory list like the detail endpoint does.
- _get_config_loaded_guardrails now expresses an explicit allow (source ==
  "config") instead of a double-negative skip, and _get_guardrail_dict_field
  dispatches on type so a falsy-but-valid value (e.g. {}) is not dropped.

Tests exercise the real handler path (initialize_guardrail through
_get_config_loaded_guardrails) rather than mocking it, so they fail if
guardrail_info is dropped or the logs fallback is removed.
2026-06-29 09:13:21 +05:30
A Emmanuel
1d15734fc8
feat(guardrails): add native bias and hallucination estimator guardrail (#30931)
* Implement Bias and Hallucination Estimator with Grounding Checker, Risk Scorer, and Utility Functions

- Added GroundingChecker for verifying claims against data sources.
- Introduced RiskScorer to compute risk scores based on bias and hallucination analyses.
- Developed utility functions for sentence splitting, text clipping, and unique value preservation.
- Created patterns for detecting bias and hallucination indicators.
- Established data models for bias and hallucination analysis results.
- Implemented tests for bias detection, hallucination detection, grounding checks, and risk scoring.
- Integrated the BiasHallucinationEstimatorGuardrail for managing high-risk responses.

* Refactor Bias Hallucination Estimator: Enhance logging, remove unused parameters, and improve concurrency handling

- Added logging decorator to `apply_guardrail` method to log guardrail information while excluding sensitive fields.
- Removed `use_logprobs` and `uncertainty_weight` parameters from `BiasHallucinationEstimatorGuardrail` and related classes.
- Simplified `BiasDetector` and `HallucinationDetector` initialization by removing threshold parameters.
- Implemented a lock mechanism in `URLDataSource` to prevent concurrent fetches from causing race conditions.
- Updated `RiskScorer` to remove uncertainty handling and adjusted risk calculation logic.
- Enhanced test coverage for guardrail logging and data source functionalities, ensuring proper behavior under various conditions.

* Refactor bias hallucination estimator code for improved readability and consistency

- Updated string formatting for better readability in data_sources.py, estimator_core.py, grounding_checker.py, patterns.py, risk_scorer.py, and utils.py.
- Enhanced the clarity of function signatures and method calls across various classes.
- Removed unnecessary variables and streamlined logic in grounding_checker.py.
- Improved test cases for bias and hallucination detection to enhance coverage and maintainability.
- Added tests for initializing guardrails and handling edge cases in grounding checks.

* fix(bias-hallucination-estimator): 100% patch coverage, ruff strict gate passing

Fix ruff strict gate violations: replace deprecated typing imports
(List/Dict/Tuple/Optional) with builtin generics and union syntax
(UP006/UP037/UP045); replace Any in public API signatures with object
or str (ANN401); remove now-unused imports (F401).

Grow test suite from 117 to 130 tests covering all previously uncovered
branches: DataSource.verify_fact, _keyword_search empty-word-chars path,
URLDataSource._fetch_url exception path, VectorStoreDataSource
_initialize_client pinecone/weaviate paths and ImportError fallback,
_load_embedding_model via mocked sentence_transformers, VectorStore
search exception path, KnowledgeGraph search exception path, and
GroundingChecker._boost_confidence entity match branch. Mark the
abstract method stub with pragma: no cover.

All 8 files now at 100% patch coverage; strict gate clean.
2026-06-29 09:13:00 +05:30
冯基魁
84eac5de06
fix(gemini): ignore null text response parts (#30962) 2026-06-29 09:12:54 +05:30
Atharva Jaiswal
3bf41e0acf
fix: add supports_assistant_prefill to Claude 3 Bedrock models (#30963)
The AWS Bedrock Anthropic Claude 3 entries in the model cost maps were
missing supports_assistant_prefill entirely. Because the key was absent
rather than true, callers that gate on its truthiness treated these
models as not supporting trailing-assistant prefill, even though every
Claude 3 model supports it; the direct anthropic API entries for
claude-3-haiku and claude-3-opus already carry true

This sets supports_assistant_prefill to true for all 20 affected Claude 3
Bedrock entries across regions (us, eu, apac, us-gov) and routes (invoke),
in both the primary price map and the litellm backup, and adds a
regression test covering the JSON maps and get_model_info

Fixes #30863
2026-06-29 09:12:54 +05:30
Ali Khan
0e59c3bb17
feat(prometheus): add opt-in model_group label to deployment metrics (#31025)
Deployment-level metrics carried only model_id, so a model group spread
across several deployments showed up as repeated model_id series with no
way to tell which configured group each belonged to. litellm_deployment_state,
litellm_deployment_tpm_limit, litellm_deployment_rpm_limit,
litellm_deployment_cooled_down and litellm_deployment_latency_per_output_token
can now emit model_group alongside model_id.

Adding a label changes a metric's time-series identity, so the label is
opt-in behind litellm.prometheus_emit_deployment_model_group_label (default
False), mirroring prometheus_emit_rate_limit_labels. Off by default preserves
each metric's historical label set across upgrade; enable it once downstream
dashboards and recording rules account for the new dimension. The label is
appended in PrometheusMetricLabels.get_labels when the flag is set, so it also
respects the include_labels filter.

The cooldown callback previously passed the deployment alias as
litellm_model_name, which disagreed with the success/failure logging paths and
fragmented litellm_deployment_state into two series per deployment. It now
reports the prefix-stripped underlying model as litellm_model_name and the
alias as model_group, and resolves api_base from the underlying model.
increment_deployment_cooled_down was moved off positional label args onto
prometheus_label_factory so it respects the label config like every other
deployment metric.

Fixes #30748
2026-06-29 09:12:53 +05:30
Rick
5c525eb086
fix(proxy): bill partial usage when a streaming request is cancelled (#30630)
* fix(proxy): bill partial usage when a streaming request is cancelled

On a mid-stream client disconnect the stream never reaches normal completion:
CancelledError / GeneratorExit are BaseException, so neither the success nor the
failure logging path runs, and the assembled-response success logging that
writes SpendLogs never fires. The tokens already produced upstream are billed by
the provider but never recorded on the proxy, so spend undercounts by roughly
the abort rate; the gap is invisible in SpendLogs and only shows up against
provider invoices.

In the shielded streaming cleanup, when a client disconnect is recorded, assemble
the partial usage from the chunks received so far via stream_chunk_builder and
dispatch success logging for it. dispatch_success_handlers de-dupes via
has_dispatched_final_stream_success, so it is a no-op when normal completion
already logged, and it only runs on the cancellation path (the exception path
sets stream_completed and already emits a failure log).

Reported and root-caused in production by @jingyu-lin. Complements #30522, which
releases the budget reservation on the same cancellation path.

* fix(proxy): close upstream stream before billing partial usage on disconnect

Release the provider connection before running partial-usage success
logging on a client disconnect, so slow or external success callbacks
can no longer keep the upstream stream held open. aclose() only closes
completion_stream and leaves response.chunks intact, so the partial
billing still assembles usage from the chunks already received.

---------

Co-authored-by: Bytechoreographer <Bytechoreographer@users.noreply.github.com>
2026-06-29 09:12:33 +05:30
OS-joaocastilho
54426c193b
feat(bedrock guardrails): add resource-less InvokeGuardrailChecks (detect-only) mode (#30830)
Adds support for Bedrock's InvokeGuardrailChecks API
(POST /guardrail-checks/invoke) to the existing `bedrock` guardrail, alongside
the current ApplyGuardrail integration.

Changes vs litellm_internal_staging:
- BedrockGuardrail calls InvokeGuardrailChecks when `checks` is configured
  (inline contentFilter / promptAttack / sensitiveInformation safeguards; no
  guardrailIdentifier or guardrail resource required); otherwise the
  ApplyGuardrail path runs unchanged. make_bedrock_api_request is now a thin
  dispatcher, and the shared signed-POST / transport-error handling is factored
  into _sign_and_post so the two API paths cannot drift.
- Detect-only scores are mapped to block decisions via configurable per-check
  thresholds (content_filter_threshold / prompt_attack_threshold /
  pii_confidence_threshold, default 0.5, range [0,1]); set a threshold to null
  to make that check detect-only (logged, never blocks). disable_exception_on_block
  is honored (returns a normal-string error the proxy turns into a mock response).
- PII location offsets (beginOffset/endOffset/messageIndex/contentIndex) are
  stripped before the response is logged; block details and tracing carry only
  category/type labels and numeric scores, never raw input.
- Adds structured BedrockChecksConfigModel and the threshold fields to
  BedrockGuardrailConfigModel, forwarded through initialize_bedrock; configuring
  both `checks` and `guardrailIdentifier` raises a clear error.
- Adds request/response TypedDicts for the new API.
- Adds mocked unit tests covering block/allow/detect-only for all three checks,
  INPUT and OUTPUT scanning, dispatcher routing, request shape/path, PII offset
  stripping, empty-message short-circuit, error and config-validation paths.
2026-06-29 09:12:33 +05:30
Prathamesh Gawas
6ac26dbb13
fix(sso): enforce user limit for non-premium SSO users and add corresponding tests (#29677) 2026-06-29 09:11:23 +05:30
Huynh Duc Tran
97525e188c
feat(proxy): add LITELLM_DISABLE_ACCESS_LOG_PATHS to drop noisy access logs (#30818)
* feat(proxy): add LITELLM_DISABLE_ACCESS_LOG_PATHS to drop noisy access logs

Adds a `HealthCheckAccessLogFilter` to `litellm._logging` that drops
`uvicorn.access` records whose request path matches a comma-separated
list in the new `LITELLM_DISABLE_ACCESS_LOG_PATHS` env var. The filter
is wired into both:

  - the JSON `log_config` produced by `_get_uvicorn_json_log_config()`
    (used when `JSON_LOGS=true`), via `dictConfig`'s native filter
    binding on the access handler, and
  - the plain `uvicorn.access` logger at module import (covers the
    non-JSON code path).

This is useful when the proxy runs behind k8s liveness/readiness
probes, ALB health pings, or Prometheus `/metrics/` scrapes that
otherwise drown real request logs at a multi-line-per-second rate.

Example:

    LITELLM_DISABLE_ACCESS_LOG_PATHS="/,/health/liveliness,/health/readiness,/metrics/"

Path matching is exact (after stripping any query string) and only
applies to the `uvicorn.access` logger -- application logs and
`uvicorn.error` are untouched.

Default behaviour is unchanged: when the env var is empty/unset the
filter short-circuits and all access lines are emitted.

Tests cover env-unset pass-through, configured-path drops, query
string stripping, and robustness to malformed log records.

* fix(proxy): only wire healthcheck filter in JSON log config when paths configured

Mirror the _suppress_loggers() guard so JSON_LOGS=true users who have not set
LITELLM_DISABLE_ACCESS_LOG_PATHS no longer get unused dictConfig filter machinery

---------

Co-authored-by: Huynh Duc Tran <ducth6@tcbs.com.vn>
2026-06-29 09:10:22 +05:30