mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-07 02:59:05 +00:00
fix(spend-logs): honor store_prompts_in_spend_logs for guardrail_information (LIT-4314) (#32688)
* fix(spend-logs): honor store_prompts_in_spend_logs for guardrail_information (LIT-4314)
_get_spend_logs_metadata passed guardrail_information entries through
verbatim, so guardrail hooks that echo the LLM request into
guardrail_response leaked the raw prompt into LiteLLM_SpendLogs.metadata
regardless of store_prompts_in_spend_logs. This mirrored the pre-existing
gap for the other prompt-carrying fields (vector_store_request_metadata,
error_information, etc.), which already sanitize via
_should_store_prompts_and_responses_in_spend_logs.
Add _sanitize_guardrail_information_for_spend_logs alongside the other
per-field sanitizers and wire it into _get_spend_logs_metadata. When the
flag is False the sanitizer replaces guardrail_request and
guardrail_response with REDACTED_BY_LITELM_STRING while preserving every
other typed field on the entry (name, provider, mode, status, timings,
action, violation_categories, risk_score, masked_entity_count, ...) so
guardrail dashboards keep working. When the flag is True (or the field
is None) the entries pass through unchanged.
Widen StandardLoggingGuardrailInformation.guardrail_request from
Optional[dict] to Optional[Union[dict, str]] so the redacted sentinel
satisfies the TypedDict without needing a cast; guardrail_response
already accepted str.
Regression tests cover the three cases (flag=False redacts,
flag=True passes through, None passes through) plus an end-to-end
get_logging_payload path that fails if the wire-in at line 139 is
reverted.
* chore(spend-logs): review nits (one-shot dict build, scrub identifier in tests)
- _redact_prompt_fields_in_guardrail_entry now returns the redacted
dict in one expression instead of seed-then-mutate (TYPE-3)
- swap the illustrative guardrail_name in the new test fixtures for
a generic 'demo-echo-guard' identifier
* chore(spend-logs): only redact guardrail prompt fields when caller supplied them
Greptile P2: the sanitizer was unconditionally writing REDACTED_BY_LITELM
into both guardrail_request and guardrail_response on the copy, so entries
that never carried one of those fields (e.g. a guardrail that only emits
a guardrail_response) came out with a phantom guardrail_request key added.
Guard both assignments with an in-check so the output shape is stable.
Add a mutation-checked regression test that fails if either guard is
removed.
* fix(spend-logs): also redact match_details and classification in guardrail_information
The initial LIT-4314 fix redacted guardrail_request and guardrail_response,
but two other typed fields on StandardLoggingGuardrailInformation also
carry raw prompt content when a first-party guardrail populates them:
- litellm_content_filter/content_filter.py:1676 sets classification =
dict(CompetitorIntentDetection), whose evidence[*].match is a substring
taken directly from the user's normalized prompt (see
litellm_content_filter/competitor_intent/base.py:184-194).
- block_code_execution/block_code_execution.py:571 sets match_details =
guardrail_response = [dict(d) for d in detections], where detections
carry the fenced-code-block content extracted from the user's message.
Reproduced live against localhost:4000 with store_prompts_in_spend_logs
false and a custom guardrail passing tracing_detail with both fields:
before this commit the raw prompt shows up in metadata.guardrail_information[0]
under match_details and classification; after, both are the sentinel.
Widen the two TypedDict fields to Optional[Union[..., str]] so the
sentinel string satisfies the schema without a cast, and consolidate
the redaction set into a tuple so future prompt-carrying additions are
one-line changes.
* fix(spend-logs): normalize non-list guardrail_information shapes in sanitizer
xecguard's logging hook (xecguard.py:246) assigns a bare dict to
standard_logging_object['guardrail_information'] instead of a list,
violating the typed contract Optional[List[StandardLoggingGuardrailInformation]].
Without defensive normalization, _sanitize_guardrail_information_for_spend_logs
iterates the dict's string keys and _redact_prompt_fields_in_guardrail_entry
raises TypeError on {**'guardrail_name'}, which get_logging_payload's
downstream update_database catches with a broad except and silently drops
the entire spend-log write for that request.
Normalize a bare-dict input to a single-item list at the sanitizer's
entry point, and skip any non-dict entries defensively (matching OTEL's
existing isinstance filter at opentelemetry.py:1751-1753 for the same
field). Downstream readers already model this defensively; make the
spend-log write path match.
The root cause is xecguard's writer, not the sanitizer. That is being
tracked as a separate ticket; this PR keeps xecguard-enabled deploys
from silently losing spend logs when store_prompts_in_spend_logs=false.
* fix(types): declare guardrail Union members str-first to avoid poisoning typing cache
CPython's typing module caches Union[...] order-insensitively (first-
construction wins), and litellm/types/utils.py has no 'from __future__
import annotations', so its unions are constructed eagerly at import
time -- before any proxy model. Declaring guardrail_request,
classification, and match_details with dict-first ordering seeds the
typing cache with a dict-first tuple, and later proxy models that
declare custom_llm_provider / model_aliases / vertex_credentials as
Optional[Union[str, dict]] pick up the same dict-first object.
Downstream, Pydantic's get_args() then reports anyOf in dict-first
order, FastAPI emits the OpenAPI accordingly, and 'npm run gen:api'
produces a schema.d.ts diff on unrelated fields, tripping the schema-
sync CI check.
Behaviorally identical in Python and at the wire; the flip only reorders
the union members so the first construction matches how the codebase
had always declared these unions, and 'npm run gen:api' now produces a
zero diff against the committed schema.d.ts.
This commit is contained in:
parent
bf02a4a47f
commit
b8bb95be8d
3 changed files with 308 additions and 5 deletions
|
|
@ -136,7 +136,7 @@ def _get_spend_logs_metadata(
|
|||
clean_metadata["vector_store_request_metadata"] = _get_vector_store_request_for_spend_logs_payload(
|
||||
vector_store_request_metadata
|
||||
)
|
||||
clean_metadata["guardrail_information"] = guardrail_information
|
||||
clean_metadata["guardrail_information"] = _sanitize_guardrail_information_for_spend_logs(guardrail_information)
|
||||
clean_metadata["usage_object"] = usage_object
|
||||
clean_metadata["model_map_information"] = model_map_information
|
||||
clean_metadata["cold_storage_object_key"] = cold_storage_object_key
|
||||
|
|
@ -868,6 +868,51 @@ def _redact_prompt_leaks_in_error_string(text: str) -> str:
|
|||
return "".join(out)
|
||||
|
||||
|
||||
def _sanitize_guardrail_information_for_spend_logs(
|
||||
guardrail_information: Optional[List[StandardLoggingGuardrailInformation]],
|
||||
) -> Optional[List[StandardLoggingGuardrailInformation]]:
|
||||
"""
|
||||
When ``store_prompts_in_spend_logs`` is False, redact prompt-carrying fields
|
||||
(``guardrail_request``, ``guardrail_response``, ``match_details``,
|
||||
``classification``) before they land in ``LiteLLM_SpendLogs.metadata``.
|
||||
|
||||
Guardrail hooks may echo the LLM request payload back into
|
||||
``guardrail_response``, and two first-party hooks
|
||||
(``block_code_execution``, ``litellm_content_filter``) inline user-prompt
|
||||
substrings into ``match_details`` / ``classification`` too, so the flag
|
||||
must cover all four fields. Every other typed field on the entry (name,
|
||||
provider, mode, status, timings, action, violation_categories, risk_score,
|
||||
masked_entity_count, ...) is preserved so guardrail dashboards keep
|
||||
working.
|
||||
|
||||
``guardrail_information`` is typed ``Optional[List[...]]`` but at least
|
||||
one writer (``xecguard``) assigns a bare dict, so normalize to a list
|
||||
here to match OTEL's defensive read pattern; otherwise iteration would
|
||||
yield the dict's keys and crash the whole spend-log write.
|
||||
"""
|
||||
if guardrail_information is None or _should_store_prompts_and_responses_in_spend_logs():
|
||||
return guardrail_information
|
||||
entries = [guardrail_information] if isinstance(guardrail_information, dict) else guardrail_information
|
||||
return [_redact_prompt_fields_in_guardrail_entry(entry) for entry in entries if isinstance(entry, dict)]
|
||||
|
||||
|
||||
_PROMPT_CARRYING_GUARDRAIL_FIELDS = (
|
||||
"guardrail_request",
|
||||
"guardrail_response",
|
||||
"match_details",
|
||||
"classification",
|
||||
)
|
||||
|
||||
|
||||
def _redact_prompt_fields_in_guardrail_entry(
|
||||
entry: StandardLoggingGuardrailInformation,
|
||||
) -> StandardLoggingGuardrailInformation:
|
||||
return {
|
||||
**entry,
|
||||
**{key: REDACTED_BY_LITELM_STRING for key in _PROMPT_CARRYING_GUARDRAIL_FIELDS if key in entry},
|
||||
}
|
||||
|
||||
|
||||
def _sanitize_error_information_for_spend_logs(
|
||||
error_information: Optional[StandardLoggingPayloadErrorInformation],
|
||||
) -> Optional[StandardLoggingPayloadErrorInformation]:
|
||||
|
|
|
|||
|
|
@ -2698,7 +2698,7 @@ class StandardLoggingGuardrailInformation(TypedDict, total=False):
|
|||
guardrail_name: Optional[str]
|
||||
guardrail_provider: Optional[str]
|
||||
guardrail_mode: Optional[Union[GuardrailEventHooks, List[GuardrailEventHooks], GuardrailMode]]
|
||||
guardrail_request: Optional[dict]
|
||||
guardrail_request: Optional[Union[str, dict]]
|
||||
guardrail_response: Optional[Union[dict, str, List[dict]]]
|
||||
guardrail_status: GuardrailStatus
|
||||
start_time: Optional[float]
|
||||
|
|
@ -2729,10 +2729,10 @@ class StandardLoggingGuardrailInformation(TypedDict, total=False):
|
|||
confidence_score: Optional[float]
|
||||
"""For LLM-judge guardrails: confidence score 0.0-1.0"""
|
||||
|
||||
classification: Optional[dict]
|
||||
classification: Optional[Union[str, dict]]
|
||||
"""For LLM-judge guardrails: structured classification output"""
|
||||
|
||||
match_details: Optional[List[dict]]
|
||||
match_details: Optional[Union[str, List[dict]]]
|
||||
"""Detailed match information for each detected pattern"""
|
||||
|
||||
patterns_checked: Optional[int]
|
||||
|
|
|
|||
|
|
@ -33,6 +33,7 @@ from litellm.proxy.spend_tracking.spend_tracking_utils import (
|
|||
_is_master_key,
|
||||
_redact_prompt_leaks_in_error_string,
|
||||
_sanitize_error_information_for_spend_logs,
|
||||
_sanitize_guardrail_information_for_spend_logs,
|
||||
_sanitize_request_body_for_spend_logs_payload,
|
||||
_should_store_prompts_and_responses_in_spend_logs,
|
||||
get_logging_payload,
|
||||
|
|
@ -1263,6 +1264,260 @@ def test_get_spend_logs_metadata_guardrail_info_fallback_from_metadata():
|
|||
assert result["guardrail_information"] is None
|
||||
|
||||
|
||||
@patch("litellm.proxy.spend_tracking.spend_tracking_utils._should_store_prompts_and_responses_in_spend_logs")
|
||||
def test_sanitize_guardrail_information_redacts_all_prompt_carrying_fields_when_flag_false(
|
||||
mock_should_store,
|
||||
):
|
||||
"""
|
||||
match_details and classification are declared as structured metadata but
|
||||
in-tree writers (litellm_content_filter, block_code_execution) inline
|
||||
raw prompt content into them, so they leak the same way
|
||||
guardrail_request/guardrail_response do. Redaction must cover all four.
|
||||
"""
|
||||
mock_should_store.return_value = False
|
||||
guardrail_info = [
|
||||
{
|
||||
"guardrail_name": "demo-echo-guard",
|
||||
"guardrail_status": "success",
|
||||
"guardrail_request": {"messages": [{"role": "user", "content": "hi"}]},
|
||||
"guardrail_response": {"evaluated_input": "hi"},
|
||||
"match_details": [{"type": "pattern", "snippet": "hi", "action_taken": "log"}],
|
||||
"classification": {"intent": "x", "evidence": [{"match": "hi"}]},
|
||||
"guardrail_action": "NONE",
|
||||
}
|
||||
]
|
||||
|
||||
result = _sanitize_guardrail_information_for_spend_logs(guardrail_info)
|
||||
|
||||
assert result is not None
|
||||
entry = result[0]
|
||||
assert entry["guardrail_request"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["guardrail_response"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["match_details"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["classification"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["guardrail_name"] == "demo-echo-guard"
|
||||
assert entry["guardrail_status"] == "success"
|
||||
assert entry["guardrail_action"] == "NONE"
|
||||
|
||||
|
||||
@patch("litellm.proxy.spend_tracking.spend_tracking_utils._should_store_prompts_and_responses_in_spend_logs")
|
||||
def test_sanitize_guardrail_information_redacts_prompt_fields_when_flag_false(
|
||||
mock_should_store,
|
||||
):
|
||||
"""
|
||||
LIT-4314 Issue A regression: with store_prompts_in_spend_logs=False,
|
||||
guardrail_request and guardrail_response must be redacted before they
|
||||
land in LiteLLM_SpendLogs.metadata, while every other field on the
|
||||
entry is preserved bit-for-bit.
|
||||
"""
|
||||
mock_should_store.return_value = False
|
||||
guardrail_info = [
|
||||
{
|
||||
"guardrail_name": "demo-echo-guard",
|
||||
"guardrail_provider": "custom",
|
||||
"guardrail_mode": "pre_call",
|
||||
"guardrail_status": "success",
|
||||
"guardrail_request": {
|
||||
"messages": [{"role": "user", "content": "Say hi in 3 words"}],
|
||||
},
|
||||
"guardrail_response": {
|
||||
"evaluated_input": "Say hi in 3 words",
|
||||
"verdict": "allow",
|
||||
},
|
||||
"start_time": 1_700_000_000.0,
|
||||
"end_time": 1_700_000_000.5,
|
||||
"duration": 0.5,
|
||||
"guardrail_id": "gd-42",
|
||||
"masked_entity_count": {"EMAIL": 1},
|
||||
"violation_categories": ["prompt_injection"],
|
||||
"risk_score": 3.5,
|
||||
"guardrail_action": "NONE",
|
||||
}
|
||||
]
|
||||
|
||||
result = _sanitize_guardrail_information_for_spend_logs(guardrail_info)
|
||||
|
||||
assert result is not None
|
||||
assert len(result) == 1
|
||||
entry = result[0]
|
||||
assert entry["guardrail_request"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["guardrail_response"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["guardrail_name"] == "demo-echo-guard"
|
||||
assert entry["guardrail_provider"] == "custom"
|
||||
assert entry["guardrail_mode"] == "pre_call"
|
||||
assert entry["guardrail_status"] == "success"
|
||||
assert entry["start_time"] == 1_700_000_000.0
|
||||
assert entry["end_time"] == 1_700_000_000.5
|
||||
assert entry["duration"] == 0.5
|
||||
assert entry["guardrail_id"] == "gd-42"
|
||||
assert entry["masked_entity_count"] == {"EMAIL": 1}
|
||||
assert entry["violation_categories"] == ["prompt_injection"]
|
||||
assert entry["risk_score"] == 3.5
|
||||
assert entry["guardrail_action"] == "NONE"
|
||||
|
||||
assert guardrail_info[0]["guardrail_request"] == {
|
||||
"messages": [{"role": "user", "content": "Say hi in 3 words"}],
|
||||
}
|
||||
assert guardrail_info[0]["guardrail_response"] == {
|
||||
"evaluated_input": "Say hi in 3 words",
|
||||
"verdict": "allow",
|
||||
}
|
||||
|
||||
|
||||
@patch("litellm.proxy.spend_tracking.spend_tracking_utils._should_store_prompts_and_responses_in_spend_logs")
|
||||
def test_sanitize_guardrail_information_passthrough_when_flag_true(
|
||||
mock_should_store,
|
||||
):
|
||||
"""
|
||||
When store_prompts_in_spend_logs=True the sanitizer must be a no-op so
|
||||
operators who explicitly opted in still see full guardrail payloads.
|
||||
"""
|
||||
mock_should_store.return_value = True
|
||||
guardrail_info = [
|
||||
{
|
||||
"guardrail_name": "content_filter",
|
||||
"guardrail_status": "success",
|
||||
"guardrail_request": {"messages": [{"role": "user", "content": "hi"}]},
|
||||
"guardrail_response": {"verdict": "allow"},
|
||||
}
|
||||
]
|
||||
|
||||
result = _sanitize_guardrail_information_for_spend_logs(guardrail_info)
|
||||
|
||||
assert result == guardrail_info
|
||||
|
||||
|
||||
@patch("litellm.proxy.spend_tracking.spend_tracking_utils._should_store_prompts_and_responses_in_spend_logs")
|
||||
def test_sanitize_guardrail_information_none_passthrough(mock_should_store):
|
||||
mock_should_store.return_value = False
|
||||
assert _sanitize_guardrail_information_for_spend_logs(None) is None
|
||||
|
||||
|
||||
@patch("litellm.proxy.spend_tracking.spend_tracking_utils._should_store_prompts_and_responses_in_spend_logs")
|
||||
def test_sanitize_guardrail_information_normalizes_bare_dict_input(mock_should_store):
|
||||
"""
|
||||
Regression: xecguard (xecguard.py:246) assigns a bare dict to
|
||||
standard_logging_object["guardrail_information"] even though the typed
|
||||
contract is Optional[List[...]]. Without defensive normalization here,
|
||||
the for-loop would iterate the dict's string keys and _redact...
|
||||
would TypeError on {**"guardrail_name"}, taking down the entire
|
||||
spend-log write via update_database's broad except.
|
||||
"""
|
||||
mock_should_store.return_value = False
|
||||
bare_dict_entry = {
|
||||
"guardrail_name": "xecguard",
|
||||
"guardrail_status": "success",
|
||||
"guardrail_response": {"decision": "SAFE", "raw_prompt": "hi"},
|
||||
"start_time": 1.0,
|
||||
"end_time": 2.0,
|
||||
"duration": 1.0,
|
||||
}
|
||||
|
||||
result = _sanitize_guardrail_information_for_spend_logs(bare_dict_entry)
|
||||
|
||||
assert result is not None
|
||||
assert isinstance(result, list)
|
||||
assert len(result) == 1
|
||||
entry = result[0]
|
||||
assert entry["guardrail_response"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["guardrail_name"] == "xecguard"
|
||||
assert entry["guardrail_status"] == "success"
|
||||
assert entry["start_time"] == 1.0
|
||||
|
||||
|
||||
@patch("litellm.proxy.spend_tracking.spend_tracking_utils._should_store_prompts_and_responses_in_spend_logs")
|
||||
def test_sanitize_guardrail_information_drops_non_dict_items_in_list(mock_should_store):
|
||||
"""
|
||||
A stray non-dict item in the list (e.g. from a buggy caller that
|
||||
accidentally appends a string) should be silently skipped instead of
|
||||
crashing the spend-log write.
|
||||
"""
|
||||
mock_should_store.return_value = False
|
||||
mixed_input = [
|
||||
{"guardrail_name": "x", "guardrail_response": {"leak": "hi"}},
|
||||
"not-a-dict",
|
||||
None,
|
||||
]
|
||||
|
||||
result = _sanitize_guardrail_information_for_spend_logs(mixed_input)
|
||||
|
||||
assert result == [{"guardrail_name": "x", "guardrail_response": REDACTED_BY_LITELM_STRING}]
|
||||
|
||||
|
||||
@patch("litellm.proxy.spend_tracking.spend_tracking_utils._should_store_prompts_and_responses_in_spend_logs")
|
||||
def test_sanitize_guardrail_information_preserves_absent_prompt_fields(mock_should_store):
|
||||
"""
|
||||
Entries that never carried guardrail_request or guardrail_response must
|
||||
not gain those keys after sanitization; consumers keying on presence
|
||||
(`"guardrail_request" in entry`) would otherwise flip from absent to
|
||||
the sentinel string.
|
||||
"""
|
||||
mock_should_store.return_value = False
|
||||
guardrail_info = [
|
||||
{
|
||||
"guardrail_name": "demo-echo-guard",
|
||||
"guardrail_status": "success",
|
||||
"guardrail_response": {"verdict": "allow", "evaluated_input": "hi"},
|
||||
}
|
||||
]
|
||||
|
||||
result = _sanitize_guardrail_information_for_spend_logs(guardrail_info)
|
||||
|
||||
assert result is not None
|
||||
entry = result[0]
|
||||
assert "guardrail_request" not in entry
|
||||
assert entry["guardrail_response"] == REDACTED_BY_LITELM_STRING
|
||||
assert entry["guardrail_name"] == "demo-echo-guard"
|
||||
assert entry["guardrail_status"] == "success"
|
||||
|
||||
|
||||
@patch("litellm.proxy.proxy_server.master_key", "sk-master")
|
||||
@patch(
|
||||
"litellm.proxy.proxy_server.general_settings",
|
||||
{"store_prompts_in_spend_logs": False},
|
||||
)
|
||||
def test_get_logging_payload_redacts_guardrail_prompt_fields_when_flag_false():
|
||||
"""
|
||||
End-to-end wire-in check: get_logging_payload -> _get_spend_logs_metadata
|
||||
-> sanitizer. Without the wire-in at line 139, the raw guardrail_response
|
||||
lands in payload["metadata"] verbatim.
|
||||
"""
|
||||
guardrail_info = [
|
||||
{
|
||||
"guardrail_name": "demo-echo-guard",
|
||||
"guardrail_provider": "custom",
|
||||
"guardrail_status": "success",
|
||||
"guardrail_request": {"messages": [{"role": "user", "content": "secret"}]},
|
||||
"guardrail_response": {"evaluated_input": "secret"},
|
||||
}
|
||||
]
|
||||
kwargs = {
|
||||
"model": "gpt-4o-mini",
|
||||
"litellm_call_id": "test-call-id",
|
||||
"litellm_params": {
|
||||
"metadata": {
|
||||
"user_api_key": "test-key",
|
||||
"standard_logging_guardrail_information": guardrail_info,
|
||||
},
|
||||
"proxy_server_request": {},
|
||||
},
|
||||
}
|
||||
|
||||
payload = get_logging_payload(
|
||||
kwargs=kwargs,
|
||||
response_obj={},
|
||||
start_time=datetime.datetime.now(tz=timezone.utc),
|
||||
end_time=datetime.datetime.now(tz=timezone.utc),
|
||||
)
|
||||
|
||||
metadata_result = json.loads(payload["metadata"])
|
||||
stored = metadata_result["guardrail_information"][0]
|
||||
assert stored["guardrail_request"] == REDACTED_BY_LITELM_STRING
|
||||
assert stored["guardrail_response"] == REDACTED_BY_LITELM_STRING
|
||||
assert stored["guardrail_name"] == "demo-echo-guard"
|
||||
assert stored["guardrail_status"] == "success"
|
||||
|
||||
|
||||
def test_get_logging_payload_guardrail_info_when_no_standard_logging_payload():
|
||||
"""
|
||||
When a guardrail blocks a request before the LLM call, the standard_logging_object
|
||||
|
|
@ -1295,7 +1550,10 @@ def test_get_logging_payload_guardrail_info_when_no_standard_logging_payload():
|
|||
}
|
||||
|
||||
with patch("litellm.proxy.proxy_server.master_key", "sk-master"):
|
||||
with patch("litellm.proxy.proxy_server.general_settings", {}):
|
||||
with patch(
|
||||
"litellm.proxy.proxy_server.general_settings",
|
||||
{"store_prompts_in_spend_logs": True},
|
||||
):
|
||||
payload = get_logging_payload(
|
||||
kwargs=kwargs,
|
||||
response_obj={},
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue