mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-05 02:41:56 +00:00
* feat(proxy): redact or drop individual batch records instead of rejecting the file A single record tripping a guardrail rejected the whole upload, which is unusable for a file holding thousands of rows. A record a guardrail rewrites is now submitted in its rewritten form, a record it blocks is left out, and the create response reports every changed record by both custom_id and line so a caller can reconcile against the file it sent. The same outcome is written to the proxy log and to request metadata, so it is not visible only to the caller. A rewritten record goes straight to a spool and only its offset is carried, so a masking guardrail touching most rows of a large upload does not build a second copy of the file on the heap, and the rewrite runs off the event loop the way the sibling full-file validation does. Both proxy-injected metadata keys are captured from the record and restored exactly, including an explicit null, so a masked row keeps the tags that decide how it is attributed. A record is dropped only when a guardrail judged its content. `GuardrailRaisedException` now carries `blocked_content` for that, because half its raise sites in the repo signal an unreachable or unparseable backend under a fail-closed policy, and treating those as blocks would turn "refuse this request" into "drop this record and submit the rest". The default is off, so a raise that does not say what it means aborts the upload instead of silently shrinking the file. * fix(proxy): only drop a batch record on a verdict the guardrail actually reached A guardrail that reports a technical failure as an HTTPException carrying a block status was read as a content block, so an unreachable backend under a fail-closed policy quietly shrank the file instead of failing the upload. Two in-tree integrations do exactly that, and one of them defaults to fail-closed, so the broken configuration was the default one. Such an exception is raised `from` the underlying error, which is a deliberate statement that something else caused it, and no content verdict in the repo is raised that way, so the chain now settles it. Implicit context is left alone, since a block raised inside an unrelated `except` would read as a failure. Two annotation errors in the same family: the one GuardrailRaisedException subclass in tree never opted into blocked_content, so a real block took the whole upload down with it, and straiker's block helper is reached both from its verdict and from its fail-closed handler, so it claimed a verdict for an outage. The helper now takes the flag from its caller. A record could also opt itself out of the chain. Guardrail selection reads a body-level `guardrails` key ahead of the proxy-injected list, and online that key can only add to the key and team selection, never replace it, so a batch record naming an empty list skipped every guardrail that was not default_on and was still reported as scanned. Every injected key is now stripped before dispatch and restored afterwards. A guardrail that reroutes a record to another model is honoured on the online path by rewriting the model, which the scan read as a rewrite and submitted in the same file, sending content to the provider the reroute existed to avoid. Every record of a batch file goes to one provider, so the upload is refused instead, naming the line. The scan spool is closed on the paths that never read it back. * fix(proxy): give the scan the metadata bag guardrails actually read, and close its spools The narrowed request metadata was installed under `litellm_metadata` only, but a record is scanned as the chat request it describes, and the guardrails that pick a policy from a request header read `metadata` instead. Noma choosing an application and Aim choosing a user both look there, so the header allowlist added for them did not reach either one and a batch record was still evaluated under the fallback policy. The scan metadata now goes into both bags, which are both stripped and restored, so neither survives into the record that ships. The scan spool was closed on the paths that abort, which are exactly the paths where it is empty, and left open on the one path where it holds the rewritten records. Nothing closed the rewrite output either, where before this feature the uploaded handle belonged to Starlette. The upload now owns both and closes them however it exits. * fix(proxy): register the scan spool before the rewrite can fail The scan spool was added to the request's cleanup list only after the rewrite returned, so a rewrite that raised, which for a spilled file can be as ordinary as the disk filling up, jumped to the handler with the list still empty and left the scan's own handle open. The rewrite also left its half-written output behind on that path, since nothing owns that handle until it is returned. Both now close. |
||
|---|---|---|
| .. | ||
| test_batch_guardrails.py | ||
| test_files_batch_file_validation.py | ||
| test_files_common_utils.py | ||
| test_files_endpoint.py | ||
| test_storage_backend_service.py | ||