Route records below WARNING to stdout (WARNING and above stay on stderr),
emit ANSI color codes only when both streams are a TTY (honoring NO_COLOR),
and parse JSON_LOGS strictly so JSON_LOGS=false no longer enables JSON logs.
* fix(presidio): chunk oversized text before /analyze so large content blocks do not fail
The Presidio PII guardrail sent each content block to the analyzer as a
single /analyze call with no size check. Analyzer deployments commonly cap
the request body (the reporting deployment rejects bodies over 1,000,000
bytes with HTTP 413), so large blocks failed closed, and analyzer latency
grew linearly with payload size.
analyze_text now splits texts larger than presidio_analyze_chunk_size_bytes
(default 500,000 UTF-8 bytes, configurable per guardrail) into overlapping
chunks, analyzes them concurrently, remaps each detection's start/end onto
the original text, and deduplicates detections from the overlap regions.
Anonymization, blocked-entity checks, score filtering, numbered-token
unmasking, telemetry, and the dashboard entity positions all consume the
remapped global offsets unchanged.
Resolves LIT-4785
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(presidio): review-round hardening for chunked analyze
- measure the chunk budget on the JSON-serialized text (non-ASCII escapes
expand beyond raw UTF-8, so a raw-byte budget could still exceed the
analyzer body limit)
- share the chunk fan-out semaphore per event loop and instance instead of
per call, so many oversized blocks cannot multiply concurrent analyzer
calls
- apply configured score thresholds and deny list per chunk BEFORE overlap
resolution, so a below-threshold span cannot displace a detection the
thresholds keep
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The /v1/messages bridge decided a Claude target could take `reasoning_effort` from
the model name, which says nothing about the params the provider in front of it
accepts. Snowflake serves Claude over the Anthropic dialect and declares `thinking`
alone, so `get_optional_params` raised `UnsupportedParamsError` before the request
reached the wire: every adaptive request carrying an effort tier turned a 200 into
a 400 for all seven of its Claude entries.
The tier is now offered only where the target declares the param, reading the same
`get_supported_openai_params` the sibling `_supports_prompt_cache_key` reads twelve
lines up. A target declaring neither carrier keeps its bare `thinking` block, which
is what this bridge sent before it carried a tier at all.
Without a resolved provider the tier stays behind rather than being offered blind.
Resolving one from the model's prefix instead would run an OAuth device flow for
github_copilot and chatgpt, blocking for minutes, and one of the two callers in that
position is a logging callback. The copilot case is pinned by a test.
The bundled-path check was an unanchored substring match, so a
user-supplied logo URL that happened to carry /assets/logos/ or
/_next/static/media/ in its path, and whose filename collided with one of
the 36 manifest entries, would pick up a dark-mode treatment meant only
for assets we ship.
assetPaths already draws this line for resolveLogoSrc, which returns an
external src untouched. Export that predicate instead of writing a second
one, and require a treated src to clear it.
The existing test only covered a remote URL with a bare filename, which
passed either way. The new ones fail without the guard.
* feat(ui): session-level cache observability in request logs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix: guard cache_hit filter against non-string defaults in direct calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): drop redundant cache_hit field comment
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A complexity-router setting placed beside complexity_router_config, or inside a
tier entry's litellm_params, is read by nobody: the router loads its settings only
from litellm_params.complexity_router_config. It does not stay inert. The
alias-marker forwarding and the per-tier param spread carry every unrecognized key
onto the outbound request, and all_litellm_params only knows the outer names, so
the key reaches the provider as an unknown body field and every call through that
model group fails with an error naming an internal config key.
Guard the whole set, derived from ComplexityRouterConfig.model_fields so a field
added later is covered, and scoped to complexity-router deployments because the
names only mean this there (embedding_model is a legitimate flat param on an
s3_vectors vector store). Scope is read from the same merged field view the naming
check is judged on, so a router named only by its default model is in scope and a
field added to the required-field table is covered without another edit. The write
endpoints reject with a 400 naming the keys and where they belong, config.yaml
refuses to start for the same reason max_agentic_loops does, and a tier entry is
judged by the config model itself.
An already-stored deployment keeps loading, so an upgrade cannot take a running
gateway down over a row that was written before the gate existed.
Two lazily loaded models changed without their generated artifacts being
regenerated, so check-ui-api-types has been red on every branch off staging.
The snapshot that /openapi.json serves for unloaded features was missing
ChatCompletionToolReferenceObject, and the dashboard types were missing
aws_external_id. The snapshot step runs first and short-circuits, so only the
first one was visible until it was fixed.
Both files are regenerated with `python -m litellm.proxy._lazy_openapi_snapshot`
and `npm run gen:api`, no hand edits.
* feat(alerting): add native Microsoft Teams alerting destination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(alerting): preserve active destinations on MS Teams save and confirm health test delivery
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): read persisted alerting destinations at MS Teams save time
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The dashboard's dark theme left a chunk of the bundled provider logos
unreadable: 25 of them are pure black marks on a transparent background,
so on a near-black surface they disappeared entirely, and another 11 are
dark multicolor marks drawn for a white page.
This adds the seam the rest of the work hangs off: a per-asset treatment
manifest in logoTreatments.ts, and a Logo component that applies the
treatment it names. Two treatments exist today. "invert" flattens a mark
to solid white with brightness(0) invert(1), which is what the vendor's
own white mark looks like for a pure-black transparent glyph. "plate"
puts a white surface behind the mark so it reads exactly as it does on a
light page.
Both are dark-only, and only assets named in the manifest are touched, so
light mode is unchanged and the other 96 bundled logos keep rendering
byte for byte as they do today. The className an untreated logo receives
is passed through verbatim rather than routed through cn(), so even the
class string is unchanged.
The split between invert and plate was measured per asset, not guessed:
luminance, saturation and alpha coverage sampled off a canvas render. Two
assets that look monochrome, aiml_api and repelloai, carry a light
knockout inside dark artwork, so inversion would flatten the knockout
into the mark and erase it. They get a plate instead, and a test pins
that.
Six assets whose artwork is an opaque dark box (aim_logo, aim_security,
deepgram, jina, lakeraai, openmeter) are deliberately left untreated. A
plate cannot show through an opaque image, so the only honest fix for
them is a replacement asset.
Two loading/empty transitions the disabled empty state got wrong.
A selection made in a range that had options survives a move to a range that
has none, and it still scopes the data below, so disabling the combobox
outright took away the only control that could clear it. Disable it only when
there is nothing selected to clear.
The customer list defaulted to an empty array while its query was in flight,
so the filter announced a range with no customers before anything had been
read. Leave it undefined until the query resolves, as the tag list now does.
/v1/messages forwarded `thinking` verbatim for a Claude-family model and then returned,
carrying `output_config.effort` only when the model string started with a Bedrock prefix.
Every other bridged provider got a bare adaptive thinking block, so the caller's effort did
nothing: max and minimal produced byte-identical upstream bodies.
Send those targets the tier as `reasoning_effort`, which is the param they take. Bedrock keeps
taking `output_config`, since the two are not interchangeable there: an application inference
profile ARN resolves to no chat config, so `reasoning_effort` is dropped and the tier vanishes,
and a provider that rebuilds `output_config` from it overwrites a caller-set `thinking.display`
on the way. The tier stays a plain string, the summary already travelling inside the forwarded
`thinking` block. Adaptive with no tier, and budgeted thinking, both stay exactly as they were.
Changing the date range left the previous range's tags in state until the
new request landed, so the filter either offered tags that range no longer
has or, when the old range was empty, stated "no tags" about a range nobody
had measured yet.
Stamp the fetched list with its range key and select it during render, the
same way the request tiles above already guard against a superseded range.
Clearing in the effect would be a render too late.
The usage page seeded its tag list as an empty array, so between first
paint and the tag request landing the filter had a resolved-looking empty
list and showed "No tags with usage in this range" for a scope nobody had
measured yet.
Seed it as null instead, which the filter already treats as unresolved and
hides, so the empty state only appears once the list has actually come back.
Kimi K3 accepts exactly low, high and max, defaults to max, and always thinks.
The map could not say that: medium and high have no supports_*_reasoning_effort
flag because every other reasoning model takes them, so the ten kimi-k3 entries
carried supports_reasoning alone and resolved to unknown. The dashboard then fell
back to a capability-blind level list that deliberately omits max, which is why a
kimi-k3 tier cannot be set to max thinking today.
Add reasoning_effort_levels, an array key in the shape the map already uses for
supported_endpoints and supported_modalities. Where present it is read first and
wins whole; every other entry keeps answering through the per-level flags,
unchanged. It is deliberately a different name from the computed
ModelGroupInfo.supported_reasoning_efforts, which stays derived from a group's
deployments and is never seeded from one deployment's model_info.
The levels are per entry rather than per model, because the deployments differ:
Moonshot, Together, Fireworks and Azure Foundry all forward the level unchanged
and get the model's own low/high/max, while Perplexity documents a six-value
enum it maps down internally and gets that. The /v1/messages degradation chain
consults the same declaration, so the level the map advertises is the level that
path forwards.
The router reload triggered by /model/new read the model table through
the read replica, so a lagging replica made the reload miss the just
committed row and fail the request with a 500 even though the write was
durable. Fixes#38556
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A non-ProxyException from the team, project or access-group lookup used to
escape the fallback loop and replace the provider's error. Treat it as a
denial and log it. Also drop the unrelated reformatting of test_router.py
and test_fallback_event_handlers.py so both diffs are additions only.