Two params were advertised for the MAI image models and dropped downstream,
so the caller got a 200 that did not match the request, or an opaque
provider 400.
n: get_supported_openai_params returns ["n", "size"], so n passes validation
and is forwarded. The MAI endpoint (/mai/v1/images/generations) has no count
field at all — its documented body is model/prompt/width/height, plus image
for edits — and ignores both `n` and the native `sampleCount`. Measured
against MAI-Image-2.5 and MAI-Image-2.5-Flash: n=2 and n=4 each return HTTP
200 with exactly one image, billed as one, with nothing in the response
saying the request was reduced. A caller balancing cost against image count
cannot see it. n=1 still passes through; n>1 now raises unless drop_params
is set, which is the existing opt-in for silently dropping a param.
size: _map_size_param's table offered five sizes, of which one is usable.
MAI requires width and height >= 768px and width*height <= 1048576, so
512x512 and 256x256 are under the per-side minimum and 1792x1024 / 1024x1792
are over the pixel budget — all four 400 at the provider with "Model does
not support request parameter value supplied: 'width' must be at least 768
pixels." Only 1024x1024 works. The bounds are now checked where the size is
mapped, so the error names the constraint instead of arriving from Azure.
width/height are deliberately left unchecked: they pass through unmapped, so
a future MAI model with different bounds stays reachable without a code
change.
Verified on a live Azure AI Foundry deployment of MAI-Image-2.5 and
MAI-Image-2.5-Flash (2026-08-17). One existing test asserted the 1792x1024
mapping; its size is changed to a size the provider accepts, keeping what it
was testing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI flagged the spec flaky twice more. Both were the same defect in different
places: the Logs drawer renders several nodes per string and the first in DOM
order is often hidden, so .first() waited 20s on an invisible element. The
entity assertions had a second problem on top, since getByText("EMAIL_ADDRESS")
substring-matched the masked prompt div, whose text contains <EMAIL_ADDRESS>,
rather than the entity chip.
Route every drawer assertion through onlyVisible, and match the entity type
and score exactly, which is what the panel renders them as: entity_type and
"Score: N.NN" each get their own span.
Verified on a live stack: 6 of 6 solo runs and the guardrails folder 5 of 5.
Mutating the analyzer to detect nothing turns the spec red on the raw address
reaching the spend log, so the assertions still carry their weight.
The Logs drawer renders the masked prompt in three nodes and the first one
in DOM order is hidden, so the previous commit's .first() traded a strict
mode violation for a locator that waits 20s on an invisible element. Local
runs against a warm stack failed on it every time, resolving the node 22
times and reporting "unexpected value hidden" each time.
Use onlyVisible, the helper this spec already uses for the playground
selectors, which filters to the visible node before taking the first. Solo
runs go 3 for 3 and the guardrails folder passes 5 of 5.
Also drop the 30s wait on the PII step added while chasing this: cold start
was never the cause, and it left the spec sitting on a dead locator longer.
The Logs drawer renders the masked prompt in three places, so matching it
without narrowing raised a strict mode violation instead of asserting
visibility. CI caught it as a flake: the spec failed its first attempt on
b0e53bfbe4 and passed on retry, which is a locator defect rather than a
timing one and would have gone red on any run that saw all three nodes.
Narrow to the first match, matching the guardrail-name assertion above it.
The leak checks below stay on toHaveCount(0), which is unaffected by
multiple matches and is what actually proves nothing raw reached the drawer.
An include entry that matches both a file next to the config that declares it and
one next to the root config now warns naming both, so a config that resolves to a
different file than it used to says so instead of quietly serving other models.
Also from reviewing that change:
- an empty root object in a bucket fails the boot again instead of coming up empty
- a YAML syntax error in a bucket object logs its own line naming the object
- an include already loaded is skipped before it is read rather than after
- reading a config out of GCS builds the plain bucket client, so it needs no
enterprise license and starts no flush loop that nothing ever cancels
The well-known index needs every stored upload repacked to publish its
digest, and both routes are unauthenticated, so each request was rebuilding
every archive on the event loop. With 21 stored skills the index took ~2s and
/health/liveliness on the same worker went from 1ms to 1.7s under two
concurrent index requests.
Repacking now runs off the event loop and each result is cached per skill
version, so a worker builds an archive once until the skill changes. The
archive route also declares application/zip in OpenAPI rather than JSON.
Both readiness loops broke out on success and fell through on timeout, so a
mock Presidio server that failed to bind left the run going with nothing
serving /analyze. The guardrail then errored at request time and the failure
surfaced as an unrelated Playwright assertion in presidioUserStory.spec.ts
rather than as the missing fixture it actually was.
Fail the local runner with the port in the message, and add the matching wait
step to both CircleCI UI jobs, which had no readiness check at all.
`include:` or `exclude:` written as `${{ ... }}` read back as a string, and
the sweep treated that as the directive being absent, so it expanded every
combination GitHub would have dropped. A job whose `name:` holds no matrix
value then looked like it repeated one name across combinations that never
run. An absent directive still means no rows; anything that is not a list
of rows now joins the names left out of the comparison
Keep reading an include left beside the root config, with a warning naming where it
was found, so a nested include written against the old rule still boots.
Also build one S3 client per config load rather than one per included object, treat an
empty included object as an empty config instead of failing the boot, and point the
error a dropped bucket include raises at the bucket error logged with it.
Adds an opt-in Agent Skills discovery index at
/.well-known/agent-skills/index.json (and the /.well-known/skills/index.json
alias) plus GET /v1/skills/{skill_id}/archive, so `npx skills add <proxy-url>`
installs skills uploaded through the Skills Gateway into any agent the CLI
supports.
The archive route repacks the stored upload so SKILL.md sits at the archive
root, with fixed entry timestamps so the SHA-256 digest published in the index
reproduces. Both routes are unauthenticated, since discovery clients send no
credentials, and stay 404 until an admin sets
`litellm_settings.public_skills_index: true`.
The prompt-management factory picks its deployment with a placeholder message. That
was inert while the pick ran on the synchronous path, which never runs the routing
plugin pipeline. Now that the pick runs the pipeline, a plugin classifying request
content would score the placeholder instead of the conversation, and the narrowing
it writes decides which deployments the real call may use.
The guardrail section of the release checklist is done by hand every cut:
create a Presidio guardrail through the wizard, send a sentence with PII from
the playground, then open Logs and check the guardrail caught it. Nothing
covered that path, so a break anywhere along it surfaced only when someone
happened to repeat the steps.
Adds a spec that walks it once and turns the eyeball checks into assertions.
The strongest of them is the leak check: it reads the request back and fails if
the stored prompt still carries the raw address or number, which is what the
manual step is really looking for.
The stack gains a Presidio stand-in that answers the two routes the guardrail
calls, detecting a fixed regex set with a Luhn check on card numbers. Real
Presidio's detection quality is Presidio's business, and pinning the UI lane to
it would mean two heavy containers with spaCy models on every CI run for a test
that is about LiteLLM's integration. The real analyzer stays covered in the
Python lane. The stand-in returns the same entities at the same spans as the
real one for the checklist's sentence, and driving the real guardrail against it
produces the same record shape, so a test written against it is written against
the product's real behavior.
Making the stand-in return the text unmasked turns the spec red on the raw
address reaching the spend log, so the leak assertion reads live data.
run_e2e.sh and both CircleCI UI jobs start the stand-in alongside the mock LLM.
Its port is overridable like the others so two checkouts can run at once.
A `name:` whose only leftover expressions read a `github.` property other
than `github.job` is filled in identically for every job of the run that
publishes it, so two jobs of one workflow carrying it land on the same
check run. Those names now compare against the other jobs of their own
file instead of sitting in the blind-spot bucket. They stay out of the
comparison across files, where two workflows can run on different events
Reading a config from a bucket ran a blocking boto3 GET straight from the
event loop for every object in the include tree, and on GCS it built a new
bucket client per object, each one starting a flush task that never ends.
S3 reads now go through a worker thread, and one bucket client serves the
whole include tree.
Two routers in one process shared a single handler, because the callback
manager dedupes on the class name plus the handler's public attributes and the
handler had none. The second router's requests were never counted. The handler
now carries the id of the cache it was built on, so routers with different
caches both register while the two selectors one router builds for its routing
groups still collapse into one.
Clamping a negative count back to zero used SET, which drops the key's TTL, so
the next write started the hour over. It uses INCRBY by the negative amount now,
which leaves the expiry alone.
The Lua script had no test that ran it, so tests/local_testing covers both the
sync and async paths against a real Redis, and the file is wired into the
CircleCI job that provides one.
An expression at `jobs.<id>.strategy` is legal on GitHub, but the model
required a mapping there, so a workflow using one made the whole file
unreadable and turned code-quality red. That job's names are now a blind
spot like any other name the sweep cannot work out offline.
A matrix whose `name:` holds no matrix value publishes that one name once
per combination, which leaves a required context just as ambiguous as two
jobs sharing a name, so it now reports instead of deduping.
A file that does not parse as one YAML document is reported the way the
module already promised, rather than escaping as a traceback.
Three ways the sweep could fail a workflow GitHub would publish fine.
`github.workflow` and `github.job` were counted as fixed for the whole run, so
two jobs naming themselves after the workflow they sit in were reported as a
collision. `runner` and `vars` were wrong the same way. Drop the exception
entirely: a name still holding an expression is one GitHub resolves per job, so
it is nothing to compare, which is what the rest of the module already does.
`format()` was resolved with Python's semantics, so an attribute lookup crashed
the script and a width specifier padded a name GitHub never pads. Fill `{0}`
holes and escaped braces, and treat anything richer as unresolved.
A matrix `include` or `exclude` row holding a value that is not a scalar lost
that key and became an empty row, which excludes every combination. Report the
row instead of quietly reshaping the matrix around it.
A job name holding an expression the sweep could not resolve was compared as
if it were the published name. Two jobs whose names differ per matrix value or
per caller input were reported as a collision, and a matrix that was itself an
expression collapsed onto the bare job id and did the same.
Model what a job publishes as known names beside the reasons the rest stay
unknown. Anything the sweep cannot work out contributes no name and is
reported as a note instead of guessed at. An expression over contexts that are
fixed for the whole run still compares, so two jobs sharing one of those are
still caught.
A config loaded from a GCS or S3 bucket skipped include processing entirely,
so every model, guardrail, and setting behind an `include` was silently
dropped. Both bucket types shared the same branch in `get_config`, which
never called `_process_includes`, and that helper only ever read from disk.
The merge now lives in one async helper that takes the loader as a
dependency, so disk and bucket configs share the same semantics: list values
extend, everything else overrides, nested includes are followed, and the
`include` key is stripped. Bucket entries resolve as object keys relative to
the config object's prefix, with a leading `/` meaning the bucket root, and
an include that cannot be read now raises instead of being skipped.
The two writes that hand the skip list to the next attempt now carry a
`# rebind-ok` reason, which is the sanctioned escape hatch for an unavoidable
parameter mutation and matches how `log_retry` already writes into the same
kwargs dict a few lines above
`get_excluded_filtered_deployments`'s docstring said returning the unfiltered
list would re-include the deployment that just failed. The retry skip does
exactly that on purpose, so the docstring now says each caller decides what an
empty result means
The reliability registry cell the new e2e test claims is marked
`fail_before_fix: proven`: the same config returns 400 at the merge base and
200 off a sibling deployment at the tip
The presidio suite proved masking happened by reading the served answer, and the
UI suite proved the wizard's "Select All & Mask" produced a row. Neither checked
the thing an operator actually looks at afterwards: the audit trail.
Adds an e2e test asserting the spend log carries the pre_call guardrail record
for a masked request: status, provider, per-entity masked counts, and the
detected-entity list the dashboard's guardrail panel renders its scores from.
It keys off the x-litellm-applied-guardrails response header rather than masked
text in the answer, because whether the model echoes the prompt back is a model
decision, not a guardrail one. Both a mutation that stops writing the record and
one that empties the entity list turn it red.
Two records land on one log, pre_call and post_call, so the assertion selects on
mode as well as name; picking by name alone could hand it the empty post_call
record depending on write order.
On the UI side, the Presidio wizard test now reads the stored guardrail back and
asserts every persisted entity carries the MASK action. A row appearing in the
table did not prove the entity selection survived the save, so a wizard that
persisted an empty pii_entities_config would have passed.
SpendLogRow gains a typed metadata field. guardrail_response is left as object
because each provider writes its own shape there (presidio a list of entities,
bedrock an assessment object, a failed run the exception string); a union narrow
enough to be useful would fail to parse the others and break every suite that
reads a spend log. The caller validates the shape it expects with a TypeAdapter.
run_e2e.sh now pins PROXY_BASE_URL to the stack's own origin. The proxy builds
its post-login redirect from that variable when it is set, so a value inherited
from a developer's .env sent the browser off the relocated stack and the suite's
login step timed out on every port but 4000.
A Redis outage read as "every deployment is idle", because batch_get_cache
swallows the failure and answers with an empty dict. batch_get_counts and its
async twin raise instead, so a worker that cannot reach Redis falls back to its
own numbers rather than routing on zeros.
The counter's TTL is now set only on a key that has none, so a +1 left behind by
a worker that died mid-request ages out an hour after the key was created. It
used to be refreshed on every touch, which kept that stuck count alive for as
long as the group took traffic.
Two least-busy groups counted the same request twice, since the pre-call list
kept a selector per group while the success list deduped by class. The selector
now goes on through add_litellm_input_callback, which dedupes the same way.
A prompt-management model picked its deployment on the synchronous path, so the
new Redis read landed on the event loop and configured routing plugins never
ran. It awaits the async selector now.
The router code coverage gate reads every function defined in router.py
and fails when no test file names it. _as_retry_skipped_deployment_ids
was only reached indirectly through the retry path, so the gate went red
on this PR's tip.
Test it directly instead: a tuple of strings survives, non-string items
inside the tuple are dropped, and every other shape a caller could send
narrows to an empty skip list.
The collision sweep read a job's name as its `name:` or bare job id, which is
wrong for a matrix job that sets no name: GitHub publishes `build (3.12)`, one
per combination. That missed real duplicates and invented ones that don't exist.
It also crossed every matrix value while ignoring `exclude`, so it checked
combinations no job ever runs.
Four smaller gaps went with it. Boolean matrix values reached a name as `True`
rather than `true`. A `format()` whose arguments cannot fill its placeholders
raised straight out of the script instead of leaving the name unresolved. A job
calling a reusable workflow only ever chained one level, and a call outside the
repo fell back to the caller's own name, which GitHub never posts. A job whose
`name:` was not a string failed validation and silently dropped every job in
that file, so the sweep now renders any scalar and reports a file it cannot read
instead of skipping it.
SPEND_LOGS_URL only diverts spend logs when db_writer_client is set, and nothing in the proxy ever assigns that global, so the queued copy was only ever skipped as a duplicate by the local insert.
The retry skip travels as a request kwarg, and the router forwards keys it
does not recognize, so a client can put _retry_skipped_deployment_ids in its
own request body. The value went straight into a pydantic TypeAdapter and
then into a set(), so an int or an object raised TypeError and a string, a
list, or a dict raised a ValidationError, each of them replacing the 400 the
provider had actually returned.
Every read now goes through one narrowing function that keeps a tuple of
strings and skips nothing otherwise, so a forged value costs the caller
nothing beyond the retry landing on the same deployment again.
disable_spend_logs has to keep meaning that no request gets logged, and the row
that makes a batch chargeable exactly once is the one row it cannot drop, so with
logging off that row now carries only what tells the retrieves apart. SPEND_LOGS_URL
deployments get their copy back too: the claim writes straight to this table, so the
row is queued as well when an external writer is the one that takes the spend logs.
A fake-streamed provider hands the adapter one chunk carrying both the
delta payload and the finish_reason, which is exactly what the combined
chunk splitter exists for, but its content check never listed the refusal.
The translation short-circuits on finish_reason, so that refusal text was
dropped and the client got `stop_reason: refusal` over an empty content
array, the symptom this PR set out to fix.
Both refusal accumulators also drop their `mutable-ok` lists for a plain
string attribute
Before excluding the deployment that just refused, the retry-skip guard asked
whether another one could still answer. It asked by re-running a single routing
filter, the order filter, while deployment selection also applies cooldowns, the
context-window pre-call check, tag routing, and routing plugins.
Any filter the guard did not replicate made it answer yes while the real pick was
left with nothing. A group narrowed to one deployment by tag routing turned the
provider's own 400 into a no-deployments 429.
The skip now runs where every filter has already been applied, and it keeps the
deployments untouched when skipping would leave none. The caller gets the
provider's error either way, and a group with one eligible deployment retries in
place as it did before.
The first-delta guard read `delta.refusal` directly, while the translation
three lines later goes through `openai_chat_refusal_text`, which also reads
the `provider_specific_fields` LiteLLM parks unrecognized fields in. A
provider that sends the refusal that way had its only refusal delta skipped
as blank, so the client got `stop_reason: refusal` over an empty content
array, which is the symptom this PR set out to fix
The /v1/messages adapter lowers a tier the entry does not accept, so dropping max from the astra
rows moves that path from Foundry's 400 to a request at xhigh. Nothing pinned that, and the guard
test's docstring named gpt-6-astra as the only gpt-5 name with an azure_ai row, which 11 rows
contradict.
The checker raised a custom exception and caught it two lines down in the
same module, which is the throw-then-catch the repo's coding guide rules
out. `main` now prints the same message and returns the exit code, so the
collision list stays a value the whole way out
The retry-skip guard checks that some other deployment could still answer
before it excludes the one that just refused, so a single-deployment group
keeps the old retry-in-place behavior. It asked that question at the group's
minimum order, but the router picks the retry's deployment at the order the
request has already escalated to.
So a group with a primary at order 1 and a backup at order 2 answered "yes,
order 1 still has a candidate" while the retry was pinned to order 2, and the
exclusion left order 2 with nothing. The caller got a no-deployments error in
place of the provider's own 400.
The helper now takes the active target order and filters by it, which is the
same value async_get_healthy_deployments reads off the request.
The guard read each matrix key's values independently and crossed them, so
a job name reading two keys off one include row published pairs no job ever
runs, which could fail a valid workflow on a required check
It now builds the combinations GitHub builds: the listed keys crossed, each
include row folded into the combinations it overwrites nothing in, and a row
that fits nowhere standing on its own
`vertex_ai/lyria-3-clip-preview` and `vertex_ai/lyria-3-pro-preview` were
registered with `supports_vision`, `supports_image_input`, and an `image`
modality, which contradicts their `gemini/lyria-3-*` siblings and makes
/model/info advertise image input on text-to-music models.
The new Lyria passthrough branch runs before the image-generation branch
and keys on the same `predictions[0].bytesBase64Encoded` shape imagen
returns, so only the cost-map lookup separates them. Cover an imagen
predict response end to end so a future change that drops that lookup
fails here instead of misbilling images as audio.
The takeover of a $0 row an older proxy left behind used to charge the batch when
the update could not reach the database. That leaves the row still reading $0, so
every later retrieve finds the same row and charges the batch again, which is the
repeat charging this PR exists to stop. The retrieve that does take the row over
is the one that charges, and a batch nobody retrieves again after that failure is
never charged, the same as one whose proxy died inside the write window.