Vertex AI Mistral models reused MistralConfig, whose reasoning_effort
advertisement checks the mistral provider entry of the cost map, so
vertex_ai/mistral-medium-3 started advertising reasoning_effort and
drop_params stopped dropping it, turning a 200 into a Vertex 400.
VertexAIMistralConfig scopes that lookup to the vertex_ai provider, and
MistralConfig now reads the provider from its custom_llm_provider
property instead of a hardcoded "mistral".
POST /config/update compared litellm_settings after lowercasing the
callback list, so a config file spelling a callback in mixed case refused
the same list sent back, and it stored every general_settings model default
next to the keys the request set. Both now use the request as sent; only
the stored callback list is lowercased.
Also drops config_data from the router settings reload callers the previous
commit left behind and teaches the legacy MockProxyConfig the ownership
check.
A merge on this branch dropped the ^ anchor and MULTILINE flag from the
doc_key_pattern, so the unanchored match started swallowing key names
into captured fields whenever a row's description cell itself contains
a pipe (Literal unions and similar). The check then reported 27
long-documented keys as undocumented. Restores the exact pattern used
on main, verified against a live litellm-docs checkout: all 58 Router
init params resolve as documented.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
GET /model/info?litellm_model_id= went through _get_proxy_model_info, a copy of _enrich_model_info_with_litellm_data that missed the unpriced check, so the same deployment read as null in the list and as 0 on the id lookup. The id lookup now delegates to the shared helper, so the two paths cannot drift again
Cancelling every in-flight job the moment shutdown reached the scheduler
dropped the rows a write job had already popped: flush_gateway_requests
drains its accumulator before committing and does not restore it on
CancelledError, and update_spend requeues its batch only after the
shutdown drain had already run.
Shutdown now waits up to JOB_FINISH_TIMEOUT_SECONDS for in-flight jobs
to finish on their own, cancels the ones still running, and does both
before the shutdown flushes so a requeued batch is still written. The
cleanup run never finishes inside the grace, so it is still cancelled
and still records outcome="aborted".
Resolves LIT-6990
Review follow-ups on #41213:
- Pause the scheduler as the first shutdown step so a job whose fire time
falls inside the shutdown window does not start only to be cancelled.
Jobs already running keep the whole window and are cancelled and
awaited before the database disconnects, as before.
- Keep the cleanup run's progress in a task-scoped ContextVar rather than
on the cleaner instance, so two runs overlapping on one cleaner
(APSCHEDULER_MAX_INSTANCES above 1 without a Redis lock) each report
their own rows and batches on cancellation.
- Drop the module docstrings the repository comment policy does not
allow; the rationale lives in the PR description.
cleanup_old_spend_logs only caught Exception, so a run cut short by
CancelledError recorded no outcome and logged nothing. Under uvicorn the
job was never cancelled at all: uvicorn re-raises the captured SIGTERM as
soon as the lifespan shutdown returns, before asyncio cancels outstanding
tasks, so an in-flight scheduler job simply died with the process.
The cleanup now handles CancelledError by logging elapsed time, rows
deleted and batch count at error level, recording outcome="aborted", and
re-raising. The lifespan shutdown stops the scheduler and awaits the jobs
it cancels while the database is still connected, so that handler runs
under uvicorn too, and the pod lock is released instead of orphaned.
Resolves LIT-6990
MutableMapping.clear pops items until the mapping is empty, but __delitem__
keeps config-owned keys, so clear spun forever on any store that had loaded a
config file. unittest.mock.patch.dict calls clear on exit, which is why the
proxy-infra and proxy-endpoints shards hung at 99 percent until the 20 minute
job timeout on every run since the store started refusing config-owned writes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The two pass-through reload tests register routes on the shared app and
never take them off, so a later allowlist test in the same xdist worker
finds routes that no component exposes. Restore the app's route list
when each of those tests finishes.
The two pass-through reload tests register real FastAPI routes on the
shared proxy app and only clean up the internal registry, so any test
running after them in the same worker sees stray /v1/kept-* and
/v1/deleted-* routes. test_component_allowlists counts those as
uncovered and fails. Snapshot app.routes and the registry up front and
restore both in a finally block.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
POST /config/update stored keys the config file owns and answered 200
while the file silently kept winning. Run the config-owned check for
general_settings, litellm_settings, and router_settings before the first
database read, the way /config/field/update already does, so a refused
request stores nothing.
The config reload re-loaded the merged settings as yaml settings, which
turned every saved router setting read-only after one tick. Read the
saved router settings row instead so database-owned values stay writable.
The integration harness seeds num_retries through /config/update instead
of the config file, which is what the effective-settings and observed
routing tests need to keep exercising a database-owned value.
Resolve dot segments in the /azure_speech endpoint path before the endpoint family and the admin-only batch guard are decided, so the guard and the forwarded upstream path agree. Bill short-audio requests for the longer of the uploaded audio duration and the recognized duration, so a NoMatch or silence response still charges for the audio Azure processed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Declare the live-verified reasoning_effort_levels on the Mistral cost-map
entries and round an undeclared request to the nearest declared level
(up to the weakest level at least as strong, down to the strongest when
the request exceeds the ceiling). Codex's default medium no longer 400s
on mistral-medium-latest, mistral-small-latest, or the vibe-cli family;
an entry that declares nothing keeps forwarding the value verbatim