Codex hides catalog entries whose supported_in_api is false when it runs
with an API key, so a stock entry the proxy serves now carries
supported_in_api true alongside its list visibility.
`lite codex` now asks the installed Codex for its own model list through
`codex debug models` before writing the catalog. A proxy model whose id
matches a stock Codex slug keeps that Codex's entry (reasoning levels,
base instructions, context window and the rest) and only its picker
position, visibility and upgrade nudge come from the proxy. Unknown slugs
still get the plain entry built from the bundled base instructions. The
catalog directory is created before the stock call so a fresh CODEX_HOME
does not make Codex refuse to run
Drop the unrequested env override on LOGGING_WORKER_TIMEOUT_SUMMARY_WINDOW_SECONDS.
The constructor still accepts a value for tests, so runtime behavior is unchanged.
Adds general_settings.custom_key_policy, one coroutine that receives the
operation ("generate", "update", "regenerate"), the existing key row, the
effective row as it will be written, and the raw request, and can deny with a
403. It runs after the request has been normalized and before the first DB
write on /key/generate, /key/service-account/generate, /key/update,
/key/bulk_update, /team/key/bulk_update and /key/{key}/regenerate. The two
legacy hooks keep running unchanged on the raw request.
* fix(proxy): attribute gate-rejected requests to their endpoint in cache analytics
Requests rejected before dispatch (bad key, blocked key, budget, rate limit, malformed body) were spend-logged with an empty call_type because the synthesized logging object never reached the failure lifter. The caching dashboard rolled all of them, plus failed calls on info routes such as /model/info, into one Unknown group.
Resolve call_type from the matched route first, falling back to body shape, and keep the synthesized logging object on request_data so the lifter sees it. Log bare auth exceptions with the 401 ProxyException the client gets so error_code is never empty. Exclude info routes from the cache analytics groups and error breakdown. The dashboard explains the Unknown group when older rows still produce one.
Resolves LIT-5884
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep the raw auth exception for failure callbacks
Record the client-facing status in the spend log through a separate client_exception argument so custom failure callbacks still receive the exception auth raised.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep the route for multi-operation endpoints and exclude info routes from cache filter options
Routes such as /v1/files map to several operations (create, list) and the
method is not available in the failure hook, so a rejected request there is
filed under its route instead of the first mapped call type. The key alias and
model filter-option queries now apply the same info-route exclusion as the
groups and error breakdown, so every offered filter value returns data. The
info-route exclusion and Unknown grouping are now covered against a real
Postgres in tests/proxy_behavior/spend/test_cache_activity.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): drop client_exception, the spend log row never used it
The DB spend row for a gate rejection is written by _ProxyDBLogger from the
original exception, so the status-bearing copy only reached the in-memory
logging payload. Live runs at the tip still recorded bare auth exceptions as
Unknown/Exception, the same as the base branch. Removing the plumbing keeps
this PR to endpoint attribution and the info-route exclusion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Codex 0.130 and 0.145 require supports_reasoning_summaries and
supports_parallel_tool_calls on every catalog entry, so a catalog written
for 0.154 made those releases exit at startup with a parse error. Every
field some release since 0.105.0 deserializes without a default is now
written, with Codex's own fallback values, and the catalog is read back
once through the installed binary (`codex debug models`) before launch.
A Codex that rejects it, or one older than 0.130 with no such command,
gets the skip notice and launches on its built-in catalog instead.
The provider_config path skipped run_realtime_guardrails for transcription
sessions to avoid sending response.create, which also dropped every
realtime_input_transcription guardrail: no violation error reached the
client and on_violation / end_session_after_n_fails never fired. Run the
guardrail for every completed transcript and only suppress response.create
when the session has no assistant turn.
Every proxy and router cost lookup goes through get_model_info, which copies
cost map keys explicitly, so the new audio cache-read branch always fell back
to the text cache-read rate there. Copy the key so models whose audio
cache-read rate differs from the text one bill cached audio correctly.
When a slow backend times out many logging callbacks at once, each timeout
hit verbose_logger.exception in _process_log_task and produced a full ERROR
traceback, clustered milliseconds apart. Count timeouts instead and arm a
debounced flush that logs one WARNING with the burst count, callback name,
timeout and cumulative total. Real programming errors keep their traceback.
Muse partials carry no turnId and belong to the most recent speechStart,
and the docs say the model may keep post processing a turn after speechEnd
until speechComplete. Releasing the active turn on speechEnd made any
partial arriving in that window raise and get dropped in ENDPOINTING mode.
The turn now stays active until its speechComplete or final transcript.