* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models
GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.
* fix(proxy): offer a Codex service tier only when every deployment of the model lists it
* fix(codex-catalog): an invalid service_tiers value offers no tier for the model
* fix(codex-catalog): read service tiers off the deployments the key's team can route to
A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them
The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default
* test(codex-catalog): drop the redundant module docstring and sort the imports
* test(integration): add the Codex catalog audit cells and the multi-worker convergence note
* test(integration): clean up every catalog test model and answer the refresh GET
* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers
Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.
* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut
The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns
The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(proxy): add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): register eager lazy routes at startup so late eager routes keep precedence
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): share one optional-feature install path between lazy and eager registration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(proxy): say eager lazy routes register at worker startup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): name the import callable passed to _install
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): prove a startup hook can drop eager lazy routes for good
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): strip the lazy-routes flag from the lazy-mode control proxies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): cover the lazy warmup route registering a feature
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(unit): run tests/unit/proxy/test__lazy_features.py in the proxy-server-core shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for LITELLM_DISABLE_LAZY_ROUTES
Flag spellings, /openapi.json at boot, the warm-up route in both modes, /mcp/proxy ahead of the /mcp mount,
route table and OpenAPI parity with a fully warmed lazy proxy, a broken optional import in both modes, every
client SDK against the completion endpoints, a boot burst with a killed worker, and a restart
* fix(proxy): keep config pass-through routes ahead of eagerly registered features
With LITELLM_DISABLE_LAZY_ROUTES set, features registered before the proxy lifespan
added config pass-through endpoints, so a pass-through overlapping a feature path
(e.g. a self-hosted /langfuse) lost to the built-in route. Restore lazy mode's
registry order once startup finishes, without bringing back routes a startup hook removed
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): edge-case matrices for malformed token limits and callback_settings shapes
Extends the integration suite so two classes of issues found by gauntlet
reviews are caught end to end against a real proxy:
- non-numeric or odd model_info token limits (from /model/new and from
config YAML) must be listed as absent on /v1/models, /models,
/v1/models/{id} and /model/info, keep sibling models listed, and still
serve chat
- every callback_settings shape (top level and per consumer) must let
the proxy boot, register the configured callbacks and serve chat
Four product bugs on main surfaced by the matrices are recorded as
BUG skips per the suite convention: chat 500 and /model_group/info 500
on non-numeric token limits, a startup crash on a non-object
callback_settings, and otel silently dropped on a non-object
callback_settings.otel
* test(integration): pin the exact coerced value for numeric-edge token limits
Addresses review feedback: the numeric-edge matrix only asserted
'int or absent'. It now asserts the listed value for each case on
/v1/models, /models and /v1/models/{id}, which also lets the listing
helper drop its optional-expectation branch.
tests/integration/conftest.py imported the bare `integration` package. Because
tests/__init__.py and tests/integration/__init__.py both exist, pytest's default
prepend import mode puts only the repo root on sys.path, so that name resolved
only under the PYTHONPATH that tests/integration/run.py injects. Every other
invocation died at conftest import with
ModuleNotFoundError: No module named 'integration' and exit 4, including the
command test_oci_integration.py documents in its own docstring.
The imports now use the tests.integration._support path that pytest actually
resolves, matching the 120 other `from tests.` imports in the suite. run.py's
PYTHONPATH still works because it already puts the repo root on the path.
tests/code_coverage_tests/test_integration_suite_imports.py collects every file
under tests/integration with PYTHONPATH scrubbed and asserts a non-zero
collection count, so an unresolvable import fails the code-quality job instead
of only the developers who run these files by hand. CI runs the three
pre-existing files through the allowlist rather than executing them, which is
why nothing caught this.