Replicas re-read cooldowns from Redis at most every 10s
(default_redis_batch_cache_expiry), so the 12s window left 2s of slack;
it is now 15s and the benched phase runs from 15s to 26s after the trip.
The reliability rows the new cells cover claimed exercised_on messages
too, but every cell drives /v1/chat/completions, so they now claim
chat_completions only. RouterSettingsOverride.timeout and
RouterCurrentValues.routing_strategy had no reader and are gone.
GET /router/settings is a management route, so the split transport now
sends it to the control plane instead of the data-plane gateway.
The rpm-1 key behind the 429 cells is spent right before the trip, after
the pair is registered, because the rate limiter's 60s window opens on
that request and the registrations' propagation waits could otherwise
outlast it. Recovery also accepts a 200 served by the benched deployment
itself, since its key's minute can be up by then.
The cooldown recovery deadline now counts from the last failure a stale
replica caused during propagation, because every failure re-arms the
cooldown TTL; the strict bench window stays anchored to the trip.