A stored null or empty list for a nested alerting field is still the value
the proxy serves when the config file leaves alerting_args alone, so the
source is db. Keying off the value rather than its presence reported those
fields as default and hid a stored setting that is genuinely in effect.
Presence in the stored row now decides, with the config file still checked
first so a config-owned key keeps reporting config. Test helpers are typed
and the router test injects a stub rather than patching a class attribute.
The /v1/responses and /v1/messages streaming wrappers only ever wrapped the
primary's stream, so a hop reached through the regular fallback chain had no
mid-stream handler: its failure re-raised, or the outer wrapper retried the
same entry with a fresh attempted set and never reached the rest of the list.
Every attempt of the chain now runs through a per-endpoint attempt function
that wraps its own stream, mirroring chat completions, and the per-request
fallback and retry overrides ride a frozen carrier so each hop's re-entry
still sees them after the retry layer pops them.
Keep sigv4 signing, eventstream decoding, smithy runtime api and types at opt-level 3 since they serve Bedrock request and streaming hot paths, and make the S3 facade binding reject region and endpoint mismatches between the projected configuration and the native handle
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The docs quickstart pipes a compose file hosted on the docs site straight
into `docker compose -f -`. That puts the content users execute in the docs
repo rather than here, and nothing lands on disk for them to read first.
Move the two-service stack (gateway + Postgres) into docker/ so it ships and
is reviewed alongside the code it starts, and pin the image to main-stable
instead of latest. The docs change to download-then-run follows separately.
Verified: `docker compose up -d` brings the stack healthy, /health/liveliness
returns 200, /v1/models returns 200 with the placeholder key and 401 without,
and /ui/ serves.
The stream chunk builder started its per-chunk accumulators at 0 and adopted only nonzero counts, then fell back to litellm's tokenizer whenever the accumulated value was falsy, so a provider that reported an explicit 0 for prompt or completion tokens was billed the estimate instead. The accumulators now start at None, a usage chunk that reports a count marks it reported (a later chunk's 0 never replaces a reported nonzero), and the estimate only runs when no chunk reported the count. The Anthropic message_start cursor reset now yields None so the estimate still covers a cancelled stream, and Ollama chat streaming only attaches usage on the done chunk when both counts are present instead of inventing 0/0 on every chunk
Optimize the aws-sdk dependency tree and cache-s3 for size in the release profile and switch to fat LTO so the native extension stays under the wheel verification gate (27.87 -> 21.88 MB)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>