- Updated MidStreamFallbackError to retrieve and maintain the original status code from the wrapped exception.
- Ensured that message, request, and response fields remain consistent after calling the parent constructor.
- Added unit tests to verify the correct propagation of status codes and attributes in various scenarios.
- compute image_generation cost from usage token metadata for vertex/gemini\n- map ImageUsage to Usage and reuse generic_cost_per_token\n- fallback to output_cost_per_image when usage metadata missing\n- add tests for token-based path and fallback path
A database timeout (httpcore.ReadTimeout) during get_config() in
_update_llm_router would propagate and prevent ALL DB models from
loading into the router.
Now get_config() failures are caught separately so model add/delete
operations still proceed. Similarly, _delete_deployment catches
get_config failures and safely skips cleanup rather than crashing the
entire sync cycle.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- add gemini-3.1-flash-image-preview + vertex_ai alias entries\n- set pricing to Gemini 3.1 Flash Image Preview rates\n- mirror updates in packaged backup model map\n- update llm cost calc regression test to cover new model
When applying a key_alias filter on the Virtual Keys page, the pagination display (total count, page count) showed stale unfiltered values. The bug was that the table component tracked total_count only from the unfiltered useKeys hook, not from the filtered API response returned by useFilterLogic.
Fixes: When selecting a key alias that matches 1 key, the UI previously showed "Showing 1-50 of 509 results" and "Page 1 of 11" instead of the correct "Showing 1-1 of 1 results" and "Page 1 of 1".
Changes:
- Added filteredTotalCount state to useFilterLogic to track the total_count from filtered API responses
- Updated VirtualKeysTable to use filteredTotalCount (when set) instead of always using the unfiltered total
- Added comprehensive tests to prevent regression of pagination display logic
Type: 🐛 Bug Fix, ✅ Test
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
- Move _is_claude_4_6_model from AnthropicConfig to AnthropicModelInfo
to eliminate duplicated logic in is_effort_used
- Use explicit effort_map dict instead of passing unknown values through
to output_config
Claude 4.6 models use output_config as a stable API feature. This commit:
- Maps reasoning_effort to output_config for 4.6 models (minimal → low)
- Restricts effort="max" to Opus 4.6 only
- Skips beta header injection for 4.6 models
- Updates docs for Claude 4.6 effort support
Phase 2 (per-worker mark_process_dead on shutdown) only ever fired
when all workers shut down together, making it redundant — Phase 1
wipes everything on next startup anyway. This aligns with the
prometheus_client docs: just wipe the directory between runs.
LiteLLM was adding a `duration` field to audio transcription responses
for internal cost tracking. The OpenAI Python SDK uses "best match
deserialization" to determine the response type from present fields —
seeing `duration` caused it to incorrectly match plain Transcription
responses as TranscriptionVerbose/TranscriptionDiarized types.
Move the internally-calculated duration to `_hidden_params` so it
remains available for cost calculation without polluting the response
body. Provider-returned duration (e.g. from verbose_json format) is
still preserved in the response as expected.
Adds TestProxyMcpStatelessBehavior to test_proxy_mcp_e2e.py with a test
that verifies two independent MCP clients can connect, initialize, and
call tools without sharing session state. This catches the regression
from PR #19809 where stateless=False broke clients that don't manage
mcp-session-id headers.
Regression test for #20242
Address Greptile review: apply type: ignore[union-attr] consistently
on all backend_ws.send(), .recv(), and .close() calls, not just the
three that CI flagged.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
CI MyPy resolves CLIENT_CONNECTION_CLASS as Optional[ClientConnection]
and flags .send() and .close() as attr-defined errors. These methods
exist at runtime on the websocket connection object. Add type: ignore
comments to unblock the linting CI.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The tests were mocking `filter_server_ids_by_ip` but the production
code in server.py now calls `filter_server_ids_by_ip_with_info` which
returns a (server_ids, blocked_count) tuple. Update all 8 mock sites
to use the correct method name and return signature.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Step-level env is not visible to the if condition — reference
secrets directly so ggshield actually runs when the key is configured.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Address github-advanced-security bot review comment by setting explicit
minimal permissions (contents: read) for the GITHUB_TOKEN.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add unit test that scans Python source for Base64 Basic Auth patterns
that would be flagged by secret scanners like GitGuardian/ggshield
- Add secret-scan job to the linting CI workflow that runs the test on
every PR and optionally runs ggshield if GITGUARDIAN_API_KEY is set
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>