Greptile 4/5 follow-up on the envelope codec:
- ``_decode_item_envelope("msg_")`` previously returned the truthy string
``"resp_"``, which leaked through truthy guards in the item_reference
resolver and propagated an empty response_id to the session handler.
The decoder now returns ``None`` when the payload after the prefix is
empty.
- ``_envelope_encode_output_item_ids`` reused the same encoded id for
every ``message`` item in a response, so parallel ``n>1`` choices
collapsed to identical ids. The encoder now takes an ``item_position``
suffix (``.{n}``) that disambiguates each item; the decoder strips
the suffix before handing the inner payload to the existing
response-id decoder, so the round-trip still returns the same
``response_id``.
- 2 new tests: empty-payload decode returns ``None``, multiple message
items receive distinct ids that all decode to the same response_id.
When the Responses API bridge fronts a chat-completions provider, message
output items inherit the upstream ``chatcmpl-*`` id directly. Clients that
follow the OpenAI Responses spec expect typed prefixes (``msg_*``, ``rs_*``)
on output item ids and break when they receive a raw ``chatcmpl-*`` id:
- Vercel AI SDK 6.x fails with "text part {id} not found" on multi-step
tool calls (issue #26529)
- Streaming bridge surfaces unregistered ``chatcmpl-`` ids in text-delta
events (issue #27671)
- Bridge responses can replay ``chatcmpl-*`` ids back into OpenAI on
cross-provider handoffs (issue #27333)
This adds an envelope codec on ``ResponsesAPIRequestUtils`` that wraps raw
upstream ids as ``msg_<base64-payload>`` using the same payload format as
the existing ``_build_responses_api_response_id`` envelope:
- ``_encode_item_envelope(raw_response_id, prefix, custom_llm_provider,
model_id)`` - reuses ``_build_responses_api_response_id`` and swaps
``resp_`` for the requested item prefix (``msg``/``rs``).
- ``_decode_item_envelope(item_id)`` - returns the ``resp_<env>`` form so
existing ``_decode_responses_api_response_id`` machinery handles the
inner payload validation.
- ``_envelope_encode_output_item_ids`` walks the response output and
rewrites only message items whose id is raw (i.e. not already prefixed
with ``msg_``, ``rs_``, or ``encitem_``). Function-call items and
already-prefixed items are left untouched.
Hooked into ``_update_responses_api_response_id_with_model_id`` so the same
post-processor that wraps ``response.id`` now also wraps message item ids,
giving downstream clients consistent typed envelopes across both fields.
- 12 unit tests covering round-trip encode/decode, malformed input
handling, selective rewriting (skip already-prefixed, skip
function_call items), empty output handling, and the full integration
via ``_update_responses_api_response_id_with_model_id``.
Add `if: github.repository == 'BerriAI/litellm'` guard to scheduled
jobs in stale.yml, codeql.yml, and create_daily_staging_branch.yml.
This matches the existing pattern in auto_update_price_and_context_window.yml
and prevents these workflows from running unnecessarily on fork repositories.
Verify that spend_logs_metadata is correctly merged into combined_metadata
and flows through to Prometheus custom labels. Tests cover: basic extraction,
precedence when keys overlap, all three metadata sources combined, and None
handling.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add spend_logs_metadata to combined_metadata in Prometheus logger so
custom metadata from x-litellm-spend-logs-metadata header can be used
in Prometheus custom labels.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Thread project_alias alongside project_id through the metadata pipeline so
callbacks receive the human-readable project name. DRY up duplicate metadata
dict construction in proxy_track_cost_callback and pass_through_endpoints by
reusing get_sanitized_user_information_from_key — future metadata fields only
need adding in one place.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Two independent bugs prevented post-call OpenAI Moderation guardrail
results from reaching downstream logging callbacks (Langfuse, Datadog).
Bug 1: process_output_response() created a throwaway request_data dict,
so guardrail info written by @log_guardrail_information was discarded.
Fixed by threading the real request_data from the unified guardrail
dispatcher through all 13 BaseTranslation handlers, with litellm_metadata
injection preserved for third-party guardrails (Zscaler, Prompt Security).
Also extended to process_output_streaming_response for consistency.
Bug 2: The @log_guardrail_information decorator collapsed the full
moderation API response (categories, scores, flagged status) to "allow".
Fixed by overriding _process_response/_process_error on
OpenAIModerationGuardrail to stash and log the full response, following
the established Model Armor pattern.
- Use Select.Option with font-medium alias + Text secondary ID to match OrganizationDropdown
- Default page size to 20
- Add useInfiniteTeams mock to AddModelForm tests
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- OldTeams: refresh table via fetchTeamsV2 after team create instead of appending
- TeamDropdown: rewrite with useInfiniteTeams for paginated fetch, scroll-to-load, and debounced search
- Update all TeamDropdown consumers to use the new self-fetching API
- Dashboard layout: switch from Sidebar2 to SidebarProvider (leftnav)
- Leftnav: add MIGRATED_PAGES routing for path-based navigation (api-reference)
- Navbar: remove chat button
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New docs page covering the HA control plane architecture where each
worker instance has its own DB, Redis, and master key. Includes a
React component diagram, setup configs, SSO notes, and local testing
instructions.