litellm/litellm/responses
devin-ai-integration[bot] f66b3ebe0d
feat(responses): honor supported_endpoints /v1/responses opt-in for OpenAI-compatible deployments (#39725)
* feat(responses): honor supported_endpoints /v1/responses opt-in for OpenAI-compatible deployments

custom_openai and other generic OpenAI-compatible deployments have no native
Responses API config, so every /v1/responses call is bridged through
/v1/chat/completions. When model_info.supported_endpoints lists /v1/responses,
resolve OpenAILikeResponsesConfig instead so the request is forwarded to
{api_base}/responses, for streaming, non-streaming and mode: responses
deployments alike. Providers with their own Responses config are unchanged.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): drop deployment supported_endpoints opt-in after cross-provider prompt swap

A prompt manager that moves the request to another provider leaves kwargs['model_info']
describing the original deployment; without this the swapped provider was sent an
OpenAI-like /responses request it does not serve.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(responses): carry prompt-swap deployment metadata as a return value instead of a kwargs marker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-05 11:40:00 -07:00
..
file_search refactor(repositories): type prisma table access with one generic protocol 2026-08-25 12:14:17 +00:00
litellm_completion_transformation merge: litellm_internal_staging into litellm_headroom_ccr_streaming_responses 2026-09-03 00:24:17 +00:00
mcp fix(responses/mcp): keep reasoning order and caller previous_response_id on stateless follow-ups 2026-09-03 13:27:48 -07:00
main.py feat(responses): honor supported_endpoints /v1/responses opt-in for OpenAI-compatible deployments (#39725) 2026-09-05 11:40:00 -07:00
sse_output_recovery.py refactor(repositories): type prisma table access with one generic protocol 2026-08-25 12:14:17 +00:00
streaming_iterator.py refactor: clear fresh tech debt from the last 24 hours (2026-09-03) 2026-09-03 08:09:51 +00:00
utils.py refactor(typing): replace Any with proven types in 42 more backend files 2026-09-02 15:35:01 +00:00