- registry_orchestrator: exact URL match for agent routing (no substring),
skip agents with empty URLs, deduplicate colliding sanitized names,
wrap a2a import in ImportError guard, hoist SemanticToolFilterHook import
out of callback loop, note user_api_key_auth auth filtering as TODO
- litellm_proxy_mcp_handler: guard FastAPI import with try/except ImportError,
use .get() on tool_server_map to handle hallucinated tool names gracefully,
add debug logging around A2A tool call dispatch
- chat_completions_handler: lazy-import RegistryOrchestrator inside function
to avoid hard proxy coupling at module load, extend semantic_filter check to
cover agent_tool_configs as well as mcp_tools, handle non-streaming ModelResponse
follow-up by emitting as synthetic final chunk instead of silently dropping
- Move A2A orchestration logic (parse, wrap, execute, semantic filter) into
RegistryOrchestrator class in litellm/proxy/agent_endpoints/registry_orchestrator.py
- Extract MCPStreamingIterator, MCPStreamWrapper, _SyncIteratorWrapper from nested
closures inside acompletion_with_mcp() to module-level classes in
chat_completions_handler.py
- Fix 7x repeated verbose_logger inline imports in MCPStreamingIterator methods
- Fix async __aiter__ (should be sync) in MCPStreamingIterator
- Remove dead rules_obj and duplicate proxy_logging_obj import from _execute_tool_calls
- Update tests to import from new locations
- server_url: 'litellm_proxy/mcp' (bare, no server suffix) now expands to ALL
registered MCP servers at inference time. Previously the code initialized
mcp_servers=[] which caused _get_allowed_mcp_servers to return nothing; fixed
by using Optional[List[str]]=None (None = all servers).
- New type='a2a_agent' tool with server_url='litellm_proxy/agents' wraps every
registered agent from global_agent_registry as an OpenAI function tool.
Agent descriptions are enriched with up to 3 skill descriptions. Names are
sanitized to ^[a-zA-Z0-9_-]{1,64}$ for OpenAI compatibility.
- A2A tool calls are executed via JSON-RPC 2.0 message/send over httpx.
Responses are parsed from result.artifacts[].parts[].text with fallback to
result.status.message.parts[].text. Both MCP and A2A calls share the same
litellm_trace_id so they appear in the same trace.
- semantic_filter: true on an MCP tool config triggers SemanticToolFilterHook
(if configured as a callback) to pre-filter tools by query relevance before
injecting into the LLM context.
- Streaming (stream=True) fully supported: MCPStreamingIterator already handles
the tool loop; agent_tool_map is threaded through the same path.
Tests (8/8 pass, no mcp package required):
test_registry_orchestration_nonstreaming - MCP + A2A in same trace
test_registry_orchestration_streaming - stream=True, both tools executed
test_bare_mcp_url_expands_to_all_servers - bare URL passes all-servers sentinel
test_agents_wrapped_as_function_tools - correct schema + name sanitization
test_parse_a2a_response_{artifacts,status_message,error} - A2A parsing
test_semantic_filter_reduces_tools - filter hook reduces injected tools
The special name check (all_team_servers, all_proxy_servers) was an elif
after the server_id-is-not-None check, making it unreachable since special
names are non-None strings. Split into separate if blocks so the special
name guard runs before the duplicate-ID check.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The team MCP manager feature was reverted in PR #24255, so the test
needs to go back to the original single auth failure test that expects
a 403 for non-admin users.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The test_create_mcp_server_auth_failure test expected a 403 for non-admin
users, but the team MCP manager feature changed the auth flow to first
check for team_id (400) before checking permissions. Split into two tests:
one for missing team_id (400) and one for non-manager rejection (403).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Fixes CI failure in test_api_docs.py which validates that all Pydantic
model fields are documented in endpoint docstrings.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add pytest.importorskip("mcp") at module level so tests skip cleanly
in CI environments without the mcp package (instead of ImportError)
- Import LiteLLM_TeamTableCachedObj into MCP_AVAILABLE block so type
annotations resolve for static analysis and get_type_hints()
- Remove string quotes from type annotations now that the import exists
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Addresses Greptile feedback about missing integration tests for PUT/DELETE
when invoked by mcp_server_manager role. Adds tests for edit success/403,
delete success with team cleanup/403, and the _remove_mcp_server_from_team
helper directly.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace raw prisma_client.db.litellm_teamtable.update with
handle_update_object_permission from team_endpoints (follows
established helper-function pattern)
- Extract _auto_assign_mcp_server_to_team and
_remove_mcp_server_from_team helpers for reuse and testability
- Update tests to mock at the correct boundaries
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
AntD v5 static message API doesn't render without an App wrapper or
useMessage() context holder. Mirrors the existing notification pattern
by adding message.useMessage() to AntdGlobalProvider and routing all
calls through a new MessageManager module.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove dead get_team_object mock in test (now reuses team_obj from _assert)
- Add test for existing_permission_id=None branch (team linkage)
- Remaining P1s are by-design per spec (admin blocked from MCP, team_id
ignored for proxy admin)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- P0: Separate try/except for auto-assign so server creation succeeds
even if team permission update fails
- P1: Clean up team permission entry on MCP server delete
- P1: Add MCP_AVAILABLE skip guard to tests
- P2: Return team_obj from _assert_can_manage_team_mcp_server to
eliminate redundant get_team_object call in create endpoint
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Cached team objects may have stale object_permission data. Bypassing
cache ensures the server-in-team check uses fresh data.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When a team has no object_permission_id yet, the auto-assign logic creates
an ObjectPermissionTable row but never linked it to the team. Now updates
the team's object_permission_id after creation.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
team_id is a request-level field, not a DB column. Excluding it prevents
Prisma from rejecting the create call.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds a control plane capability that enables a central admin instance
to manage multiple regional worker proxies from a single UI.
Backend:
- Worker registry loaded from YAML config (worker_id, name, url)
- /.well-known/litellm-ui-config exposes is_control_plane and workers list
- /v3/login + /v3/login/exchange: opaque code exchange for cross-origin
username/password auth (JWT never in URL/logs, single-use 60s TTL)
- SSO cookie handoff with return_to → opaque code → exchange
- _validate_return_to: full origin validation (scheme+hostname+port)
- Startup warning when control_plane_url set without Redis
- Both /v3 endpoints gated behind control_plane_url config
Frontend:
- Worker selector dropdown on login page (gated behind is_control_plane)
- Cross-origin SSO code exchange handling on callback
- switchToWorkerUrl: localStorage-persisted worker URL for API calls
- useWorker hook: shared worker state management
- WorkerDropdown in navbar for switching workers
- Logout/switch clears worker state from localStorage
Tests:
- 7 tests for /v3/login + /v3/login/exchange
- 10 tests for _validate_return_to
- 2 tests for control plane discovery endpoint
Migrate the OldTeams table from Tremor to Ant Design components, matching the
Access Groups page pattern. Switch from /team/list to /v2/team/list for
server-side pagination, filtering, and sorting.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- clearChatHistory: use functional setChatHistory updater so blob URL
revocation operates on the latest snapshot, not a stale closure capture.
- Simplified mode: skip sessionStorage hydration and persistence for
messageTraceId, responsesSessionId, and useApiSessionManagement so
embedded widgets don't cross-contaminate the full playground session.
- Debounce race: skip re-writing empty chatHistory to sessionStorage
after clearChatHistory already removed the key.
- Added 5 new tests covering these fixes (39 total).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>