diff --git a/docs/my-website/docs/proxy/cursor.md b/docs/my-website/docs/proxy/cursor.md new file mode 100644 index 00000000000..d01c1e62036 --- /dev/null +++ b/docs/my-website/docs/proxy/cursor.md @@ -0,0 +1,108 @@ +--- +id: cursor +title: /cursor/chat/completions - Cursor Endpoint +description: Accept Responses API input from Cursor and return OpenAI Chat Completions output +--- + +LiteLLM provides a Cursor-specific endpoint to make Cursor IDE work seamlessly with the LiteLLM Proxy when using BYOK + custom `base_url`. + +- Accepts Requests in OpenAI Responses API input format (Cursor sends this) +- Returns Responses in OpenAI Chat Completions format (Cursor expects this) +- Supports streaming and non‑streaming + +## Endpoint + +- Path: `/cursor/chat/completions` +- Auth: Standard LiteLLM Proxy auth (`Authorization: Bearer `) +- Behavior: Internally routes to LiteLLM `/responses` flow and transforms output to Chat Completions + +## Why this exists + +When setting up Cursor with BYOK against a custom `base_url`, Cursor sends requests to the Chat Completions endpoint but in the OpenAI Responses API input shape. Without translation, Cursor won’t display streamed output. This endpoint bridges the formats: + +- Input: Responses API (`input`, tool calls, etc.) +- Output: Chat Completions (`choices`, `delta`, `finish_reason`, etc.) + +## Usage + +### Non-streaming + +```bash +curl -X POST https://litellm-internal/cursor/chat/completions \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer sk-1234" \ + -d '{ + "model": "gpt-4o", + "input": [{"role": "user", "content": "Hello"}] + }' +``` + +Example response (shape): + +```json +{ + "id": "chatcmpl-123", + "object": "chat.completion", + "created": 1733333333, + "model": "gpt-4o", + "choices": [ + { + "index": 0, + "message": { + "role": "assistant", + "content": "Hello! How can I help you?" + }, + "finish_reason": "stop" + } + ], + "usage": { + "prompt_tokens": 10, + "completion_tokens": 8, + "total_tokens": 18 + } +} +``` + +### Streaming + +```bash +curl -N -X POST https://litellm-internal/cursor/chat/completions \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer sk-1234" \ + -d '{ + "model": "gpt-4o", + "input": [{"role": "user", "content": "Hello"}], + "stream": true + }' +``` + +- Server-Sent Events (SSE) +- Emits `chat.completion.chunk` deltas (`choices[].delta`) and ends with `data: [DONE]` + +## Configuration + +### Base URL Setup + +**Important**: When configuring Cursor IDE to use this endpoint, you must include `/cursor` in the base URL. + +Cursor automatically appends `/chat/completions` to the base URL you provide. To ensure requests go to `/cursor/chat/completions`, configure your base URL in Cursor as: + +``` +Base URL: https://litellm-internal/cursor +``` + +This way, when Cursor appends `/chat/completions`, the full path becomes `/cursor/chat/completions`, which is the correct endpoint. + +**Example**: If your LiteLLM Proxy is running at `https://litellm-internal`, set the base URL in Cursor to `https://litellm-internal/cursor` (not just `https://litellm-internal`). + +### General Setup + +No special configuration is required beyond your normal LiteLLM Proxy setup. Ensure that: + +- Your `config.yaml` includes the models you want to call via this endpoint +- Your Cursor project uses your LiteLLM Proxy `base_url` (with `/cursor` included) and a valid API key + +## Notes +- This endpoint is intended specifically for Cursor’s request/response expectations. Other clients should continue to use `/v1/chat/completions` or `/v1/responses` as appropriate. + + diff --git a/docs/my-website/docs/tutorials/cursor_integration.md b/docs/my-website/docs/tutorials/cursor_integration.md new file mode 100644 index 00000000000..f0d87b050cf --- /dev/null +++ b/docs/my-website/docs/tutorials/cursor_integration.md @@ -0,0 +1,226 @@ +--- +sidebar_label: "Cursor IDE" +--- + +import Tabs from '@theme/Tabs'; +import TabItem from '@theme/TabItem'; + +# Cursor IDE Integration with LiteLLM + +This tutorial shows you how to integrate Cursor IDE with LiteLLM Proxy, allowing you to use any LiteLLM-supported model through Cursor's interface with BYOK (Bring Your Own Key) and custom base URL. + +## Benefits of using Cursor with LiteLLM + +When you use Cursor IDE with LiteLLM you get the following benefits: + +**Developer Benefits:** +- Universal Model Access: Use any LiteLLM supported model (Anthropic, OpenAI, Vertex AI, Bedrock, etc.) through the Cursor IDE interface. +- Higher Rate Limits & Reliability: Load balance across multiple models and providers to avoid hitting individual provider limits, with fallbacks to ensure you get responses even if one provider fails. +- Streaming Support: Full streaming support with proper response transformation for Cursor's expected format. + +**Proxy Admin Benefits:** +- Centralized Management: Control access to all models through a single LiteLLM proxy instance without giving your developers API Keys to each provider. +- Budget Controls: Set spending limits and track costs across all Cursor usage. +- Request Logging: Track all requests made through Cursor for debugging and monitoring. + +## Prerequisites + +Before you begin, ensure you have: +- Cursor IDE installed +- A running LiteLLM Proxy instance with **HTTPS enabled** (HTTP is not supported) +- A valid LiteLLM Proxy API key +- An HTTPS domain for your LiteLLM Proxy (required by Cursor) + +## Quick Start Guide + +### Step 1: Install LiteLLM + +Install LiteLLM with proxy support: + +```bash +pip install litellm[proxy] +``` + +### Step 2: Configure LiteLLM Proxy + +Create a `config.yaml` file with your model configurations: + +```yaml showLineNumbers title="config.yaml" +model_list: + - model_name: gpt-4o + litellm_params: + model: gpt-4o + api_key: os.environ/OPENAI_API_KEY + + - model_name: claude-3-5-sonnet + litellm_params: + model: anthropic/claude-3-5-sonnet-20241022 + api_key: os.environ/ANTHROPIC_API_KEY + +general_settings: + master_key: sk-1234567890 # Change this to a secure key +``` + +### Step 3: Start LiteLLM Proxy + +Start the proxy server with HTTPS enabled: + +```bash +litellm --config config.yaml --port 4000 +``` + +:::warning HTTPS Required + +**Important**: Cursor IDE requires HTTPS connections. HTTP (`http://`) will not work. You must: +- Deploy your LiteLLM Proxy with HTTPS enabled +- Use a valid SSL certificate +- Access the proxy via an HTTPS domain (e.g., `https://your-proxy-domain.com`) + +For local development, you'll need to set up HTTPS (e.g., using a reverse proxy like nginx with SSL, or deploying to a cloud service with HTTPS). + +::: + +### Step 4: Configure Cursor IDE + +Configure Cursor IDE to use your LiteLLM proxy with the `/cursor/chat/completions` endpoint: + +1. Open Cursor IDE +2. Go to **Settings** → **Features** → **AI** +3. Enable **"Use Custom API"** or **"Bring Your Own Key"** +4. Set the following: + - **Base URL**: `https://your-proxy-domain.com/cursor` (⚠️ **Important**: Must use HTTPS and include `/cursor`) + - **API Key**: Your LiteLLM Proxy API key (e.g., `sk-1234567890`) + +:::warning HTTPS Required + +Cursor IDE **requires HTTPS** connections. HTTP (`http://`) will not work. You must: +- Use an HTTPS URL for your base URL (e.g., `https://your-proxy-domain.com/cursor`) +- Ensure your LiteLLM Proxy is accessible via HTTPS +- Have a valid SSL certificate configured + +::: + +**Example Configuration:** + +``` +Base URL: https://your-proxy-domain.com/cursor +API Key: sk-1234567890 +``` + +Replace `your-proxy-domain.com` with your actual HTTPS domain where LiteLLM Proxy is running. + +:::info Why `/cursor` in the base URL? + +Cursor automatically appends `/chat/completions` to the base URL you provide. By setting the base URL to `https://your-proxy-domain.com/cursor`, Cursor will send requests to `/cursor/chat/completions`, which is the special endpoint that handles Cursor's Responses API input format and transforms it to Chat Completions output format. + +If you set the base URL to just `https://your-proxy-domain.com`, Cursor would send requests to `/chat/completions`, which won't work correctly with Cursor's request format. + + +::: + +### Step 5: Test the Integration + +1. Restart Cursor IDE to apply the settings +2. Open a code file and try using Cursor's AI features (completions, chat, etc.) +3. Your requests will now be routed through LiteLLM Proxy + +You can verify it's working by: +- Checking the LiteLLM Proxy logs for incoming requests +- Using Cursor's chat feature and seeing responses stream correctly +- Checking your LiteLLM dashboard for request logs and cost tracking + +## How It Works + +The `/cursor/chat/completions` endpoint is specifically designed to handle Cursor's unique request format: + +1. **Input**: Cursor sends requests in OpenAI Responses API format (with `input` field) +2. **Processing**: LiteLLM processes the request through its internal `/responses` flow +3. **Output**: The response is transformed to OpenAI Chat Completions format (with `choices` field) that Cursor expects + +This transformation happens automatically for both streaming and non-streaming responses. + +## Advanced Configuration + +### Using Different Models + +You can configure Cursor to use different models by updating your `config.yaml`: + +```yaml showLineNumbers title="config.yaml" +model_list: + - model_name: gpt-4o + litellm_params: + model: gpt-4o + api_key: os.environ/OPENAI_API_KEY + + - model_name: claude-3-5-sonnet + litellm_params: + model: anthropic/claude-3-5-sonnet-20241022 + api_key: os.environ/ANTHROPIC_API_KEY + + - model_name: gemini-pro + litellm_params: + model: gemini/gemini-1.5-pro + api_key: os.environ/GEMINI_API_KEY +``` + +Then in Cursor, you can specify which model to use in your requests. + +### Rate Limiting and Budgets + +Set up rate limits and budgets in your `config.yaml`: + +```yaml showLineNumbers title="config.yaml" +general_settings: + master_key: sk-1234567890 + +litellm_settings: + # Set max budget per user + max_budget: 100.0 + + # Set rate limits + rate_limit: 100 # requests per minute +``` + +### Request Logging + +All requests from Cursor will be logged by LiteLLM Proxy. You can: +- View logs in the LiteLLM Admin UI +- Export logs to your preferred logging service +- Track costs per user/team + +## Troubleshooting + +### Cursor shows no output + +- **Check base URL**: Ensure it uses HTTPS and includes `/cursor` (e.g., `https://your-proxy-domain.com/cursor`, not `http://` or without `/cursor`) +- **Verify HTTPS**: Cursor requires HTTPS - HTTP connections will not work +- **Check API key**: Verify your LiteLLM Proxy API key is correct +- **Check proxy logs**: Look for errors in the LiteLLM Proxy logs + +### Requests failing + +- **Verify HTTPS is enabled**: Cursor requires HTTPS connections. Ensure your LiteLLM Proxy is accessible via HTTPS with a valid SSL certificate +- **Verify proxy is running**: Check that LiteLLM Proxy is accessible at your HTTPS base URL +- **Check SSL certificate**: Ensure your SSL certificate is valid and not expired +- **Check model configuration**: Ensure the model you're trying to use is configured in `config.yaml` +- **Check API keys**: Verify provider API keys are set correctly in environment variables + +### HTTP not working + +If you're trying to use HTTP (`http://`) and it's not working: +- **This is expected**: Cursor IDE requires HTTPS connections +- **Solution**: Deploy your LiteLLM Proxy with HTTPS enabled (use a reverse proxy like nginx, or deploy to a cloud service that provides HTTPS) + +### Streaming not working + +The `/cursor/chat/completions` endpoint automatically handles streaming. If streaming isn't working: +- Check that your model supports streaming +- Verify the proxy logs for any transformation errors +- Ensure Cursor IDE is up to date + +## Related Documentation + +- [Cursor Endpoint Documentation](/docs/proxy/cursor) - Detailed endpoint documentation +- [LiteLLM Proxy Setup](/docs/proxy/quick_start) - General proxy setup guide +- [Model Configuration](/docs/proxy/configs) - How to configure models + diff --git a/docs/my-website/sidebars.js b/docs/my-website/sidebars.js index b7143490dea..073f2f5d288 100644 --- a/docs/my-website/sidebars.js +++ b/docs/my-website/sidebars.js @@ -105,6 +105,7 @@ const sidebars = { items: [ "tutorials/claude_responses_api", "tutorials/cost_tracking_coding", + "tutorials/cursor_integration", "tutorials/github_copilot_integration", "tutorials/litellm_gemini_cli", "tutorials/litellm_qwen_code_cli", @@ -129,16 +130,6 @@ const sidebars = { }, items: [ "proxy/docker_quick_start", - { - type: "link", - label: "A2A Agent Gateway", - href: "https://docs.litellm.ai/docs/a2a", - }, - { - type: "link", - label: "MCP Gateway", - href: "https://docs.litellm.ai/docs/mcp", - }, { "type": "category", "label": "Config.yaml", @@ -195,7 +186,6 @@ const sidebars = { label: "Architecture", items: [ "proxy/architecture", - "proxy/multi_tenant_architecture", "proxy/control_plane_and_data_plane", "proxy/db_deadlocks", "proxy/db_info", @@ -327,14 +317,6 @@ const sidebars = { slug: "/supported_endpoints", }, items: [ - { - type: "category", - label: "/a2a - A2A Agent Gateway", - items: [ - "a2a", - "a2a_agent_permissions", - ], - }, "assistants", { type: "category", @@ -490,11 +472,6 @@ const sidebars = { id: "provider_registration/index", label: "Integrate as a Model Provider", }, - { - type: "doc", - id: "contributing/adding_openai_compatible_providers", - label: "Add OpenAI-Compatible Provider (JSON)", - }, { type: "doc", id: "provider_registration/add_model_pricing", @@ -820,7 +797,6 @@ const sidebars = { type: "category", label: "Adding Providers", items: [ - "contributing/adding_openai_compatible_providers", "adding_provider/directory_structure", "adding_provider/new_rerank_provider", ] diff --git a/litellm/proxy/response_api_endpoints/endpoints.py b/litellm/proxy/response_api_endpoints/endpoints.py index 26d10c1ac47..7736c37a809 100644 --- a/litellm/proxy/response_api_endpoints/endpoints.py +++ b/litellm/proxy/response_api_endpoints/endpoints.py @@ -85,6 +85,150 @@ async def responses_api( ) +@router.post( + "/cursor/chat/completions", + dependencies=[Depends(user_api_key_auth)], + tags=["responses"], +) +async def cursor_chat_completions( + request: Request, + fastapi_response: Response, + user_api_key_dict: UserAPIKeyAuth = Depends(user_api_key_auth), +): + """ + Cursor-specific endpoint that accepts Responses API input format but returns chat completions format. + + This endpoint handles requests from Cursor IDE which sends Responses API format (`input` field) + but expects chat completions format response (`choices`, `messages`, etc.). + + ```bash + curl -X POST http://localhost:4000/cursor/chat/completions \ + -H "Content-Type: application/json" \ + -H "Authorization: Bearer sk-1234" \ + -d '{ + "model": "gpt-4o", + "input": [{"role": "user", "content": "Hello"}] + }' + Responds back in chat completions format. + ``` + """ + from litellm.completion_extras.litellm_responses_transformation.handler import ( + responses_api_bridge, + ) + from litellm.litellm_core_utils.streaming_handler import CustomStreamWrapper + from litellm.proxy.proxy_server import ( + _read_request_body, + async_data_generator, + general_settings, + llm_router, + proxy_config, + proxy_logging_obj, + user_api_base, + user_max_tokens, + user_model, + user_request_timeout, + user_temperature, + version, + ) + from litellm.responses.streaming_iterator import BaseResponsesAPIStreamingIterator + from litellm.types.llms.openai import ResponsesAPIResponse + + data = await _read_request_body(request=request) + processor = ProxyBaseLLMRequestProcessing(data=data) + + def cursor_data_generator(response, user_api_key_dict, request_data): + """ + Custom generator that transforms Responses API streaming chunks to chat completion chunks. + + This generator is used for the cursor endpoint to convert Responses API format responses + to chat completion format that Cursor IDE expects. + + Args: + response: The streaming response (BaseResponsesAPIStreamingIterator or other) + user_api_key_dict: User API key authentication dict + request_data: Request data containing model, logging_obj, etc. + + Returns: + Async generator that yields SSE-formatted chat completion chunks + """ + # If response is a BaseResponsesAPIStreamingIterator, transform it first + if isinstance(response, BaseResponsesAPIStreamingIterator): + # Transform Responses API iterator to chat completion iterator + completion_stream = responses_api_bridge.transformation_handler.get_model_response_iterator( + streaming_response=response, + sync_stream=False, + json_mode=False, + ) + # Wrap in CustomStreamWrapper to get the async generator + logging_obj = request_data.get("litellm_logging_obj") + streamwrapper = CustomStreamWrapper( + completion_stream=completion_stream, + model=request_data.get("model", ""), + custom_llm_provider=None, + logging_obj=logging_obj, + ) + # Use async_data_generator to format as SSE + return async_data_generator( + response=streamwrapper, + user_api_key_dict=user_api_key_dict, + request_data=request_data, + ) + # Otherwise, use the default generator + return async_data_generator( + response=response, + user_api_key_dict=user_api_key_dict, + request_data=request_data, + ) + + try: + response = await processor.base_process_llm_request( + request=request, + fastapi_response=fastapi_response, + user_api_key_dict=user_api_key_dict, + route_type="aresponses", + proxy_logging_obj=proxy_logging_obj, + llm_router=llm_router, + general_settings=general_settings, + proxy_config=proxy_config, + select_data_generator=cursor_data_generator, + model=None, + user_model=user_model, + user_temperature=user_temperature, + user_request_timeout=user_request_timeout, + user_max_tokens=user_max_tokens, + user_api_base=user_api_base, + version=version, + ) + + # Transform non-streaming Responses API response to chat completions format + if isinstance(response, ResponsesAPIResponse): + logging_obj = processor.data.get("litellm_logging_obj") + transformed_response = responses_api_bridge.transformation_handler.transform_response( + model=processor.data.get("model", ""), + raw_response=response, + model_response=None, + logging_obj=logging_obj, + request_data=processor.data, + messages=processor.data.get("input", []), + optional_params={}, + litellm_params={}, + encoding=None, + api_key=None, + json_mode=None, + ) + return transformed_response + + # Streaming responses are already transformed by cursor_select_data_generator + return response + except Exception as e: + raise await processor._handle_llm_api_exception( + e=e, + user_api_key_dict=user_api_key_dict, + proxy_logging_obj=proxy_logging_obj, + version=version, + ) + + @router.get( "/v1/responses/{response_id}", dependencies=[Depends(user_api_key_auth)], diff --git a/tests/test_litellm/proxy/response_api_endpoints/test_endpoints.py b/tests/test_litellm/proxy/response_api_endpoints/test_endpoints.py index bca0944aeac..4bbbf87edb8 100644 --- a/tests/test_litellm/proxy/response_api_endpoints/test_endpoints.py +++ b/tests/test_litellm/proxy/response_api_endpoints/test_endpoints.py @@ -51,3 +51,66 @@ class TestResponsesAPIEndpoints(unittest.TestCase): assert response.status_code in [200, 401, 500] + @pytest.mark.asyncio + @patch("litellm.proxy.proxy_server.llm_router") + @patch("litellm.proxy.proxy_server.user_api_key_auth") + async def test_cursor_chat_completions_route(self, mock_auth, mock_router): + """ + Test that /cursor/chat/completions endpoint: + 1. Accepts Responses API input format + 2. Returns chat completions format response + 3. Transforms streaming responses correctly + """ + from litellm.types.llms.openai import ResponsesAPIResponse + from litellm.types.utils import ResponseOutputMessage, ResponseOutputText + + mock_auth.return_value = MagicMock( + token="test_token", + user_id="test_user", + team_id=None, + ) + + # Mock a Responses API response + mock_responses_response = ResponsesAPIResponse( + id="resp_cursor123", + created_at=1234567890, + model="gpt-4o", + object="response", + output=[ + ResponseOutputMessage( + type="message", + role="assistant", + content=[ + ResponseOutputText(type="output_text", text="Hello from Cursor!") + ], + ) + ], + ) + + mock_router.aresponses = AsyncMock(return_value=mock_responses_response) + + client = TestClient(app) + + # Test with Responses API input format (what Cursor sends) + test_data = { + "model": "gpt-4o", + "input": [{"role": "user", "content": "Hello"}], + } + + response = client.post( + "/cursor/chat/completions", + json=test_data, + headers={"Authorization": "Bearer sk-1234"}, + ) + + # Should return 200 (or 401/500 if auth fails) + assert response.status_code in [200, 401, 500] + + # If successful, verify it returns chat completions format + if response.status_code == 200: + response_data = response.json() + # Should have chat completion structure + assert "choices" in response_data or "id" in response_data + # Should not have Responses API structure + assert "output" not in response_data or "status" not in response_data +