Merge pull request #17519 from BerriAI/litellm_cursor_integration

Add support for cursor BYOK with its own configuration
This commit is contained in:
Sameer Kankute 2025-12-05 22:23:45 +05:30 • committed by GitHub
commit 558c8f92d1
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
5 changed files with 542 additions and 25 deletions

View file

@ -0,0 +1,108 @@
---
id: cursor
title: /cursor/chat/completions - Cursor Endpoint
description: Accept Responses API input from Cursor and return OpenAI Chat Completions output
---
LiteLLM provides a Cursor-specific endpoint to make Cursor IDE work seamlessly with the LiteLLM Proxy when using BYOK + custom `base_url`.
- Accepts Requests in OpenAI Responses API input format (Cursor sends this)
- Returns Responses in OpenAI Chat Completions format (Cursor expects this)
- Supports streaming and non‑streaming
## Endpoint
- Path: `/cursor/chat/completions`
- Auth: Standard LiteLLM Proxy auth (`Authorization: Bearer <key>`)
- Behavior: Internally routes to LiteLLM `/responses` flow and transforms output to Chat Completions
## Why this exists
When setting up Cursor with BYOK against a custom `base_url`, Cursor sends requests to the Chat Completions endpoint but in the OpenAI Responses API input shape. Without translation, Cursor won’t display streamed output. This endpoint bridges the formats:
- Input: Responses API (`input`, tool calls, etc.)
- Output: Chat Completions (`choices`, `delta`, `finish_reason`, etc.)
## Usage
### Non-streaming
```bash
curl -X POST https://litellm-internal/cursor/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-1234" \
-d '{
"model": "gpt-4o",
"input": [{"role": "user", "content": "Hello"}]
}'
```
Example response (shape):
```json
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1733333333,
"model": "gpt-4o",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 8,
"total_tokens": 18
}
}
```
### Streaming
```bash
curl -N -X POST https://litellm-internal/cursor/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-1234" \
-d '{
"model": "gpt-4o",
"input": [{"role": "user", "content": "Hello"}],
"stream": true
}'
```
- Server-Sent Events (SSE)
- Emits `chat.completion.chunk` deltas (`choices[].delta`) and ends with `data: [DONE]`
## Configuration
### Base URL Setup
**Important**: When configuring Cursor IDE to use this endpoint, you must include `/cursor` in the base URL.
Cursor automatically appends `/chat/completions` to the base URL you provide. To ensure requests go to `/cursor/chat/completions`, configure your base URL in Cursor as:
```
Base URL: https://litellm-internal/cursor
```
This way, when Cursor appends `/chat/completions`, the full path becomes `/cursor/chat/completions`, which is the correct endpoint.
**Example**: If your LiteLLM Proxy is running at `https://litellm-internal`, set the base URL in Cursor to `https://litellm-internal/cursor` (not just `https://litellm-internal`).
### General Setup
No special configuration is required beyond your normal LiteLLM Proxy setup. Ensure that:
- Your `config.yaml` includes the models you want to call via this endpoint
- Your Cursor project uses your LiteLLM Proxy `base_url` (with `/cursor` included) and a valid API key
## Notes
- This endpoint is intended specifically for Cursor’s request/response expectations. Other clients should continue to use `/v1/chat/completions` or `/v1/responses` as appropriate.

View file

@ -0,0 +1,226 @@
---
sidebar_label: "Cursor IDE"
---
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
# Cursor IDE Integration with LiteLLM
This tutorial shows you how to integrate Cursor IDE with LiteLLM Proxy, allowing you to use any LiteLLM-supported model through Cursor's interface with BYOK (Bring Your Own Key) and custom base URL.
## Benefits of using Cursor with LiteLLM
When you use Cursor IDE with LiteLLM you get the following benefits:
**Developer Benefits:**
- Universal Model Access: Use any LiteLLM supported model (Anthropic, OpenAI, Vertex AI, Bedrock, etc.) through the Cursor IDE interface.
- Higher Rate Limits & Reliability: Load balance across multiple models and providers to avoid hitting individual provider limits, with fallbacks to ensure you get responses even if one provider fails.
- Streaming Support: Full streaming support with proper response transformation for Cursor's expected format.
**Proxy Admin Benefits:**
- Centralized Management: Control access to all models through a single LiteLLM proxy instance without giving your developers API Keys to each provider.
- Budget Controls: Set spending limits and track costs across all Cursor usage.
- Request Logging: Track all requests made through Cursor for debugging and monitoring.
## Prerequisites
Before you begin, ensure you have:
- Cursor IDE installed
- A running LiteLLM Proxy instance with **HTTPS enabled** (HTTP is not supported)
- A valid LiteLLM Proxy API key
- An HTTPS domain for your LiteLLM Proxy (required by Cursor)
## Quick Start Guide
### Step 1: Install LiteLLM
Install LiteLLM with proxy support:
```bash
pip install litellm[proxy]
```
### Step 2: Configure LiteLLM Proxy
Create a `config.yaml` file with your model configurations:
```yaml showLineNumbers title="config.yaml"
model_list:
- model_name: gpt-4o
litellm_params:
model: gpt-4o
api_key: os.environ/OPENAI_API_KEY
- model_name: claude-3-5-sonnet
litellm_params:
model: anthropic/claude-3-5-sonnet-20241022
api_key: os.environ/ANTHROPIC_API_KEY
general_settings:
master_key: sk-1234567890 # Change this to a secure key
```
### Step 3: Start LiteLLM Proxy
Start the proxy server with HTTPS enabled:
```bash
litellm --config config.yaml --port 4000
```
:::warning HTTPS Required
**Important**: Cursor IDE requires HTTPS connections. HTTP (`http://`) will not work. You must:
- Deploy your LiteLLM Proxy with HTTPS enabled
- Use a valid SSL certificate
- Access the proxy via an HTTPS domain (e.g., `https://your-proxy-domain.com`)
For local development, you'll need to set up HTTPS (e.g., using a reverse proxy like nginx with SSL, or deploying to a cloud service with HTTPS).
:::
### Step 4: Configure Cursor IDE
Configure Cursor IDE to use your LiteLLM proxy with the `/cursor/chat/completions` endpoint:
1. Open Cursor IDE
2. Go to **Settings** → **Features** → **AI**
3. Enable **"Use Custom API"** or **"Bring Your Own Key"**
4. Set the following:
- **Base URL**: `https://your-proxy-domain.com/cursor` (⚠️ **Important**: Must use HTTPS and include `/cursor`)
- **API Key**: Your LiteLLM Proxy API key (e.g., `sk-1234567890`)
:::warning HTTPS Required
Cursor IDE **requires HTTPS** connections. HTTP (`http://`) will not work. You must:
- Use an HTTPS URL for your base URL (e.g., `https://your-proxy-domain.com/cursor`)
- Ensure your LiteLLM Proxy is accessible via HTTPS
- Have a valid SSL certificate configured
:::
**Example Configuration:**
```
Base URL: https://your-proxy-domain.com/cursor
API Key: sk-1234567890
```
Replace `your-proxy-domain.com` with your actual HTTPS domain where LiteLLM Proxy is running.
:::info Why `/cursor` in the base URL?
Cursor automatically appends `/chat/completions` to the base URL you provide. By setting the base URL to `https://your-proxy-domain.com/cursor`, Cursor will send requests to `/cursor/chat/completions`, which is the special endpoint that handles Cursor's Responses API input format and transforms it to Chat Completions output format.
If you set the base URL to just `https://your-proxy-domain.com`, Cursor would send requests to `/chat/completions`, which won't work correctly with Cursor's request format.
:::
### Step 5: Test the Integration
1. Restart Cursor IDE to apply the settings
2. Open a code file and try using Cursor's AI features (completions, chat, etc.)
3. Your requests will now be routed through LiteLLM Proxy
You can verify it's working by:
- Checking the LiteLLM Proxy logs for incoming requests
- Using Cursor's chat feature and seeing responses stream correctly
- Checking your LiteLLM dashboard for request logs and cost tracking
## How It Works
The `/cursor/chat/completions` endpoint is specifically designed to handle Cursor's unique request format:
1. **Input**: Cursor sends requests in OpenAI Responses API format (with `input` field)
2. **Processing**: LiteLLM processes the request through its internal `/responses` flow
3. **Output**: The response is transformed to OpenAI Chat Completions format (with `choices` field) that Cursor expects
This transformation happens automatically for both streaming and non-streaming responses.
## Advanced Configuration
### Using Different Models
You can configure Cursor to use different models by updating your `config.yaml`:
```yaml showLineNumbers title="config.yaml"
model_list:
- model_name: gpt-4o
litellm_params:
model: gpt-4o
api_key: os.environ/OPENAI_API_KEY
- model_name: claude-3-5-sonnet
litellm_params:
model: anthropic/claude-3-5-sonnet-20241022
api_key: os.environ/ANTHROPIC_API_KEY
- model_name: gemini-pro
litellm_params:
model: gemini/gemini-1.5-pro
api_key: os.environ/GEMINI_API_KEY
```
Then in Cursor, you can specify which model to use in your requests.
### Rate Limiting and Budgets
Set up rate limits and budgets in your `config.yaml`:
```yaml showLineNumbers title="config.yaml"
general_settings:
master_key: sk-1234567890
litellm_settings:
# Set max budget per user
max_budget: 100.0
# Set rate limits
rate_limit: 100 # requests per minute
```
### Request Logging
All requests from Cursor will be logged by LiteLLM Proxy. You can:
- View logs in the LiteLLM Admin UI
- Export logs to your preferred logging service
- Track costs per user/team
## Troubleshooting
### Cursor shows no output
- **Check base URL**: Ensure it uses HTTPS and includes `/cursor` (e.g., `https://your-proxy-domain.com/cursor`, not `http://` or without `/cursor`)
- **Verify HTTPS**: Cursor requires HTTPS - HTTP connections will not work
- **Check API key**: Verify your LiteLLM Proxy API key is correct
- **Check proxy logs**: Look for errors in the LiteLLM Proxy logs
### Requests failing
- **Verify HTTPS is enabled**: Cursor requires HTTPS connections. Ensure your LiteLLM Proxy is accessible via HTTPS with a valid SSL certificate
- **Verify proxy is running**: Check that LiteLLM Proxy is accessible at your HTTPS base URL
- **Check SSL certificate**: Ensure your SSL certificate is valid and not expired
- **Check model configuration**: Ensure the model you're trying to use is configured in `config.yaml`
- **Check API keys**: Verify provider API keys are set correctly in environment variables
### HTTP not working
If you're trying to use HTTP (`http://`) and it's not working:
- **This is expected**: Cursor IDE requires HTTPS connections
- **Solution**: Deploy your LiteLLM Proxy with HTTPS enabled (use a reverse proxy like nginx, or deploy to a cloud service that provides HTTPS)
### Streaming not working
The `/cursor/chat/completions` endpoint automatically handles streaming. If streaming isn't working:
- Check that your model supports streaming
- Verify the proxy logs for any transformation errors
- Ensure Cursor IDE is up to date
## Related Documentation
- [Cursor Endpoint Documentation](/docs/proxy/cursor) - Detailed endpoint documentation
- [LiteLLM Proxy Setup](/docs/proxy/quick_start) - General proxy setup guide
- [Model Configuration](/docs/proxy/configs) - How to configure models

View file

@ -105,6 +105,7 @@ const sidebars = {
items: [
"tutorials/claude_responses_api",
"tutorials/cost_tracking_coding",
"tutorials/cursor_integration",
"tutorials/github_copilot_integration",
"tutorials/litellm_gemini_cli",
"tutorials/litellm_qwen_code_cli",
@ -129,16 +130,6 @@ const sidebars = {
},
items: [
"proxy/docker_quick_start",
{
type: "link",
label: "A2A Agent Gateway",
href: "https://docs.litellm.ai/docs/a2a",
},
{
type: "link",
label: "MCP Gateway",
href: "https://docs.litellm.ai/docs/mcp",
},
{
"type": "category",
"label": "Config.yaml",
@ -195,7 +186,6 @@ const sidebars = {
label: "Architecture",
items: [
"proxy/architecture",
"proxy/multi_tenant_architecture",
"proxy/control_plane_and_data_plane",
"proxy/db_deadlocks",
"proxy/db_info",
@ -327,14 +317,6 @@ const sidebars = {
slug: "/supported_endpoints",
},
items: [
{
type: "category",
label: "/a2a - A2A Agent Gateway",
items: [
"a2a",
"a2a_agent_permissions",
],
},
"assistants",
{
type: "category",
@ -490,11 +472,6 @@ const sidebars = {
id: "provider_registration/index",
label: "Integrate as a Model Provider",
},
{
type: "doc",
id: "contributing/adding_openai_compatible_providers",
label: "Add OpenAI-Compatible Provider (JSON)",
},
{
type: "doc",
id: "provider_registration/add_model_pricing",
@ -820,7 +797,6 @@ const sidebars = {
type: "category",
label: "Adding Providers",
items: [
"contributing/adding_openai_compatible_providers",
"adding_provider/directory_structure",
"adding_provider/new_rerank_provider",
]

View file

@ -85,6 +85,150 @@ async def responses_api(
)
@router.post(
"/cursor/chat/completions",
dependencies=[Depends(user_api_key_auth)],
tags=["responses"],
)
async def cursor_chat_completions(
request: Request,
fastapi_response: Response,
user_api_key_dict: UserAPIKeyAuth = Depends(user_api_key_auth),
):
"""
Cursor-specific endpoint that accepts Responses API input format but returns chat completions format.
This endpoint handles requests from Cursor IDE which sends Responses API format (`input` field)
but expects chat completions format response (`choices`, `messages`, etc.).
```bash
curl -X POST http://localhost:4000/cursor/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-1234" \
-d '{
"model": "gpt-4o",
"input": [{"role": "user", "content": "Hello"}]
}'
Responds back in chat completions format.
```
"""
from litellm.completion_extras.litellm_responses_transformation.handler import (
responses_api_bridge,
)
from litellm.litellm_core_utils.streaming_handler import CustomStreamWrapper
from litellm.proxy.proxy_server import (
_read_request_body,
async_data_generator,
general_settings,
llm_router,
proxy_config,
proxy_logging_obj,
user_api_base,
user_max_tokens,
user_model,
user_request_timeout,
user_temperature,
version,
)
from litellm.responses.streaming_iterator import BaseResponsesAPIStreamingIterator
from litellm.types.llms.openai import ResponsesAPIResponse
data = await _read_request_body(request=request)
processor = ProxyBaseLLMRequestProcessing(data=data)
def cursor_data_generator(response, user_api_key_dict, request_data):
"""
Custom generator that transforms Responses API streaming chunks to chat completion chunks.
This generator is used for the cursor endpoint to convert Responses API format responses
to chat completion format that Cursor IDE expects.
Args:
response: The streaming response (BaseResponsesAPIStreamingIterator or other)
user_api_key_dict: User API key authentication dict
request_data: Request data containing model, logging_obj, etc.
Returns:
Async generator that yields SSE-formatted chat completion chunks
"""
# If response is a BaseResponsesAPIStreamingIterator, transform it first
if isinstance(response, BaseResponsesAPIStreamingIterator):
# Transform Responses API iterator to chat completion iterator
completion_stream = responses_api_bridge.transformation_handler.get_model_response_iterator(
streaming_response=response,
sync_stream=False,
json_mode=False,
)
# Wrap in CustomStreamWrapper to get the async generator
logging_obj = request_data.get("litellm_logging_obj")
streamwrapper = CustomStreamWrapper(
completion_stream=completion_stream,
model=request_data.get("model", ""),
custom_llm_provider=None,
logging_obj=logging_obj,
)
# Use async_data_generator to format as SSE
return async_data_generator(
response=streamwrapper,
user_api_key_dict=user_api_key_dict,
request_data=request_data,
)
# Otherwise, use the default generator
return async_data_generator(
response=response,
user_api_key_dict=user_api_key_dict,
request_data=request_data,
)
try:
response = await processor.base_process_llm_request(
request=request,
fastapi_response=fastapi_response,
user_api_key_dict=user_api_key_dict,
route_type="aresponses",
proxy_logging_obj=proxy_logging_obj,
llm_router=llm_router,
general_settings=general_settings,
proxy_config=proxy_config,
select_data_generator=cursor_data_generator,
model=None,
user_model=user_model,
user_temperature=user_temperature,
user_request_timeout=user_request_timeout,
user_max_tokens=user_max_tokens,
user_api_base=user_api_base,
version=version,
)
# Transform non-streaming Responses API response to chat completions format
if isinstance(response, ResponsesAPIResponse):
logging_obj = processor.data.get("litellm_logging_obj")
transformed_response = responses_api_bridge.transformation_handler.transform_response(
model=processor.data.get("model", ""),
raw_response=response,
model_response=None,
logging_obj=logging_obj,
request_data=processor.data,
messages=processor.data.get("input", []),
optional_params={},
litellm_params={},
encoding=None,
api_key=None,
json_mode=None,
)
return transformed_response
# Streaming responses are already transformed by cursor_select_data_generator
return response
except Exception as e:
raise await processor._handle_llm_api_exception(
e=e,
user_api_key_dict=user_api_key_dict,
proxy_logging_obj=proxy_logging_obj,
version=version,
)
@router.get(
"/v1/responses/{response_id}",
dependencies=[Depends(user_api_key_auth)],

View file

@ -51,3 +51,66 @@ class TestResponsesAPIEndpoints(unittest.TestCase):
assert response.status_code in [200, 401, 500]
@pytest.mark.asyncio
@patch("litellm.proxy.proxy_server.llm_router")
@patch("litellm.proxy.proxy_server.user_api_key_auth")
async def test_cursor_chat_completions_route(self, mock_auth, mock_router):
"""
Test that /cursor/chat/completions endpoint:
1. Accepts Responses API input format
2. Returns chat completions format response
3. Transforms streaming responses correctly
"""
from litellm.types.llms.openai import ResponsesAPIResponse
from litellm.types.utils import ResponseOutputMessage, ResponseOutputText
mock_auth.return_value = MagicMock(
token="test_token",
user_id="test_user",
team_id=None,
)
# Mock a Responses API response
mock_responses_response = ResponsesAPIResponse(
id="resp_cursor123",
created_at=1234567890,
model="gpt-4o",
object="response",
output=[
ResponseOutputMessage(
type="message",
role="assistant",
content=[
ResponseOutputText(type="output_text", text="Hello from Cursor!")
],
)
],
)
mock_router.aresponses = AsyncMock(return_value=mock_responses_response)
client = TestClient(app)
# Test with Responses API input format (what Cursor sends)
test_data = {
"model": "gpt-4o",
"input": [{"role": "user", "content": "Hello"}],
}
response = client.post(
"/cursor/chat/completions",
json=test_data,
headers={"Authorization": "Bearer sk-1234"},
)
# Should return 200 (or 401/500 if auth fails)
assert response.status_code in [200, 401, 500]
# If successful, verify it returns chat completions format
if response.status_code == 200:
response_data = response.json()
# Should have chat completion structure
assert "choices" in response_data or "id" in response_data
# Should not have Responses API structure
assert "output" not in response_data or "status" not in response_data