diff --git a/docs/my-website/docs/completion/knowledgebase.md b/docs/my-website/docs/completion/knowledgebase.md index 35bd1942b9d..ee0e3086785 100644 --- a/docs/my-website/docs/completion/knowledgebase.md +++ b/docs/my-website/docs/completion/knowledgebase.md @@ -160,6 +160,129 @@ print(response.choices[0].message.content) +## Provider Specific Guides + +This section covers how to add your vector stores to LiteLLM. If you want support for a new provider, please file an issue [here](https://github.com/BerriAI/litellm/issues). + +### Bedrock Knowledge Bases + +**1. Set up your Bedrock Knowledge Base** + +Ensure you have a Bedrock Knowledge Base created in your AWS account with the appropriate permissions configured. + +**2. Add to LiteLLM UI** + +1. Navigate to **Tools > Vector Stores > "Add new vector store"** +2. Select **"Bedrock"** as the provider +3. Enter your Bedrock Knowledge Base ID in the **"Vector Store ID"** field + + + + +### Vertex AI RAG Engine + +**1. Get your Vertex AI RAG Engine ID** + +1. Navigate to your RAG Engine Corpus in the [Google Cloud Console](https://console.cloud.google.com/vertex-ai/rag/corpus) +2. Select the **RAG Engine** you want to integrate with LiteLLM + +
+ +
+ +3. Click the **"Details"** button and copy the UUID for the RAG Engine +4. The ID should look like: `6917529027641081856` + +
+ +
+ +**2. Add to LiteLLM UI** + +1. Navigate to **Tools > Vector Stores > "Add new vector store"** +2. Select **"Vertex AI RAG Engine"** as the provider +3. Enter your Vertex AI RAG Engine ID in the **"Vector Store ID"** field + +
+ +
+ +### PG Vector + +**1. Deploy the litellm-pg-vector-store connector** + +LiteLLM provides a server that exposes OpenAI-compatible `vector_store` endpoints for PG Vector. The LiteLLM Proxy server connects to your deployed service and uses it as a vector store when querying. + +1. Follow the deployment instructions for the litellm-pg-vector-store connector [here](https://github.com/BerriAI/litellm-pgvector) +2. For detailed configuration options, see the [configuration guide](https://github.com/BerriAI/litellm-pgvector?tab=readme-ov-file#configuration) + +**Example .env configuration for deploying litellm-pg-vector-store:** + +```env +DATABASE_URL="postgresql://neondb_owner:xxxx" +SERVER_API_KEY="sk-1234" +HOST="0.0.0.0" +PORT=8001 +EMBEDDING__MODEL="text-embedding-ada-002" +EMBEDDING__BASE_URL="http://localhost:4000" +EMBEDDING__API_KEY="sk-1234" +EMBEDDING__DIMENSIONS=1536 +DB_FIELDS__ID_FIELD="id" +DB_FIELDS__CONTENT_FIELD="content" +DB_FIELDS__METADATA_FIELD="metadata" +DB_FIELDS__EMBEDDING_FIELD="embedding" +DB_FIELDS__VECTOR_STORE_ID_FIELD="vector_store_id" +DB_FIELDS__CREATED_AT_FIELD="created_at" +``` + +**2. Add to LiteLLM UI** + +Once your litellm-pg-vector-store is deployed: + +1. Navigate to **Tools > Vector Stores > "Add new vector store"** +2. Select **"PG Vector"** as the provider +3. Enter your **API Base URL** and **API Key** for your `litellm-pg-vector-store` container + - The API Key field corresponds to the `SERVER_API_KEY` from your .env configuration + +
+ +
+ +### OpenAI Vector Stores + +**1. Set up your OpenAI Vector Store** + +1. Create your Vector Store on the [OpenAI platform](https://platform.openai.com/storage/vector_stores) +2. Note your Vector Store ID (format: `vs_687ae3b2439881918b433cb99d10662e`) + +**2. Add to LiteLLM UI** + +1. Navigate to **Tools > Vector Stores > "Add new vector store"** +2. Select **"OpenAI"** as the provider +3. Enter your **Vector Store ID** in the corresponding field +4. Enter your **OpenAI API Key** in the API Key field + +
+ +
diff --git a/docs/my-website/docs/vector_stores/search.md b/docs/my-website/docs/vector_stores/search.md index c1dda06a608..5c3d02be3da 100644 --- a/docs/my-website/docs/vector_stores/search.md +++ b/docs/my-website/docs/vector_stores/search.md @@ -1,7 +1,7 @@ import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; -# /vector_stores/\{vector_store_id\}/search - Search Vector Store +# /vector_stores/search - Search Vector Store Search a vector store for relevant chunks based on a query and file attributes filter. This is useful for retrieval-augmented generation (RAG) use cases. @@ -175,194 +175,14 @@ curl -L -X POST 'http://0.0.0.0:4000/v1/vector_stores/vs_abc123/search' \ -### OpenAI SDK (Standalone) +## Setting Up Vector Stores - - +To use vector store search, configure your vector stores in the `vector_store_registry`. See the [Vector Store Configuration Guide](../completion/knowledgebase.md) for: -```python showLineNumbers title="OpenAI SDK Direct" -from openai import OpenAI +- Provider-specific configuration (Bedrock, OpenAI, Azure, Vertex AI, PG Vector) +- Python SDK and Proxy setup examples +- Authentication and credential management -client = OpenAI(api_key="your-openai-api-key") +## Using Vector Stores with Chat Completions -search_results = client.beta.vector_stores.search( - vector_store_id="vs_abc123", - query="What is the capital of France?", - max_num_results=5 -) -print(search_results) -``` - - - - -## Request Format - -The request body follows OpenAI's vector stores search API format. - -#### Example request body - -```json -{ - "query": "What is the capital of France?", - "filters": { - "file_ids": ["file-abc123", "file-def456"] - }, - "max_num_results": 5, - "ranking_options": { - "score_threshold": 0.7 - }, - "rewrite_query": true -} -``` - -#### Required Fields -- **query** (string or array of strings): A query string or array for the search. The query is used to find relevant chunks in the vector store. - -#### Optional Fields -- **filters** (object): Optional filter to apply based on file attributes. - - **file_ids** (array of strings): Filter chunks based on specific file IDs. -- **max_num_results** (integer): Maximum number of results to return. Must be between 1 and 50. Default is 10. -- **ranking_options** (object): Optional ranking options for search. - - **score_threshold** (number): Minimum similarity score threshold for results. -- **rewrite_query** (boolean): Whether to rewrite the natural language query for vector search optimization. Default is true. - -## Response Format - -#### Example Response - -```json -{ - "object": "vector_store.search_results.page", - "search_query": "What is the capital of France?", - "data": [ - { - "score": 0.95, - "content": [ - { - "type": "text", - "text": "Paris is the capital and most populous city of France. With an official estimated population of 2,102,650 residents as of 1 January 2023 in an area of more than 105 km², Paris is the fourth-most populated city in the European Union and the 30th most densely populated city in the world in 2022." - } - ] - }, - { - "score": 0.87, - "content": [ - { - "type": "text", - "text": "France, officially the French Republic, is a country located primarily in Western Europe. Its capital is Paris, one of the most important cultural and economic centers in Europe." - } - ] - } - ] -} -``` - -#### Response Fields - -- **object** (string): The object type, which is always `vector_store.search_results.page`. -- **search_query** (string): The query that was used for the search. -- **data** (array): An array of search result objects. - - **score** (number): The similarity score of the search result, typically between 0 and 1, where 1 is the most similar. - - **content** (array): Array of content objects containing the retrieved text. - - **type** (string): The type of content, typically `text`. - - **text** (string): The actual text content that was retrieved from the vector store. - -## Mock Response Testing - -For testing purposes, you can use mock responses: - -```python showLineNumbers title="Mock Response Example" -import litellm - -# Mock response for testing -mock_results = [ - { - "score": 0.95, - "content": [ - { - "text": "Paris is the capital of France.", - "type": "text" - } - ] - }, - { - "score": 0.87, - "content": [ - { - "text": "France is a country in Western Europe.", - "type": "text" - } - ] - } -] - -response = await litellm.vector_stores.asearch( - vector_store_id="vs_abc123", - query="What is the capital of France?", - mock_response=mock_results -) -print(response) -``` - -## Error Handling - -Common errors you might encounter: - -```python showLineNumbers title="Error Handling Example" -import litellm - -try: - response = await litellm.vector_stores.asearch( - vector_store_id="vs_invalid", - query="What is the capital of France?" - ) -except litellm.NotFoundError as e: - print(f"Vector store not found: {e}") -except litellm.RateLimitError as e: - print(f"Rate limit exceeded: {e}") -except Exception as e: - print(f"Unexpected error: {e}") -``` - -## Best Practices - -1. **Query Optimization**: Use clear, specific queries for better search results. -2. **Result Filtering**: Use file_ids filter to limit search scope when needed. -3. **Score Thresholds**: Set appropriate score thresholds to filter out irrelevant results. -4. **Batch Queries**: Use array queries when searching for multiple related topics. -5. **Error Handling**: Always implement proper error handling for production use. - -```python showLineNumbers title="Best Practices Example" -import litellm - -async def search_documents(vector_store_id: str, user_query: str): - """ - Search documents with best practices applied - """ - try: - response = await litellm.vector_stores.asearch( - vector_store_id=vector_store_id, - query=user_query, - max_num_results=5, - ranking_options={ - "score_threshold": 0.7 # Filter out low-relevance results - }, - rewrite_query=True # Optimize query for vector search - ) - - # Filter results by score for additional quality control - high_quality_results = [ - result for result in response.data - if result.score >= 0.8 - ] - - return high_quality_results - - except Exception as e: - print(f"Search failed: {e}") - return [] - -# Usage -results = await search_documents("vs_abc123", "What is the capital of France?") -``` \ No newline at end of file +Pass `vector_store_ids` in chat completion requests to automatically retrieve relevant context. See [Using Vector Stores with Chat Completions](../completion/knowledgebase.md#2-make-a-request-with-vector_store_ids-parameter) for implementation details. \ No newline at end of file diff --git a/docs/my-website/img/kb_openai1.png b/docs/my-website/img/kb_openai1.png new file mode 100644 index 00000000000..8b5b92b7940 Binary files /dev/null and b/docs/my-website/img/kb_openai1.png differ diff --git a/docs/my-website/img/kb_pg1.png b/docs/my-website/img/kb_pg1.png new file mode 100644 index 00000000000..c5d7331f6a6 Binary files /dev/null and b/docs/my-website/img/kb_pg1.png differ diff --git a/docs/my-website/img/kb_vertex1.png b/docs/my-website/img/kb_vertex1.png new file mode 100644 index 00000000000..16dbb4b992f Binary files /dev/null and b/docs/my-website/img/kb_vertex1.png differ diff --git a/docs/my-website/img/kb_vertex2.png b/docs/my-website/img/kb_vertex2.png new file mode 100644 index 00000000000..4606008091b Binary files /dev/null and b/docs/my-website/img/kb_vertex2.png differ diff --git a/docs/my-website/img/kb_vertex3.png b/docs/my-website/img/kb_vertex3.png new file mode 100644 index 00000000000..1329c47433f Binary files /dev/null and b/docs/my-website/img/kb_vertex3.png differ diff --git a/docs/my-website/sidebars.js b/docs/my-website/sidebars.js index fb70bcc3695..3065ef4fb8d 100644 --- a/docs/my-website/sidebars.js +++ b/docs/my-website/sidebars.js @@ -282,7 +282,6 @@ const sidebars = { type: "category", label: "/vector_stores", items: [ - "vector_stores/create", "vector_stores/search", ] },