docs - vector stores (#12781)

* docs vertex vector store

* guide for using other non openai providers

* docs polish

* docs KBs

* docs search endpoint

* docs vector stores
This commit is contained in:
Ishaan Jaff 2025-07-19 17:07:44 -07:00 • committed by GitHub
parent eb7e50b1f2
commit 2cf4d164fc
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
8 changed files with 131 additions and 189 deletions

View file

@ -160,6 +160,129 @@ print(response.choices[0].message.content)
</TabItem>
</Tabs>
## Provider Specific Guides
This section covers how to add your vector stores to LiteLLM. If you want support for a new provider, please file an issue [here](https://github.com/BerriAI/litellm/issues).
### Bedrock Knowledge Bases
**1. Set up your Bedrock Knowledge Base**
Ensure you have a Bedrock Knowledge Base created in your AWS account with the appropriate permissions configured.
**2. Add to LiteLLM UI**
1. Navigate to **Tools > Vector Stores > "Add new vector store"**
2. Select **"Bedrock"** as the provider
3. Enter your Bedrock Knowledge Base ID in the **"Vector Store ID"** field
<Image
img={require('../../img/kb_2.png')}
style={{width: '60%', display: 'block'}}
/>
### Vertex AI RAG Engine
**1. Get your Vertex AI RAG Engine ID**
1. Navigate to your RAG Engine Corpus in the [Google Cloud Console](https://console.cloud.google.com/vertex-ai/rag/corpus)
2. Select the **RAG Engine** you want to integrate with LiteLLM
<div style={{margin: '20px 0', padding: '10px', border: '1px solid #ddd', borderRadius: '8px', display: 'inline-block', boxShadow: '0 2px 8px rgba(0,0,0,0.1)'}}>
<Image
img={require('../../img/kb_vertex1.png')}
style={{width: '60%', display: 'block'}}
/>
</div>
3. Click the **"Details"** button and copy the UUID for the RAG Engine
4. The ID should look like: `6917529027641081856`
<div style={{margin: '20px 0', padding: '10px', border: '1px solid #ddd', borderRadius: '8px', display: 'inline-block', boxShadow: '0 2px 8px rgba(0,0,0,0.1)'}}>
<Image
img={require('../../img/kb_vertex2.png')}
style={{width: '60%', display: 'block'}}
/>
</div>
**2. Add to LiteLLM UI**
1. Navigate to **Tools > Vector Stores > "Add new vector store"**
2. Select **"Vertex AI RAG Engine"** as the provider
3. Enter your Vertex AI RAG Engine ID in the **"Vector Store ID"** field
<div style={{margin: '20px 0', padding: '10px', border: '1px solid #ddd', borderRadius: '8px', display: 'inline-block', boxShadow: '0 2px 8px rgba(0,0,0,0.1)'}}>
<Image
img={require('../../img/kb_vertex3.png')}
style={{width: '60%', display: 'block'}}
/>
</div>
### PG Vector
**1. Deploy the litellm-pg-vector-store connector**
LiteLLM provides a server that exposes OpenAI-compatible `vector_store` endpoints for PG Vector. The LiteLLM Proxy server connects to your deployed service and uses it as a vector store when querying.
1. Follow the deployment instructions for the litellm-pg-vector-store connector [here](https://github.com/BerriAI/litellm-pgvector)
2. For detailed configuration options, see the [configuration guide](https://github.com/BerriAI/litellm-pgvector?tab=readme-ov-file#configuration)
**Example .env configuration for deploying litellm-pg-vector-store:**
```env
DATABASE_URL="postgresql://neondb_owner:xxxx"
SERVER_API_KEY="sk-1234"
HOST="0.0.0.0"
PORT=8001
EMBEDDING__MODEL="text-embedding-ada-002"
EMBEDDING__BASE_URL="http://localhost:4000"
EMBEDDING__API_KEY="sk-1234"
EMBEDDING__DIMENSIONS=1536
DB_FIELDS__ID_FIELD="id"
DB_FIELDS__CONTENT_FIELD="content"
DB_FIELDS__METADATA_FIELD="metadata"
DB_FIELDS__EMBEDDING_FIELD="embedding"
DB_FIELDS__VECTOR_STORE_ID_FIELD="vector_store_id"
DB_FIELDS__CREATED_AT_FIELD="created_at"
```
**2. Add to LiteLLM UI**
Once your litellm-pg-vector-store is deployed:
1. Navigate to **Tools > Vector Stores > "Add new vector store"**
2. Select **"PG Vector"** as the provider
3. Enter your **API Base URL** and **API Key** for your `litellm-pg-vector-store` container
- The API Key field corresponds to the `SERVER_API_KEY` from your .env configuration
<div style={{margin: '20px 0', padding: '10px', border: '1px solid #ddd', borderRadius: '8px', display: 'inline-block', boxShadow: '0 2px 8px rgba(0,0,0,0.1)'}}>
<Image
img={require('../../img/kb_pg1.png')}
style={{width: '60%', display: 'block'}}
/>
</div>
### OpenAI Vector Stores
**1. Set up your OpenAI Vector Store**
1. Create your Vector Store on the [OpenAI platform](https://platform.openai.com/storage/vector_stores)
2. Note your Vector Store ID (format: `vs_687ae3b2439881918b433cb99d10662e`)
**2. Add to LiteLLM UI**
1. Navigate to **Tools > Vector Stores > "Add new vector store"**
2. Select **"OpenAI"** as the provider
3. Enter your **Vector Store ID** in the corresponding field
4. Enter your **OpenAI API Key** in the API Key field
<div style={{margin: '20px 0', padding: '10px', border: '1px solid #ddd', borderRadius: '8px', display: 'inline-block', boxShadow: '0 2px 8px rgba(0,0,0,0.1)'}}>
<Image
img={require('../../img/kb_openai1.png')}
style={{width: '60%', display: 'block'}}
/>
</div>

View file

@ -1,7 +1,7 @@
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
# /vector_stores/\{vector_store_id\}/search - Search Vector Store
# /vector_stores/search - Search Vector Store
Search a vector store for relevant chunks based on a query and file attributes filter. This is useful for retrieval-augmented generation (RAG) use cases.
@ -175,194 +175,14 @@ curl -L -X POST 'http://0.0.0.0:4000/v1/vector_stores/vs_abc123/search' \
</TabItem>
</Tabs>
### OpenAI SDK (Standalone)
## Setting Up Vector Stores
<Tabs>
<TabItem value="openai-direct" label="Direct OpenAI Usage">
To use vector store search, configure your vector stores in the `vector_store_registry`. See the [Vector Store Configuration Guide](../completion/knowledgebase.md) for:
```python showLineNumbers title="OpenAI SDK Direct"
from openai import OpenAI
- Provider-specific configuration (Bedrock, OpenAI, Azure, Vertex AI, PG Vector)
- Python SDK and Proxy setup examples
- Authentication and credential management
client = OpenAI(api_key="your-openai-api-key")
## Using Vector Stores with Chat Completions
search_results = client.beta.vector_stores.search(
vector_store_id="vs_abc123",
query="What is the capital of France?",
max_num_results=5
)
print(search_results)
```
</TabItem>
</Tabs>
## Request Format
The request body follows OpenAI's vector stores search API format.
#### Example request body
```json
{
"query": "What is the capital of France?",
"filters": {
"file_ids": ["file-abc123", "file-def456"]
},
"max_num_results": 5,
"ranking_options": {
"score_threshold": 0.7
},
"rewrite_query": true
}
```
#### Required Fields
- **query** (string or array of strings): A query string or array for the search. The query is used to find relevant chunks in the vector store.
#### Optional Fields
- **filters** (object): Optional filter to apply based on file attributes.
- **file_ids** (array of strings): Filter chunks based on specific file IDs.
- **max_num_results** (integer): Maximum number of results to return. Must be between 1 and 50. Default is 10.
- **ranking_options** (object): Optional ranking options for search.
- **score_threshold** (number): Minimum similarity score threshold for results.
- **rewrite_query** (boolean): Whether to rewrite the natural language query for vector search optimization. Default is true.
## Response Format
#### Example Response
```json
{
"object": "vector_store.search_results.page",
"search_query": "What is the capital of France?",
"data": [
{
"score": 0.95,
"content": [
{
"type": "text",
"text": "Paris is the capital and most populous city of France. With an official estimated population of 2,102,650 residents as of 1 January 2023 in an area of more than 105 km², Paris is the fourth-most populated city in the European Union and the 30th most densely populated city in the world in 2022."
}
]
},
{
"score": 0.87,
"content": [
{
"type": "text",
"text": "France, officially the French Republic, is a country located primarily in Western Europe. Its capital is Paris, one of the most important cultural and economic centers in Europe."
}
]
}
]
}
```
#### Response Fields
- **object** (string): The object type, which is always `vector_store.search_results.page`.
- **search_query** (string): The query that was used for the search.
- **data** (array): An array of search result objects.
- **score** (number): The similarity score of the search result, typically between 0 and 1, where 1 is the most similar.
- **content** (array): Array of content objects containing the retrieved text.
- **type** (string): The type of content, typically `text`.
- **text** (string): The actual text content that was retrieved from the vector store.
## Mock Response Testing
For testing purposes, you can use mock responses:
```python showLineNumbers title="Mock Response Example"
import litellm
# Mock response for testing
mock_results = [
{
"score": 0.95,
"content": [
{
"text": "Paris is the capital of France.",
"type": "text"
}
]
},
{
"score": 0.87,
"content": [
{
"text": "France is a country in Western Europe.",
"type": "text"
}
]
}
]
response = await litellm.vector_stores.asearch(
vector_store_id="vs_abc123",
query="What is the capital of France?",
mock_response=mock_results
)
print(response)
```
## Error Handling
Common errors you might encounter:
```python showLineNumbers title="Error Handling Example"
import litellm
try:
response = await litellm.vector_stores.asearch(
vector_store_id="vs_invalid",
query="What is the capital of France?"
)
except litellm.NotFoundError as e:
print(f"Vector store not found: {e}")
except litellm.RateLimitError as e:
print(f"Rate limit exceeded: {e}")
except Exception as e:
print(f"Unexpected error: {e}")
```
## Best Practices
1. **Query Optimization**: Use clear, specific queries for better search results.
2. **Result Filtering**: Use file_ids filter to limit search scope when needed.
3. **Score Thresholds**: Set appropriate score thresholds to filter out irrelevant results.
4. **Batch Queries**: Use array queries when searching for multiple related topics.
5. **Error Handling**: Always implement proper error handling for production use.
```python showLineNumbers title="Best Practices Example"
import litellm
async def search_documents(vector_store_id: str, user_query: str):
"""
Search documents with best practices applied
"""
try:
response = await litellm.vector_stores.asearch(
vector_store_id=vector_store_id,
query=user_query,
max_num_results=5,
ranking_options={
"score_threshold": 0.7 # Filter out low-relevance results
},
rewrite_query=True # Optimize query for vector search
)
# Filter results by score for additional quality control
high_quality_results = [
result for result in response.data
if result.score >= 0.8
]
return high_quality_results
except Exception as e:
print(f"Search failed: {e}")
return []
# Usage
results = await search_documents("vs_abc123", "What is the capital of France?")
```
Pass `vector_store_ids` in chat completion requests to automatically retrieve relevant context. See [Using Vector Stores with Chat Completions](../completion/knowledgebase.md#2-make-a-request-with-vector_store_ids-parameter) for implementation details.

Binary file not shown.

After

Width:  |  Height:  |  Size: 198 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 718 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 683 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 509 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 375 KiB

View file

@ -282,7 +282,6 @@ const sidebars = {
type: "category",
label: "/vector_stores",
items: [
"vector_stores/create",
"vector_stores/search",
]
},