diff --git a/docs/my-website/docs/providers/anthropic.md b/docs/my-website/docs/providers/anthropic.md
index 27c12232c56..792e57fc03f 100644
--- a/docs/my-website/docs/providers/anthropic.md
+++ b/docs/my-website/docs/providers/anthropic.md
@@ -60,11 +60,30 @@ export ANTHROPIC_API_KEY="your-api-key"
### 2. Start the proxy
+
+
+
```bash
$ litellm --model claude-3-opus-20240229
# Server running on http://0.0.0.0:4000
```
+
+
+
+```yaml
+model_list:
+ - model_name: claude-3 ### RECEIVED MODEL NAME ###
+ litellm_params: # all params accepted by litellm.completion() - https://docs.litellm.ai/docs/completion/input
+ model: claude-3-opus-20240229 ### MODEL NAME sent to `litellm.completion()` ###
+ api_key: "os.environ/ANTHROPIC_API_KEY" # does os.getenv("AZURE_API_KEY_EU")
+```
+
+```bash
+litellm --config /path/to/config.yaml
+```
+
+
### 3. Test it
@@ -76,7 +95,7 @@ $ litellm --model claude-3-opus-20240229
curl --location 'http://0.0.0.0:4000/chat/completions' \
--header 'Content-Type: application/json' \
--data ' {
- "model": "gpt-3.5-turbo",
+ "model": "claude-3",
"messages": [
{
"role": "user",
@@ -97,7 +116,7 @@ client = openai.OpenAI(
)
# request sent to model set on litellm proxy, `litellm --model`
-response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
+response = client.chat.completions.create(model="claude-3", messages = [
{
"role": "user",
"content": "this is a test request, write a short poem"
@@ -121,7 +140,7 @@ from langchain.schema import HumanMessage, SystemMessage
chat = ChatOpenAI(
openai_api_base="http://0.0.0.0:4000", # set openai_api_base to the LiteLLM Proxy
- model = "gpt-3.5-turbo",
+ model = "claude-3",
temperature=0.1
)
@@ -238,7 +257,7 @@ resp = litellm.completion(
print(f"\nResponse: {resp}")
```
-### Usage - "Assistant Pre-fill"
+## Usage - "Assistant Pre-fill"
You can "put words in Claude's mouth" by including an `assistant` role message as the last item in the `messages` array.
@@ -271,8 +290,8 @@ Human: How do you say 'Hello' in German? Return your answer as a JSON object, li
Assistant: {
```
-### Usage - "System" messages
-If you're using Anthropic's Claude 2.1 with Bedrock, `system` role messages are properly formatted for you.
+## Usage - "System" messages
+If you're using Anthropic's Claude 2.1, `system` role messages are properly formatted for you.
```python
import os
diff --git a/docs/my-website/docs/proxy/virtual_keys.md b/docs/my-website/docs/proxy/virtual_keys.md
index 525843cfd71..6ea101c5ce6 100644
--- a/docs/my-website/docs/proxy/virtual_keys.md
+++ b/docs/my-website/docs/proxy/virtual_keys.md
@@ -1,14 +1,14 @@
-# 🔑 Virtual Keys, Users
-Track Spend, Set budgets and create virtual keys for the proxy
-
-Grant other's temporary access to your proxy, with keys that expire after a set duration.
+import Tabs from '@theme/Tabs';
+import TabItem from '@theme/TabItem';
+# 🔑 Virtual Keys
+Track Spend, and control model access via virtual keys for the proxy
:::info
- 🔑 [UI to Generate, Edit, Delete Keys (with SSO)](https://docs.litellm.ai/docs/proxy/ui)
- [Deploy LiteLLM Proxy with Key Management](https://docs.litellm.ai/docs/proxy/deploy#deploy-with-database)
-- Dockerfile.database for LiteLLM Proxy + Key Management [here](https://github.com/BerriAI/litellm/blob/main/Dockerfile.database)
+- [Dockerfile.database for LiteLLM Proxy + Key Management](https://github.com/BerriAI/litellm/blob/main/Dockerfile.database)
:::
@@ -30,7 +30,7 @@ export DATABASE_URL=postgresql://:@:/
```
-You can then generate temporary keys by hitting the `/key/generate` endpoint.
+You can then generate keys by hitting the `/key/generate` endpoint.
[**See code**](https://github.com/BerriAI/litellm/blob/7a669a36d2689c7f7890bc9c93e04ff3c2641299/litellm/proxy/proxy_server.py#L672)
@@ -46,8 +46,8 @@ model_list:
model: ollama/llama2
general_settings:
- master_key: sk-1234 # [OPTIONAL] if set all calls to proxy will require either this key or a valid generated token
- database_url: "postgresql://:@:/"
+ master_key: sk-1234
+ database_url: "postgresql://:@:/" # 👈 KEY CHANGE
```
**Step 2: Start litellm**
@@ -56,62 +56,220 @@ general_settings:
litellm --config /path/to/config.yaml
```
-**Step 3: Generate temporary keys**
+**Step 3: Generate keys**
```shell
curl 'http://0.0.0.0:4000/key/generate' \
--header 'Authorization: Bearer ' \
--header 'Content-Type: application/json' \
---data-raw '{"models": ["gpt-3.5-turbo", "gpt-4", "claude-2"], "duration": "20m","metadata": {"user": "ishaan@berri.ai"}}'
+--data-raw '{"models": ["gpt-3.5-turbo", "gpt-4"], "metadata": {"user": "ishaan@berri.ai"}}'
```
+## Advanced - Spend Tracking
-## /key/generate
+Get spend per:
+- key - via `/key/info` [Swagger](https://litellm-api.up.railway.app/#/key%20management/info_key_fn_key_info_get)
+- user - via `/user/info` [Swagger](https://litellm-api.up.railway.app/#/user%20management/user_info_user_info_get)
+- team - via `/team/info` [Swagger](https://litellm-api.up.railway.app/#/team%20management/team_info_team_info_get)
+- ⏳ end-users - via `/end_user/info` - [Comment on this issue for end-user cost tracking](https://github.com/BerriAI/litellm/issues/2633)
-### Request
-```shell
-curl 'http://0.0.0.0:4000/key/generate' \
---header 'Authorization: Bearer ' \
---header 'Content-Type: application/json' \
---data-raw '{
- "models": ["gpt-3.5-turbo", "gpt-4", "claude-2"],
- "duration": "20m",
- "metadata": {"user": "ishaan@berri.ai"},
- "team_id": "core-infra",
- "max_budget": 10,
- "soft_budget": 5,
-}'
+**How is it calculated?**
+
+The cost per model is stored [here](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json) and calculated by the [`completion_cost`](https://github.com/BerriAI/litellm/blob/db7974f9f216ee50b53c53120d1e3fc064173b60/litellm/utils.py#L3771) function.
+
+**How is it tracking?**
+
+Spend is automatically tracked for the key in the "LiteLLM_VerificationTokenTable". If the key has an attached 'user_id' or 'team_id', the spend for that user is tracked in the "LiteLLM_UserTable", and team in the "LiteLLM_TeamTable".
+
+
+
+
+You can get spend for a key by using the `/key/info` endpoint.
+
+```bash
+curl 'http://0.0.0.0:4000/key/info?key=' \
+ -X GET \
+ -H 'Authorization: Bearer '
```
+This is automatically updated (in USD) when calls are made to /completions, /chat/completions, /embeddings using litellm's completion_cost() function. [**See Code**](https://github.com/BerriAI/litellm/blob/1a6ea20a0bb66491968907c2bfaabb7fe45fc064/litellm/utils.py#L1654).
-Request Params:
-
-- `duration`: *Optional[str]* - Specify the length of time the token is valid for. You can set duration as seconds ("30s"), minutes ("30m"), hours ("30h"), days ("30d").
-- `key_alias`: *Optional[str]* - User defined key alias
-- `team_id`: *Optional[str]* - The team id of the user
-- `models`: *Optional[list]* - Model_name's a user is allowed to call. (if empty, key is allowed to call all models)
-- `aliases`: *Optional[dict]* - Any alias mappings, on top of anything in the config.yaml model list. - https://docs.litellm.ai/docs/proxy/virtual_keys#managing-auth---upgradedowngrade-models
-- `config`: *Optional[dict]* - any key-specific configs, overrides config in config.yaml
-- `spend`: *Optional[int]* - Amount spent by key. Default is 0. Will be updated by proxy whenever key is used. https://docs.litellm.ai/docs/proxy/virtual_keys#managing-auth---tracking-spend
-- `max_budget`: *Optional[float]* - Specify max budget for a given key.
-- `soft_budget`: *Optional[float]* - Specify soft limit budget for a given key. Get Alerts when key hits its soft budget
-- `model_max_budget`: *Optional[dict[str, float]]* - Specify max budget for each model, `model_max_budget={"gpt4": 0.5, "gpt-5": 0.01}`
-- `max_parallel_requests`: *Optional[int]* - Rate limit a user based on the number of parallel requests. Raises 429 error, if user's parallel requests > x.
-- `metadata`: *Optional[dict]* - Metadata for key, store information for key. Example metadata = {"team": "core-infra", "app": "app2", "email": "ishaan@berri.ai" }
-
-
-### Response
+**Sample response**
```python
{
- "key": "sk-kdEXbIqZRwEeEiHwdg7sFA", # Bearer token
- "expires": "2023-11-19T01:38:25.834000+00:00" # datetime object
- "key_name": "sk-...7sFA" # abbreviated key string, ONLY stored in db if `allow_user_auth: true` set - [see](./ui.md)
- ...
+ "key": "sk-tXL0wt5-lOOVK9sfY2UacA",
+ "info": {
+ "token": "sk-tXL0wt5-lOOVK9sfY2UacA",
+ "spend": 0.0001065, # 👈 SPEND
+ "expires": "2023-11-24T23:19:11.131000Z",
+ "models": [
+ "gpt-3.5-turbo",
+ "gpt-4",
+ "claude-2"
+ ],
+ "aliases": {
+ "mistral-7b": "gpt-3.5-turbo"
+ },
+ "config": {}
+ }
}
```
-### Upgrade/Downgrade Models
+
+
+
+**1. Create a user**
+
+```bash
+curl --location 'http://localhost:4000/user/new' \
+--header 'Authorization: Bearer ' \
+--header 'Content-Type: application/json' \
+--data-raw '{user_email: "krrish@berri.ai"}'
+```
+
+**Expected Response**
+
+```bash
+{
+ ...
+ "expires": "2023-12-22T09:53:13.861000Z",
+ "user_id": "my-unique-id", # 👈 unique id
+ "max_budget": 0.0
+}
+```
+
+**2. Create a key for that user**
+
+```bash
+curl 'http://0.0.0.0:4000/key/generate' \
+--header 'Authorization: Bearer ' \
+--header 'Content-Type: application/json' \
+--data-raw '{"models": ["gpt-3.5-turbo", "gpt-4"], "user_id": "my-unique-id"}'
+```
+
+Returns a key - `sk-...`.
+
+**3. See spend for user**
+
+```bash
+curl 'http://0.0.0.0:4000/user/info?user_id=my-unique-id' \
+ -X GET \
+ -H 'Authorization: Bearer '
+```
+
+Expected Response
+
+```bash
+{
+ ...
+ "spend": 0 # 👈 SPEND
+}
+```
+
+
+
+
+Use teams, if you want keys to be owned by multiple people (e.g. for a production app).
+
+**1. Create a team**
+
+```bash
+curl --location 'http://localhost:4000/team/new' \
+--header 'Authorization: Bearer ' \
+--header 'Content-Type: application/json' \
+--data-raw '{"team_alias": "my-awesome-team"}'
+```
+
+**Expected Response**
+
+```bash
+{
+ ...
+ "expires": "2023-12-22T09:53:13.861000Z",
+ "team_id": "my-unique-id", # 👈 unique id
+ "max_budget": 0.0
+}
+```
+
+**2. Create a key for that team**
+
+```bash
+curl 'http://0.0.0.0:4000/key/generate' \
+--header 'Authorization: Bearer ' \
+--header 'Content-Type: application/json' \
+--data-raw '{"models": ["gpt-3.5-turbo", "gpt-4"], "team_id": "my-unique-id"}'
+```
+
+Returns a key - `sk-...`.
+
+**3. See spend for team**
+
+```bash
+curl 'http://0.0.0.0:4000/team/info?team_id=my-unique-id' \
+ -X GET \
+ -H 'Authorization: Bearer '
+```
+
+Expected Response
+
+```bash
+{
+ ...
+ "spend": 0 # 👈 SPEND
+}
+```
+
+
+
+
+## Advanced - Model Access
+
+### Restrict models by `team_id`
+`litellm-dev` can only access `azure-gpt-3.5`
+
+**1. Create a team via `/team/new`**
+```shell
+curl --location 'http://localhost:4000/team/new' \
+--header 'Authorization: Bearer ' \
+--header 'Content-Type: application/json' \
+--data-raw '{
+ "team_alias": "litellm-dev",
+ "models": ["azure-gpt-3.5"]
+}'
+
+# returns {...,"team_id": "my-unique-id"}
+```
+
+**2. Create a key for team**
+```shell
+curl --location 'http://localhost:4000/key/generate' \
+--header 'Authorization: Bearer sk-1234' \
+--header 'Content-Type: application/json' \
+--data-raw '{"team_id": "my-unique-id"}'
+```
+
+**3. Test it**
+```shell
+curl --location 'http://0.0.0.0:4000/chat/completions' \
+ --header 'Content-Type: application/json' \
+ --header 'Authorization: Bearer sk-qo992IjKOC2CHKZGRoJIGA' \
+ --data '{
+ "model": "BEDROCK_GROUP",
+ "messages": [
+ {
+ "role": "user",
+ "content": "hi"
+ }
+ ]
+ }'
+```
+
+```shell
+{"error":{"message":"Invalid model for team litellm-dev: BEDROCK_GROUP. Valid models for team are: ['azure-gpt-3.5']\n\n\nTraceback (most recent call last):\n File \"/Users/ishaanjaffer/Github/litellm/litellm/proxy/proxy_server.py\", line 2298, in chat_completion\n _is_valid_team_configs(\n File \"/Users/ishaanjaffer/Github/litellm/litellm/proxy/utils.py\", line 1296, in _is_valid_team_configs\n raise Exception(\nException: Invalid model for team litellm-dev: BEDROCK_GROUP. Valid models for team are: ['azure-gpt-3.5']\n\n","type":"None","param":"None","code":500}}%
+```
+
+### Model Aliases
If a user is expected to use a given model (i.e. gpt3-5), and you want to:
@@ -189,421 +347,9 @@ curl --location 'http://localhost:4000/key/generate' \
"max_budget": 0,}'
```
+## Advanced - Custom Auth
-## /key/info
-
-### Request
-```shell
-curl -X GET "http://0.0.0.0:4000/key/info?key=sk-02Wr4IAlN3NvPXvL5JVvDA" \
--H "Authorization: Bearer sk-1234"
-```
-
-Request Params:
-- key: str - The key you want the info for
-
-### Response
-
-`token` is the hashed key (The DB stores the hashed key for security)
-```json
-{
- "key": "sk-02Wr4IAlN3NvPXvL5JVvDA",
- "info": {
- "token": "80321a12d03412c527f2bd9db5fabd746abead2e1d50b435a534432fbaca9ef5",
- "spend": 0.0,
- "expires": "2024-01-18T23:52:09.125000+00:00",
- "models": ["azure-gpt-3.5", "azure-embedding-model"],
- "aliases": {},
- "config": {},
- "user_id": "ishaan2@berri.ai",
- "team_id": "None",
- "max_parallel_requests": null,
- "metadata": {}
- }
-}
-
-
-```
-
-## /key/update
-
-### Request
-```shell
-curl 'http://0.0.0.0:4000/key/update' \
---header 'Authorization: Bearer ' \
---header 'Content-Type: application/json' \
---data-raw '{
- "key": "sk-kdEXbIqZRwEeEiHwdg7sFA",
- "models": ["gpt-3.5-turbo", "gpt-4", "claude-2"],
- "metadata": {"user": "ishaan@berri.ai"},
- "team_id": "core-infra"
-}'
-```
-
-Request Params:
-- key: str - The key that needs to be updated.
-
-- models: list or null (optional) - Specify the models a token has access to. If null, then the token has access to all models on the server.
-
-- metadata: dict or null (optional) - Pass metadata for the updated token. If null, defaults to an empty dictionary.
-
-- team_id: str or null (optional) - Specify the team_id for the associated key.
-
-### Response
-
-```json
-{
- "key": "sk-kdEXbIqZRwEeEiHwdg7sFA",
- "models": ["gpt-3.5-turbo", "gpt-4", "claude-2"],
- "metadata": {
- "user": "ishaan@berri.ai"
- }
-}
-
-```
-
-
-## /key/delete
-
-### Request
-```shell
-curl 'http://0.0.0.0:4000/key/delete' \
---header 'Authorization: Bearer ' \
---header 'Content-Type: application/json' \
---data-raw '{
- "keys": ["sk-kdEXbIqZRwEeEiHwdg7sFA"]
-}'
-```
-
-Request Params:
-- keys: List[str] - List of keys to delete
-
-### Response
-
-```json
-{
- "deleted_keys": ["sk-kdEXbIqZRwEeEiHwdg7sFA"]
-}
-```
-
-## /user/new
-
-### Request
-
-All [key/generate params supported](#keygenerate) for creating a user
-```shell
-curl 'http://0.0.0.0:4000/user/new' \
---header 'Authorization: Bearer sk-1234' \
---header 'Content-Type: application/json' \
---data-raw '{
- "user_id": "ishaan1",
- "user_email": "ishaan@litellm.ai",
- "user_role": "admin",
- "team_id": "cto-team",
- "max_budget": 20,
- "budget_duration": "1h"
-
-}'
-```
-
-Request Params:
-
-- user_id: str (optional - defaults to uuid) - The unique identifier for the user.
-- user_email: str (optional - defaults to "") - The email address associated with the user.
-- user_role: str (optional - defaults to "app_user") - The role assigned to the user. Can be "admin", "app_owner", "app_user"
-
-**Possible `user_role` values**
-```
-"admin" - Maintaining the proxy and owning the overall budget
-"app_owner" - employees maintaining the apps, each owner may own more than one app
-"app_user" - users who know nothing about the proxy. These users get created when you pass `user` to /chat/completions
-```
-- team_id: str (optional - defaults to "") - The identifier for the team to which the user belongs.
-- max_budget: float (optional - defaults to `null`) - The maximum budget allocated for the user. No budget checks done if `max_budget==null`
-- budget_duration: str (optional - defaults to `null`) - The duration for which the budget is valid, e.g., "1h", "1d"
-
-### Response
-A key will be generated for the new user created
-
-```shell
-{
- "models": [],
- "spend": 0.0,
- "max_budget": null,
- "user_id": "ishaan1",
- "team_id": null,
- "max_parallel_requests": null,
- "metadata": {},
- "tpm_limit": null,
- "rpm_limit": null,
- "budget_duration": null,
- "allowed_cache_controls": [],
- "key_alias": null,
- "duration": null,
- "aliases": {},
- "config": {},
- "key": "sk-JflB33ucTqc2NYvNAgiBCA",
- "key_name": null,
- "expires": null
-}
-```
-
-
-## /user/info
-
-### Request
-
-#### View all Users
-If you're trying to view all users, we recommend using pagination with the following args
-- `view_all=true`
-- `page=0` Optional(int) min = 0, default=0
-- `page_size=25` Optional(int) min = 1, default = 25
-```shell
-curl -X GET "http://0.0.0.0:4000/user/info?view_all=true&page=0&page_size=25" -H "Authorization: Bearer sk-1234"
-```
-
-#### View specific user_id
-```shell
-curl -X GET "http://0.0.0.0:4000/user/info?user_id=228da235-eef0-4c30-bf53-5d6ac0d278c2" -H "Authorization: Bearer sk-1234"
-```
-
-### Response
-View user spend, budget, models, keys and teams
-
-```json
-{
- "user_id": "228da235-eef0-4c30-bf53-5d6ac0d278c2",
- "user_info": {
- "user_id": "228da235-eef0-4c30-bf53-5d6ac0d278c2",
- "team_id": null,
- "teams": [],
- "user_role": "app_user",
- "max_budget": null,
- "spend": 200000.0,
- "user_email": null,
- "models": [],
- "max_parallel_requests": null,
- "tpm_limit": null,
- "rpm_limit": null,
- "budget_duration": null,
- "budget_reset_at": null,
- "allowed_cache_controls": [],
- "model_spend": {
- "chatgpt-v-2": 200000
- },
- "model_max_budget": {}
- },
- "keys": [
- {
- "token": "16c337f9df00a0e6472627e39a2ed02e67bc9a8a760c983c4e9b8cad7954f3c0",
- "key_name": null,
- "key_alias": null,
- "spend": 200000.0,
- "expires": null,
- "models": [],
- "aliases": {},
- "config": {},
- "user_id": "228da235-eef0-4c30-bf53-5d6ac0d278c2",
- "team_id": null,
- "permissions": {},
- "max_parallel_requests": null,
- "metadata": {},
- "tpm_limit": null,
- "rpm_limit": null,
- "max_budget": null,
- "budget_duration": null,
- "budget_reset_at": null,
- "allowed_cache_controls": [],
- "model_spend": {
- "chatgpt-v-2": 200000
- },
- "model_max_budget": {}
- }
- ],
- "teams": []
-}
-
-```
-
-## Advanced
-### Upperbound /key/generate params
-Use this, if you need to control the upperbound that users can use for `max_budget`, `budget_duration` or any `key/generate` param per key.
-
-Set `litellm_settings:upperbound_key_generate_params`:
-```yaml
-litellm_settings:
- upperbound_key_generate_params:
- max_budget: 100 # upperbound of $100, for all /key/generate requests
- duration: "30d" # upperbound of 30 days for all /key/generate requests
-```
-
-** Expected Behavior **
-
-- Send a `/key/generate` request with `max_budget=200`
-- Key will be created with `max_budget=100` since 100 is the upper bound
-
-### Default /key/generate params
-Use this, if you need to control the default `max_budget` or any `key/generate` param per key.
-
-When a `/key/generate` request does not specify `max_budget`, it will use the `max_budget` specified in `default_key_generate_params`
-
-Set `litellm_settings:default_key_generate_params`:
-```yaml
-litellm_settings:
- default_key_generate_params:
- max_budget: 1.5000
- models: ["azure-gpt-3.5"]
- duration: # blank means `null`
- metadata: {"setting":"default"}
- team_id: "core-infra"
-```
-
-### Restrict models by `team_id`
-`litellm-dev` can only access `azure-gpt-3.5`
-
-```yaml
-litellm_settings:
- default_team_settings:
- - team_id: litellm-dev
- models: ["azure-gpt-3.5"]
-```
-
-#### Create key with team_id="litellm-dev"
-```shell
-curl --location 'http://localhost:4000/key/generate' \
---header 'Authorization: Bearer sk-1234' \
---header 'Content-Type: application/json' \
---data-raw '{"team_id": "litellm-dev"}'
-```
-
-#### Use Key to call invalid model - Fails
-```shell
-curl --location 'http://0.0.0.0:4000/chat/completions' \
- --header 'Content-Type: application/json' \
- --header 'Authorization: Bearer sk-qo992IjKOC2CHKZGRoJIGA' \
- --data '{
- "model": "BEDROCK_GROUP",
- "messages": [
- {
- "role": "user",
- "content": "hi"
- }
- ]
- }'
-```
-
-```shell
-{"error":{"message":"Invalid model for team litellm-dev: BEDROCK_GROUP. Valid models for team are: ['azure-gpt-3.5']\n\n\nTraceback (most recent call last):\n File \"/Users/ishaanjaffer/Github/litellm/litellm/proxy/proxy_server.py\", line 2298, in chat_completion\n _is_valid_team_configs(\n File \"/Users/ishaanjaffer/Github/litellm/litellm/proxy/utils.py\", line 1296, in _is_valid_team_configs\n raise Exception(\nException: Invalid model for team litellm-dev: BEDROCK_GROUP. Valid models for team are: ['azure-gpt-3.5']\n\n","type":"None","param":"None","code":500}}%
-```
-
-### Set Budgets - Per Key
-
-Set `max_budget` in (USD $) param in the `key/generate` request. By default the `max_budget` is set to `null` and is not checked for keys
-
-```shell
-curl 'http://0.0.0.0:4000/key/generate' \
---header 'Authorization: Bearer ' \
---header 'Content-Type: application/json' \
---data-raw '{
- "metadata": {"user": "ishaan@berri.ai"},
- "team_id": "core-infra",
- "max_budget": 10,
-}'
-```
-
-#### Expected Behaviour
-- Costs Per key get auto-populated in `LiteLLM_VerificationToken` Table
-- After the key crosses it's `max_budget`, requests fail
-
-Example Request to `/chat/completions` when key has crossed budget
-
-```shell
-curl --location 'http://0.0.0.0:4000/chat/completions' \
- --header 'Content-Type: application/json' \
- --header 'Authorization: Bearer sk-ULl_IKCVFy2EZRzQB16RUA' \
- --data ' {
- "model": "azure-gpt-3.5",
- "user": "e09b4da8-ed80-4b05-ac93-e16d9eb56fca",
- "messages": [
- {
- "role": "user",
- "content": "respond in 50 lines"
- }
- ],
-}'
-```
-
-
-Expected Response from `/chat/completions` when key has crossed budget
-```shell
-{
- "detail":"Authentication Error, ExceededTokenBudget: Current spend for token: 7.2e-05; Max Budget for Token: 2e-07"
-}
-```
-
-
-### Set Budgets - Per User
-
-LiteLLM exposes a `/user/new` endpoint to create budgets for users, that persist across multiple keys.
-
-This is documented in the swagger (live on your server root endpoint - e.g. `http://0.0.0.0:4000/`). Here's an example request.
-
-```shell
-curl --location 'http://localhost:4000/user/new' \
---header 'Authorization: Bearer ' \
---header 'Content-Type: application/json' \
---data-raw '{"models": ["azure-models"], "max_budget": 0, "user_id": "krrish3@berri.ai"}'
-```
-The request is a normal `/key/generate` request body + a `max_budget` field.
-
-**Sample Response**
-
-```shell
-{
- "key": "sk-YF2OxDbrgd1y2KgwxmEA2w",
- "expires": "2023-12-22T09:53:13.861000Z",
- "user_id": "krrish3@berri.ai",
- "max_budget": 0.0
-}
-```
-
-### Tracking Spend
-
-You can get spend for a key by using the `/key/info` endpoint.
-
-```bash
-curl 'http://0.0.0.0:4000/key/info?key=' \
- -X GET \
- -H 'Authorization: Bearer '
-```
-
-This is automatically updated (in USD) when calls are made to /completions, /chat/completions, /embeddings using litellm's completion_cost() function. [**See Code**](https://github.com/BerriAI/litellm/blob/1a6ea20a0bb66491968907c2bfaabb7fe45fc064/litellm/utils.py#L1654).
-
-**Sample response**
-
-```python
-{
- "key": "sk-tXL0wt5-lOOVK9sfY2UacA",
- "info": {
- "token": "sk-tXL0wt5-lOOVK9sfY2UacA",
- "spend": 0.0001065,
- "expires": "2023-11-24T23:19:11.131000Z",
- "models": [
- "gpt-3.5-turbo",
- "gpt-4",
- "claude-2"
- ],
- "aliases": {
- "mistral-7b": "gpt-3.5-turbo"
- },
- "config": {}
- }
-}
-```
-
-
-### Custom Auth
-
-You can now override the default api key auth.
+You can now override the default api key auth.
Here's how:
@@ -737,4 +483,56 @@ litellm_settings:
general_settings:
custom_key_generate: custom_auth.custom_generate_key_fn
-```
\ No newline at end of file
+```
+
+
+## Upperbound /key/generate params
+Use this, if you need to set default upperbounds for `max_budget`, `budget_duration` or any `key/generate` param per key.
+
+Set `litellm_settings:upperbound_key_generate_params`:
+```yaml
+litellm_settings:
+ upperbound_key_generate_params:
+ max_budget: 100 # upperbound of $100, for all /key/generate requests
+ duration: "30d" # upperbound of 30 days for all /key/generate requests
+```
+
+** Expected Behavior **
+
+- Send a `/key/generate` request with `max_budget=200`
+- Key will be created with `max_budget=100` since 100 is the upper bound
+
+## Default /key/generate params
+Use this, if you need to control the default `max_budget` or any `key/generate` param per key.
+
+When a `/key/generate` request does not specify `max_budget`, it will use the `max_budget` specified in `default_key_generate_params`
+
+Set `litellm_settings:default_key_generate_params`:
+```yaml
+litellm_settings:
+ default_key_generate_params:
+ max_budget: 1.5000
+ models: ["azure-gpt-3.5"]
+ duration: # blank means `null`
+ metadata: {"setting":"default"}
+ team_id: "core-infra"
+```
+
+## Endpoints
+
+### Keys
+
+#### [**👉 API REFERENCE DOCS**](https://litellm-api.up.railway.app/#/key%20management/)
+
+### Users
+
+#### [**👉 API REFERENCE DOCS**](https://litellm-api.up.railway.app/#/user%20management/)
+
+
+### Teams
+
+#### [**👉 API REFERENCE DOCS**](https://litellm-api.up.railway.app/#/team%20management)
+
+
+
+