mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-02 02:11:58 +00:00
* feat(langfuse): migrate the sdk callback to langfuse v4
Replace the v2 trace()/generation()/span() calls with SDK v4 observations exported over OpenTelemetry, with one isolated tracer provider per Langfuse credential set, a discarding exporter for mock mode, and v4 trace and observation id normalization. Keeps the session-header trace provenance logic from main so each call under a session alias still gets its own trace
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): drop the always-true prompt client check now that v4 get_prompt is non-optional
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langfuse): isolate the e2e sync test from cached clients and log the real sdk major
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): type the slack trace-url lookup and drop dead v2 test shims
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(slack): cover the langfuse trace url built from the logger host
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* build(docker): pin langfuse to the locked 4.15.2 in the pip image
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): hash all-zero trace and observation ids instead of passing them through
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(langfuse): honour caller generation ids and assert v4 OTLP exports in legacy tests
v2 accepted generation(id=...). v4 derives the observation id from the OTel
span id, so the isolated tracer provider now carries an id generator that
hands out the id start_generation asked for through a context variable, and
the callback passes the resolved generation_id metadata into it.
The legacy e2e suite patched httpx.Client.post and compared v2 ingestion
batches; it now patches requests.Session.post, decodes the OTLP protobuf
and compares the exported generation against regenerated fixtures. The
local readback test replaces the removed get_generations() with
api.observations.get_many() and polls Langfuse Cloud instead of sleeping.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): read the sdk version header from package metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): propagate trace_metadata as trace-level attributes in v4
v2 wrote trace(metadata=...) onto the trace object. In v4 the trace only
carries what the observations propagate, so a continuation request with
update_trace_keys=["trace_metadata"] updated the generation's metadata
while the trace kept its stale values. Coerce each entry to the SDK's
string limit and hand it to propagate_attributes(metadata=...).
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): propagate interrupts raised during deferred client teardown
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): honor ssl_verify=False and SSL_VERIFY on the v4 OTLP exporter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): fall back to the default CA when the configured bundle path is missing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): renew the client when eviction lands before the callback lease
The cache can evict a logger between handing it to the callback and the callback taking its
lease. Such a lease now hands back a fresh client acquired through the same parameters, so that
callback exports through a live tracer provider instead of one teardown already shut down.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): emit litellm_call_id and response_id as generation metadata
v2 put the provider response id inside the generation id. v4 observation ids are 16 hex chars derived from that string, so the ids move to generation metadata to keep generations searchable by response id
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): read the response id through a typed protocol
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): do not claim trace root when continuing an existing trace
Langfuse derives a trace's name and I/O from any observation flagged
langfuse.internal.as_root, so a request carrying existing_trace_id
renamed the trace to the generation name and replaced the trace input
and output on every continuation. v2 only updated the keys listed in
update_trace_keys. Continuations now export as plain children of the
remote parent and keep the explicit langfuse.trace.* attributes for the
fields they do want changed.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): iterate lease renewal instead of recursing, monkeypatch update_trace_keys flag in tests
The recursive lease fallback tripped tests/code_coverage_tests/recursive_detector.py; the renewal
candidates are now walked with itertools.chain. The six update_trace_keys tests set the litellm
global through pytest monkeypatch so the TQ008 budget stays within its ceiling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): retry raised OTLP exports and honor LANGFUSE_TIMEOUT
The OTLP http exporter only retries 429 and 5xx; a connect or read timeout
propagates and BatchSpanProcessor drops the batch. Wrap the exporter in
RetryingSpanExporter (three backoff retries, as the v2 consumer did) and
build it on every path so the default and private-CA deployments share the
same channel, timeout and retry behaviour
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): sample on a hash of the full trace id and tolerate bad LANGFUSE_SAMPLE_RATE
TraceIdRatioBased reads the low 64 bits of the trace id. litellm trace ids are
UUIDs, whose variant bits sit at the top of that word, so every fractional rate
up to 0.5 dropped all traces. A SHA-256 of the full id gives an unbiased,
deterministic decision. Values outside [0, 1] or non numeric now warn and export
everything instead of raising during callback construction, which surfaced as a
500 on the first request of each worker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): put the Langfuse trace link back into Slack alerts
The proxy registers LangfusePromptManagement for callbacks: ["langfuse"], so the alert helper never saw the literal "langfuse" string and returned before looking up the trace id, and the prompt management logger never stored the trace id it got back from log_event_on_langfuse. Recognize LangFuseLogger instances in the callback list, record the returned trace id in the shared service trace id cache, and skip the link when no trace id arrives
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(deps): relock langfuse 4.15.2 and opentelemetry 1.33.1 on current main
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(langfuse): mark the deliberate blind except in client teardown for the strict ruff gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): pass the resource attributes mapping straight to Resource.create
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): warn about ignored UPSTREAM_LANGFUSE_* on the shared client init path too
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): normalise the OTLP export path so a trailing host slash never yields a double slash
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): nest guardrail and grounding spans under the generation
Langfuse v4 derives the trace name and I/O from every observation marked as_root, and the one with the latest start time wins. Guardrail and grounding spans used to claim root next to the generation, so a post_call guardrail could replace the model's request and response on the trace with its own. Only the generation claims root now; the sibling spans become its children
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): rebuild the cached bundle when mock mode or sample rate changes
The SDK keys resource bundles on the public key alone, so a bundle built with the discarding exporter for LANGFUSE_MOCK, or with an earlier LANGFUSE_SAMPLE_RATE, was handed back to a client that asked for a live exporter or a different rate. Compare both when deciding whether the cached bundle is still valid
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): keep trace_public true when a guardrail span is exported
Langfuse folds langfuse.trace.public across every observation in the trace and reads a missing attribute as false, so a guardrail child span without the flag turned a trace_public: true request private on Langfuse Cloud. Child spans now repeat the generation's value
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): emit observations as plain OTel spans, keep the SDK for prompts and auth
The callback now owns an isolated TracerProvider and OTLP exporter and builds generation and child spans with public OpenTelemetry APIs plus the LangfuseOtelSpanAttributes constants. Caller trace ids, generation ids, parent observation ids and historical start and end times are honoured through the OTel id generator, remote SpanContext and explicit span timestamps, so no private Langfuse SDK tracing handle is used any more. The Langfuse client stays only for get_prompt and auth_check
This also resolves the gauntlet findings on the previous draft: fresh traces start from an empty context so caller application spans are never stamped, the Slack trace link is read from the request logging state instead of constructing a logger per alert, a truthy non-mapping trace_metadata is serialized instead of raising, trace_input and trace_output land on the root generation, discarding a cached client is done under the lock, and the prompt cache no longer leaks a task manager because the client cache no longer tears down shared providers
Fixtures under tests/logging_callback_tests lose the SDK-private langfuse.internal.as_root marker; every other exported attribute is unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): hand the SDK client a validated sample rate so an unusable LANGFUSE_SAMPLE_RATE no longer breaks the callback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): gate the SDK version before importing the OTel module in prompt management
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): flush every export channel on proxy shutdown and use the callback's host in Slack trace links
The shutdown hook imported litellm.utils.langFuseLogger, a global the callback registry never assigns, so a graceful restart dropped the spans still queued in the batch processors. Shutdown now calls flush_langfuse_tracing, which force-flushes every acquired channel. The Slack alert link falls back to the registered LangFuseLogger's langfuse_host when the request carries no dynamic host, and the export endpoint tests pin that scheme-relative or absolute LANGFUSE_OTEL_TRACES_EXPORT_PATH values stay on the configured host
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): store resolved credentials on LangfusePromptManagement
The Slack alert trace link reads langfuse_host from every registered LangFuseLogger. Prompt management subclasses it without calling the parent constructor, so it never set the attribute and the alerting handler crashed before posting
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): flush every export channel concurrently under one shutdown deadline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): flush export channels on daemon threads so a stuck channel cannot hold up interpreter exit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): own the tracer config and drop the SDK client for prompts and auth
The callback's TracerProvider now sets its sampler, span limits and id generator explicitly so unrelated OTEL_* variables no longer change what Langfuse receives, and trace metadata is written once on the trace instead of folded into the generation, which kept input and output under the attribute cap. Spans are emitted under the langfuse-sdk scope so Langfuse renders them natively, the batch processor queues 100k spans and honors LANGFUSE_FLUSH_AT, and the proxy shutdown flush runs off the event loop with a 10s deadline and logs a miss.
Prompts, auth_check and the project id now go through LangfuseAPI directly with a litellm-owned TTL cache, so no Langfuse() client is built and a host application's client on the same public key is left alone. Dead attributes, the unreachable exporter branch and the export list are cleaned up, and the client-budget eviction behavior is documented.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): export OTLP spans and fetch prompts through litellm's HTTPHandler instead of a private requests session
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): gate the SDK version before importing the tracing module and retire unheld export channels
An installed v2 SDK used to fail inside the langfuse_sdk import and surface as "Langfuse not installed"; the version check now runs first so v2 users get the upgrade message, and only PackageNotFoundError means the package is missing
Export channels are now leased per credential set: acquire adds a holder, LangFuseLogger.stop (called by DynamicLoggingCache on expiry) releases one, and a channel with no holders is flushed and shut down after a 60 s grace, so rotating key or team credentials no longer grows one batch thread per credential set for the life of the process
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): end the generation when a child span fails, take the client slot last, keep prompt cache keys structured
Generation spans now end in a finally block so a bad guardrail or provider entry cannot strand the trace. The logger acquires its export channel and REST client before counting a client slot and releases the channel synchronously if the REST client fails to build, so retries after a bad config do not exhaust the budget. LANGFUSE_TIMEOUT accepts decimals for the REST client like it already did for OTLP export. The prompt cache keys on (name, version, label) so a missing label and the literal label None stay apart
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): claim the cache entry before releasing its slot and channel hold on eviction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): coerce generation names, keep v2 release, timeout and retry defaults, refresh stale prompts off the loop
A non-string metadata generation_name reached the OTLP encoder and took the whole batch down; it is now exported as its text and the exporter drops only the span the encoder rejects. LANGFUSE_RELEASE falls back to the deploy platform's commit variable again, the export deadline is back to the v2 default of 20 s and LANGFUSE_MAX_RETRIES sizes the retry ladder. An expired prompt is served at once while one background thread refreshes it, a re-acquired export channel cancels the pending retire timer, and flush reports delivery rather than a drained queue
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): assert the current Langfuse shutdown flush warning
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): keep host OTel resource out, carry big metadata ints, tolerate bad flush and TTL env, stamp trace I/O under a parent
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): name a malformed prompt cache TTL before the SDK import, keep metadata ints JSON safe, retry every 5xx export
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): name the auth check failure, split a 413 export, wire LANGFUSE_DEBUG, stamp error output under a parent, send the ingestion version header
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): honor LANGFUSE_DEBUG on the callbacks path, cap retry backoff, name the auth failure status and body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): cap LANGFUSE_MAX_RETRIES at 1000 so an absurd value cannot stall callback init
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): fold 413 halving into bounded rounds instead of recursion
The code-quality recursive-function gate flagged LangfuseSpanExporter.export. A batch of n spans settles within n.bit_length() halving rounds, so the split is a reduce over a frozen round state with the same posts, logs and results. The TTL gate test now asserts the gate returns without raising
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): truncate a single oversized span like v2 instead of dropping it, no retries on REST auth and project lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): write the metadata truncation marker under a flattened key so Langfuse keeps it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langfuse): patch the HTTPHandler export path and sync the metadata fixture and lease registry with the v4 callback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(langfuse): give the 413 split helpers a single explicit return path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): url-encode prompt names and fetch cold prompts without client retries
A cold get_prompt runs inline on the event loop; the generated v4 client's default two retries slept through
Retry-After (up to 60 s per attempt) and held the loop. The wrapper also passed the raw name into
api/public/v2/prompts/{name}, so 'what?' fetched prompt 'what' and folder names left the route. Quote the
name with safe='' like the v4 SDK's own get_prompt and pass max_retries=0 like the projects.get calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): retry a cold prompt miss once and drop upstream headers from prompt errors
A cold prompt fetch makes one immediate second attempt after a 5xx or a
transport failure, as the v2 client did, still with the generated client's
sleeping retries and Retry-After handling off so the event loop never stalls.
A failed fetch raises LangfusePromptError carrying only the status and body,
so the proxy no longer forwards Langfuse's response headers to its client
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langfuse): stub the logger in the health auth_check test instead of dialing a closed port
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(langfuse): integration test for OTLP v4 delivery and prompt fetch through a real proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* build(docker): keep the pip image's langfuse and otel pins on the v2 line its litellm 1.83.0 wheel expects
The image validates the published PyPI artifact, whose langfuse callback still
reads langfuse.version, so the 4.15.2 pin broke that callback. The pins move
together with the next LITELLM_VERSION bump. Also rewords the trace_version
precedence test docstring: v2 carried two version fields, v4 has one per span
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2162 lines
98 KiB
Python
2162 lines
98 KiB
Python
import os
|
||
import sys
|
||
from types import MappingProxyType
|
||
from typing import Final, Literal
|
||
|
||
from litellm.litellm_core_utils.env_utils import get_env_int, get_env_int_in_range, get_env_int_or_none
|
||
|
||
DEFAULT_HEALTH_CHECK_PROMPT: Final = str(os.getenv("DEFAULT_HEALTH_CHECK_PROMPT", "test from litellm"))
|
||
AZURE_DEFAULT_RESPONSES_API_VERSION: Final = str(os.getenv("AZURE_DEFAULT_RESPONSES_API_VERSION", "preview"))
|
||
AZURE_OPENAI_AUDIO_PROVIDERS: Final = frozenset({"azure", "azure_ai"})
|
||
ROUTER_MAX_FALLBACKS: Final = int(os.getenv("ROUTER_MAX_FALLBACKS", 5))
|
||
ROUTER_FALLBACK_ERROR_DETAIL_MAX_CHARS: Final = 2000
|
||
RUNTIME_UPDATABLE_ROUTER_SETTINGS: Final[frozenset[str]] = frozenset(
|
||
{
|
||
"routing_strategy_args",
|
||
"routing_strategy",
|
||
"routing_groups",
|
||
"allowed_fails",
|
||
"cooldown_time",
|
||
"num_retries",
|
||
"timeout",
|
||
"max_retries",
|
||
"retry_after",
|
||
"fallbacks",
|
||
"context_window_fallbacks",
|
||
"retry_policy",
|
||
"model_group_retry_policy",
|
||
"model_group_alias",
|
||
"enable_weighted_failover",
|
||
"enable_tag_filtering",
|
||
"tag_routing_prefix",
|
||
"optional_pre_call_checks",
|
||
}
|
||
)
|
||
ROUTER_SETTINGS_MANAGED_OUTSIDE_CONFIG: Final[frozenset[str]] = frozenset(
|
||
{
|
||
"model_list",
|
||
"search_tools",
|
||
"assistants_config",
|
||
"router_general_settings",
|
||
"ignore_invalid_deployments",
|
||
"fallback_access_check",
|
||
"fallback_budget_check",
|
||
"auto_router_capability_limit",
|
||
}
|
||
)
|
||
DEFAULT_BATCH_SIZE: Final = int(os.getenv("DEFAULT_BATCH_SIZE", 512))
|
||
DEFAULT_FLUSH_INTERVAL_SECONDS: Final = int(os.getenv("DEFAULT_FLUSH_INTERVAL_SECONDS", 5))
|
||
DEFAULT_S3_FLUSH_INTERVAL_SECONDS: Final = int(os.getenv("DEFAULT_S3_FLUSH_INTERVAL_SECONDS", 10))
|
||
DEFAULT_S3_BATCH_SIZE: Final = int(os.getenv("DEFAULT_S3_BATCH_SIZE", 512))
|
||
DEFAULT_S3_MAX_CONCURRENT_UPLOADS: Final = int(os.getenv("DEFAULT_S3_MAX_CONCURRENT_UPLOADS", "16"))
|
||
# https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-keys.html
|
||
MAX_S3_OBJECT_KEY_BYTES: Final = 1024
|
||
S3_BOUNDED_OBJECT_KEY_HEAD_BYTES: Final = 64
|
||
S3_PREFIX_DIGEST_CHARS: Final = 16
|
||
# s3 allows 2048 bytes of combined metadata headers, which Content-Disposition counts against
|
||
MAX_S3_OBJECT_DOWNLOAD_FILENAME_BYTES: Final = 1024
|
||
S3_LOG_PROMPTS_ONLY_ENV_VAR: Final = "S3_LOG_PROMPTS_ONLY"
|
||
MAX_FILE_LIST_LIMIT: Final = 10000
|
||
DEFAULT_SQS_FLUSH_INTERVAL_SECONDS: Final = int(os.getenv("DEFAULT_SQS_FLUSH_INTERVAL_SECONDS", 10))
|
||
DEFAULT_NUM_WORKERS_LITELLM_PROXY: Final = int(os.getenv("DEFAULT_NUM_WORKERS_LITELLM_PROXY", 1))
|
||
budget_reservation_disabled_info_emitted = False
|
||
DYNAMIC_RATE_LIMIT_ERROR_THRESHOLD_PER_MINUTE = int(os.getenv("DYNAMIC_RATE_LIMIT_ERROR_THRESHOLD_PER_MINUTE", 1))
|
||
DEFAULT_SQS_BATCH_SIZE: Final = int(os.getenv("DEFAULT_SQS_BATCH_SIZE", 512))
|
||
SQS_SEND_MESSAGE_ACTION: Final = "SendMessage"
|
||
SQS_API_VERSION: Final = "2012-11-05"
|
||
DEFAULT_MAX_RETRIES: Final = int(os.getenv("DEFAULT_MAX_RETRIES", 2))
|
||
# Max records accepted in one POST /v1/callbacks/logs batch. Bounds the blast
|
||
# radius: each record fans out to spend logs + every callback integration.
|
||
MAX_CALLBACK_LOG_RECORDS: Final = 1000
|
||
DEFAULT_MAX_RECURSE_DEPTH: Final = int(os.getenv("DEFAULT_MAX_RECURSE_DEPTH", 100))
|
||
DEFAULT_MAX_RECURSE_DEPTH_SENSITIVE_DATA_MASKER = int(os.getenv("DEFAULT_MAX_RECURSE_DEPTH_SENSITIVE_DATA_MASKER", 10))
|
||
DEFAULT_FAILURE_THRESHOLD_PERCENT: Final = float(
|
||
os.getenv("DEFAULT_FAILURE_THRESHOLD_PERCENT", 0.5)
|
||
) # default cooldown a deployment if 50% of requests fail in a given minute
|
||
DEFAULT_MAX_TOKENS: Final = int(os.getenv("DEFAULT_MAX_TOKENS", 4096))
|
||
DEFAULT_ALLOWED_FAILS: Final = int(os.getenv("DEFAULT_ALLOWED_FAILS", 3))
|
||
DEFAULT_REDIS_SYNC_INTERVAL: Final = int(os.getenv("DEFAULT_REDIS_SYNC_INTERVAL", 1))
|
||
DEFAULT_COOLDOWN_TIME_SECONDS: Final = int(os.getenv("DEFAULT_COOLDOWN_TIME_SECONDS", 5))
|
||
DEFAULT_COOLDOWN_REDIS_READ_INTERVAL_SECONDS: Final = float(
|
||
os.getenv("DEFAULT_COOLDOWN_REDIS_READ_INTERVAL_SECONDS", "1")
|
||
)
|
||
DEFAULT_REPLICATE_POLLING_RETRIES: Final = int(os.getenv("DEFAULT_REPLICATE_POLLING_RETRIES", 5))
|
||
DEFAULT_REPLICATE_POLLING_DELAY_SECONDS: Final = int(os.getenv("DEFAULT_REPLICATE_POLLING_DELAY_SECONDS", 1))
|
||
DEFAULT_IMAGE_TOKEN_COUNT: Final = int(os.getenv("DEFAULT_IMAGE_TOKEN_COUNT", 250))
|
||
HF_CONFIG_FETCH_TIMEOUT_SECONDS: Final = 10.0
|
||
|
||
# Maximum wall-clock seconds a streaming response is allowed to run.
|
||
# Streams exceeding this duration are terminated with a Timeout error.
|
||
# None (default) = no limit. Set env var to a number of seconds to enable globally.
|
||
_max_stream_duration_env: Final = os.getenv("LITELLM_MAX_STREAMING_DURATION_SECONDS", None)
|
||
LITELLM_MAX_STREAMING_DURATION_SECONDS: Final = (
|
||
float(_max_stream_duration_env) if _max_stream_duration_env is not None else None
|
||
)
|
||
|
||
# Maximum number of base64 characters to keep in logging payloads.
|
||
# Data URIs exceeding this are replaced with a size placeholder.
|
||
# Set to 0 to disable truncation.
|
||
MAX_BASE64_LENGTH_FOR_LOGGING: Final = int(os.getenv("MAX_BASE64_LENGTH_FOR_LOGGING", 64))
|
||
BASE64_TRUNCATION_OFFLOAD_THRESHOLD_CHARS: Final = 256 * 1024
|
||
REDACTED_BY_LITELLM: Final = "redacted-by-litellm"
|
||
# in-memory stand-in handed to provider converters for redacted arguments; never stored
|
||
REDACTED_TOOL_CALL_ARGUMENTS_PLACEHOLDER: Final = "{}"
|
||
|
||
MAX_STRING_LENGTH_STDOUT_LOG: Final = get_env_int("MAX_STRING_LENGTH_STDOUT_LOG", 4096)
|
||
MAX_BASE64_LENGTH_STDOUT_LOG: Final = get_env_int("MAX_BASE64_LENGTH_STDOUT_LOG", 4096)
|
||
|
||
# When true, adds detailed per-phase timing breakdown headers to responses.
|
||
# Headers: x-litellm-timing-{pre-processing,llm-api,post-processing,message-copy}-ms
|
||
LITELLM_DETAILED_TIMING: Final = os.getenv("LITELLM_DETAILED_TIMING", "false").lower() == "true"
|
||
|
||
# Model cost map validation constants
|
||
MODEL_COST_MAP_MIN_MODEL_COUNT: Final = int(
|
||
os.getenv("MODEL_COST_MAP_MIN_MODEL_COUNT", 50)
|
||
) # Minimum number of models a fetched cost map must contain to be considered valid
|
||
MODEL_COST_MAP_MAX_SHRINK_RATIO: Final = float(
|
||
os.getenv("MODEL_COST_MAP_MAX_SHRINK_RATIO", 0.5)
|
||
) # Maximum allowed shrinkage ratio vs local backup (0.5 = reject if fetched map is <50% of backup)
|
||
DEFAULT_IMAGE_WIDTH: Final = int(os.getenv("DEFAULT_IMAGE_WIDTH", 300))
|
||
DEFAULT_IMAGE_HEIGHT: Final = int(os.getenv("DEFAULT_IMAGE_HEIGHT", 300))
|
||
# Maximum size for image URL downloads in MB (default 50MB, set to 0 to disable limit)
|
||
# This prevents memory issues from downloading very large images
|
||
# Maps to OpenAI's 50 MB payload limit - requests with images exceeding this size will be rejected
|
||
# Set MAX_IMAGE_URL_DOWNLOAD_SIZE_MB=0 to disable image URL handling entirely
|
||
MAX_IMAGE_URL_DOWNLOAD_SIZE_MB: Final = float(os.getenv("MAX_IMAGE_URL_DOWNLOAD_SIZE_MB", 50))
|
||
MAX_SIZE_PER_ITEM_IN_MEMORY_CACHE_IN_KB: Final = int(
|
||
os.getenv("MAX_SIZE_PER_ITEM_IN_MEMORY_CACHE_IN_KB", 1024)
|
||
) # 1MB = 1024KB
|
||
# Surrogate-repair fallback in _read_request_body runs two full-body re.sub passes
|
||
# that block the event loop on multi-MB malformed bodies. Skip the repair above this
|
||
# size and raise the existing 400 immediately. Set to 0 to disable the cap.
|
||
MAX_REQUEST_BODY_SIZE_TO_REPAIR_MB: Final = get_env_int("MAX_REQUEST_BODY_SIZE_TO_REPAIR_MB", 1)
|
||
SINGLE_DEPLOYMENT_TRAFFIC_FAILURE_THRESHOLD: Final = int(
|
||
os.getenv("SINGLE_DEPLOYMENT_TRAFFIC_FAILURE_THRESHOLD", 1000)
|
||
) # Minimum number of requests to consider "reasonable traffic". Used for single-deployment cooldown logic.
|
||
DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS: Final = int(
|
||
os.getenv("DEFAULT_FAILURE_THRESHOLD_MINIMUM_REQUESTS", 5)
|
||
) # Minimum number of requests before applying error rate cooldown. Prevents cooldown from triggering on first failure.
|
||
|
||
DEFAULT_REASONING_EFFORT_DISABLE_THINKING_BUDGET = int(os.getenv("DEFAULT_REASONING_EFFORT_DISABLE_THINKING_BUDGET", 0))
|
||
|
||
# MCP Semantic Tool Filter Defaults
|
||
DEFAULT_MCP_SEMANTIC_FILTER_EMBEDDING_MODEL: Final = str(
|
||
os.getenv("DEFAULT_MCP_SEMANTIC_FILTER_EMBEDDING_MODEL", "text-embedding-3-small")
|
||
)
|
||
DEFAULT_MCP_SEMANTIC_FILTER_TOP_K: Final = int(os.getenv("DEFAULT_MCP_SEMANTIC_FILTER_TOP_K", 10))
|
||
DEFAULT_MCP_SEMANTIC_FILTER_SIMILARITY_THRESHOLD: Final = float(
|
||
os.getenv("DEFAULT_MCP_SEMANTIC_FILTER_SIMILARITY_THRESHOLD", 0.3)
|
||
)
|
||
MAX_MCP_SEMANTIC_FILTER_TOOLS_HEADER_LENGTH: Final = int(os.getenv("MAX_MCP_SEMANTIC_FILTER_TOOLS_HEADER_LENGTH", 150))
|
||
MAX_LITELLM_CALL_ID_LENGTH: Final = 256
|
||
MAX_GUARDRAIL_SCAN_METADATA_HEADER_LENGTH: Final = 2048
|
||
|
||
DEFAULT_AUTO_ROUTER_MAX_INPUT_CHARS: Final = 2000
|
||
|
||
# Semantic Guard Defaults
|
||
DEFAULT_SEMANTIC_GUARD_EMBEDDING_MODEL: Final = str(
|
||
os.getenv("DEFAULT_SEMANTIC_GUARD_EMBEDDING_MODEL", "text-embedding-3-small")
|
||
)
|
||
DEFAULT_SEMANTIC_GUARD_SIMILARITY_THRESHOLD = float(os.getenv("DEFAULT_SEMANTIC_GUARD_SIMILARITY_THRESHOLD", 0.75))
|
||
|
||
DEFAULT_OPENAI_MODERATIONS_MODEL: Final = "omni-moderation-latest"
|
||
|
||
# MCP OAuth2 Client Credentials Defaults
|
||
MCP_OAUTH2_TOKEN_EXPIRY_BUFFER_SECONDS: Final = int(os.getenv("MCP_OAUTH2_TOKEN_EXPIRY_BUFFER_SECONDS", "60"))
|
||
MCP_OAUTH2_TOKEN_CACHE_MAX_SIZE: Final = int(os.getenv("MCP_OAUTH2_TOKEN_CACHE_MAX_SIZE", "200"))
|
||
MCP_OAUTH2_TOKEN_CACHE_DEFAULT_TTL: Final = int(os.getenv("MCP_OAUTH2_TOKEN_CACHE_DEFAULT_TTL", "3600"))
|
||
MCP_SSO_ASSERTION_CACHE_TTL_SECONDS: Final = int(os.getenv("MCP_SSO_ASSERTION_CACHE_TTL_SECONDS", "60"))
|
||
|
||
# mcp_tool_permissions entry that grants every current and future tool on a server
|
||
MCP_ALL_TOOLS_WILDCARD: Final = "*"
|
||
|
||
# Default npm cache directory for STDIO MCP servers.
|
||
# npm/npx needs a writable cache dir; in containers the default (~/.npm)
|
||
# may not exist or be read-only. /tmp is always writable.
|
||
MCP_NPM_CACHE_DIR: Final = os.getenv("MCP_NPM_CACHE_DIR", "/tmp/.npm_mcp_cache")
|
||
MCP_OAUTH2_TOKEN_CACHE_MIN_TTL: Final = int(os.getenv("MCP_OAUTH2_TOKEN_CACHE_MIN_TTL", "10"))
|
||
|
||
# Per-user OAuth token Redis cache (for server-side token storage)
|
||
MCP_PER_USER_TOKEN_REDIS_KEY_PREFIX: Final = "mcp:per_user_token"
|
||
MCP_PER_USER_TOKEN_DEFAULT_TTL: Final = int(
|
||
os.getenv("MCP_PER_USER_TOKEN_DEFAULT_TTL", "43200") # 12 hours
|
||
)
|
||
MCP_PER_USER_TOKEN_EXPIRY_BUFFER_SECONDS: Final = int(os.getenv("MCP_PER_USER_TOKEN_EXPIRY_BUFFER_SECONDS", "60"))
|
||
|
||
# MCP timeout defaults (seconds). Override via env vars for slow/custom MCP servers.
|
||
MCP_CLIENT_TIMEOUT: Final = float(os.getenv("LITELLM_MCP_CLIENT_TIMEOUT", "60.0"))
|
||
MCP_TOOL_LISTING_TIMEOUT: Final = float(os.getenv("LITELLM_MCP_TOOL_LISTING_TIMEOUT", "30.0"))
|
||
MCP_METADATA_TIMEOUT: Final = float(os.getenv("LITELLM_MCP_METADATA_TIMEOUT", "10.0"))
|
||
MCP_HEALTH_CHECK_TIMEOUT: Final = float(os.getenv("LITELLM_MCP_HEALTH_CHECK_TIMEOUT", "10.0"))
|
||
MCP_TOOL_LISTING_MAX_PAGES: Final = 1000
|
||
MCP_GATEWAY_SESSION_ID_PREFIX_LENGTH: Final = 8
|
||
MCP_BYOK_CREDENTIAL_CACHE_TTL_SECONDS: Final = 60
|
||
MCP_BYOK_CREDENTIAL_CACHE_MAX_SIZE: Final = 4096
|
||
|
||
# Allowlist of commands permitted for MCP stdio transport.
|
||
# Prevents arbitrary command execution via /mcp-rest/test/* endpoints or server creation.
|
||
# Note: allowlisted runtimes can still execute code via args (e.g. python -c "...").
|
||
# This is an accepted residual risk since these endpoints require PROXY_ADMIN.
|
||
# Extend via LITELLM_MCP_STDIO_EXTRA_COMMANDS env var (comma-separated).
|
||
_MCP_STDIO_EXTRA_COMMANDS: Final = os.getenv("LITELLM_MCP_STDIO_EXTRA_COMMANDS", "")
|
||
MCP_STDIO_ALLOWED_COMMANDS: Final[frozenset] = frozenset(
|
||
{"npx", "uvx", "python", "python3", "node", "docker", "deno"} | (set(_MCP_STDIO_EXTRA_COMMANDS.split(",")) - {""})
|
||
)
|
||
|
||
# MCP OAuth2 Token Exchange (OBO) Defaults
|
||
MCP_TOKEN_EXCHANGE_CACHE_MAX_SIZE: Final = int(os.getenv("MCP_TOKEN_EXCHANGE_CACHE_MAX_SIZE", "500"))
|
||
|
||
LITELLM_UI_ALLOW_HEADERS: Final = [
|
||
"x-litellm-semantic-filter",
|
||
"x-litellm-semantic-filter-tools",
|
||
"x-litellm-adaptive-router-model",
|
||
"x-litellm-applied-guardrails",
|
||
"x-litellm-guardrail-scan-id",
|
||
"x-litellm-guardrail-scan-metadata",
|
||
"x-litellm-cache-key",
|
||
]
|
||
|
||
# Gemini model-specific minimal thinking budget constants
|
||
DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET_GEMINI_2_5_FLASH: Final = int(
|
||
os.getenv("DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET_GEMINI_2_5_FLASH", 1)
|
||
)
|
||
DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET_GEMINI_2_5_PRO: Final = int(
|
||
os.getenv("DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET_GEMINI_2_5_PRO", 128)
|
||
)
|
||
DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET_GEMINI_2_5_FLASH_LITE: Final = int(
|
||
os.getenv("DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET_GEMINI_2_5_FLASH_LITE", 512)
|
||
)
|
||
|
||
# Maximum number of callbacks that can be registered
|
||
# This prevents callbacks from exponentially growing and consuming CPU resources
|
||
# Override with LITELLM_MAX_CALLBACKS env var for large deployments (e.g., many teams with guardrails)
|
||
MAX_CALLBACKS: Final = get_env_int("LITELLM_MAX_CALLBACKS", 100)
|
||
|
||
# Metadata key recording which pre_call guardrails the proxy loop already ran,
|
||
# so the deployment-level hook does not re-run them for the same request
|
||
PRE_CALL_EXECUTED_GUARDRAILS_KEY: Final = "_pre_call_executed_guardrails"
|
||
|
||
# Attribute stamped on log_guardrail_information wrappers so __init_subclass__ does not wrap them again
|
||
LOGS_GUARDRAIL_INFORMATION_MARKER: Final = "_litellm_logs_guardrail_information"
|
||
|
||
# llm_provider stamped on proxy-side rate limit errors when the model resolves to no deployment
|
||
PROXY_LLM_PROVIDER_FALLBACK: Final = "litellm_proxy"
|
||
|
||
# litellm_params flag on failure logs for requests the proxy rejected before routing to a deployment
|
||
PROXY_REJECTED_BEFORE_ROUTING_KEY: Final = "proxy_rejected_before_routing"
|
||
|
||
# Generic fallback for unknown models
|
||
DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET: Final = int(
|
||
os.getenv("DEFAULT_REASONING_EFFORT_MINIMAL_THINKING_BUDGET", 128)
|
||
)
|
||
|
||
# Provider-specific API base URLs
|
||
XAI_API_BASE: Final = "https://api.x.ai/v1"
|
||
OPEN_SANDBOX_API_BASE_ENV_VAR: Final = "OPEN_SANDBOX_API_BASE"
|
||
OPEN_SANDBOX_API_KEY_ENV_VAR: Final = "OPEN_SANDBOX_API_KEY"
|
||
OPEN_SANDBOX_DEFAULT_TEMPLATE: Final = "opensandbox/code-interpreter:v1.1.0"
|
||
_OPEN_SANDBOX_FALLBACK_ENTRYPOINT: Final = "/opt/code-interpreter/code-interpreter.sh"
|
||
OPEN_SANDBOX_DEFAULT_ENTRYPOINT: Final = (_OPEN_SANDBOX_FALLBACK_ENTRYPOINT,)
|
||
OPEN_SANDBOX_DEFAULT_LANGUAGE: Final = "python"
|
||
OPEN_SANDBOX_DEFAULT_CPU_LIMIT: Final = "1"
|
||
OPEN_SANDBOX_DEFAULT_MEMORY_LIMIT: Final = "2Gi"
|
||
OPEN_SANDBOX_EXECD_PORT: Final = 44772
|
||
OPEN_SANDBOX_DEFAULT_TIMEOUT: Final = 300
|
||
OPEN_SANDBOX_READY_TIMEOUT: Final = 30.0
|
||
OPEN_SANDBOX_POLL_INTERVAL: Final = 0.2
|
||
|
||
DEFAULT_REASONING_EFFORT_LOW_THINKING_BUDGET = int(os.getenv("DEFAULT_REASONING_EFFORT_LOW_THINKING_BUDGET", 1024))
|
||
DEFAULT_REASONING_EFFORT_MEDIUM_THINKING_BUDGET: Final = int(
|
||
os.getenv("DEFAULT_REASONING_EFFORT_MEDIUM_THINKING_BUDGET", 2048)
|
||
)
|
||
DEFAULT_REASONING_EFFORT_HIGH_THINKING_BUDGET = int(os.getenv("DEFAULT_REASONING_EFFORT_HIGH_THINKING_BUDGET", 4096))
|
||
DEFAULT_REASONING_EFFORT_XHIGH_THINKING_BUDGET = int(os.getenv("DEFAULT_REASONING_EFFORT_XHIGH_THINKING_BUDGET", 8192))
|
||
DEFAULT_REASONING_EFFORT_MAX_THINKING_BUDGET = int(os.getenv("DEFAULT_REASONING_EFFORT_MAX_THINKING_BUDGET", 16384))
|
||
MAX_TOKEN_TRIMMING_ATTEMPTS: Final = int(
|
||
os.getenv("MAX_TOKEN_TRIMMING_ATTEMPTS", 10)
|
||
) # Maximum number of attempts to trim the message
|
||
|
||
RUNWAYML_DEFAULT_API_VERSION: Final = str(os.getenv("RUNWAYML_DEFAULT_API_VERSION", "2024-11-06"))
|
||
RUNWAYML_POLLING_TIMEOUT = int(os.getenv("RUNWAYML_POLLING_TIMEOUT", 600)) # 10 minutes default for image generation
|
||
|
||
########## Networking constants ##############################################################
|
||
_DEFAULT_TTL_FOR_HTTPX_CLIENTS: Final = 3600 # 1 hour, re-use the same httpx client for 1 hour
|
||
|
||
# The earliest an evicted, litellm-created client may be closed. A request handed the
|
||
# client just before eviction is still using it, so nothing is closed inside this window;
|
||
# past it, the client is closed once it reports no connection in flight.
|
||
EVICTED_LLM_CLIENT_CLOSE_GRACE_SECONDS: Final = 900
|
||
|
||
# How many evicted clients may be queued for closing at once. Past this, an evicted client
|
||
# is left to the collector rather than letting a cache-churning workload grow the queue
|
||
# without bound. Each queued entry is ~100 bytes and comes due within one grace window.
|
||
EVICTED_LLM_CLIENT_CLOSE_MAX_PENDING: Final = 10_000
|
||
|
||
# Aiohttp connection pooling - prevents memory leaks from unbounded connection growth
|
||
# Set to 0 for unlimited (not recommended for production)
|
||
AIOHTTP_CONNECTOR_LIMIT: Final = int(os.getenv("AIOHTTP_CONNECTOR_LIMIT", 1000))
|
||
AIOHTTP_CONNECTOR_LIMIT_PER_HOST: Final = int(os.getenv("AIOHTTP_CONNECTOR_LIMIT_PER_HOST", 500))
|
||
AIOHTTP_KEEPALIVE_TIMEOUT: Final = int(os.getenv("AIOHTTP_KEEPALIVE_TIMEOUT", 120))
|
||
AIOHTTP_TTL_DNS_CACHE: Final = int(os.getenv("AIOHTTP_TTL_DNS_CACHE", 300))
|
||
# TCP keep-alive (SO_KEEPALIVE) — opt-in. Required when running behind NAT/LBs
|
||
# whose idle timeout is shorter than provider response timeouts (e.g. AWS NAT
|
||
# Gateway: 350s vs OpenAI/Azure: 600s). Without this, the kernel sends nothing
|
||
# during a long provider call and the NAT reaps the flow before the response
|
||
# arrives. Enabling SO_KEEPALIVE makes the kernel emit TCP probes that reset
|
||
# the NAT idle timer.
|
||
AIOHTTP_SO_KEEPALIVE: Final = os.getenv("AIOHTTP_SO_KEEPALIVE", "False").lower() == "true"
|
||
AIOHTTP_TCP_KEEPIDLE: Final = int(os.getenv("AIOHTTP_TCP_KEEPIDLE", 60))
|
||
AIOHTTP_TCP_KEEPINTVL: Final = int(os.getenv("AIOHTTP_TCP_KEEPINTVL", 30))
|
||
AIOHTTP_TCP_KEEPCNT: Final = int(os.getenv("AIOHTTP_TCP_KEEPCNT", 5))
|
||
# enable_cleanup_closed is only needed for Python versions with the SSL leak bug
|
||
# Fixed in Python 3.12.7+ and 3.13.1+ (see https://github.com/python/cpython/pull/118960)
|
||
# Reference: https://github.com/aio-libs/aiohttp/blob/master/aiohttp/connector.py#L74-L78
|
||
AIOHTTP_NEEDS_CLEANUP_CLOSED: Final = (3, 13, 0) <= sys.version_info < (
|
||
3,
|
||
13,
|
||
1,
|
||
) or sys.version_info < (3, 12, 7)
|
||
|
||
# WebSocket constants
|
||
# Default to None (unlimited) to match OpenAI's official agents SDK behavior
|
||
# https://github.com/openai/openai-agents-python/blob/cf1b933660e44fd37b4350c41febab8221801409/src/agents/realtime/openai_realtime.py#L235
|
||
_max_size_env: Final = os.getenv("REALTIME_WEBSOCKET_MAX_MESSAGE_SIZE_BYTES")
|
||
REALTIME_WEBSOCKET_MAX_MESSAGE_SIZE_BYTES: Final = int(_max_size_env) if _max_size_env is not None else None
|
||
REALTIME_CREDENTIAL_RESOLUTION_TIMEOUT_SECONDS: Final = float(
|
||
os.getenv("REALTIME_CREDENTIAL_RESOLUTION_TIMEOUT_SECONDS", "20.0")
|
||
)
|
||
|
||
# RFC 6455 caps the close frame payload at 125 bytes, 2 of which carry the status code
|
||
WEBSOCKET_CLOSE_REASON_MAX_BYTES: Final = 123
|
||
|
||
DEEPGRAM_DEFAULT_API_BASE: Final = "https://api.deepgram.com/v1"
|
||
NADIR_DEFAULT_API_BASE: Final = "https://api.getnadir.com/v1"
|
||
DEEPGRAM_LISTEN_DEFAULT_MODEL: Final = "nova-3"
|
||
|
||
BEDROCK_REALTIME_PENDING_SESSION_UPDATE_SCOPE_KEY: Final = "litellm.bedrock_realtime.pending_session_update"
|
||
BEDROCK_REALTIME_SESSION_COMMITTED_SCOPE_KEY: Final = "litellm.bedrock_realtime.session_committed"
|
||
BEDROCK_REALTIME_COMMITTED_FAILURE_SCOPE_KEY: Final = "litellm.bedrock_realtime.committed_failure"
|
||
BEDROCK_REALTIME_SDK_DISTRIBUTION: Final = "aws-sdk-bedrock-runtime"
|
||
BEDROCK_REALTIME_SDK_SUPPORTED_RANGE: Final = ">=0.10.0,<0.12.0"
|
||
CLIENT_REQUESTED_MODEL_SCOPE_KEY: Final = "litellm.client_requested_model"
|
||
MODEL_GROUP_ALIAS_RESOLVED_SCOPE_KEY: Final = "litellm.model_group_alias_resolved"
|
||
REALTIME_SESSION_SUCCESS_LOGGED_KEY: Final = "realtime_session_success_logged"
|
||
REALTIME_SESSION_FAILURE_LOGGED_KEY: Final = "realtime_session_failure_logged"
|
||
|
||
# SSL/TLS cipher configuration for faster handshakes
|
||
# Strategy: Strongly prefer fast modern ciphers, but allow fallback to commonly supported ones
|
||
# This balances performance with broad compatibility
|
||
DEFAULT_SSL_CIPHERS: Final = os.getenv(
|
||
"LITELLM_SSL_CIPHERS",
|
||
# Priority 1: TLS 1.3 ciphers (fastest, ~50ms handshake)
|
||
"TLS_AES_256_GCM_SHA384:" # Fastest observed in testing
|
||
"TLS_AES_128_GCM_SHA256:" # Slightly faster than 256-bit
|
||
"TLS_CHACHA20_POLY1305_SHA256:" # Fast on ARM/mobile
|
||
# Priority 2: TLS 1.2 ECDHE+GCM (fast, ~100ms handshake, widely supported)
|
||
"ECDHE-RSA-AES256-GCM-SHA384:"
|
||
"ECDHE-RSA-AES128-GCM-SHA256:"
|
||
"ECDHE-ECDSA-AES256-GCM-SHA384:"
|
||
"ECDHE-ECDSA-AES128-GCM-SHA256:"
|
||
# Priority 3: Additional modern ciphers (good balance)
|
||
"ECDHE-RSA-CHACHA20-POLY1305:"
|
||
"ECDHE-ECDSA-CHACHA20-POLY1305:"
|
||
# Priority 4: Widely compatible fallbacks (slower but universally supported)
|
||
"ECDHE-RSA-AES256-SHA384:" # Common fallback
|
||
"ECDHE-RSA-AES128-SHA256:" # Very widely supported
|
||
"AES256-GCM-SHA384:" # Non-PFS fallback (compatibility)
|
||
"AES128-GCM-SHA256", # Last resort (maximum compatibility)
|
||
)
|
||
|
||
########### v2 Architecture constants for managing writing updates to the database ###########
|
||
REDIS_UPDATE_BUFFER_KEY: Final = "litellm_spend_update_buffer"
|
||
REDIS_GATEWAY_REQUESTS_BUFFER_KEY: Final = "litellm_gateway_requests_buffer"
|
||
REDIS_DAILY_SPEND_UPDATE_BUFFER_KEY: Final = "litellm_daily_spend_update_buffer"
|
||
REDIS_DAILY_TEAM_SPEND_UPDATE_BUFFER_KEY: Final = "litellm_daily_team_spend_update_buffer"
|
||
REDIS_DAILY_ORG_SPEND_UPDATE_BUFFER_KEY: Final = "litellm_daily_org_spend_update_buffer"
|
||
REDIS_DAILY_END_USER_SPEND_UPDATE_BUFFER_KEY: Final = "litellm_daily_end_user_spend_update_buffer"
|
||
REDIS_DAILY_AGENT_SPEND_UPDATE_BUFFER_KEY: Final = "litellm_daily_agent_spend_update_buffer"
|
||
REDIS_DAILY_TAG_SPEND_UPDATE_BUFFER_KEY: Final = "litellm_daily_tag_spend_update_buffer"
|
||
REDIS_WINDOW_SPEND_UPDATE_BUFFER_KEY: Final = "litellm_window_spend_update_buffer"
|
||
MAX_REDIS_BUFFER_DEQUEUE_COUNT: Final = int(os.getenv("MAX_REDIS_BUFFER_DEQUEUE_COUNT", 100))
|
||
REDIS_SPEND_LOGS_BUFFER_KEY: Final = "litellm_spend_logs_buffer"
|
||
REDIS_SPEND_LOGS_BUFFER_MAX_ROWS: Final = 100000
|
||
REDIS_SPEND_LOGS_BUFFER_DEQUEUE_COUNT: Final = 1000
|
||
# Bounds asyncio.Queue() instances (log queues, spend update queues, etc.) to prevent unbounded memory growth
|
||
LITELLM_ASYNCIO_QUEUE_MAXSIZE: Final = int(os.getenv("LITELLM_ASYNCIO_QUEUE_MAXSIZE", 1000))
|
||
TOOL_POLICY_CACHE_TTL_SECONDS: Final = int(os.getenv("TOOL_POLICY_CACHE_TTL_SECONDS", 60))
|
||
GUARDRAIL_SCANNED_MESSAGES_CACHE_TTL_SECONDS: Final = int(
|
||
os.getenv("GUARDRAIL_SCANNED_MESSAGES_CACHE_TTL_SECONDS", 24 * 60 * 60)
|
||
)
|
||
BEDROCK_APPLY_GUARDRAIL_CHUNK_BUDGET_CHARS: Final = 25_000
|
||
CONTENT_FILTER_STREAMING_HOLDBACK_CHARS: Final = 50
|
||
CONTENT_FILTER_STREAMING_SCAN_CONTEXT_CHARS: Final = 512
|
||
DEFAULT_PRESIDIO_ANALYZE_CHUNK_SIZE_BYTES: Final = 500_000
|
||
PRESIDIO_ANALYZE_CHUNK_OVERLAP_CHARS: Final = 4096
|
||
PRESIDIO_ANALYZE_CHUNK_CONCURRENCY: Final = 8
|
||
# Aggregation threshold: default to 80% of the asyncio queue maxsize so the check can always trigger.
|
||
# Must be < LITELLM_ASYNCIO_QUEUE_MAXSIZE; if set higher the aggregation logic will never fire.
|
||
MAX_SIZE_IN_MEMORY_QUEUE: Final = int(os.getenv("MAX_SIZE_IN_MEMORY_QUEUE", int(LITELLM_ASYNCIO_QUEUE_MAXSIZE * 0.8)))
|
||
MAX_IN_MEMORY_QUEUE_FLUSH_COUNT: Final = int(os.getenv("MAX_IN_MEMORY_QUEUE_FLUSH_COUNT", 1000))
|
||
###############################################################################################
|
||
# Providers will not cache a prefix below a minimum size. That minimum is per-model, not global:
|
||
# Anthropic's ranges from 512 to 4096 depending on the model, and can differ per platform for the
|
||
# same model. The real minimum is resolved from `prompt_cache_min_tokens` in the model cost map;
|
||
# this value is only the fallback for models the cost map has no entry for, and doubles as a global
|
||
# escape hatch when `MINIMUM_PROMPT_CACHE_TOKEN_COUNT` is explicitly set.
|
||
MINIMUM_PROMPT_CACHE_TOKEN_COUNT_OVERRIDE: Final[int | None] = get_env_int_or_none("MINIMUM_PROMPT_CACHE_TOKEN_COUNT")
|
||
DEFAULT_MINIMUM_PROMPT_CACHE_TOKEN_COUNT: Final = 1024
|
||
MINIMUM_PROMPT_CACHE_TOKEN_COUNT: Final = (
|
||
MINIMUM_PROMPT_CACHE_TOKEN_COUNT_OVERRIDE
|
||
if MINIMUM_PROMPT_CACHE_TOKEN_COUNT_OVERRIDE is not None
|
||
else DEFAULT_MINIMUM_PROMPT_CACHE_TOKEN_COUNT
|
||
)
|
||
PROMPT_CACHE_LOOKBACK_POSITIONS: Final = 20
|
||
DEFAULT_TRIM_RATIO: Final = float(
|
||
os.getenv("DEFAULT_TRIM_RATIO", 0.75)
|
||
) # default ratio of tokens to trim from the end of a prompt
|
||
HOURS_IN_A_DAY: Final = int(os.getenv("HOURS_IN_A_DAY", 24))
|
||
DAYS_IN_A_WEEK: Final = int(os.getenv("DAYS_IN_A_WEEK", 7))
|
||
DAYS_IN_A_MONTH: Final = int(os.getenv("DAYS_IN_A_MONTH", 28))
|
||
DAYS_IN_A_YEAR: Final = int(os.getenv("DAYS_IN_A_YEAR", 365))
|
||
REPLICATE_MODEL_NAME_WITH_ID_LENGTH: Final = int(os.getenv("REPLICATE_MODEL_NAME_WITH_ID_LENGTH", 64))
|
||
#### TOKEN COUNTING ####
|
||
FUNCTION_DEFINITION_TOKEN_COUNT: Final = int(os.getenv("FUNCTION_DEFINITION_TOKEN_COUNT", 9))
|
||
SYSTEM_MESSAGE_TOKEN_COUNT: Final = int(os.getenv("SYSTEM_MESSAGE_TOKEN_COUNT", 4))
|
||
TOOL_CHOICE_OBJECT_TOKEN_COUNT: Final = int(os.getenv("TOOL_CHOICE_OBJECT_TOKEN_COUNT", 4))
|
||
DEFAULT_MOCK_RESPONSE_PROMPT_TOKEN_COUNT: Final = int(os.getenv("DEFAULT_MOCK_RESPONSE_PROMPT_TOKEN_COUNT", 10))
|
||
DEFAULT_MOCK_RESPONSE_COMPLETION_TOKEN_COUNT: Final = int(os.getenv("DEFAULT_MOCK_RESPONSE_COMPLETION_TOKEN_COUNT", 20))
|
||
MAX_SHORT_SIDE_FOR_IMAGE_HIGH_RES: Final = int(os.getenv("MAX_SHORT_SIDE_FOR_IMAGE_HIGH_RES", 768))
|
||
MAX_LONG_SIDE_FOR_IMAGE_HIGH_RES: Final = int(os.getenv("MAX_LONG_SIDE_FOR_IMAGE_HIGH_RES", 2000))
|
||
# tiktoken's BPE merge loop is quadratic in the length of a single regex piece, so a long run of one
|
||
# repeated character (dot leaders, whitespace, zero-padded base64) can take minutes on a multi-MB payload.
|
||
# Encoding in chunks makes the cost linear, at a drift of at most ~1 token per chunk boundary. The upper
|
||
# bound keeps a misconfigured chunk size from restoring the quadratic cost this exists to remove.
|
||
TIKTOKEN_ENCODE_MAX_CHUNK_SIZE_CHARS: Final = 4096
|
||
TIKTOKEN_ENCODE_CHUNK_SIZE_CHARS: Final = get_env_int_in_range(
|
||
"TIKTOKEN_ENCODE_CHUNK_SIZE_CHARS",
|
||
default=1024,
|
||
minimum=1,
|
||
maximum=TIKTOKEN_ENCODE_MAX_CHUNK_SIZE_CHARS,
|
||
)
|
||
TOKEN_COUNTER_MAX_EXACT_CHARS: Final = get_env_int_in_range(
|
||
"TOKEN_COUNTER_MAX_EXACT_CHARS",
|
||
default=4_000_000,
|
||
minimum=1,
|
||
maximum=1_000_000_000,
|
||
)
|
||
TOKEN_COUNTER_MAX_CONCURRENT_COUNTS: Final = get_env_int_in_range(
|
||
"TOKEN_COUNTER_MAX_CONCURRENT_COUNTS",
|
||
default=4,
|
||
minimum=1,
|
||
maximum=256,
|
||
)
|
||
MAX_TILE_WIDTH: Final = int(os.getenv("MAX_TILE_WIDTH", 512))
|
||
MAX_TILE_HEIGHT: Final = int(os.getenv("MAX_TILE_HEIGHT", 512))
|
||
OPENAI_FILE_SEARCH_COST_PER_1K_CALLS: Final = float(os.getenv("OPENAI_FILE_SEARCH_COST_PER_1K_CALLS", 2.5 / 1000))
|
||
GROQ_BROWSER_VISIT_WEBSITE_COST_PER_CALL: Final = 1.0 / 1000
|
||
# Azure OpenAI Assistants feature costs
|
||
# Source: https://azure.microsoft.com/en-us/pricing/details/cognitive-services/openai-service/
|
||
AZURE_FILE_SEARCH_COST_PER_GB_PER_DAY: Final = float(
|
||
os.getenv("AZURE_FILE_SEARCH_COST_PER_GB_PER_DAY", 0.1) # $0.1 USD per 1 GB/Day
|
||
)
|
||
AZURE_COMPUTER_USE_INPUT_COST_PER_1K_TOKENS: Final = float(
|
||
os.getenv("AZURE_COMPUTER_USE_INPUT_COST_PER_1K_TOKENS", 3.0) # $0.003 USD per 1K Tokens
|
||
)
|
||
AZURE_COMPUTER_USE_OUTPUT_COST_PER_1K_TOKENS: Final = float(
|
||
os.getenv("AZURE_COMPUTER_USE_OUTPUT_COST_PER_1K_TOKENS", 12.0) # $0.012 USD per 1K Tokens
|
||
)
|
||
AZURE_VECTOR_STORE_COST_PER_GB_PER_DAY: Final = float(
|
||
os.getenv("AZURE_VECTOR_STORE_COST_PER_GB_PER_DAY", 0.1) # $0.1 USD per 1 GB/Day (same as file search)
|
||
)
|
||
MIN_NON_ZERO_TEMPERATURE: Final = float(os.getenv("MIN_NON_ZERO_TEMPERATURE", 0.0001))
|
||
#### RELIABILITY ####
|
||
REPEATED_STREAMING_CHUNK_LIMIT: Final = int(
|
||
os.getenv("REPEATED_STREAMING_CHUNK_LIMIT", 100)
|
||
) # catch if model starts looping the same chunk while streaming. Uses high default to prevent false positives.
|
||
# Shared maxsize for functools.lru_cache usage across hot paths.
|
||
# Defaulted to 64 to avoid cache thrash in multi-model production workloads.
|
||
DEFAULT_MAX_LRU_CACHE_SIZE: Final = int(os.getenv("DEFAULT_MAX_LRU_CACHE_SIZE", 64))
|
||
_REALTIME_BODY_CACHE_SIZE = 1000 # Keep realtime helper caches bounded; workloads rarely exceed 1k models/intents
|
||
INITIAL_RETRY_DELAY: Final = float(os.getenv("INITIAL_RETRY_DELAY", 0.5))
|
||
MAX_RETRY_DELAY: Final = float(os.getenv("MAX_RETRY_DELAY", 8.0))
|
||
JITTER: Final = float(os.getenv("JITTER", 0.75))
|
||
DEFAULT_IN_MEMORY_TTL = int(os.getenv("DEFAULT_IN_MEMORY_TTL", 5)) # default time to live for the in-memory cache
|
||
DEFAULT_MAX_REDIS_BATCH_CACHE_SIZE: Final = int(
|
||
os.getenv("DEFAULT_MAX_REDIS_BATCH_CACHE_SIZE", 1000)
|
||
) # default max size for redis batch cache
|
||
DEFAULT_POLLING_INTERVAL: Final = float(
|
||
os.getenv("DEFAULT_POLLING_INTERVAL", 0.03)
|
||
) # default polling interval for the scheduler
|
||
AZURE_OPERATION_POLLING_TIMEOUT: Final = int(os.getenv("AZURE_OPERATION_POLLING_TIMEOUT", 120))
|
||
AZURE_DOCUMENT_INTELLIGENCE_API_VERSION: Final = str(os.getenv("AZURE_DOCUMENT_INTELLIGENCE_API_VERSION", "2024-11-30"))
|
||
AZURE_DOCUMENT_INTELLIGENCE_DEFAULT_DPI: Final = int(os.getenv("AZURE_DOCUMENT_INTELLIGENCE_DEFAULT_DPI", 96))
|
||
REDIS_SOCKET_TIMEOUT: Final = float(os.getenv("REDIS_SOCKET_TIMEOUT", 0.1))
|
||
CACHE_WRITE_SHUTDOWN_FLUSH_TIMEOUT_SECONDS: Final[float] = 5.0
|
||
REDIS_CONNECTION_POOL_TIMEOUT: Final = int(os.getenv("REDIS_CONNECTION_POOL_TIMEOUT", 5))
|
||
REDIS_CIRCUIT_BREAKER_FAILURE_THRESHOLD: Final = int(os.getenv("REDIS_CIRCUIT_BREAKER_FAILURE_THRESHOLD", 5))
|
||
REDIS_CIRCUIT_BREAKER_RECOVERY_TIMEOUT: Final = int(os.getenv("REDIS_CIRCUIT_BREAKER_RECOVERY_TIMEOUT", 60))
|
||
REDIS_CIRCUIT_BREAKER_ENABLED: Final = os.getenv("REDIS_CIRCUIT_BREAKER_ENABLED", "true").lower() == "true"
|
||
# minimum seconds a timeout-only failure streak must span before it can open the breaker,
|
||
# so one event-loop stall timing out many queued calls at once does not trip it
|
||
REDIS_CIRCUIT_BREAKER_TIMEOUT_MIN_DURATION: Final = float(os.getenv("REDIS_CIRCUIT_BREAKER_TIMEOUT_MIN_DURATION", 5.0))
|
||
REDIS_TIMEOUT_LOG_INTERVAL: Final = float(os.getenv("REDIS_TIMEOUT_LOG_INTERVAL", "5.0"))
|
||
# Seconds of idle before a Redis cluster connection is validated with a PING and
|
||
# reconnected if dead, so a connection silently dropped by a cluster restart
|
||
# (e.g. ElastiCache Serverless maintenance) is not reused while broken
|
||
REDIS_CLUSTER_HEALTH_CHECK_INTERVAL: Final = 25
|
||
# Default Redis major version to assume when version cannot be determined
|
||
# Using 7 as it's the modern version that supports LPOP with count parameter
|
||
DEFAULT_REDIS_MAJOR_VERSION: Final = int(os.getenv("DEFAULT_REDIS_MAJOR_VERSION", 7))
|
||
NON_LLM_CONNECTION_TIMEOUT: Final = int(
|
||
os.getenv("NON_LLM_CONNECTION_TIMEOUT", 15)
|
||
) # timeout for adjacent services (e.g. jwt auth)
|
||
MAX_EXCEPTION_MESSAGE_LENGTH: Final = int(os.getenv("MAX_EXCEPTION_MESSAGE_LENGTH", 2000))
|
||
MAX_STRING_LENGTH_PROMPT_IN_DB: Final = int(os.getenv("MAX_STRING_LENGTH_PROMPT_IN_DB", 2048))
|
||
BEDROCK_MAX_POLICY_SIZE: Final = int(os.getenv("BEDROCK_MAX_POLICY_SIZE", 75))
|
||
# One entry per distinct AWS credential-argument set. Per-user cost attribution passes the attributed
|
||
# identity as aws_session_name, so this bounds how many attributed identities keep a cached STS session.
|
||
BEDROCK_IAM_CACHE_MAX_ENTRIES: Final = 1000
|
||
# Single-flight lock stripes over that cache. Only keys landing on the same stripe wait for each
|
||
# other, so a burst of distinct identities still resolves its credentials in parallel.
|
||
BEDROCK_IAM_CACHE_FETCH_LOCK_STRIPES: Final = 64
|
||
# Retire a cached STS credential this many seconds before AWS expires it, so a request that reads it
|
||
# still has a usable credential for the whole call.
|
||
STS_CREDENTIAL_EXPIRY_SAFETY_MARGIN_SECONDS: Final = 60
|
||
BEDROCK_MIN_THINKING_BUDGET_TOKENS: Final = int(os.getenv("BEDROCK_MIN_THINKING_BUDGET_TOKENS", 1024))
|
||
# Anthropic's Messages API rejects thinking.budget_tokens < 1024.
|
||
ANTHROPIC_MIN_THINKING_BUDGET_TOKENS: Final = 1024
|
||
REPLICATE_POLLING_DELAY_SECONDS: Final = float(os.getenv("REPLICATE_POLLING_DELAY_SECONDS", 0.5))
|
||
DEFAULT_ANTHROPIC_CHAT_MAX_TOKENS: Final = int(os.getenv("DEFAULT_ANTHROPIC_CHAT_MAX_TOKENS", 4096))
|
||
DEFAULT_OCI_CHAT_MAX_TOKENS: Final = 4096
|
||
TOGETHER_AI_4_B: Final = int(os.getenv("TOGETHER_AI_4_B", 4))
|
||
TOGETHER_AI_8_B: Final = int(os.getenv("TOGETHER_AI_8_B", 8))
|
||
TOGETHER_AI_21_B: Final = int(os.getenv("TOGETHER_AI_21_B", 21))
|
||
TOGETHER_AI_41_B: Final = int(os.getenv("TOGETHER_AI_41_B", 41))
|
||
TOGETHER_AI_80_B: Final = int(os.getenv("TOGETHER_AI_80_B", 80))
|
||
TOGETHER_AI_110_B: Final = int(os.getenv("TOGETHER_AI_110_B", 110))
|
||
TOGETHER_AI_EMBEDDING_150_M: Final = int(os.getenv("TOGETHER_AI_EMBEDDING_150_M", 150))
|
||
TOGETHER_AI_EMBEDDING_350_M: Final = int(os.getenv("TOGETHER_AI_EMBEDDING_350_M", 350))
|
||
QDRANT_SCALAR_QUANTILE: Final = float(os.getenv("QDRANT_SCALAR_QUANTILE", 0.99))
|
||
QDRANT_VECTOR_SIZE: Final = int(os.getenv("QDRANT_VECTOR_SIZE", 1536))
|
||
CACHED_STREAMING_CHUNK_DELAY: Final = float(os.getenv("CACHED_STREAMING_CHUNK_DELAY", 0.02))
|
||
AUDIO_SPEECH_CHUNK_SIZE: Final = int(
|
||
os.getenv("AUDIO_SPEECH_CHUNK_SIZE", 8192)
|
||
) # chunk_size for audio speech streaming. Balance between latency and memory usage
|
||
DEFAULT_MAX_TOKENS_FOR_TRITON: Final = int(os.getenv("DEFAULT_MAX_TOKENS_FOR_TRITON", 2000))
|
||
#### Networking settings ####
|
||
# Sentinel used when `REQUEST_TIMEOUT` is unset: `litellm.request_timeout` keeps this
|
||
# value so longer-running surfaces (Router `timeout or litellm.request_timeout`,
|
||
# speech/TTS, responses, vector stores, etc.) get a long HTTP deadline. Chat
|
||
# `completion()` maps this sentinel down to 600s when the caller did not set a
|
||
# per-request/model timeout—see ``CompletionTimeout.resolve`` in completion_timeout.py. MCP uses
|
||
# dedicated timeouts (e.g. `MCP_CLIENT_TIMEOUT`), not `request_timeout`.
|
||
DEFAULT_REQUEST_TIMEOUT_SECONDS: Final[float] = 6000.0
|
||
# Pair used for default httpx clients when no custom timeout is passed: read/write
|
||
# deadline and connect handshake (see ``http_handler`` cached handler paths).
|
||
COMPLETION_HTTP_FALLBACK_SECONDS: Final[float] = 600.0
|
||
HTTP_HANDLER_CONNECT_TIMEOUT_SECONDS: Final[float] = 5.0
|
||
SEMANTIC_CACHE_EMBEDDING_TIMEOUT_SECONDS: Final[float] = float(
|
||
os.getenv("SEMANTIC_CACHE_EMBEDDING_TIMEOUT_SECONDS", "5.0")
|
||
)
|
||
request_timeout: float = float(os.getenv("REQUEST_TIMEOUT", str(int(DEFAULT_REQUEST_TIMEOUT_SECONDS))))
|
||
request_timeout_explicitly_set: bool = "REQUEST_TIMEOUT" in os.environ
|
||
DEFAULT_A2A_AGENT_TIMEOUT: Final[float] = float(os.getenv("DEFAULT_A2A_AGENT_TIMEOUT", 6000)) # 10 minutes
|
||
AGENT_KILL_SWITCH_TIMEOUT_SECONDS: Final = 10.0
|
||
AGENT_KILL_SWITCH_RESPONSE_BODY_MAX_CHARS: Final = 2000
|
||
# Patterns that indicate a localhost/internal URL in A2A agent cards that should be
|
||
# replaced with the original base_url. This is a common misconfiguration where
|
||
# developers deploy agents with development URLs in their agent cards.
|
||
LOCALHOST_URL_PATTERNS: Final[list[str]] = [
|
||
"localhost",
|
||
"127.0.0.1",
|
||
"0.0.0.0",
|
||
"[::1]", # IPv6 localhost
|
||
]
|
||
# Patterns in error messages that indicate a connection failure
|
||
CONNECTION_ERROR_PATTERNS: Final[list[str]] = [
|
||
"connect",
|
||
"connection",
|
||
"network",
|
||
"refused",
|
||
]
|
||
STREAM_SSE_DONE_STRING: Final[str] = "[DONE]"
|
||
STREAM_SSE_DATA_PREFIX: Final[str] = "data: "
|
||
STREAM_SSE_KEEPALIVE_PING_CHUNK: Final[str] = 'event: ping\ndata: {"type": "ping"}\n\n'
|
||
STREAM_SSE_KEEPALIVE_PING_BYTES: Final[bytes] = STREAM_SSE_KEEPALIVE_PING_CHUNK.encode("utf-8")
|
||
### SPEND TRACKING ###
|
||
DEFAULT_REPLICATE_GPU_PRICE_PER_SECOND: Final = float(
|
||
os.getenv("DEFAULT_REPLICATE_GPU_PRICE_PER_SECOND", 0.001400)
|
||
) # price per second for a100 80GB
|
||
FIREWORKS_AI_56_B_MOE: Final = int(os.getenv("FIREWORKS_AI_56_B_MOE", 56))
|
||
FIREWORKS_AI_176_B_MOE: Final = int(os.getenv("FIREWORKS_AI_176_B_MOE", 176))
|
||
FIREWORKS_AI_4_B: Final = int(os.getenv("FIREWORKS_AI_4_B", 4))
|
||
FIREWORKS_AI_16_B: Final = int(os.getenv("FIREWORKS_AI_16_B", 16))
|
||
FIREWORKS_AI_80_B: Final = int(os.getenv("FIREWORKS_AI_80_B", 80))
|
||
# https://docs.fireworks.ai/guides/prompt-caching (accessed 2026-09-19): serverless cached prompt tokens
|
||
# default to a 50% discount off the input rate
|
||
FIREWORKS_AI_DEFAULT_CACHE_READ_RATE_RATIO: Final = 0.5
|
||
#### Logging callback constants ####
|
||
REDACTED_BY_LITELM_STRING: Final = "REDACTED_BY_LITELM"
|
||
MAX_LANGFUSE_INITIALIZED_CLIENTS: Final = int(os.getenv("MAX_LANGFUSE_INITIALIZED_CLIENTS", 50))
|
||
LANGFUSE_SHUTDOWN_FLUSH_TIMEOUT_MILLIS: Final = 10_000
|
||
# Backpressure + lifetime bounds for the /v1/messages streaming relay (see
|
||
# BaseAnthropicMessagesStreamingIterator.async_sse_wrapper). The relay queue is
|
||
# bounded so a slow client throttles the upstream pump instead of letting it
|
||
# buffer the whole response in memory; the detached-drain cap bounds how many
|
||
# post-disconnect drains may run concurrently so client behavior can't create
|
||
# unbounded worker state.
|
||
ANTHROPIC_MESSAGES_STREAM_RELAY_QUEUE_MAXSIZE: Final = int(
|
||
os.getenv("ANTHROPIC_MESSAGES_STREAM_RELAY_QUEUE_MAXSIZE", "1024")
|
||
)
|
||
# Setting this to 0 disables detached draining entirely: every post-disconnect
|
||
# pump bills whatever partial output it has already collected and aborts the
|
||
# upstream stream immediately, instead of continuing to drain for the real
|
||
# terminal usage.
|
||
ANTHROPIC_MESSAGES_MAX_DETACHED_STREAM_DRAINS: Final = int(
|
||
os.getenv("ANTHROPIC_MESSAGES_MAX_DETACHED_STREAM_DRAINS", "100")
|
||
)
|
||
LOGGING_WORKER_CONCURRENCY: Final = int(os.getenv("LOGGING_WORKER_CONCURRENCY", 100)) # Must be above 0
|
||
LOGGING_WORKER_MAX_QUEUE_SIZE: Final = int(os.getenv("LOGGING_WORKER_MAX_QUEUE_SIZE", 50_000))
|
||
LOGGING_WORKER_MAX_TIME_PER_COROUTINE: Final = float(os.getenv("LOGGING_WORKER_MAX_TIME_PER_COROUTINE", 20.0))
|
||
LOGGING_WORKER_TIMEOUT_SUMMARY_WINDOW_SECONDS: Final = 5.0
|
||
LOGGING_WORKER_CLEAR_PERCENTAGE: Final = int(
|
||
os.getenv("LOGGING_WORKER_CLEAR_PERCENTAGE", 50)
|
||
) # Percentage of queue to clear (default: 50%)
|
||
MAX_ITERATIONS_TO_CLEAR_QUEUE: Final = int(os.getenv("MAX_ITERATIONS_TO_CLEAR_QUEUE", 200))
|
||
MAX_TIME_TO_CLEAR_QUEUE: Final = float(os.getenv("MAX_TIME_TO_CLEAR_QUEUE", 5.0))
|
||
LOGGING_WORKER_AGGRESSIVE_CLEAR_COOLDOWN_SECONDS: Final = float(
|
||
os.getenv("LOGGING_WORKER_AGGRESSIVE_CLEAR_COOLDOWN_SECONDS", 0.5)
|
||
) # Cooldown time in seconds before allowing another aggressive clear (default: 0.5s)
|
||
LOGGING_EXECUTOR_MAX_THREADS: Final = get_env_int("LOGGING_EXECUTOR_MAX_THREADS", 100)
|
||
LOGGING_EXECUTOR_MAX_PENDING_TASKS: Final = get_env_int("LOGGING_EXECUTOR_MAX_PENDING_TASKS", 10_000)
|
||
LOGGING_EXECUTOR_DROPPED_TASK_LOG_INTERVAL_SECONDS: Final = 30.0
|
||
AWS_SIGNING_MAX_THREADS: Final = 16
|
||
PROMPT_INJECTION_HEURISTICS_MAX_THREADS: Final = max(1, get_env_int("PROMPT_INJECTION_HEURISTICS_MAX_THREADS", 1))
|
||
DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE: Final = os.getenv(
|
||
"DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE", "streaming.chunk.yield"
|
||
)
|
||
|
||
LITELLM_HTTP_STATUS_CLIENT_DISCONNECTED: Final = 499
|
||
|
||
EMAIL_BUDGET_ALERT_TTL: Final = int(os.getenv("EMAIL_BUDGET_ALERT_TTL", 24 * 60 * 60)) # 24 hours in seconds
|
||
EMAIL_BUDGET_ALERT_MAX_SPEND_ALERT_PERCENTAGE: Final = float(
|
||
os.getenv("EMAIL_BUDGET_ALERT_MAX_SPEND_ALERT_PERCENTAGE", 0.8)
|
||
) # 80% of max budget
|
||
############### LLM Provider Constants ###############
|
||
### ANTHROPIC CONSTANTS ###
|
||
ANTHROPIC_TOKEN_COUNTING_BETA_VERSION = os.getenv("ANTHROPIC_TOKEN_COUNTING_BETA_VERSION", "token-counting-2024-11-01")
|
||
ANTHROPIC_SKILLS_API_BETA_VERSION: Final = "skills-2025-10-02"
|
||
ANTHROPIC_BATCHES_ROUTE: Final = "/v1/messages/batches"
|
||
VERTEX_BATCH_PREDICTION_JOBS_ROUTE: Final = "batchPredictionJobs"
|
||
ANTHROPIC_WEB_SEARCH_TOOL_MAX_USES: Final = {
|
||
"low": 1,
|
||
"medium": 5,
|
||
"high": 10,
|
||
}
|
||
|
||
# LiteLLM standard web search tool name
|
||
# Used for web search interception across providers
|
||
LITELLM_WEB_SEARCH_TOOL_NAME: Final = "litellm_web_search"
|
||
|
||
DEFAULT_IMAGE_ENDPOINT_MODEL: Final = "dall-e-2"
|
||
DEFAULT_VIDEO_ENDPOINT_MODEL: Final = "sora-2"
|
||
|
||
DEFAULT_GOOGLE_VIDEO_DURATION_SECONDS: Final = int(os.getenv("DEFAULT_GOOGLE_VIDEO_DURATION_SECONDS", 8))
|
||
|
||
### DATAFORSEO CONSTANTS ###
|
||
DEFAULT_DATAFORSEO_LOCATION_CODE: Final = int(
|
||
os.getenv("DEFAULT_DATAFORSEO_LOCATION_CODE", 2250)
|
||
) # Default to France (2250) - lower number, commonly used location
|
||
|
||
LITELLM_CHAT_PROVIDERS: Final = [
|
||
"openai",
|
||
"openai_like",
|
||
"bytez",
|
||
"gdc",
|
||
"xai",
|
||
"custom_openai",
|
||
"text-completion-openai",
|
||
"cohere",
|
||
"cohere_chat",
|
||
"clarifai",
|
||
"anthropic",
|
||
"anthropic_text",
|
||
"replicate",
|
||
"huggingface",
|
||
"together_ai",
|
||
"datarobot",
|
||
"helicone",
|
||
"openrouter",
|
||
"cometapi",
|
||
"vertex_ai",
|
||
"vertex_ai_beta",
|
||
"gemini",
|
||
"ai21",
|
||
"baseten",
|
||
"azure",
|
||
"azure_text",
|
||
"azure_ai",
|
||
"sagemaker",
|
||
"sagemaker_chat",
|
||
"sagemaker_nova",
|
||
"bedrock",
|
||
"vllm",
|
||
"nlp_cloud",
|
||
"petals",
|
||
"oobabooga",
|
||
"ollama",
|
||
"ollama_chat",
|
||
"deepinfra",
|
||
"perplexity",
|
||
"mistral",
|
||
"groq",
|
||
"gigachat",
|
||
"nvidia_nim",
|
||
"cerebras",
|
||
"nadir",
|
||
"baseten",
|
||
"ai21_chat",
|
||
"volcengine",
|
||
"codestral",
|
||
"text-completion-codestral",
|
||
"text-completion-inception",
|
||
"deepseek",
|
||
"tencent",
|
||
"sambanova",
|
||
"maritalk",
|
||
"cloudflare",
|
||
"fireworks_ai",
|
||
"friendliai",
|
||
"watsonx",
|
||
"watsonx_text",
|
||
"triton",
|
||
"predibase",
|
||
"databricks",
|
||
"empower",
|
||
"github",
|
||
"custom",
|
||
"litellm_proxy",
|
||
"hosted_vllm",
|
||
"llamafile",
|
||
"lm_studio",
|
||
"galadriel",
|
||
"gradient_ai",
|
||
"github_copilot", # GitHub Copilot Chat API
|
||
"chatgpt", # ChatGPT subscription API
|
||
"novita",
|
||
"meta_llama",
|
||
"featherless_ai",
|
||
"nscale",
|
||
"nebius",
|
||
"dashscope",
|
||
"qwencloud",
|
||
"qwen_ai_platform",
|
||
"modelscope",
|
||
"moonshot",
|
||
"publicai",
|
||
"v0",
|
||
"heroku",
|
||
"oci",
|
||
"morph",
|
||
"lambda_ai",
|
||
"inception",
|
||
"vercel_ai_gateway",
|
||
"wandb",
|
||
"edenai",
|
||
"ovhcloud",
|
||
"lemonade",
|
||
"docker_model_runner",
|
||
"amazon_nova",
|
||
]
|
||
|
||
# Resolving these providers runs an OAuth device flow (their provider info IS the login), so any
|
||
# metadata or capability lookup against them can block for minutes waiting on a human.
|
||
PROVIDERS_THAT_AUTHENTICATE_ON_PROVIDER_INFO: Final = frozenset(
|
||
{
|
||
"github_copilot",
|
||
"chatgpt",
|
||
}
|
||
)
|
||
|
||
LITELLM_EMBEDDING_PROVIDERS_SUPPORTING_INPUT_ARRAY_OF_TOKENS: Final = [
|
||
"openai",
|
||
"azure",
|
||
"hosted_vllm",
|
||
"nebius",
|
||
]
|
||
|
||
|
||
OPENAI_CHAT_COMPLETION_PARAMS: Final = [
|
||
"functions",
|
||
"function_call",
|
||
"temperature",
|
||
"temperature",
|
||
"top_p",
|
||
"n",
|
||
"stream",
|
||
"stream_options",
|
||
"stop",
|
||
"max_completion_tokens",
|
||
"modalities",
|
||
"prediction",
|
||
"audio",
|
||
"max_tokens",
|
||
"presence_penalty",
|
||
"frequency_penalty",
|
||
"logit_bias",
|
||
"user",
|
||
"request_timeout",
|
||
"api_base",
|
||
"api_version",
|
||
"api_key",
|
||
"deployment_id",
|
||
"organization",
|
||
"base_url",
|
||
"default_headers",
|
||
"timeout",
|
||
"response_format",
|
||
"seed",
|
||
"tools",
|
||
"tool_choice",
|
||
"max_retries",
|
||
"parallel_tool_calls",
|
||
"logprobs",
|
||
"top_logprobs",
|
||
"reasoning_effort",
|
||
"extra_headers",
|
||
"thinking",
|
||
"web_search_options",
|
||
"include_server_side_tool_invocations",
|
||
"service_tier",
|
||
"prompt_cache_key",
|
||
"prompt_cache_retention",
|
||
"safety_identifier",
|
||
"verbosity",
|
||
"store",
|
||
]
|
||
|
||
OPENAI_TRANSCRIPTION_PARAMS: Final = [
|
||
"language",
|
||
"response_format",
|
||
"timestamp_granularities",
|
||
]
|
||
|
||
OPENAI_EMBEDDING_PARAMS: Final = ["dimensions", "encoding_format", "user"]
|
||
|
||
DEFAULT_EMBEDDING_PARAM_VALUES: Final = {
|
||
**{k: None for k in OPENAI_EMBEDDING_PARAMS},
|
||
"model": None,
|
||
"custom_llm_provider": "",
|
||
"input": None,
|
||
}
|
||
|
||
DEFAULT_CHAT_COMPLETION_PARAM_VALUES: Final = {
|
||
"functions": None,
|
||
"function_call": None,
|
||
"temperature": None,
|
||
"top_p": None,
|
||
"n": None,
|
||
"stream": None,
|
||
"stream_options": None,
|
||
"stop": None,
|
||
"max_tokens": None,
|
||
"max_completion_tokens": None,
|
||
"modalities": None,
|
||
"prediction": None,
|
||
"audio": None,
|
||
"presence_penalty": None,
|
||
"frequency_penalty": None,
|
||
"logit_bias": None,
|
||
"user": None,
|
||
"model": None,
|
||
"custom_llm_provider": "",
|
||
"response_format": None,
|
||
"seed": None,
|
||
"tools": None,
|
||
"tool_choice": None,
|
||
"max_retries": None,
|
||
"logprobs": None,
|
||
"top_logprobs": None,
|
||
"extra_headers": None,
|
||
"api_version": None,
|
||
"parallel_tool_calls": None,
|
||
"drop_params": None,
|
||
"allowed_openai_params": None,
|
||
"additional_drop_params": None,
|
||
"messages": None,
|
||
"reasoning_effort": None,
|
||
"verbosity": None,
|
||
"thinking": None,
|
||
"web_search_options": None,
|
||
"include_server_side_tool_invocations": None,
|
||
"service_tier": None,
|
||
"safety_identifier": None,
|
||
"prompt_cache_key": None,
|
||
"prompt_cache_retention": None,
|
||
"store": None,
|
||
"metadata": None,
|
||
"context_management": None,
|
||
}
|
||
|
||
openai_compatible_endpoints: Final[list] = [
|
||
"api.perplexity.ai",
|
||
"api.endpoints.anyscale.com/v1",
|
||
"api.deepinfra.com/v1/openai",
|
||
"api.mistral.ai/v1",
|
||
"codestral.mistral.ai/v1/chat/completions",
|
||
"codestral.mistral.ai/v1/fim/completions",
|
||
"api.groq.com/openai/v1",
|
||
"https://integrate.api.nvidia.com/v1",
|
||
NADIR_DEFAULT_API_BASE,
|
||
"api.deepseek.com/v1",
|
||
"api.together.ai/v1",
|
||
"api.together.xyz/v1",
|
||
"app.empower.dev/api/v1",
|
||
"https://api.friendli.ai/serverless/v1",
|
||
"api.sambanova.ai/v1",
|
||
"api.x.ai/v1",
|
||
"ollama.com",
|
||
"api.galadriel.ai/v1",
|
||
"api.llama.com/compat/v1/",
|
||
"api.featherless.ai/v1",
|
||
"inference.api.nscale.com/v1",
|
||
"api.studio.nebius.ai/v1",
|
||
"https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
|
||
"https://dashscope.aliyuncs.com/compatible-mode/v1",
|
||
"https://api-inference.modelscope.cn/v1",
|
||
"https://api.moonshot.ai/v1",
|
||
"https://api.publicai.co/v1",
|
||
"https://api.synthetic.new/openai/v1",
|
||
"https://serverless.tensormesh.ai/v1",
|
||
"https://api.stima.tech/v1",
|
||
"https://nano-gpt.com/api/v1",
|
||
"https://api.poe.com/v1",
|
||
"https://llm.chutes.ai/v1/",
|
||
"https://api.v0.dev/v1",
|
||
"https://api.morphllm.com/v1",
|
||
"https://api.lambda.ai/v1",
|
||
"https://api.inceptionlabs.ai/v1",
|
||
"https://api.hyperbolic.xyz/v1",
|
||
"https://ai-gateway.helicone.ai/",
|
||
"https://ai-gateway.vercel.sh/v1",
|
||
"https://api.edenai.run/v3",
|
||
"https://api.inference.wandb.ai/v1",
|
||
"https://api.clarifai.com/v2/ext/openai/v1",
|
||
"https://api.libertai.io/v1",
|
||
"https://pinstripes.io/v1",
|
||
"https://api.meta.ai/v1",
|
||
"https://api.cognition.ai/v1",
|
||
"https://api.scx.ai/v1",
|
||
"https://gigachat.devices.sberbank.ru/api/v1",
|
||
]
|
||
|
||
|
||
openai_compatible_providers: Final[list] = [
|
||
"anyscale",
|
||
"groq",
|
||
"nvidia_nim",
|
||
"cerebras",
|
||
"baseten",
|
||
"sambanova",
|
||
"ai21_chat",
|
||
"ai21",
|
||
"volcengine",
|
||
"codestral",
|
||
"deepseek",
|
||
"tencent",
|
||
"deepinfra",
|
||
"perplexity",
|
||
"xinference",
|
||
"xai",
|
||
"zai",
|
||
"together_ai",
|
||
"fireworks_ai",
|
||
"empower",
|
||
"friendliai",
|
||
"azure_ai",
|
||
"github",
|
||
"litellm_proxy",
|
||
"hosted_vllm",
|
||
"llamafile",
|
||
"lm_studio",
|
||
"galadriel",
|
||
"github_copilot", # GitHub Copilot Chat API
|
||
"chatgpt", # ChatGPT subscription API
|
||
"novita",
|
||
"meta_llama",
|
||
"publicai", # PublicAI - JSON-configured provider
|
||
"synthetic", # Synthetic - JSON-configured provider
|
||
"tensormesh", # Tensormesh - JSON-configured provider
|
||
"apertis", # Apertis - JSON-configured provider
|
||
"nano-gpt", # Nano-GPT - JSON-configured provider
|
||
"poe", # Poe - JSON-configured provider
|
||
"chutes", # Chutes - JSON-configured provider
|
||
"parasail", # Parasail - JSON-configured provider
|
||
"libertai", # LibertAI - JSON-configured provider
|
||
"featherless_ai",
|
||
"nscale",
|
||
"nebius",
|
||
"dashscope",
|
||
"qwencloud",
|
||
"qwen_ai_platform",
|
||
"modelscope",
|
||
"moonshot",
|
||
"v0",
|
||
"helicone",
|
||
"morph",
|
||
"lambda_ai",
|
||
"inception",
|
||
"hyperbolic",
|
||
"vercel_ai_gateway",
|
||
"aiml",
|
||
"edenai",
|
||
"wandb",
|
||
"cometapi",
|
||
"clarifai",
|
||
"docker_model_runner",
|
||
"ragflow",
|
||
"pinstripes", # Pinstripes - JSON-configured provider
|
||
"darkbloom",
|
||
"meta", # Meta Model API (Muse Spark) - JSON-configured provider
|
||
"cognition",
|
||
"scx-ai",
|
||
]
|
||
|
||
OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS: Final = frozenset({"openai"} | frozenset(openai_compatible_providers))
|
||
|
||
openai_text_completion_compatible_providers: Final[list] = [ # providers that support `/v1/completions`
|
||
"together_ai",
|
||
"fireworks_ai",
|
||
"hosted_vllm",
|
||
"meta_llama",
|
||
"llamafile",
|
||
"featherless_ai",
|
||
"nebius",
|
||
"dashscope",
|
||
"qwencloud",
|
||
"qwen_ai_platform",
|
||
"modelscope",
|
||
"moonshot",
|
||
"publicai",
|
||
"synthetic",
|
||
"tensormesh",
|
||
"apertis",
|
||
"nano-gpt",
|
||
"poe",
|
||
"chutes",
|
||
"v0",
|
||
"lambda_ai",
|
||
"hyperbolic",
|
||
"wandb",
|
||
]
|
||
_openai_like_providers: Final[list] = [
|
||
"predibase",
|
||
"databricks",
|
||
"lemonade",
|
||
"watsonx",
|
||
] # private helper. similar to openai but require some custom auth / endpoint handling, so can't use the openai sdk
|
||
# well supported replicate llms
|
||
replicate_models: Final[set] = set(
|
||
[
|
||
# llama replicate supported LLMs
|
||
"replicate/llama-2-70b-chat:2796ee9483c3fd7aa2e171d38f4ca12251a30609463dcfd4cd76703f22e96cdf",
|
||
"a16z-infra/llama-2-13b-chat:2a7f981751ec7fdf87b5b91ad4db53683a98082e9ff7bfd12c8cd5ea85980a52",
|
||
"meta/codellama-13b:1c914d844307b0588599b8393480a3ba917b660c7e9dfae681542b5325f228db",
|
||
# Vicuna
|
||
"replicate/vicuna-13b:6282abe6a492de4145d7bb601023762212f9ddbbe78278bd6771c8b3b2f2a13b",
|
||
"joehoover/instructblip-vicuna13b:c4c54e3c8c97cd50c2d2fec9be3b6065563ccf7d43787fb99f84151b867178fe",
|
||
# Flan T-5
|
||
"daanelson/flan-t5-large:ce962b3f6792a57074a601d3979db5839697add2e4e02696b3ced4c022d4767f",
|
||
# Others
|
||
"replicate/dolly-v2-12b:ef0e1aefc61f8e096ebe4db6b2bacc297daf2ef6899f0f7e001ec445893500e5",
|
||
"replit/replit-code-v1-3b:b84f4c074b807211cd75e3e8b1589b6399052125b4c27106e43d47189e8415ad",
|
||
]
|
||
)
|
||
|
||
clarifai_models: Final[set] = set(
|
||
[
|
||
"clarifai/openai.chat-completion.gpt-oss-20b",
|
||
"clarifai/qwen.qwenLM.Qwen3-30B-A3B-Instruct-2507",
|
||
"clarifai/qwen.qwen3.qwen3-next-80B-A3B-Thinking",
|
||
"clarifai/openai.chat-completion.gpt-oss-120b",
|
||
"clarifai/qwen.qwenLM.Qwen3-30B-A3B-Thinking-2507clarifai/openai.chat-completion.gpt-5-nano",
|
||
"clarifai/openai.chat-completion.gpt-4o",
|
||
"clarifai/gcp.generate.gemini-2_5-pro",
|
||
"clarifai/anthropic.completion.claude-sonnet-4",
|
||
"clarifai/xai.chat-completion.grok-2-vision-1212",
|
||
"clarifai/openbmb.miniCPM.MiniCPM-o-2_6-language",
|
||
"clarifai/microsoft.text-generation.Phi-4-reasoning-plus",
|
||
"clarifai/openbmb.miniCPM.MiniCPM3-4B",
|
||
"clarifai/openbmb.miniCPM.MiniCPM4-8B",
|
||
"clarifai/xai.chat-completion.grok-2-1212",
|
||
"clarifai/anthropic.completion.claude-opus-4",
|
||
"clarifai/xai.chat-completion.grok-code-fast-1",
|
||
"clarifai/qwen.qwenCoder.Qwen3-Coder-30B-A3B-Instruct",
|
||
"clarifai/deepseek-ai.deepseek-chat.DeepSeek-R1-0528-Qwen3-8B",
|
||
"clarifai/openai.chat-completion.gpt-5-mini",
|
||
"clarifai/microsoft.text-generation.phi-4",
|
||
"clarifai/openai.chat-completion.gpt-5",
|
||
"clarifai/meta.Llama-3.Llama-3_2-3B-Instruct",
|
||
"clarifai/xai.image-generation.grok-2-image-1212",
|
||
"clarifai/xai.chat-completion.grok-3",
|
||
"clarifai/openai.chat-completion.o3",
|
||
"clarifai/qwen.qwen-VL.Qwen2_5-VL-7B-Instruct",
|
||
"clarifai/qwen.qwenLM.Qwen3-14B",
|
||
"clarifai/qwen.qwenLM.QwQ-32B-AWQ",
|
||
"clarifai/anthropic.completion.claude-3_5-haiku",
|
||
"clarifai/anthropic.completion.claude-3_7-sonnet",
|
||
]
|
||
)
|
||
|
||
|
||
huggingface_models: Final[set] = set(
|
||
[
|
||
"meta-llama/Llama-2-7b-hf",
|
||
"meta-llama/Llama-2-7b-chat-hf",
|
||
"meta-llama/Llama-2-13b-hf",
|
||
"meta-llama/Llama-2-13b-chat-hf",
|
||
"meta-llama/Llama-2-70b-hf",
|
||
"meta-llama/Llama-2-70b-chat-hf",
|
||
"meta-llama/Llama-2-7b",
|
||
"meta-llama/Llama-2-7b-chat",
|
||
"meta-llama/Llama-2-13b",
|
||
"meta-llama/Llama-2-13b-chat",
|
||
"meta-llama/Llama-2-70b",
|
||
"meta-llama/Llama-2-70b-chat",
|
||
]
|
||
) # these have been tested on extensively. But by default all text2text-generation and text-generation models are supported by liteLLM. - https://docs.litellm.ai/docs/providers
|
||
empower_models: Final = set(
|
||
[
|
||
"empower/empower-functions",
|
||
"empower/empower-functions-small",
|
||
]
|
||
)
|
||
|
||
together_ai_models: Final[set] = set(
|
||
[
|
||
# llama llms - chat
|
||
"togethercomputer/llama-2-70b-chat",
|
||
# llama llms - language / instruct
|
||
"togethercomputer/llama-2-70b",
|
||
"togethercomputer/LLaMA-2-7B-32K",
|
||
"togethercomputer/Llama-2-7B-32K-Instruct",
|
||
"togethercomputer/llama-2-7b",
|
||
# falcon llms
|
||
"togethercomputer/falcon-40b-instruct",
|
||
"togethercomputer/falcon-7b-instruct",
|
||
# alpaca
|
||
"togethercomputer/alpaca-7b",
|
||
# chat llms
|
||
"HuggingFaceH4/starchat-alpha",
|
||
# code llms
|
||
"togethercomputer/CodeLlama-34b",
|
||
"togethercomputer/CodeLlama-34b-Instruct",
|
||
"togethercomputer/CodeLlama-34b-Python",
|
||
"defog/sqlcoder",
|
||
"NumbersStation/nsql-llama-2-7B",
|
||
"WizardLM/WizardCoder-15B-V1.0",
|
||
"WizardLM/WizardCoder-Python-34B-V1.0",
|
||
# language llms
|
||
"NousResearch/Nous-Hermes-Llama2-13b",
|
||
"Austism/chronos-hermes-13b",
|
||
"upstage/SOLAR-0-70b-16bit",
|
||
"WizardLM/WizardLM-70B-V1.0",
|
||
]
|
||
)
|
||
# supports all together ai models, just pass in the model id e.g. completion(model="together_computer/replit_code_3b",...)
|
||
|
||
|
||
baseten_models: Final[set] = set(
|
||
[
|
||
"qvv0xeq",
|
||
"q841o8w",
|
||
"31dxrj3",
|
||
]
|
||
) # FALCON 7B # WizardLM # Mosaic ML
|
||
|
||
featherless_ai_models: Final[set] = set(
|
||
[
|
||
"featherless-ai/Qwerky-72B",
|
||
"featherless-ai/Qwerky-QwQ-32B",
|
||
"Qwen/Qwen2.5-72B-Instruct",
|
||
"all-hands/openhands-lm-32b-v0.1",
|
||
"Qwen/Qwen2.5-Coder-32B-Instruct",
|
||
"deepseek-ai/DeepSeek-V3-0324",
|
||
"mistralai/Mistral-Small-24B-Instruct-2501",
|
||
"mistralai/Mistral-Nemo-Instruct-2407",
|
||
"ProdeusUnity/Stellar-Odyssey-12b-v0.0",
|
||
]
|
||
)
|
||
|
||
nebius_models: Final[set] = set(
|
||
[
|
||
# deepseek models
|
||
"deepseek-ai/DeepSeek-R1-0528",
|
||
"deepseek-ai/DeepSeek-V3-0324",
|
||
"deepseek-ai/DeepSeek-V3",
|
||
"deepseek-ai/DeepSeek-R1",
|
||
"deepseek-ai/DeepSeek-R1-Distill-Llama-70B",
|
||
# google models
|
||
"google/gemma-2-2b-it",
|
||
"google/gemma-2-9b-it-fast",
|
||
# llama models
|
||
"meta-llama/Llama-3.3-70B-Instruct",
|
||
"meta-llama/Meta-Llama-3.1-70B-Instruct",
|
||
"meta-llama/Meta-Llama-3.1-8B-Instruct",
|
||
"meta-llama/Meta-Llama-3.1-405B-Instruct",
|
||
"NousResearch/Hermes-3-Llama-405B",
|
||
# microsoft models
|
||
"microsoft/phi-4",
|
||
# mistral models
|
||
"mistralai/Mistral-Nemo-Instruct-2407",
|
||
"mistralai/Devstral-Small-2505",
|
||
# moonshot models
|
||
"moonshotai/Kimi-K2-Instruct",
|
||
# nvidia models
|
||
"nvidia/Llama-3_1-Nemotron-Ultra-253B-v1",
|
||
"nvidia/Llama-3_3-Nemotron-Super-49B-v1",
|
||
# openai models
|
||
"openai/gpt-oss-120b",
|
||
"openai/gpt-oss-20b",
|
||
# qwen models
|
||
"Qwen/Qwen3-Coder-480B-A35B-Instruct",
|
||
"Qwen/Qwen3-235B-A22B-Instruct-2507",
|
||
"Qwen/Qwen3-235B-A22B",
|
||
"Qwen/Qwen3-30B-A3B",
|
||
"Qwen/Qwen3-32B",
|
||
"Qwen/Qwen3-14B",
|
||
"Qwen/Qwen3-4B-fast",
|
||
"Qwen/Qwen2.5-Coder-7B",
|
||
"Qwen/Qwen2.5-Coder-32B-Instruct",
|
||
"Qwen/Qwen2.5-72B-Instruct",
|
||
"Qwen/QwQ-32B",
|
||
"Qwen/Qwen3-30B-A3B-Thinking-2507",
|
||
"Qwen/Qwen3-30B-A3B-Instruct-2507",
|
||
# zai models
|
||
"zai-org/GLM-4.5",
|
||
"zai-org/GLM-4.5-Air",
|
||
# other models
|
||
"aaditya/Llama3-OpenBioLLM-70B",
|
||
"ProdeusUnity/Stellar-Odyssey-12b-v0.0",
|
||
"all-hands/openhands-lm-32b-v0.1",
|
||
]
|
||
)
|
||
|
||
dashscope_models: Final[frozenset] = frozenset(
|
||
[
|
||
"qwen-turbo",
|
||
"qwen-plus",
|
||
"qwen-max",
|
||
"qwen-turbo-latest",
|
||
"qwen-plus-latest",
|
||
"qwen-max-latest",
|
||
"qwq-32b",
|
||
"qwen3-235b-a22b",
|
||
"qwen3-32b",
|
||
"qwen3-30b-a3b",
|
||
]
|
||
)
|
||
|
||
qwencloud_models: Final[frozenset] = frozenset(dashscope_models)
|
||
|
||
qwen_ai_platform_models: Final[frozenset] = frozenset(dashscope_models)
|
||
|
||
nebius_embedding_models: Final[set] = set(
|
||
[
|
||
"BAAI/bge-en-icl",
|
||
"BAAI/bge-multilingual-gemma2",
|
||
"intfloat/e5-mistral-7b-instruct",
|
||
]
|
||
)
|
||
|
||
WANDB_MODELS: Final[set] = set(
|
||
[
|
||
# openai models
|
||
"openai/gpt-oss-120b",
|
||
"openai/gpt-oss-20b",
|
||
# zai-org models
|
||
"zai-org/GLM-4.5",
|
||
# Qwen models
|
||
"Qwen/Qwen3-235B-A22B-Instruct-2507",
|
||
"Qwen/Qwen3-Coder-480B-A35B-Instruct",
|
||
"Qwen/Qwen3-235B-A22B-Thinking-2507",
|
||
# moonshotai
|
||
"moonshotai/Kimi-K2-Instruct",
|
||
"moonshotai/Kimi-K2.5",
|
||
# MiniMaxAI
|
||
"MiniMaxAI/MiniMax-M2.5",
|
||
# meta models
|
||
"meta-llama/Llama-3.1-8B-Instruct",
|
||
"meta-llama/Llama-3.3-70B-Instruct",
|
||
"meta-llama/Llama-4-Scout-17B-16E-Instruct",
|
||
# deepseek-ai
|
||
"deepseek-ai/DeepSeek-V3.1",
|
||
"deepseek-ai/DeepSeek-R1-0528",
|
||
"deepseek-ai/DeepSeek-V3-0324",
|
||
# microsoft
|
||
"microsoft/Phi-4-mini-instruct",
|
||
]
|
||
)
|
||
|
||
modelscope_models: Final[set] = set(
|
||
[
|
||
# Qwen series models
|
||
"Qwen/Qwen3-0.6B",
|
||
"Qwen/Qwen3-1.7B",
|
||
"Qwen/Qwen3-4B",
|
||
"Qwen/Qwen3-8B",
|
||
"Qwen/Qwen3-14B",
|
||
"Qwen/Qwen3-30B-A3B",
|
||
"Qwen/Qwen3-32B",
|
||
"Qwen/Qwen3-235B-A22B",
|
||
"Qwen/Qwen3-235B-A22B-Instruct-2507",
|
||
"Qwen/Qwen3-235B-A22B-Thinking-2507",
|
||
"Qwen/Qwen3-30B-A3B-Thinking-2507",
|
||
"Qwen/Qwen3-Coder-30B-A3B-Instruct",
|
||
"Qwen/Qwen3-Coder-480B-A35B-Instruct",
|
||
"Qwen/Qwen3-Next-80B-A3B-Instruct",
|
||
"Qwen/Qwen3-Next-80B-A3B-Thinking",
|
||
"Qwen/Qwen3-VL-235B-A22B-Instruct",
|
||
"Qwen/Qwen3-VL-8B-Instruct",
|
||
"Qwen/Qwen3-VL-8B-Thinking",
|
||
"Qwen/Qwen3.5-122B-A10B",
|
||
"Qwen/Qwen3.5-27B",
|
||
"Qwen/Qwen3.5-35B-A3B",
|
||
"Qwen/Qwen3.5-397B-A17B",
|
||
"Qwen/QwQ-32B",
|
||
"Qwen/QwQ-32B-Preview",
|
||
"Qwen/QVQ-72B-Preview",
|
||
"Qwen/Qwen-Image-Edit",
|
||
# DeepSeek series models
|
||
"deepseek-ai/DeepSeek-R1-0528",
|
||
"deepseek-ai/DeepSeek-R1-Distill-Llama-70B",
|
||
"deepseek-ai/DeepSeek-R1-Distill-Llama-8B",
|
||
"deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B",
|
||
"deepseek-ai/DeepSeek-R1-Distill-Qwen-14B",
|
||
"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
|
||
"deepseek-ai/DeepSeek-R1-Distill-Qwen-7B",
|
||
"deepseek-ai/DeepSeek-V3.2",
|
||
"deepseek-ai/DeepSeek-V4-Flash",
|
||
]
|
||
)
|
||
|
||
BEDROCK_INVOKE_PROVIDERS_LITERAL = Literal[
|
||
"cohere",
|
||
"anthropic",
|
||
"mistral",
|
||
"amazon",
|
||
"meta",
|
||
"llama",
|
||
"ai21",
|
||
"nova",
|
||
"deepseek_r1",
|
||
"qwen3",
|
||
"qwen2",
|
||
"twelvelabs",
|
||
"openai",
|
||
"stability",
|
||
"moonshot",
|
||
]
|
||
|
||
BEDROCK_EMBEDDING_PROVIDERS_LITERAL = Literal[
|
||
"cohere",
|
||
"amazon",
|
||
"twelvelabs",
|
||
"nova",
|
||
]
|
||
|
||
BEDROCK_CONVERSE_MODELS: Final = [
|
||
"qwen.qwen3-coder-480b-a35b-v1:0",
|
||
"qwen.qwen3-coder-next",
|
||
"qwen.qwen3-235b-a22b-2507-v1:0",
|
||
"qwen.qwen3-coder-30b-a3b-v1:0",
|
||
"qwen.qwen3-32b-v1:0",
|
||
"deepseek.v3-v1:0",
|
||
"deepseek.v3.2",
|
||
"openai.gpt-oss-20b-1:0",
|
||
"openai.gpt-oss-120b-1:0",
|
||
"anthropic.claude-haiku-4-5-20251001-v1:0",
|
||
"anthropic.claude-sonnet-4-5-20250929-v1:0",
|
||
"anthropic.claude-fable-5-1",
|
||
"anthropic.claude-fable-5",
|
||
"anthropic.claude-sonnet-5",
|
||
"anthropic.claude-opus-5",
|
||
"anthropic.claude-opus-4-8",
|
||
"anthropic.claude-opus-4-7",
|
||
"anthropic.claude-opus-4-6-v1:0",
|
||
"anthropic.claude-opus-4-6-v1",
|
||
"anthropic.claude-sonnet-4-6",
|
||
"anthropic.claude-opus-4-1-20250805-v1:0",
|
||
"anthropic.claude-opus-4-20250514-v1:0",
|
||
"anthropic.claude-sonnet-4-20250514-v1:0",
|
||
"anthropic.claude-3-7-sonnet-20250219-v1:0",
|
||
"anthropic.claude-3-5-haiku-20241022-v1:0",
|
||
"anthropic.claude-3-5-sonnet-20241022-v2:0",
|
||
"anthropic.claude-3-5-sonnet-20240620-v1:0",
|
||
"anthropic.claude-3-opus-20240229-v1:0",
|
||
"anthropic.claude-3-sonnet-20240229-v1:0",
|
||
"anthropic.claude-3-haiku-20240307-v1:0",
|
||
"anthropic.claude-v2",
|
||
"anthropic.claude-v2:1",
|
||
"anthropic.claude-v1",
|
||
"anthropic.claude-instant-v1",
|
||
"ai21.jamba-instruct-v1:0",
|
||
"ai21.jamba-1-5-mini-v1:0",
|
||
"ai21.jamba-1-5-large-v1:0",
|
||
"meta.llama3-70b-instruct-v1:0",
|
||
"meta.llama3-8b-instruct-v1:0",
|
||
"meta.llama3-1-8b-instruct-v1:0",
|
||
"meta.llama3-1-70b-instruct-v1:0",
|
||
"meta.llama3-1-405b-instruct-v1:0",
|
||
"meta.llama3-70b-instruct-v1:0",
|
||
"mistral.mistral-large-2407-v1:0",
|
||
"mistral.mistral-large-2402-v1:0",
|
||
"mistral.mistral-small-2402-v1:0",
|
||
"meta.llama3-2-1b-instruct-v1:0",
|
||
"meta.llama3-2-3b-instruct-v1:0",
|
||
"meta.llama3-2-11b-instruct-v1:0",
|
||
"meta.llama3-2-90b-instruct-v1:0",
|
||
"amazon.nova-lite-v1:0",
|
||
"amazon.nova-2-lite-v1:0",
|
||
"amazon.nova-2-pro-preview-20251202-v1:0",
|
||
"amazon.nova-pro-v1:0",
|
||
"writer.palmyra-x4-v1:0",
|
||
"writer.palmyra-x5-v1:0",
|
||
"minimax.minimax-m2.1",
|
||
"moonshotai.kimi-k2.5",
|
||
]
|
||
|
||
|
||
open_ai_embedding_models: Final[set] = set(["text-embedding-ada-002"])
|
||
cohere_embedding_models: Final[set] = set(
|
||
[
|
||
"embed-v4.0",
|
||
"embed-english-v3.0",
|
||
"embed-english-light-v3.0",
|
||
"embed-multilingual-v3.0",
|
||
"embed-english-v2.0",
|
||
"embed-english-light-v2.0",
|
||
"embed-multilingual-v2.0",
|
||
]
|
||
)
|
||
bedrock_embedding_models: Final[set] = set(
|
||
[
|
||
"amazon.titan-embed-text-v1",
|
||
"amazon.nova-2-multimodal-embeddings-v1:0",
|
||
"cohere.embed-english-v3",
|
||
"cohere.embed-multilingual-v3",
|
||
"cohere.embed-v4:0",
|
||
"twelvelabs.marengo-embed-2-7-v1:0",
|
||
"twelvelabs.marengo-embed-3-0-v1:0",
|
||
]
|
||
)
|
||
|
||
known_tokenizer_config: Final = {
|
||
"mistralai/Mistral-7B-Instruct-v0.1": {
|
||
"tokenizer": {
|
||
"chat_template": "{{ bos_token }}{% for message in messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if message['role'] == 'user' %}{{ '[INST] ' + message['content'] + ' [/INST]' }}{% elif message['role'] == 'assistant' %}{{ message['content'] + eos_token + ' ' }}{% else %}{{ raise_exception('Only user and assistant roles are supported!') }}{% endif %}{% endfor %}",
|
||
"bos_token": "<s>",
|
||
"eos_token": "</s>",
|
||
},
|
||
"status": "success",
|
||
},
|
||
"meta-llama/Meta-Llama-3-8B-Instruct": {
|
||
"tokenizer": {
|
||
"chat_template": "{% set loop_messages = messages %}{% for message in loop_messages %}{% set content = '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n'+ message['content'] | trim + '<|eot_id|>' %}{% if loop.index0 == 0 %}{% set content = bos_token + content %}{% endif %}{{ content }}{% endfor %}{{ '<|start_header_id|>assistant<|end_header_id|>\n\n' }}",
|
||
"bos_token": "<|begin_of_text|>",
|
||
"eos_token": "",
|
||
},
|
||
"status": "success",
|
||
},
|
||
"deepseek-r1/deepseek-r1-7b-instruct": {
|
||
"tokenizer": {
|
||
"add_bos_token": True,
|
||
"add_eos_token": False,
|
||
"bos_token": {
|
||
"__type": "AddedToken",
|
||
"content": "<|begin▁of▁sentence|>",
|
||
"lstrip": False,
|
||
"normalized": True,
|
||
"rstrip": False,
|
||
"single_word": False,
|
||
},
|
||
"clean_up_tokenization_spaces": False,
|
||
"eos_token": {
|
||
"__type": "AddedToken",
|
||
"content": "<|end▁of▁sentence|>",
|
||
"lstrip": False,
|
||
"normalized": True,
|
||
"rstrip": False,
|
||
"single_word": False,
|
||
},
|
||
"legacy": True,
|
||
"model_max_length": 16384,
|
||
"pad_token": {
|
||
"__type": "AddedToken",
|
||
"content": "<|end▁of▁sentence|>",
|
||
"lstrip": False,
|
||
"normalized": True,
|
||
"rstrip": False,
|
||
"single_word": False,
|
||
},
|
||
"sp_model_kwargs": {},
|
||
"unk_token": None,
|
||
"tokenizer_class": "LlamaTokenizerFast",
|
||
"chat_template": "{% if not add_generation_prompt is defined %}{% set add_generation_prompt = false %}{% endif %}{% set ns = namespace(is_first=false, is_tool=false, is_output_first=true, system_prompt='') %}{%- for message in messages %}{%- if message['role'] == 'system' %}{% set ns.system_prompt = message['content'] %}{%- endif %}{%- endfor %}{{bos_token}}{{ns.system_prompt}}{%- for message in messages %}{%- if message['role'] == 'user' %}{%- set ns.is_tool = false -%}{{'<|User|>' + message['content']}}{%- endif %}{%- if message['role'] == 'assistant' and message['content'] is none %}{%- set ns.is_tool = false -%}{%- for tool in message['tool_calls']%}{%- if not ns.is_first %}{{'<|Assistant|><|tool▁calls▁begin|><|tool▁call▁begin|>' + tool['type'] + '<|tool▁sep|>' + tool['function']['name'] + '\\n' + '```json' + '\\n' + tool['function']['arguments'] + '\\n' + '```' + '<|tool▁call▁end|>'}}{%- set ns.is_first = true -%}{%- else %}{{'\\n' + '<|tool▁call▁begin|>' + tool['type'] + '<|tool▁sep|>' + tool['function']['name'] + '\\n' + '```json' + '\\n' + tool['function']['arguments'] + '\\n' + '```' + '<|tool▁call▁end|>'}}{{'<|tool▁calls▁end|><|end▁of▁sentence|>'}}{%- endif %}{%- endfor %}{%- endif %}{%- if message['role'] == 'assistant' and message['content'] is not none %}{%- if ns.is_tool %}{{'<|tool▁outputs▁end|>' + message['content'] + '<|end▁of▁sentence|>'}}{%- set ns.is_tool = false -%}{%- else %}{% set content = message['content'] %}{% if '</think>' in content %}{% set content = content.split('</think>')[-1] %}{% endif %}{{'<|Assistant|>' + content + '<|end▁of▁sentence|>'}}{%- endif %}{%- endif %}{%- if message['role'] == 'tool' %}{%- set ns.is_tool = true -%}{%- if ns.is_output_first %}{{'<|tool▁outputs▁begin|><|tool▁output▁begin|>' + message['content'] + '<|tool▁output▁end|>'}}{%- set ns.is_output_first = false %}{%- else %}{{'\\n<|tool▁output▁begin|>' + message['content'] + '<|tool▁output▁end|>'}}{%- endif %}{%- endif %}{%- endfor -%}{% if ns.is_tool %}{{'<|tool▁outputs▁end|>'}}{% endif %}{% if add_generation_prompt and not ns.is_tool %}{{'<|Assistant|><think>\\n'}}{% endif %}",
|
||
},
|
||
"status": "success",
|
||
},
|
||
}
|
||
|
||
|
||
OPENAI_FINISH_REASONS: Final = [
|
||
"stop",
|
||
"length",
|
||
"function_call",
|
||
"tool_calls",
|
||
"content_filter",
|
||
]
|
||
HUMANLOOP_PROMPT_CACHE_TTL_SECONDS: Final = int(os.getenv("HUMANLOOP_PROMPT_CACHE_TTL_SECONDS", 60)) # 1 minute
|
||
RESPONSE_FORMAT_TOOL_NAME = "json_tool_call" # default tool name used when converting response format to tool call
|
||
|
||
########################### Logging Callback Constants ###########################
|
||
AZURE_STORAGE_MSFT_VERSION: Final = "2019-07-07"
|
||
AZURE_STORAGE_DEFAULT_ENDPOINT_SUFFIX: Final = "core.windows.net"
|
||
PROMETHEUS_BUDGET_METRICS_REFRESH_INTERVAL_MINUTES: Final = int(
|
||
os.getenv("PROMETHEUS_BUDGET_METRICS_REFRESH_INTERVAL_MINUTES", 5)
|
||
)
|
||
CLOUDZERO_EXPORT_INTERVAL_MINUTES: Final = int(os.getenv("CLOUDZERO_EXPORT_INTERVAL_MINUTES", 60))
|
||
MCP_TOOL_NAME_PREFIX: Final = "mcp_tool"
|
||
MAXIMUM_TRACEBACK_LINES_TO_LOG: Final = int(os.getenv("MAXIMUM_TRACEBACK_LINES_TO_LOG", 100))
|
||
PASSTHROUGH_UPSTREAM_ERROR_BODY_MAX_LOG_CHARS: Final = 4096
|
||
|
||
# Headers to control callbacks
|
||
X_LITELLM_DISABLE_CALLBACKS: Final = "x-litellm-disable-callbacks"
|
||
LITELLM_METADATA_FIELD: Final = "litellm_metadata"
|
||
OLD_LITELLM_METADATA_FIELD: Final = "metadata"
|
||
RETURN_RAW_MODEL_NAME_METADATA_KEY: Final = "_complexity_router_return_raw_model_name"
|
||
SESSION_DEPLOYMENT_AFFINITY_TTL_METADATA_KEY: Final = "_session_deployment_affinity_ttl"
|
||
OUTPUT_TOKEN_CEILING_PARAMS: Final = frozenset({"max_tokens", "max_completion_tokens", "max_output_tokens"})
|
||
CLIENT_OUTPUT_CEILING_METADATA_KEY: Final = "_client_output_ceiling"
|
||
CONSUMED_REQUEST_TAGS_METADATA_KEY: Final = "_consumed_request_tags"
|
||
ROUTING_REQUEST_TAGS_METADATA_KEY: Final = "_routing_request_tags"
|
||
ROUTER_USAGE_COUNTED_TOKENS_METADATA_KEY: Final = "_litellm_router_usage_counted_tokens"
|
||
INTERNAL_CALL_ORIGIN_METADATA_KEY: Final = "internal_call_origin"
|
||
SESSION_ID_GENERATED_METADATA_KEY: Final = "litellm_session_id_generated"
|
||
SESSION_ID_OMITTED_METADATA_KEY: Final = "litellm_session_id_omitted"
|
||
LITELLM_TRUNCATED_PAYLOAD_FIELD: Final = "litellm_truncated"
|
||
LITELLM_TRUNCATION_DB_SAFEGUARD_NOTE: Final = (
|
||
"Truncation is a DB storage safeguard. "
|
||
"Full, untruncated data is logged to logging callbacks (OTEL, Datadog, etc.). "
|
||
"To increase the truncation limit, set `MAX_STRING_LENGTH_PROMPT_IN_DB` in your env."
|
||
)
|
||
LITELLM_TRUNCATION_STDOUT_SAFEGUARD_NOTE: Final = (
|
||
"Truncation is a stdout logging safeguard. "
|
||
"Full, untruncated data is logged to logging callbacks (OTEL, Datadog, etc.) and at DEBUG level. "
|
||
"To increase the truncation limit, set `MAX_STRING_LENGTH_STDOUT_LOG` in your env."
|
||
)
|
||
|
||
########################### LiteLLM Proxy Specific Constants ###########################
|
||
########################################################################################
|
||
|
||
# Standard headers that are always checked for customer/end-user ID (no configuration required)
|
||
# These headers work out-of-the-box for tools like Claude Code that support custom headers
|
||
STANDARD_CUSTOMER_ID_HEADERS: Final = [
|
||
"x-litellm-customer-id",
|
||
"x-litellm-end-user-id",
|
||
]
|
||
MAX_SPENDLOG_ROWS_TO_QUERY: Final = int(
|
||
os.getenv("MAX_SPENDLOG_ROWS_TO_QUERY", 1_000_000)
|
||
) # if spendLogs has more than 1M rows, do not query the DB
|
||
DEFAULT_SOFT_BUDGET: Final = float(
|
||
os.getenv("DEFAULT_SOFT_BUDGET", 50.0)
|
||
) # by default all litellm proxy keys have a soft budget of 50.0
|
||
# makes it clear this is a rate limit error for a litellm virtual key
|
||
RATE_LIMIT_ERROR_MESSAGE_FOR_VIRTUAL_KEY: Final = "LiteLLM Virtual Key user_api_key_hash"
|
||
# Prefix of the 401 raised when a submitted virtual key is not shaped like one.
|
||
INVALID_VIRTUAL_KEY_ERROR_MESSAGE: Final = "LiteLLM Virtual Key expected"
|
||
# Attribute stamped on that 401 at its raise site so log routing recognises it by
|
||
# provenance. Message text is caller-influenceable on other 401s, so it must not
|
||
# be used to classify.
|
||
INVALID_VIRTUAL_KEY_ERROR_MARKER: Final = "_litellm_invalid_virtual_key_error"
|
||
|
||
# Python garbage collection threshold configuration
|
||
# Format: "gen0,gen1,gen2" e.g., "1000,50,50"
|
||
PYTHON_GC_THRESHOLD: Final = os.getenv("PYTHON_GC_THRESHOLD")
|
||
|
||
# pass through route constansts
|
||
BEDROCK_AGENT_RUNTIME_PASS_THROUGH_ROUTES: Final = [
|
||
"agents/",
|
||
"knowledgebases/",
|
||
"flows/",
|
||
"retrieveAndGenerate/",
|
||
"rerank/",
|
||
"generateQuery/",
|
||
"optimize-prompt/",
|
||
]
|
||
|
||
|
||
# Headers that are safe to forward from incoming requests to Vertex AI
|
||
# Using an allowlist approach for security - only forward headers we explicitly trust
|
||
ALLOWED_VERTEX_AI_PASSTHROUGH_HEADERS: Final = {
|
||
"anthropic-beta", # Required for Anthropic features like extended context windows
|
||
"content-type", # Required for request body parsing
|
||
}
|
||
|
||
# Prefix for headers that should be forwarded to the provider with the prefix stripped
|
||
# e.g., 'x-pass-anthropic-beta: value' becomes 'anthropic-beta: value'
|
||
# Works for all LLM pass-through endpoints (Vertex AI, Anthropic, Bedrock, etc.)
|
||
PASS_THROUGH_HEADER_PREFIX: Final = "x-pass-"
|
||
|
||
AZURE_SPEECH_CUSTOM_LLM_PROVIDER: Final = "azure_speech"
|
||
AZURE_SPEECH_PASS_THROUGH_ROUTE_PREFIX: Final = "/azure_speech"
|
||
AZURE_SPEECH_SHORT_AUDIO_PATH_PREFIX: Final = "/speech/"
|
||
AZURE_SPEECH_BATCH_PATH_PREFIX: Final = "/speechtotext/"
|
||
AZURE_SPEECH_FAST_TRANSCRIPTION_PATH: Final = "/speechtotext/transcriptions:transcribe"
|
||
AZURE_SPEECH_STT_DOMAIN: Final = "stt.speech.microsoft.com"
|
||
AZURE_SPEECH_COGNITIVE_SERVICES_DOMAIN: Final = "api.cognitive.microsoft.com"
|
||
AZURE_SPEECH_SUBSCRIPTION_KEY_HEADER: Final = "Ocp-Apim-Subscription-Key"
|
||
AZURE_SPEECH_SHORT_AUDIO_MODEL: Final = "short-audio"
|
||
AZURE_SPEECH_BATCH_MODEL: Final = "batch-transcription"
|
||
AZURE_SPEECH_FAST_TRANSCRIPTION_MODEL: Final = "fast-transcription"
|
||
AZURE_SPEECH_PRICING_MODEL: Final = "azure/speech/azure-stt"
|
||
AZURE_SPEECH_TICKS_PER_SECOND: Final = 10_000_000
|
||
AZURE_SPEECH_MILLISECONDS_PER_SECOND: Final = 1_000
|
||
|
||
BASE_MCP_ROUTE: Final = "/mcp"
|
||
|
||
TRANSCRIBE_JOB_POLLING_INTERVAL_SECONDS: Final = 10.0
|
||
TRANSCRIBE_JOB_MAX_POLLING_ATTEMPTS: Final = 720 # 2 hours
|
||
TRANSCRIBE_MAX_MEDIA_DURATION_SECONDS: Final = 28800 # Amazon Transcribe quota: maximum audio file length
|
||
TRANSCRIBE_MAX_MEDIA_BYTES: Final = 2 * 1024**3 # Amazon Transcribe quota: maximum audio file size
|
||
TRANSCRIBE_MEDIA_DOWNLOAD_CONCURRENCY: Final = 1
|
||
TRANSCRIBE_MEDIA_FETCH_ATTEMPTS: Final = 3
|
||
TRANSCRIBE_MEDIA_LAST_MODIFIED_TOLERANCE_SECONDS: Final = 1.0 # S3 Last-Modified carries whole seconds only
|
||
TRANSCRIBE_MEASURABLE_MEDIA_FORMATS: Final = frozenset({"flac", "mp3", "ogg", "wav"}) # what libsndfile can read
|
||
|
||
BATCH_STATUS_POLL_INTERVAL_SECONDS: Final = int(os.getenv("BATCH_STATUS_POLL_INTERVAL_SECONDS", 3600)) # 1 hour
|
||
BATCH_STATUS_POLL_MAX_ATTEMPTS: Final = int(os.getenv("BATCH_STATUS_POLL_MAX_ATTEMPTS", 24)) # for 24 hours
|
||
BATCH_TPD_WINDOW_SECONDS: Final = 86400
|
||
BATCH_TPD_DESCRIPTOR_SUFFIX: Final = "_tpd"
|
||
|
||
HEALTH_CHECK_TIMEOUT_SECONDS: Final = int(os.getenv("HEALTH_CHECK_TIMEOUT_SECONDS", 60)) # 60 seconds
|
||
_background_health_check_max_tokens_env: Final = os.getenv("BACKGROUND_HEALTH_CHECK_MAX_TOKENS")
|
||
try:
|
||
_raw_background_health_check_max_tokens: Final = (
|
||
_background_health_check_max_tokens_env.strip() if _background_health_check_max_tokens_env is not None else ""
|
||
)
|
||
BACKGROUND_HEALTH_CHECK_MAX_TOKENS: int | None = (
|
||
int(_raw_background_health_check_max_tokens) if _raw_background_health_check_max_tokens else None
|
||
)
|
||
except (ValueError, TypeError):
|
||
BACKGROUND_HEALTH_CHECK_MAX_TOKENS = None
|
||
|
||
|
||
_background_health_check_max_tokens_reasoning_env: Final = os.getenv("BACKGROUND_HEALTH_CHECK_MAX_TOKENS_REASONING")
|
||
try:
|
||
_raw_background_health_check_max_tokens_reasoning: Final = (
|
||
_background_health_check_max_tokens_reasoning_env.strip()
|
||
if _background_health_check_max_tokens_reasoning_env is not None
|
||
else ""
|
||
)
|
||
BACKGROUND_HEALTH_CHECK_MAX_TOKENS_REASONING: int | None = (
|
||
int(_raw_background_health_check_max_tokens_reasoning)
|
||
if _raw_background_health_check_max_tokens_reasoning
|
||
else None
|
||
)
|
||
except (ValueError, TypeError):
|
||
BACKGROUND_HEALTH_CHECK_MAX_TOKENS_REASONING = None
|
||
|
||
LITTELM_INTERNAL_HEALTH_SERVICE_ACCOUNT_NAME: Final = "litellm-internal-health-check"
|
||
LITTELM_CLI_SERVICE_ACCOUNT_NAME: Final = "litellm-cli"
|
||
LITELLM_INTERNAL_JOBS_SERVICE_ACCOUNT_NAME: Final = "litellm_internal_jobs"
|
||
# Stable identifier substituted in place of the master key on UserAPIKeyAuth
|
||
# objects so the master key (or its hash) never propagates to spend logs,
|
||
# Prometheus metrics, audit trails, or any other downstream consumer.
|
||
LITELLM_PROXY_MASTER_KEY_ALIAS: Final = "litellm_proxy_master_key"
|
||
|
||
# Marker placed in ``model_call_details`` on a synthetic ``Logging`` object that
|
||
# records a proxy-gate error (auth/rate-limit rejection) for a request that never
|
||
# reached an upstream provider. Tracing callbacks key off it to avoid fabricating
|
||
# an LLM-call span for a call that did not happen. See
|
||
# ``ProxyLogging._handle_logging_proxy_only_error``.
|
||
LITELLM_LOGGING_NO_UPSTREAM_LLM_CALL: Final = "litellm_no_upstream_llm_call"
|
||
|
||
# Key/team metadata fields naming the OTel Resource ``service.name``, highest
|
||
# precedence first. Shared between the OTel v2 tenant router (which reads them
|
||
# out of ``user_api_key_auth_metadata``) and proxy request setup (which re-applies
|
||
# the key's values after the team metadata merge so a key outranks its team).
|
||
OTEL_SERVICE_NAME_METADATA_KEYS: Final = ("otel_service_name_override", "otel_service_name")
|
||
|
||
# Key Rotation Constants
|
||
LITELLM_KEY_ROTATION_ENABLED: Final = os.getenv("LITELLM_KEY_ROTATION_ENABLED", "false")
|
||
LITELLM_KEY_ROTATION_CHECK_INTERVAL_SECONDS: Final = int(
|
||
os.getenv("LITELLM_KEY_ROTATION_CHECK_INTERVAL_SECONDS", 86400)
|
||
) # 24 hours default
|
||
LITELLM_KEY_ROTATION_GRACE_PERIOD: Final[str] = os.getenv(
|
||
"LITELLM_KEY_ROTATION_GRACE_PERIOD", ""
|
||
) # Duration to keep old key valid after rotation (e.g. "24h", "2d"); empty = immediate revoke (default)
|
||
LITELLM_KEY_ROTATION_LOCK_TTL_SECONDS: Final = int(
|
||
os.getenv("LITELLM_KEY_ROTATION_LOCK_TTL_SECONDS", 600)
|
||
) # 10 minutes default — caps the deadlock window if a pod crashes mid-rotation
|
||
UI_SESSION_TOKEN_TEAM_ID: Final = "litellm-dashboard"
|
||
LITELLM_EXPIRED_UI_SESSION_KEY_CLEANUP_ENABLED = os.getenv("LITELLM_EXPIRED_UI_SESSION_KEY_CLEANUP_ENABLED", "false")
|
||
LITELLM_EXPIRED_UI_SESSION_KEY_CLEANUP_INTERVAL_SECONDS: Final = int(
|
||
os.getenv("LITELLM_EXPIRED_UI_SESSION_KEY_CLEANUP_INTERVAL_SECONDS", 86400)
|
||
) # 24 hours default
|
||
LITELLM_EXPIRED_UI_SESSION_KEY_CLEANUP_BATCH_SIZE: Final = int(
|
||
os.getenv("LITELLM_EXPIRED_UI_SESSION_KEY_CLEANUP_BATCH_SIZE", 1000)
|
||
)
|
||
LOGIN_THROTTLE_CACHE_KEY_PREFIX: Final = "login_fail"
|
||
LOGIN_THROTTLE_UNKNOWN_SOURCE: Final = "unknown"
|
||
LOGIN_THROTTLE_MAX_TRACKED_COUNTERS: Final = 20_000
|
||
LOGIN_THROTTLE_MAX_TRACKED_BLOCKS: Final = 10_000
|
||
LOGIN_THROTTLE_NOT_BLOCKED: Final = (0, 0)
|
||
LITELLM_PROXY_ADMIN_NAME: Final = "default_user_id"
|
||
LITELLM_PROXY_BUDGET_NAME: Final = "litellm-proxy-budget"
|
||
GLOBAL_PROXY_SPEND_CACHE_KEY: Final = f"{LITELLM_PROXY_ADMIN_NAME}:spend"
|
||
LITELLM_EXECUTED_BATCH_CONCURRENCY: Final = max(1, int(os.getenv("LITELLM_EXECUTED_BATCH_CONCURRENCY", "4")))
|
||
|
||
########################### CLI SSO AUTHENTICATION CONSTANTS ###########################
|
||
LITELLM_CLI_SOURCE_IDENTIFIER: Final = "litellm-cli"
|
||
LITELLM_CLI_SESSION_TOKEN_PREFIX: Final = "litellm-session-token"
|
||
CLI_SSO_SESSION_CACHE_KEY_PREFIX: Final = "cli_sso_session"
|
||
CLI_SSO_SESSION_TTL_SECONDS: Final = 600
|
||
CLI_SESSION_KEY_PREFIX: Final = "cli-session"
|
||
# Support both CLI_JWT_EXPIRATION_HOURS and LITELLM_CLI_JWT_EXPIRATION_HOURS for backwards compatibility
|
||
CLI_JWT_EXPIRATION_HOURS: Final = int(
|
||
os.getenv("CLI_JWT_EXPIRATION_HOURS") or os.getenv("LITELLM_CLI_JWT_EXPIRATION_HOURS") or 24
|
||
)
|
||
# Comma-separated allowlisted OIDC claim map for CLI SSO polling, e.g.
|
||
# "employment_type->acme_employment_type,org_info.department->department"
|
||
CLI_SSO_CLAIM_MAP: Final = os.getenv("CLI_SSO_CLAIM_MAP") or os.getenv("LITELLM_CLI_SSO_CLAIM_MAP") or ""
|
||
CLI_SSO_CLAIM_MAX_SCALAR_LENGTH: Final = 1024
|
||
|
||
########################### UI SESSION DURATION ###########################
|
||
# Duration for UI login session (username/password, SSO, invitation links). Format: "30s", "30m", "24h", "7d"
|
||
# Does NOT apply to EXPERIMENTAL_UI_LOGIN flow, which intentionally uses a fixed 10-minute expiry for security.
|
||
LITELLM_UI_SESSION_DURATION: Final = os.getenv("LITELLM_UI_SESSION_DURATION", "24h")
|
||
|
||
########################### DB CRON JOB NAMES ###########################
|
||
DB_SPEND_UPDATE_JOB_NAME: Final = "db_spend_update_job"
|
||
DB_DAILY_TAG_SPEND_UPDATE_JOB_NAME: Final = "db_daily_tag_spend_update_job"
|
||
PROMETHEUS_EMIT_BUDGET_METRICS_JOB_NAME: Final = "prometheus_emit_budget_metrics"
|
||
CLOUDZERO_EXPORT_USAGE_DATA_JOB_NAME: Final = "cloudzero_export_usage_data"
|
||
MAVVRIK_FOCUS_EXPORT_JOB_NAME: Final = "mavvrik_focus_export_usage_data"
|
||
CLOUDZERO_MAX_FETCHED_DATA_RECORDS: Final = int(os.getenv("CLOUDZERO_MAX_FETCHED_DATA_RECORDS", 50000))
|
||
SPEND_LOG_CLEANUP_JOB_NAME: Final = "spend_log_cleanup"
|
||
BACKGROUND_HEALTH_CHECK_DB_SAVE_JOB_NAME: Final = "background_health_check_db_save"
|
||
KEY_ROTATION_JOB_NAME: Final = "litellm_key_rotation_job"
|
||
EXPIRED_UI_SESSION_KEY_CLEANUP_JOB_NAME: Final = "litellm_expired_ui_session_key_cleanup_job"
|
||
WEEKLY_SPEND_REPORT_JOB_ID: Final = "weekly_spend_report_job"
|
||
MONTHLY_SPEND_REPORT_JOB_ID: Final = "monthly_spend_report_job"
|
||
USER_SPEND_ALERTS_JOB_ID: Final = "user_spend_alerts_job"
|
||
PROMETHEUS_FALLBACK_STATS_JOB_ID: Final = "prometheus_fallback_stats_job"
|
||
SLACK_DAILY_REPORT_LOCK_ID: Final = "slack_daily_report"
|
||
SLACK_MODEL_DEPRECATION_LOCK_ID: Final = "slack_model_deprecation_warning"
|
||
SPEND_LOG_RUN_LOOPS: Final = int(os.getenv("SPEND_LOG_RUN_LOOPS", 500))
|
||
SPEND_LOG_CLEANUP_BATCH_SIZE: Final = int(os.getenv("SPEND_LOG_CLEANUP_BATCH_SIZE", 1000))
|
||
SPEND_LOG_CLEANUP_MAX_CONSECUTIVE_BATCH_FAILURES = int(os.getenv("SPEND_LOG_CLEANUP_MAX_CONSECUTIVE_BATCH_FAILURES", 3))
|
||
SPEND_LOG_CLEANUP_BATCH_FAILURE_BACKOFF_SECONDS: Final = float(
|
||
os.getenv("SPEND_LOG_CLEANUP_BATCH_FAILURE_BACKOFF_SECONDS", 0.5)
|
||
)
|
||
SPEND_LOG_CLEANUP_RUN_BUDGET_SECONDS: Final = float(os.getenv("SPEND_LOG_CLEANUP_RUN_BUDGET_SECONDS", "300"))
|
||
SPEND_LOG_CLEANUP_BATCH_TIMEOUT_SECONDS: Final = float(os.getenv("SPEND_LOG_CLEANUP_BATCH_TIMEOUT_SECONDS", "30"))
|
||
SPEND_LOG_CLEANUP_REMAINING_COUNT_CAP: Final = int(os.getenv("SPEND_LOG_CLEANUP_REMAINING_COUNT_CAP", "100000"))
|
||
SCHEDULED_JOB_SHUTDOWN_FINISH_TIMEOUT_SECONDS: Final = float(
|
||
os.getenv("SCHEDULED_JOB_SHUTDOWN_FINISH_TIMEOUT_SECONDS", "5")
|
||
)
|
||
SCHEDULED_JOB_SHUTDOWN_CANCEL_TIMEOUT_SECONDS: Final = float(
|
||
os.getenv("SCHEDULED_JOB_SHUTDOWN_CANCEL_TIMEOUT_SECONDS", "5")
|
||
)
|
||
TOOL_SPEND_TOP_TOOLS: Final = 100
|
||
SPEND_LOG_PARTITION_INTERVAL: Final = os.getenv("SPEND_LOG_PARTITION_INTERVAL", "day")
|
||
SPEND_LOG_PARTITION_PRECREATE_AHEAD: Final = int(os.getenv("SPEND_LOG_PARTITION_PRECREATE_AHEAD", 7))
|
||
SPEND_LOG_WRITE_BATCH_MAX_BYTES: Final = max(1, int(os.getenv("SPEND_LOG_WRITE_BATCH_MAX_BYTES", 2_000_000)))
|
||
SPEND_LOG_WRITE_BATCH_MAX_ROWS: Final = max(1, int(os.getenv("SPEND_LOG_WRITE_BATCH_MAX_ROWS", "100")))
|
||
SPEND_LOG_QUEUE_SIZE_THRESHOLD: Final = int(os.getenv("SPEND_LOG_QUEUE_SIZE_THRESHOLD", 100))
|
||
SPEND_LOG_QUEUE_MAX_BYTES: Final = max(1, int(os.getenv("SPEND_LOG_QUEUE_MAX_BYTES", "64000000")))
|
||
SPEND_LOG_QUEUE_POLL_INTERVAL: Final = float(os.getenv("SPEND_LOG_QUEUE_POLL_INTERVAL", 2.0))
|
||
RESPONSES_SESSION_LOOKUP_MAX_ATTEMPTS: Final = max(1, int(os.getenv("RESPONSES_SESSION_LOOKUP_MAX_ATTEMPTS", "3")))
|
||
RESPONSES_SESSION_LOOKUP_RETRY_INTERVAL: Final = float(os.getenv("RESPONSES_SESSION_LOOKUP_RETRY_INTERVAL", "0.2"))
|
||
SPEND_COUNTER_RESEED_LOCKS_MAX_SIZE: Final = int(os.getenv("SPEND_COUNTER_RESEED_LOCKS_MAX_SIZE", 10000))
|
||
PROXY_DB_LOOKUP_MAX_CONCURRENCY: Final = max(1, int(os.getenv("PROXY_DB_LOOKUP_MAX_CONCURRENCY", "25")))
|
||
PROXY_DB_LOOKUP_DEADLINE_SECONDS: Final = max(0.1, float(os.getenv("PROXY_DB_LOOKUP_DEADLINE_SECONDS", "10")))
|
||
PROXY_DB_LOOKUP_STALL_WINDOW_SECONDS: Final = max(0.0, float(os.getenv("PROXY_DB_LOOKUP_STALL_WINDOW_SECONDS", "30")))
|
||
DEFAULT_CRON_JOB_LOCK_TTL_SECONDS: Final = int(os.getenv("DEFAULT_CRON_JOB_LOCK_TTL_SECONDS", 60)) # 1 minute
|
||
PROXY_BUDGET_RESCHEDULER_MIN_TIME: Final = int(os.getenv("PROXY_BUDGET_RESCHEDULER_MIN_TIME", 597))
|
||
RESET_BUDGET_JOB_BATCH_SIZE: Final = max(1, int(os.getenv("RESET_BUDGET_JOB_BATCH_SIZE", "500")))
|
||
RESET_BUDGET_JOB_MAX_CHUNKS_PER_RUN: Final = max(1, int(os.getenv("RESET_BUDGET_JOB_MAX_CHUNKS_PER_RUN", "100")))
|
||
RESET_BUDGET_JOB_NAME: Final = "reset_budget_job"
|
||
# Comfortably longer than one PROXY_BUDGET_RESCHEDULER_MIN_TIME tick, so a healthy
|
||
# leader keeps the lease across its own run, and a crashed one strands the sweep for
|
||
# at most a single tick.
|
||
RESET_BUDGET_JOB_LOCK_TTL_SECONDS: Final[int] = 900
|
||
PROXY_BATCH_POLLING_INTERVAL: Final = int(os.getenv("PROXY_BATCH_POLLING_INTERVAL", 3600))
|
||
MAX_OBJECTS_PER_POLL_CYCLE: Final = max(1, int(os.getenv("MAX_OBJECTS_PER_POLL_CYCLE", 50)))
|
||
MANAGED_OBJECT_STALENESS_CUTOFF_DAYS: Final = max(1, int(os.getenv("MANAGED_OBJECT_STALENESS_CUTOFF_DAYS", 7)))
|
||
STALE_OBJECT_CLEANUP_BATCH_SIZE: Final = max(1, int(os.getenv("STALE_OBJECT_CLEANUP_BATCH_SIZE", 1000)))
|
||
# Set PROXY_BATCH_POLLING_ENABLED=false to disable the CheckBatchCost and
|
||
# CheckResponsesCost background polling jobs entirely (e.g. to avoid DB load on
|
||
# installations with large numbers of stale managed objects).
|
||
_batch_polling_env: Final = os.getenv("PROXY_BATCH_POLLING_ENABLED", "true").lower()
|
||
PROXY_BATCH_POLLING_ENABLED: Final = _batch_polling_env == "true"
|
||
BACKGROUND_INTERACTION_COST_POLL_INITIAL_INTERVAL_SECONDS: Final = float(
|
||
os.getenv("BACKGROUND_INTERACTION_COST_POLL_INITIAL_INTERVAL_SECONDS", "5")
|
||
)
|
||
BACKGROUND_INTERACTION_COST_POLL_MAX_INTERVAL_SECONDS: Final = float(
|
||
os.getenv("BACKGROUND_INTERACTION_COST_POLL_MAX_INTERVAL_SECONDS", "60")
|
||
)
|
||
BACKGROUND_INTERACTION_COST_POLL_TIMEOUT_SECONDS: Final = float(
|
||
os.getenv("BACKGROUND_INTERACTION_COST_POLL_TIMEOUT_SECONDS", "3600")
|
||
)
|
||
_background_interaction_cost_polling_env: Final = os.getenv(
|
||
"BACKGROUND_INTERACTION_COST_POLLING_ENABLED", "true"
|
||
).lower()
|
||
BACKGROUND_INTERACTION_COST_POLLING_ENABLED: Final = _background_interaction_cost_polling_env == "true"
|
||
PROXY_BUDGET_RESCHEDULER_MAX_TIME: Final = int(os.getenv("PROXY_BUDGET_RESCHEDULER_MAX_TIME", 605))
|
||
PROXY_BATCH_WRITE_AT: Final = int(os.getenv("PROXY_BATCH_WRITE_AT", 10)) # in seconds, increased from 10
|
||
PROXY_CONFIG_RELOAD_INTERVAL_SECONDS: Final = get_env_int("PROXY_CONFIG_RELOAD_INTERVAL_SECONDS", 30)
|
||
|
||
# APScheduler Configuration - MEMORY LEAK FIX
|
||
# These settings prevent memory leaks in APScheduler's normalize() and _apply_jitter() functions
|
||
APSCHEDULER_COALESCE: Final = os.getenv("APSCHEDULER_COALESCE", "True").lower() in [
|
||
"true",
|
||
"1",
|
||
] # collapse many missed runs into one
|
||
APSCHEDULER_MISFIRE_GRACE_TIME: Final = int(
|
||
os.getenv("APSCHEDULER_MISFIRE_GRACE_TIME", 3600)
|
||
) # ignore runs older than 1 hour (was 120)
|
||
APSCHEDULER_MAX_INSTANCES: Final = int(os.getenv("APSCHEDULER_MAX_INSTANCES", 1)) # prevent concurrent job instances
|
||
APSCHEDULER_REPLACE_EXISTING: Final = os.getenv("APSCHEDULER_REPLACE_EXISTING", "True").lower() in [
|
||
"true",
|
||
"1",
|
||
] # always replace existing jobs
|
||
|
||
# Width of the window scheduled background jobs are spread across, so they do not all fire
|
||
# on one instant on every replica. Tunable per deployment via general_settings.
|
||
DEFAULT_STAGGER_WINDOW_SECONDS: Final = 300
|
||
|
||
# The number of tag entries are higher than number of user, team entries. This leads to a higher QPS.
|
||
# This will run tag spcific tasks at a later time to smooth QPS
|
||
DAILY_TAG_SPEND_BATCH_MULTIPLIER: Final = 2.3
|
||
|
||
DEFAULT_HEALTH_CHECK_INTERVAL: Final = int(os.getenv("DEFAULT_HEALTH_CHECK_INTERVAL", 300)) # 5 minutes
|
||
DEFAULT_SHARED_HEALTH_CHECK_TTL: Final = int(
|
||
os.getenv("DEFAULT_SHARED_HEALTH_CHECK_TTL", 300)
|
||
) # 5 minutes - TTL for cached health check results
|
||
DEFAULT_SHARED_HEALTH_CHECK_LOCK_TTL: Final = int(
|
||
os.getenv("DEFAULT_SHARED_HEALTH_CHECK_LOCK_TTL", 60)
|
||
) # 1 minute - TTL for health check lock
|
||
DEFAULT_HEALTH_CHECK_STALENESS_MULTIPLIER: Final = 2 # health state is stale after interval * this
|
||
PROMETHEUS_FALLBACK_STATS_SEND_TIME_HOURS: Final = int(os.getenv("PROMETHEUS_FALLBACK_STATS_SEND_TIME_HOURS", 9))
|
||
DEFAULT_MODEL_CREATED_AT_TIME: Final = int(
|
||
os.getenv("DEFAULT_MODEL_CREATED_AT_TIME", 1677610602)
|
||
) # returns on `/models` endpoint
|
||
DEFAULT_SLACK_ALERTING_THRESHOLD: Final = int(os.getenv("DEFAULT_SLACK_ALERTING_THRESHOLD", 300))
|
||
MAX_TEAM_LIST_LIMIT: Final = int(os.getenv("MAX_TEAM_LIST_LIMIT", 20))
|
||
MAX_POLICY_ESTIMATE_IMPACT_ROWS: Final = int(os.getenv("MAX_POLICY_ESTIMATE_IMPACT_ROWS", 1000))
|
||
DEFAULT_PROMPT_INJECTION_SIMILARITY_THRESHOLD = float(os.getenv("DEFAULT_PROMPT_INJECTION_SIMILARITY_THRESHOLD", 0.7))
|
||
LENGTH_OF_LITELLM_GENERATED_KEY: Final = int(os.getenv("LENGTH_OF_LITELLM_GENERATED_KEY", 16))
|
||
MINIMUM_CUSTOM_KEY_LENGTH: Final = int(os.getenv("MINIMUM_CUSTOM_KEY_LENGTH", 16))
|
||
SECRET_MANAGER_REFRESH_INTERVAL: Final = int(os.getenv("SECRET_MANAGER_REFRESH_INTERVAL", 86400))
|
||
OPENAI_SYSTEM_MESSAGES_FIRST_PROVIDERS: Final = frozenset({"openai", "azure"})
|
||
LITELLM_SETTINGS_SAFE_DB_OVERRIDES: Final = [
|
||
"default_internal_user_params",
|
||
"default_team_params",
|
||
"public_mcp_servers",
|
||
"public_agent_groups",
|
||
"public_model_groups",
|
||
"public_model_groups_links",
|
||
"cost_discount_config",
|
||
"cost_margin_config",
|
||
"block_requests_for_models_without_pricing",
|
||
"budget_exceeded_throttle_percentage",
|
||
# Every field editable from the Admin UI (proxy_server._GENERAL_SETTINGS_UI_LITELLM_FIELDS)
|
||
# must be listed here so a DB write from one worker overrides the live litellm attribute on
|
||
# the others when config reloads; otherwise peer workers stay on their startup value.
|
||
# test_general_settings_ui_fields_are_db_overridable enforces that pairing.
|
||
"enable_anthropic_prompt_caching",
|
||
"anthropic_prompt_caching_ttl",
|
||
"openai_system_messages_first",
|
||
"max_ui_session_budget",
|
||
"budget_rollover",
|
||
"mcp_tool_search",
|
||
"turn_off_message_logging",
|
||
"datadog_params",
|
||
"datadog_llm_observability_params",
|
||
"newrelic_params",
|
||
"pointfive_params",
|
||
"aws_sqs_callback_params",
|
||
]
|
||
SPECIAL_LITELLM_AUTH_TOKEN: Final = ["ui-token"]
|
||
DEFAULT_MANAGEMENT_OBJECT_IN_MEMORY_CACHE_TTL = int(os.getenv("DEFAULT_MANAGEMENT_OBJECT_IN_MEMORY_CACHE_TTL", 60))
|
||
DEFAULT_ACCESS_GROUP_CACHE_TTL: Final = int(os.getenv("DEFAULT_ACCESS_GROUP_CACHE_TTL", 600))
|
||
SPEND_LOG_KEY_METADATA_CACHE_TTL: Final = 600
|
||
SPEND_LOG_KEY_METADATA_MISS_CACHE_TTL: Final = 30
|
||
SPEND_LOG_KEY_METADATA_CACHE_MAX_ITEMS: Final = 10000
|
||
SPEND_LOG_KEY_METADATA_QUERY_TIMEOUT_MS: Final = 5000
|
||
# Short TTL for negative MCP access-group existence lookups. Keeps unauthenticated
|
||
# callers from forcing a DB query per request for unknown names, while bounding
|
||
# staleness so a transient DB error (which surfaces as an empty list) cannot
|
||
# hide a real group for long.
|
||
DEFAULT_MCP_ACCESS_GROUP_NEGATIVE_CACHE_TTL: Final = 10
|
||
# Maximum number of comma-separated MCP server / access-group tokens accepted
|
||
# in a single ``/{name1,name2,...}/mcp`` URL. Bounds the per-request DB / cache
|
||
# fan-out an authenticated caller can trigger by stuffing the path with tokens.
|
||
DEFAULT_MCP_NAMESPACE_CSV_MAX_TOKENS: Final = 16
|
||
# Ceilings on the cached auth registries; larger tables fall back to per-row lookups
|
||
# instead of holding an unbounded id set in every worker.
|
||
TAG_REGISTRY_MAX_SIZE: Final = 5000
|
||
MODEL_ACCESS_GROUP_REGISTRY_MAX_SIZE: Final = 5000
|
||
END_USER_RESTRICTED_REGISTRY_MAX_SIZE: Final = 5000
|
||
# How long a failed registry load is remembered as "unusable", so a degraded Postgres
|
||
# is not re-scanned on every request on top of the per-id lookups it falls back to.
|
||
REGISTRY_ERROR_NEGATIVE_CACHE_TTL: Final = 30
|
||
|
||
# Sentry Scrubbing Configuration
|
||
SENTRY_DENYLIST: Final = [
|
||
# API Keys and Tokens
|
||
"api_key",
|
||
"token",
|
||
"key",
|
||
"secret",
|
||
"password",
|
||
"auth",
|
||
"credential",
|
||
"OPENAI_API_KEY",
|
||
"ANTHROPIC_API_KEY",
|
||
"ANTHROPIC_AUTH_TOKEN",
|
||
"AZURE_API_KEY",
|
||
"COHERE_API_KEY",
|
||
"REPLICATE_API_KEY",
|
||
"HUGGINGFACE_API_KEY",
|
||
"TOGETHERAI_API_KEY",
|
||
"CLOUDFLARE_API_KEY",
|
||
"BASETEN_KEY",
|
||
"OPENROUTER_KEY",
|
||
"COMETAPI_KEY",
|
||
"DATAROBOT_API_TOKEN",
|
||
"FIREWORKS_API_KEY",
|
||
"FIREWORKS_AI_API_KEY",
|
||
"FIREWORKSAI_API_KEY",
|
||
"OVHCLOUD_API_KEY",
|
||
"CLARIFAI_API_KEY",
|
||
# Database and Connection Strings
|
||
"database_url",
|
||
"redis_url",
|
||
"connection_string",
|
||
# Authentication and Security
|
||
"master_key",
|
||
"LITELLM_MASTER_KEY",
|
||
"auth_token",
|
||
"jwt_token",
|
||
"private_key",
|
||
"SLACK_WEBHOOK_URL",
|
||
"ALERTING_WEBHOOK_URL",
|
||
"webhook_url",
|
||
"LANGFUSE_SECRET_KEY",
|
||
# Email Configuration
|
||
"SMTP_PASSWORD",
|
||
"SMTP_USERNAME",
|
||
"email_password",
|
||
# Cloud Provider Credentials
|
||
"aws_access_key",
|
||
"aws_secret_key",
|
||
"gcp_credentials",
|
||
"azure_credentials",
|
||
"HCP_VAULT_TOKEN",
|
||
"CIRCLE_OIDC_TOKEN",
|
||
# Proxy and Environment Settings
|
||
"proxy_url",
|
||
"proxy_key",
|
||
"environment_variables",
|
||
]
|
||
SENTRY_PII_DENYLIST: Final = [
|
||
"user_id",
|
||
"email",
|
||
"phone",
|
||
"address",
|
||
"ip_address",
|
||
"SMTP_SENDER_EMAIL",
|
||
"TEST_EMAIL_ADDRESS",
|
||
]
|
||
|
||
# CoroutineChecker cache configuration
|
||
COROUTINE_CHECKER_MAX_SIZE_IN_MEMORY: Final = int(os.getenv("COROUTINE_CHECKER_MAX_SIZE_IN_MEMORY", 1000))
|
||
|
||
########################### RAG Text Splitter Constants ###########################
|
||
DEFAULT_CHUNK_SIZE: Final = int(os.getenv("DEFAULT_CHUNK_SIZE", 1000))
|
||
DEFAULT_CHUNK_OVERLAP: Final = int(os.getenv("DEFAULT_CHUNK_OVERLAP", 200))
|
||
|
||
########################### S3 Vectors RAG Constants ###########################
|
||
S3_VECTORS_DEFAULT_DIMENSION: Final = int(os.getenv("S3_VECTORS_DEFAULT_DIMENSION", 1024))
|
||
S3_VECTORS_DEFAULT_DISTANCE_METRIC: Final = str(os.getenv("S3_VECTORS_DEFAULT_DISTANCE_METRIC", "cosine"))
|
||
S3_VECTORS_DEFAULT_NON_FILTERABLE_METADATA_KEYS: Final = ["source_text"]
|
||
|
||
########################### Microsoft SSO Constants ###########################
|
||
MICROSOFT_USER_EMAIL_ATTRIBUTE: Final = str(os.getenv("MICROSOFT_USER_EMAIL_ATTRIBUTE", "userPrincipalName"))
|
||
MICROSOFT_USER_DISPLAY_NAME_ATTRIBUTE: Final = str(os.getenv("MICROSOFT_USER_DISPLAY_NAME_ATTRIBUTE", "displayName"))
|
||
MICROSOFT_USER_ID_ATTRIBUTE: Final = str(os.getenv("MICROSOFT_USER_ID_ATTRIBUTE", "id"))
|
||
MICROSOFT_USER_FIRST_NAME_ATTRIBUTE: Final = str(os.getenv("MICROSOFT_USER_FIRST_NAME_ATTRIBUTE", "givenName"))
|
||
MICROSOFT_USER_LAST_NAME_ATTRIBUTE: Final = str(os.getenv("MICROSOFT_USER_LAST_NAME_ATTRIBUTE", "surname"))
|
||
|
||
# Maximum payload size (in bytes) to fully serialize for DEBUG logging.
|
||
# Payloads larger than this are truncated to avoid multi-second json.dumps blocking the response.
|
||
MAX_PAYLOAD_SIZE_FOR_DEBUG_LOG: Final = int(os.getenv("MAX_PAYLOAD_SIZE_FOR_DEBUG_LOG", 102400)) # 100 KB
|
||
|
||
# Policy template enrichment
|
||
MAX_COMPETITOR_NAMES: Final = int(os.getenv("MAX_COMPETITOR_NAMES", 100))
|
||
COMPETITOR_LLM_TEMPERATURE: Final = float(os.getenv("COMPETITOR_LLM_TEMPERATURE", 0.3))
|
||
DEFAULT_COMPETITOR_DISCOVERY_MODEL: Final = "gpt-4o-mini"
|
||
|
||
# Advisor tool orchestration
|
||
# Providers that support advisor_20260301 natively (no LiteLLM orchestration needed).
|
||
# Add vertex_ai here once verified.
|
||
ADVISOR_NATIVE_PROVIDERS: Final[frozenset] = frozenset({"anthropic"})
|
||
# Hard cap on advisor iterations per request to prevent runaway loops.
|
||
ADVISOR_MAX_USES: Final[int] = 5
|
||
# Description injected into the synthetic advisor tool definition sent to non-native providers.
|
||
ADVISOR_TOOL_DESCRIPTION: Final[str] = (
|
||
"Consult a highly intelligent advisor model when you need expert guidance, "
|
||
"want to verify your reasoning, or face a complex decision. "
|
||
"Describe your question or challenge clearly in the 'question' field."
|
||
)
|
||
|
||
# Headers that must be stripped from a provider exception before it's forwarded as
|
||
# the proxy's own HTTP response, or they conflict with the framing the proxy sets.
|
||
HTTP_FRAMING_HEADERS: Final[frozenset[str]] = frozenset(
|
||
{
|
||
"content-length",
|
||
"transfer-encoding",
|
||
"content-encoding",
|
||
"content-type",
|
||
"set-cookie",
|
||
"cookie",
|
||
"proxy-authenticate",
|
||
"proxy-authorization",
|
||
}
|
||
)
|
||
|
||
PROVIDER_REQUEST_ID_HEADERS: Final[tuple[str, ...]] = (
|
||
"x-amzn-requestid",
|
||
"x-request-id",
|
||
"request-id",
|
||
"x-ms-request-id",
|
||
"apim-request-id",
|
||
"x-goog-request-id",
|
||
"cf-ray",
|
||
)
|
||
|
||
# Browser-facing security headers that a malicious or misconfigured upstream
|
||
# provider must not be able to set on the proxy's own response.
|
||
BROWSER_SECURITY_HEADERS: Final[frozenset[str]] = frozenset(
|
||
{
|
||
"access-control-allow-origin",
|
||
"access-control-allow-credentials",
|
||
"access-control-allow-methods",
|
||
"access-control-allow-headers",
|
||
"access-control-expose-headers",
|
||
"content-security-policy",
|
||
"content-security-policy-report-only",
|
||
"clear-site-data",
|
||
"strict-transport-security",
|
||
"x-frame-options",
|
||
"cross-origin-opener-policy",
|
||
"cross-origin-embedder-policy",
|
||
"cross-origin-resource-policy",
|
||
}
|
||
)
|
||
|
||
UNSAFE_PROXY_RESPONSE_HEADERS: Final[frozenset[str]] = HTTP_FRAMING_HEADERS | BROWSER_SECURITY_HEADERS
|
||
|
||
STRINGIFIED_NONE: Final[str] = "None"
|
||
|
||
# A retrieved response replays the usage of the call that created it, so pricing these
|
||
# read/management routes like inference bills the same tokens twice.
|
||
NON_INFERENCE_CALL_TYPES: Final[frozenset[str]] = frozenset(
|
||
{
|
||
"get_responses",
|
||
"aget_responses",
|
||
"delete_responses",
|
||
"adelete_responses",
|
||
"cancel_responses",
|
||
"acancel_responses",
|
||
"list_input_items",
|
||
"alist_input_items",
|
||
"vector_store_create",
|
||
"avector_store_create",
|
||
"vector_store_retrieve",
|
||
"avector_store_retrieve",
|
||
"vector_store_list",
|
||
"avector_store_list",
|
||
"vector_store_update",
|
||
"avector_store_update",
|
||
"vector_store_delete",
|
||
"avector_store_delete",
|
||
"vector_store_file_create",
|
||
"avector_store_file_create",
|
||
"vector_store_file_list",
|
||
"avector_store_file_list",
|
||
"vector_store_file_retrieve",
|
||
"avector_store_file_retrieve",
|
||
"vector_store_file_content",
|
||
"avector_store_file_content",
|
||
"vector_store_file_update",
|
||
"avector_store_file_update",
|
||
"vector_store_file_delete",
|
||
"avector_store_file_delete",
|
||
}
|
||
)
|
||
|
||
UNKNOWN_MODEL_SPEND_LOG_MODEL: Final[str] = "unknown-model"
|
||
MAX_SPEND_LOG_MODEL_NAME_LENGTH: Final[int] = 256
|
||
MCP_SPEND_LOG_MODEL_PREFIX: Final[str] = "MCP: "
|
||
|
||
# PTU reservation rollup writes rows to LiteLLM_DailyTeamSpend with this
|
||
# sentinel api_key so PTU flat cost stays distinguishable from real per-request
|
||
# spend under the table's composite unique constraint.
|
||
PTU_SENTINEL_API_KEY: Final[str] = "__ptu_flat_cost__"
|
||
PTU_ROLLUP_JOB_ID: Final[str] = "ptu_flat_cost_rollup_job"
|
||
PTU_ROLLUP_LOCK_TTL_SECONDS: Final[int] = 900
|
||
USAGE_TOP_API_KEYS_LIMIT: Final[int] = int(os.getenv("USAGE_TOP_API_KEYS_LIMIT", "100"))
|
||
# Furthest back the catch-up pass looks for unpriced PTU days when a deployment
|
||
# declares no ptu_effective_from, bounding the scan for an open-ended window.
|
||
PTU_ROLLUP_MAX_BACKFILL_DAYS: Final[int] = 90
|
||
# Deployments named in the lapsed-window alert before it is truncated, so a fleet-wide
|
||
# expiry cannot produce an alert too large for the channel delivering it.
|
||
PTU_LAPSED_ALERT_LIMIT: Final[int] = 10
|
||
DAILY_GLOBAL_SPEND_RECONCILE_JOB_ID: Final[str] = "daily_global_spend_reconcile_job"
|
||
DAILY_GLOBAL_SPEND_RECONCILE_LOCK_TTL_SECONDS: Final[int] = 3600
|
||
DAILY_GLOBAL_SPEND_RECONCILED_THROUGH_PARAM: Final[str] = "daily_global_spend_reconciled_through"
|
||
SPEND_CAPTURE_RATE_CHECK_JOB_ID: Final[str] = "spend_capture_rate_check_job"
|
||
SPEND_CAPTURE_RATE_CHECK_LOCK_TTL_SECONDS: Final[int] = 900
|
||
SPEND_CAPTURE_RATE_MAX_RANGE_DAYS: Final[int] = 180
|
||
SPEND_CAPTURE_RATE_DOCS_URL: Final[str] = "https://docs.litellm.ai/docs/proxy/spend_capture_rate"
|
||
OPENAI_ORGANIZATION_COSTS_URL: Final[str] = "https://api.openai.com/v1/organization/costs"
|
||
# Buckets per page the OpenAI costs endpoint allows (1 to 180, default 7), 2026-09-24
|
||
OPENAI_ORGANIZATION_COSTS_PAGE_LIMIT: Final[int] = 180
|
||
PROVIDER_BILLING_TIMEOUT_SECONDS: Final[float] = 30.0
|
||
# Slack allowed when deciding a sentinel row is stale. The row's updated_at and the
|
||
# run's cutoff are stamped by different hosts, so clock skew between them must not let
|
||
# one run delete a charge another just wrote. A stale row is hours old and a concurrent
|
||
# one is seconds old, so a few minutes separates them.
|
||
PTU_PRUNE_SKEW_GRACE_SECONDS: Final[int] = 300
|
||
|
||
# How long enqueued-token reservations for batches live without a refund. Providers
|
||
# complete or expire batches within their completion window (24h for OpenAI), so a
|
||
# reservation still unrefunded after 8 days belongs to a batch whose terminal state
|
||
# was never observed (e.g. proxy restart); expiry returns the tokens to the caller.
|
||
BATCH_ENQUEUED_TOKEN_TTL_SECONDS: Final[int] = 8 * 24 * 60 * 60
|
||
|
||
# Key/team metadata field that opts batches into enqueued-token limiting. Only proxy
|
||
# admins may write it: when present it replaces the standard RPM/TPM checks for
|
||
# batch submissions.
|
||
BATCH_ENQUEUED_TOKEN_LIMIT_METADATA_KEY: Final = "batch_enqueued_token_limit"
|
||
|
||
# Shared read-only empty mapping, for defaulting optional Mapping parameters without
|
||
# constructing a fresh mutable dict at each call site.
|
||
EMPTY_MAPPING: Final = MappingProxyType({})
|
||
|
||
# API endpoint for breached password k-anonymity search
|
||
HIBP_RANGE_API_BASE: Final = "https://api.pwnedpasswords.com/range"
|