mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
14 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
75b45e39c9
|
fix(proxy): enforce internal-user model creation prohibition (#44438)
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com> |
||
|
|
3286782dea
|
fix(caching): key response cache by router model group in litellm_metadata (#44542)
* test: align integration fixtures with current behavior Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): wait for the spend flush before asserting its trace placement Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(caching): key response-cache entries by model group from litellm_metadata /v1/responses routes through the router with model_group in litellm_metadata, which the cache key ignored, so identical requests to different model groups sharing one underlying model hit each other's cached responses Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0ed1c08f02
|
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one commit on top of litellm_internal_staging without the dashboard changes. Deployments on anthropic/ without a static api_key can exchange an OIDC workload assertion for a short-lived sk-ant-oat01 token through a shared RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file, an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment, per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation fields are server-owned: refused inline in request bodies and on POST /model/new, proxy-admin only on credentials, and the token exchange is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS adds a host. GET /credentials/{name}/jwks exports the public key set of a LiteLLM-signed credential for the Claude Console. The OpenAI federation trio from #39613 rides along on the backend side with the same server-owned handling. Fixes #28607 Resolves LIT-6107 Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com> * fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries The files handler enabled workload identity on batch-result downloads but never received the deployment's litellm_params, so a deployment authenticating through a named credential could only mint from process-wide env vars. It now threads litellm_params through to the auth header the way the batch retrieve path already does. LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme, so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the gateway. Entries are now parsed as network locations whether or not they carry a scheme. * fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle * test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries * fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly * fix(proxy): decrypt stored litellm_params before the WIF write gate * fix(proxy): hide WIF secret references from /health output * fix(proxy): keep the proxy error shape on credential endpoint refusals * fix(proxy): hide identity token file paths from /health output * fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working The Bedrock Claude Platform route already reads anthropic_workspace_id from optional_params, so banning that spelling as a server-owned federation parameter broke a pre-existing client capability. The federation field is now anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID), which restores the base branch's behavior for Bedrock callers, drops the Bedrock-specific hint from the refusal message, and deletes the unconditional ban constant that no longer had a reader * fix(auth): share one exchanged token across workers reading the same assertion Anthropic accepts each identity assertion exactly once, so two uvicorn workers reading the same token file both minting from it means the second exchange is denied with jti_reused. Minted tokens now land in a per-user 0700 cache directory guarded by a file lock, so workers on the same host reuse one exchange until the token expires or the assertion rotates. A 401 is only retried when the re-read assertion actually differs, and the denial hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache and an empty value disables it * fix: keep anthropic federation from being shadowed or leaked An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated deployment sent an empty x-api-key on every call instead of minting a token. Blank values now read as unset, and a real static key on a federated deployment logs once that it outranks federation and nothing is being federated. The exchange-host allowlist matched hostnames only, so a second process on another port of an allowed host was trusted with the workload's identity token. An entry that names a port now trusts that port alone, while a bare host still trusts every port. The shared token store exists so the workers reading one projected token file do not each spend its single-use jti. A source that mints its own assertion per exchange shares nothing with another worker, so it no longer writes a live token to disk for a lookup that can never hit. * fix: unlink a staged token file a failed write leaves behind The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads. * refactor: move anthropic jwks derivation behind a provider-owned tagged union * fix: unlink the staged token file when its write fails at close A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token. * fix(anthropic): close the staging descriptor before writing the shared token file * fix(wif): judge federation writes by what they set, not what is stored The admin gate read the stored deployment, so a team admin lost edit, delete and Test Connection on any deployment carrying federation params. It now returns early unless the submitted fields touch the federation surface, and a Test Connection probe that points the deployment at its own api_base is still refused, with the 403 no longer wrapped into a 500 The rest of the same review pass: POST /model/new refuses only a blocking value of `blocked`, so a client that always sends `blocked: false` is not turned away; a request body can no longer pick which federated identity to mint as by naming a stored credential; an advisory refresh the executor refuses disarms the entry instead of wedging the identity until the follower timeout; the static-key shadow warning resolves its env fallback inside the cache instead of once per request; credential writes drop nulls before storing them; the token exchange validates the endpoint URL before reading an assertion and keeps refusing redirects across a client heal; /health hides every server-owned federation field from non-admins; and the async create_file and create_batch paths say which setting is missing when the provider resolves no URL * fix(proxy): let a deployment write name a federated credential reject_federated_credential_reference runs from is_request_body_safe, which pre_db_read_auth_checks calls on every route, so it also fired on POST /model/new, /model/update, /model/{id}/update and /health/test_connection. A proxy admin could no longer attach a federated credential to a deployment over the API or the Admin UI, leaving a static config.yaml entry as the only way to configure the feature the rejection told the caller to go configure, and _reject_non_admin_wif_write never got to make the call it exists to make. is_request_body_safe now takes the route and skips only the credential-reference check on the routes that reach can_user_make_model_call. Federation fields typed inline into a body stay refused everywhere, and a call naming a federated credential still cannot pick the identity it mints as. * refactor(proxy): derive health display policy from the federation key sets The health check module hand-copied the five workload identity fields whose value is a credential, so a shared proxy surface named provider-specific parameters and a newly added secret-bearing field would have gone on being displayed until someone remembered both places WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of, types/utils derives secret_bearing_wif_litellm_params from it, and the health layer splats that tuple the same way it already splats the admin-only one * fix(anthropic_wif): treat blank identity-source fields as unset * test(proxy): classify the federation params in the credential slot registry main's registry test (#43298) now fails the build for any credential-named deployment param without a classification. The five federation fields that carry a token, a token file path, or a signing or client secret reference are Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak settings name a URL, a client id, an auth method, or a scope and are NotSecret * fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics Register the 18 Anthropic and 3 OpenAI federation params as frozen ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel and the request-body ban list read one declaration. Pass the deployment api_base through to the count-tokens handler instead of a pre-suffixed URL, which doubled the /count_tokens path on main's prompt-cache predictor. Add the five litellm_anthropic_wif_* families to the all-metrics Grafana dashboard. * fix(credentials): gate PATCH on WIF fields resolved from model_id The credential PATCH handler checked server-owned workload identity federation fields only on the values the caller sent, while a body that named a deployment through model_id had its credential values resolved after that check. A non-admin could therefore copy a federated deployment's WIF fields onto an ordinary credential. Resolve the incoming values first and run the non-admin gate on them, matching the POST path * fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header Count-tokens walked its own credential ladder: a static key, else skip minting when ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the same deployment authenticated with that token. The handler now takes the auth header that AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills use, and merges the oauth beta a minted or consumer token carries with the token-counting beta --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com> Co-authored-by: mateo-berri <happymvw@gmail.com> |
||
|
|
fe683ea139
|
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models
GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.
* fix(proxy): offer a Codex service tier only when every deployment of the model lists it
* fix(codex-catalog): an invalid service_tiers value offers no tier for the model
* fix(codex-catalog): read service tiers off the deployments the key's team can route to
A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them
The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default
* test(codex-catalog): drop the redundant module docstring and sort the imports
* test(integration): add the Codex catalog audit cells and the multi-worker convergence note
* test(integration): clean up every catalog test model and answer the refresh GET
* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers
Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.
* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut
The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns
The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
|
||
|
|
e340e546e2
|
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures * chore(dev): use OpenAI model in tracing config * chore(dev): align tracing credentials with UI E2E * fix(dev): update fixture seeder query scope * feat(dev): seed linked tracing and spend fixtures * chore(dev): use OpenAI model in tracing config * chore(dev): align tracing credentials with UI E2E * fix(dev): update fixture seeder query scope * wip * wip * wip * chore(trace): checkpoint ongoing Rust migration * refactor(trace): group Python bridge under trace package * refactor(traces): read span conventions through a Convention trait Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under normalize/convention/ as a unit struct implementing Convention, owning both its detection and its extraction. Precedence is one ordered registry instead of an if-chain in mod.rs that reached into each module differently. The modules now share one way to read attributes: present() for the first non-empty key and Payload for a text that also reports the key it consumed, replacing three different idioms and the &mut Vec threaded through payload readers. Instrumentation::adjust returns a new Extraction instead of mutating one, with each SDK rule as its own function, and the LangChain middleware suffix list exists once. * feat(trace): export Rust-owned wire schemas and enforce contract bounds * fix(trace): bound quoted counts in ClickHouse wire schemas * feat(trace): generate Python wire contracts with datamodel-code-generator * test(trace): validate migrated callers and generated contracts at the native boundary * refactor(traces): rename normalization convention to format * fix(traces): reconcile spend evidence and preserve unknown costs * feat(traces): normalize additional telemetry formats * test(traces): cover captured normalization fixtures * refactor(traces): isolate SDK normalization rules * feat(tracing): seed all trace exports for local dashboard * fix(clickhouse): preserve custom LiteLLM request metadata * docs(traces): define normalization module boundaries * docs(traces): define resolution and OTLP boundaries * fix(ui): normalize nullable trace message names * refactor(traces): split resolver modules and cover resolution behavior * test(traces): replace normalization snapshots with behavior assertions * fix(ui): align dashboard API contracts with generated types * refactor(traces): type normalization and storage boundaries * fix(traces): seed captured SDK spend and preserve provider identities * wip * test(traces): verify guide discovery and content ordering --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
7b432d78d2
|
fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model (#44341)
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): streamed alias matching a capability rule bills the deployment price Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): assert every streamed chunk carries the client alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(logging): log the client alias on the priced streamed response, the same as non-streamed Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
12d4b75b7a
|
fix(search_tools): encrypt search tool litellm_params at rest (#43631)
* fix(search_tools): encrypt search tool litellm_params at rest Encrypt every string value of a search tool's litellm_params on create and update and decrypt on every DB read, so legacy plaintext rows load unchanged. Include the table in master key rotation, LITELLM_MIGRATE_FROM_MASTER_KEY and the migrate-encryption scan. * fix(search_tools): keep edits made while the master key rotates Write each rotated search tool row only if it still holds the litellm_params that were read, and re-read and rotate it again if it was edited in between, so a PUT that lands during /key/regenerate is not overwritten. * fix(search_tools): retry rotation writes until the row stops changing Rotate a search tool row again for as long as it keeps being edited instead of giving up after five attempts, and stop with a warning only when the conditional write fails on an unchanged row. Build the decrypted read result without mutating it in place. * refactor(search_tools): rotate edited rows in a loop, drop the step comment Retry the conditional rotation write in a loop instead of recursion so sustained edits cannot deepen the call stack, drop the step comment on the rotation call, and stop mutating local state in the rotation tests. * test(search_tools): drop the rotation test docstring * Store search tool params as written when no encryption key is configured * Rotate search tools under the salt key, keep non-ciphertext values and loaded tools that do not decrypt * Treat a search tool as undecryptable only when its provider is ciphertext-length * Drop suppressions the type discipline gate on main now reports as unused * Show the loaded search tool in the admin list and info views when its DB params do not decrypt * Keep the DB row's other fields when the admin views substitute loaded params |
||
|
|
9b5562f89b
|
test: repair stale and polluting tests red on scheduled main CI (#44229)
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context * test(secret-detection): give the hand-built redaction request an ASGI path Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path * test(integration): isolate litellm callback lists per sdk test usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards * test(integration): keep the owner-lookup fault proxy off the shared read replica The owned proxy points DATABASE_URL at a scratch database but inherited DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared database and rejected the freshly created key with token_not_found_in_db. Drop the replica variable like the other scratch-database owned proxies * test(integration): request every seeded key in the team owner breakdown The aggregated team activity endpoint now caps breakdown.api_keys at the top 100 keys by default (#43398), so the 300 seeded keys came back as 100 rows. The test guarantees each key is reported with its own owner, so ask for an api_key_limit that covers all seeded keys * test(integration): give every owned Redis its own port in the redis-cache container On CircleCI every owned Redis ran on the fixed port 16379 inside the shared redis-cache container. When an earlier server still held that port, the new one failed to bind, readiness pinged the old server, the pidfile read failed and cleanup then reported "Owned Redis still serves after shutdown" Reserve an ephemeral port for the docker-exec path the same way the local binary path already does, and refuse to start when something already serves the chosen port so the failure names the real cause * test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail * test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to * test(e2e): check only stored message content for a leaked card number The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there * test(e2e): assert the proxy decodes token-array embeddings for titan The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer * test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5 * test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed * test(e2e): cite the tokenizer and date behind the titan token-array fixture * test(e2e): let migration seed replicas finish their request-log indexes before cloning Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema * test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope #43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake * test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly |
||
|
|
2b85808011
|
feat(mcp)!: disable stdio MCP servers by default (#44066)
* feat(mcp)!: disable stdio MCP servers by default stdio MCP servers now only run when the proxy is started with LITELLM_ENABLE_MCP_STDIO=true. While it is off, existing stdio servers stay registered but never start: tool listings skip them quietly, direct tool calls and health checks return a 403 naming the env var, and creating or updating a stdio server is rejected. The flag is read from the process environment only, so DB-stored environment_variables cannot turn it on. The UI reads mcp_stdio_enabled from /.well-known/litellm-ui-config to grey out the stdio transport, show a banner on stdio forms, and badge stdio server cards. BREAKING CHANGE: stdio MCP servers are off by default. Set LITELLM_ENABLE_MCP_STDIO=true in the proxy environment and restart to keep using them. * fix(mcp): ignore stdio flag from config file and read UI flag from the selected worker LITELLM_ENABLE_MCP_STDIO set under environment_variables in config.yaml is now skipped like the DB-stored value, so only the process environment can enable stdio. The dashboard reads mcp_stdio_enabled from the proxy it is managing, so a control plane shows each worker's own setting. * test(mcp): cover non-mapping payloads in the shared transport validator * fix(mcp): skip blocked stdio servers quietly in every listing and keep the UI unchanged until the flag loads Prompt, resource and resource-template listings now skip a blocked stdio server at debug level like tool listing does, instead of logging a warning per server on every call. The dashboard only treats stdio as disabled once the proxy explicitly reports mcp_stdio_enabled false, so a proxy with the flag on, or an older one without the field, renders exactly as before with no flicker while loading. * fix(mcp): route blocked stdio tool calls to the flag error and warn once per server A gateway tools/call naming a blocked stdio server's tool now returns the LITELLM_ENABLE_MCP_STDIO message instead of "Tool not found". The "will not start" warning moves out of build_mcp_server_from_table, which DB reload re-runs on every cycle for rows with a NULL updated_at and which drafts and test-connection also call. It now fires when a row first enters the registry or changes transport. * fix(ui): explain on the server detail page why a stdio server is inert The Overview and MCP Tools tabs showed "No tools available" with no reason while stdio is disabled. The detail page now shows the same warning banner as the edit form, and hands off to the form's banner once editing starts. * refactor(ui): name the stdio banner conditions on the server detail page Keeps local/no-long-condition-chain within its budget * fix(proxy): log the ignored DB-stored LITELLM_ENABLE_MCP_STDIO warning once The DB config sync re-reads environment_variables on every cycle, so a stored flag logged the warning on each sync per worker |
||
|
|
276fc9c63a
|
fix(tracing): unify ClickHouse storage configuration (#43941)
* fix(tracing): use ClickHouse URL for reads by default * fix(tracing): unify ClickHouse storage configuration * fix(tracing): update dashboard setup copy for one URL * test(tracing): make tests/unit/tracing a package Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(config): drop legacy string tracing store variant Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(tracing): own ClickHouse defaults in constants and reject unset env references Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(tracing): use raw regex patterns in config tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(tracing): read ClickHouse env defaults when tracing config resolves Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ui): split audit log query guard to fit condition-chain budget Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
aa601ce4e8
|
refactor(repositories): daily activity repository with centralized bounded usage queries (#43398)
Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
ec605826d4
|
feat: improve trace ingestion and trace details (#43975)
* refactor: separate OTLP HTTP decoding from trace codec * feat: complete trace ingestion and read paths * fix: encode OTLP protobuf errors in Rust * fix: raise OTLP body limit to 16 MiB * test: cover OTLP auth body parsing boundary * refactor: parse OTLP media type into enum * fix: enforce OTLP body size at HTTP boundary * perf: preserve shared OTLP metadata across ingestion * bench: compare owned and shared trace resource fanout * refactor: extract shared storage and Python conversion caches * refactor: keep shared storage owned by traces * test: keep trace loopback coverage in Rust * test(proxy): adapt trace coverage to injected access context * fix(tracing): satisfy stacked branch lint checks * refactor(tracing): use immutable ingestion payloads * fix(tracing): declare native error encoder export * test(proxy): resolve trace access through dependency * fix(tracing): align merged normalizer types and bridge tests * fix(tracing): address ingestion and diagnostic review findings * fix(proxy): preserve body parsing for partial request scopes * test(proxy): use valid HTTP scopes in request fixtures * test(proxy): complete auth request flow scopes |
||
|
|
be67fce26a
|
refactor(proxy): inject tracing receiver and access context (#44035)
* refactor(proxy): inject tracing receiver and access context * refactor(proxy): own tracing resources through FastAPI lifespan * test(proxy): pass tracing dependency in Lens lifecycle * refactor(proxy): stop tracing logger cooperatively * refactor(proxy): derive tracing permissions in one place * refactor(proxy): compose application lifespan state * refactor(proxy): give Lens tracing storage directly * refactor(tracing): name shared ClickHouse storage explicitly * refactor(tracing): extract shared ClickHouse storage crate * test(proxy): isolate db push timeout from Lens safety check * fix(tracing): drain spend retries during shutdown |
||
|
|
24584d3d3d
|
test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy (#44012)
* test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): keep tuple identity in proxy state restore and fix misc target paths Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yuneng <yuneng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |