Commit graph

53578 commits

Author SHA1 Message Date
Daniel JB Clark
0e756bf0f8
test(clinepass): expect context-window and content-policy recognition
clinepass is mapped by _map_openai_exception, the same path as openai,
mistral and runwayml, so a full context window and a content-policy block
reach the caller as ContextWindowExceededError / ContentPolicyViolationError.
The per-provider tests added upstream pin that set, and clinepass was
missing from it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-03 13:58:26 -04:00
Daniel JB Clark
6d3a53f773
fix(clinepass): register the provider with litellm's own guards
Two upstream guards did not know about ClinePass, and both caught something real.

`test_every_cost_map_provider_is_registered[main|backup]`: its failure message
says exactly what a new provider owes — "a `<provider>_models` set in
litellm/__init__.py, filled in _populate_provider_model_sets and listed in
_build_models_by_provider". Adding the cost-map entry without that leaves the
provider's models unreachable through `models_by_provider`. Wired at all four
sites, mirroring `deepseek`. Verified with LITELLM_LOCAL_MODEL_COST_MAP=True,
since `model_cost` is fetched remotely by default and the remote map naturally
has no clinepass yet: `clinepass_models` and `models_by_provider["clinepass"]`
both resolve to `clinepass/deepseek-v4-flash`.

`test_a_provider_without_a_handler_maps_by_the_upstream_status[clinepass-403]`:
expected PermissionDeniedError, got APIError. The test derives its
no-handler list as `LlmProviders - PROVIDERS_WITH_A_HANDLER - aliases -
openai_compatible_providers`. Removing clinepass from
`openai_compatible_providers` (the P1 fix) dropped it into the no-handler bucket
while the explicit `_map_openai_exception` registration gave it a handler, so the
test's model of reality went stale rather than the code being wrong. `clinepass`
now sits in `PROVIDERS_WITH_A_HANDLER` beside `mistral` and `runwayml` — the two
providers whose registration shape this change copied.

Not addressed, and not mine: `proxy-infra`'s
test_every_model_in_the_prisma_schema_is_a_renderable_span_table. This branch
touches no prisma schema, span, or proxy-db file — the full changed-file list is
the clinepass provider, its tests, the two cost maps, the provider-support JSONs,
the UI helpers and README.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:58 -04:00
Daniel JB Clark
cb9aeab01d
fix(clinepass): type logging_obj and encoding instead of Any
The `lint` job's failing step was not plain ruff but
`scripts/ruff_strict_gate.py`, a strict-rule budget gate that compares totals
against the base commit: "ANN401: total 361 over limit 119 (this change added
2)", pointing at transformation.py:192 and :197.

Those two were `logging_obj: Any` and `encoding: Any`. The base class already
types them (`litellm/llms/base_llm/chat/transformation.py:350,355`), so they now
use `LiteLLMLoggingObj` and `"Tokenizer | None"`, imported under TYPE_CHECKING
with an `Any` fallback exactly as litellm/llms/openai/chat/gpt_transformation.py
does -- importing litellm_logging at runtime from a provider module risks a
circular import, which is presumably why that pattern exists.

This was also codex review finding #6, which said to "use the existing
logging/tokenizer types as current CometAPI and Perplexity transformations do".
I had read that as satisfied by the builtin-generics work and it was not.

`scripts/ruff_strict_gate.py --base upstream/main` now reports "OK: every strict
rule is within its codebase ceiling"; ruff clean under both CI configurations;
43 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:58 -04:00
Daniel JB Clark
d33e84b6e0
fix(clinepass): satisfy ruff — drop a redundant pass, narrow two raises
Two lint failures, both mine, found by CI rather than locally because I had been
running ruff from the repo root while the job runs it from `litellm/` and with
`ruff-tests.toml` for the test tree.

- PIE790: `ClinePassException` had a `pass` after its docstring.
- PT011 x2: `pytest.raises(Exception)` is too broad. Replaced with an explicit
  try/except that is also more precise about the contract: which exception
  litellm raises for an unsupported endpoint is its business and may change,
  whereas "nothing was transmitted" is what the test exists to prove. The
  fixture's own AssertionError (meaning the network WAS reached) is re-raised
  rather than swallowed as "some exception happened", which the previous form
  only caught via an isinstance check afterwards.

43 tests still pass; `ruff check` clean under both of the configurations CI uses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:58 -04:00
Daniel JB Clark
51214ed26e
fix(clinepass): address CI — mirrored cost map, logoless set, formatting
Three failures from the first CI run on the PR, all genuine and all local:

- `cost-map-guard`: litellm/model_prices_and_context_window_backup.json is a
  mirror of the root cost map and must be copied over whenever the root changes.
  I did not know that second copy existed.
- `ui-unit-tests`: provider_info_helpers.test.tsx asserts every provider maps to
  a bundled logo *except* a known-logoless set. ClinePass intentionally ships no
  logo -- the `<Logo>` component falls back to a first-letter circle rather than
  an invented asset -- so it belongs in that set, not in the logo map.
- `lint`: `ruff format` on litellm/main.py and the endpoint-guard test.

Not fixed here, because it cannot be: `documentation` and `code-quality` both
fail with "Environment variables read under ./litellm but mentioned nowhere in
the docs: ['CLINEPASS_API_BASE', 'CLINEPASS_API_KEY']". That check resolves docs
by checking out BerriAI/litellm-docs into docs/my-website (the test raises
"check out BerriAI/litellm-docs into docs/my-website" when the directory is
absent), so those two keys stay undocumented until a companion PR lands there.
It is a genuine cross-repo ordering dependency, not something this branch can
satisfy alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:57 -04:00
Daniel JB Clark
a6c9f01b39
test(clinepass): cover streaming, usage, auth headers and 429
Closes the last external-review finding (codex P2 #9), which noted the only
streaming test checked concatenated text and that there was no async-stream,
tool-call-fragment, usage-preservation or outgoing-Authorization assertion, and
asked for the registry/`__dict__` introspection to become behaviour checks.

Added: an async streaming test; tool-call fragments split across SSE chunks
reassembling intact; usage (prompt/completion/total) surviving the envelope
unwrap into ModelResponse.usage; the request actually carrying
`Authorization: Bearer <clinepass key>`; an explicit `api_key=` argument taking
precedence over CLINEPASS_API_KEY in the environment; and an upstream 429
mapping to RateLimitError, mirroring the existing 401 test.

Replaced three introspection assertions with observable behaviour: the JSON
provider-registry check now proves the transform layer is actually reached, and
the async-transform check proves `acompletion` evaluates the synchronous
`transform_request` rather than reading `__dict__`.

Kept the structural membership assertion rather than replacing it. The
behavioural test proves the consequence; the one-line structural test proves the
cause, with no mocking that could itself drift into passing for the wrong
reason. Both are wanted, and its docstring carries why the membership was
dangerous: the list is read by the speech branch (which pairs the provider's
api_base with OPENAI_API_KEY), by the transcription set derived from it, by
image generation, and by the extra_body wrapping that BaseLLMHTTPHandler never
unwraps.

36 -> 43 passing; ruff check and format green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:52 -04:00
Daniel JB Clark
edf31323fd
fix(clinepass): report the upstream finish reason instead of inferring truncation
`_correct_truncated_finish_reason` rewrote a single-choice response's
`finish_reason` from `stop` to `length` whenever completion usage reached the
requested cap. Two independent reviewers flagged it, and it is unsound:

* usage reaching the cap is a token count, not a reason for stopping. A natural
  completion, or a stop-sequence match, can land exactly on the cap, and the
  heuristic then mislabels a successful response as truncated -- which can
  provoke spurious continuation requests in callers;
* it read no explicit provider truncation signal, so it was pure inference;
* it was inconsistent. Real streaming never reaches `transform_response` --
  `BaseLLMHTTPHandler` delegates to `get_model_response_iterator()` -- so a
  streamed response kept the provider's `stop` while the identical non-streamed
  response was rewritten to `length`.

The original incident is recorded in a comment on `transform_response` so it is
not lost: ClinePass was once observed returning `stop` on a completion cut off
at 4000 tokens, and follow-up probes on 2026-08-22 did not reproduce it.

Tests: the three end-to-end rewrite tests become regression tests asserting an
explicit `stop` survives at and above the cap, including through the
`max_completion_tokens` -> `max_tokens` mapping path. The streaming test now
appends a terminal finish-reason chunk and asserts it survives, which closes the
streaming/non-streaming asymmetry that motivated the change. Ten tests that only
exercised the removed helper's internals (non-positive cap filtering, missing
usage, multi-choice guard, float coercion) are deleted with it: 46 -> 36 passing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:52 -04:00
Daniel JB Clark
69ecca0c7b
feat(clinepass): register the model in the price/context catalogue
`model_prices_and_context_window.json` had no ClinePass entry, so
provider-specific model info could not resolve: a DeepSeek row declares
`"litellm_provider": "deepseek"`, which `_check_provider_match` rejects when the
requested provider is `clinepass`.

Prices and limits are deliberately omitted rather than guessed. ClinePass is a
flat-rate monthly subscription, so inventing per-token rates would report an
invented allocation as provider spend; `null` prices are not schema-valid, and
the schema requires only `litellm_provider`. Omitting them follows the existing
subscription precedent -- several `github_copilot/*` entries carry metadata with
no price fields. Context and output limits are left unset because nothing in
this repo establishes them, and DeepSeek's limits are not evidence for
ClinePass's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:51 -04:00
Daniel JB Clark
bf18f22f9c
fix(clinepass): address external review -- logging, unicode, dead code, UI, README
Fixes from two independent reviews of the provider addition:

* `main.py`: drop the provider's own `logging.post_call`. The shared
  `BaseLLMHTTPHandler` already logs post-call, so this double-fired callbacks and
  overwrote raw-response metadata; on the async path `response` is still a
  coroutine there, so callbacks ran before the request completed. Also adopt the
  `Final` annotations and `_dispatch_client_http(ctx)` used by the sibling
  `_complete_*` blocks.
* `transformation.py`: `json.dumps(..., ensure_ascii=False)` when rebuilding the
  unwrapped body -- the default escapes non-ASCII to `\uXXXX`, and
  `OpenAIGPTConfig.transform_response` hands `raw_response.text` straight to
  `logging.post_call`, so observability backends received mangled text.
* `transformation.py`: remove two unreachable branches. `httpx.StreamError` cannot
  arrive, because `BaseLLMHTTPHandler` returns `get_model_response_iterator()`
  without calling `transform_response` when streaming (the only paths that do call
  it set `stream=False`, so the body is fully buffered). The
  `max_completion_tokens` fallback cannot fire either, because `request_data` is
  the post-`map_openai_params` payload and the mapping has already rewritten that
  key to `max_tokens`.
* `transformation.py`: modernise typing to builtin generics and PEP 604 unions,
  matching the repo's UP006/UP007/UP035/UP045 rules and the CometAPI house style.
* UI: add `CLINEPASS` to the `Providers` enum and `provider_map`.
  `CredentialModal` builds its options from `Object.entries(Providers)`, so
  backend credential metadata alone left the provider unselectable.
* README: add the provider table row, ticking only the endpoints
  `provider_endpoints_support.json` declares.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:32 -04:00
Daniel JB Clark
b9d159fcc9
fix(clinepass): stop unsupported endpoints reaching the provider with the OpenAI key
ClinePass implements chat completions only, but it was listed in
`openai_compatible_providers` to reach `_map_openai_exception`. That list is not
inert for routing:

* the speech branch in `main.py` matched it, and sends the request to the
  provider's own `api_base` while reading the credential from `OPENAI_API_KEY` --
  so `litellm.speech(model="clinepass/...")` POSTed the caller's OpenAI key to
  the Cline host for an endpoint that does not exist there;
* `OPENAI_AUDIO_TRANSCRIPTION_PROVIDERS` is derived from the same list, opening
  the identical hole for transcription;
* `litellm/images/main.py` consults it for image generation;
* `_add_provider_specific_params` packs unknown kwargs into `extra_body`, an
  OpenAI *SDK* concept the SDK unwraps client-side. ClinePass dispatches through
  `BaseLLMHTTPHandler`, which serialises optional params straight into the JSON
  body -- so membership put a literal `"extra_body": {}` on the wire on every
  chat request and buried genuine vendor kwargs one level deep.

Drop the membership and register ClinePass explicitly for `_map_openai_exception`
alongside `mistral` and `runwayml`, which is the established shape for a provider
with its own module. Exception mapping is unchanged and still covered by
`test_upstream_401_maps_to_authentication_error`.

Verified: speech, transcription and image generation now make zero outbound
requests and transmit no credential; unknown chat kwargs are flattened instead of
wrapped. Image generation returns an empty `ImageResponse` rather than raising,
which is pre-existing upstream behaviour for a provider with no image support --
`mistral` behaves identically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-03 13:55:31 -04:00
Daniel JB Clark
7d9ef1546d
test(clinepass): move the provider tests into the CI-selected tree
The tests were under tests/test_litellm/llms/clinepass/chat/, but
.circleci/scripts/unit_selection.sh:62 builds the llm-other-providers shard
with `find tests/unit/llms -name 'test_*.py'`, and GitHub Actions consumes the
same script. Nothing errored -- the 482 lines of tests simply never ran in CI,
so the PR would have looked green on tests that never executed.

Moved to tests/unit/llms/clinepass/chat/, matching the layout cometapi and
deepseek already use for an OpenAI-compatible provider, including the
__init__.py files those directories carry, and renamed to
test_clinepass_chat_transformation.py to match that same convention.

41 passed at the new location, where tests/unit/conftest.py applies rather
than tests/test_litellm/conftest.py. The CI selector now lists the file.
2026-10-03 13:55:29 -04:00
Daniel JB Clark
9a91da39d3
feat(clinepass): add Admin UI credential fields entry
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 890301fa4fc25b5df46c735c7f2ca2c024c8ac63)
2026-10-03 13:55:27 -04:00
Daniel JB Clark
e688210c3f
fix(clinepass): correct the model namespace, harden truncation detection
Addresses findings from three independent code reviews (Claude Opus 5,
GPT-5.6-sol xhigh via codex, Grok 4.6 xhigh via cursor) ahead of proposing
this branch upstream.

Correctness:

* The restored model qualifier was `clinepass/`, a namespace that does not
  exist. The catalog namespace is `cline-pass/` (hyphenated); `clinepass/` is
  LiteLLM's own routing prefix, which is stripped before the request is built.
  This failed silently because the API validates only the *shape* of a model
  id -- `totallybogus/deepseek-v4-flash` also returns HTTP 200. It is not
  inert, though: verified against the live API, `cline-pass/deepseek-v4-flash`
  resolves to `deepseek/deepseek-v4-flash` while any unrecognised namespace
  falls back to the date-pinned `deepseek/deepseek-v4-flash-0731`.

* `_correct_truncated_finish_reason` compared an aggregate
  `usage.completion_tokens` against a per-choice `max_tokens`. With `n > 1`
  that relabels naturally-finished choices as truncated: two 60-token choices
  under a cap of 100 aggregate to 120 and both became `length`. Restricted to
  single-choice responses, where the inference is sound.

* Zero, negative and `bool` caps are now rejected. `bool` subclasses `int`, so
  `max_tokens=True` was read as a cap of 1 and would relabel everything.
  Float usage/caps are now accepted rather than silently skipped.

* `_unwrap_response_envelope` reached for the private `_request` attribute
  because `httpx.Response.request` raises `RuntimeError` instead of returning
  `None`. Ask for the public attribute defensively instead.

* `get_models()` is overridden to return an empty catalog. ClinePass has no
  `/models` endpoint (404), and the inherited OpenAI implementation asked for
  it at the wrong path.

* Dropped the `async_transform_request` override. `BaseLLMHTTPHandler` builds
  the body with the synchronous `transform_request` on both the sync and async
  paths, so it was dead code -- the same shape of bug this provider shipped
  once already. Pinned with a test.

Packaging / CI:

* Added the `clinepass` entry to `provider_endpoints_support.json` and its
  backup. Without it `check_provider_folders_documented.py` fails, which is a
  required Code Quality job -- verified failing before, passing after.

* `transformation.py` did not satisfy `ruff format` under the repo's pinned
  ruff 0.15.3, which the changed-file CI gate runs. Reformatted.

* This commit replaces the previous branch tip, which had accidentally swept
  in ~435 regenerated Next.js artifacts under `litellm/proxy/_experimental/`.
  The branch is now +883/-0 across 12 files.

Honesty note on the truncation fix: re-probing the live API (streaming and
non-streaming, caps of 2000 and 4000, both namespaces) could NOT reproduce the
`stop`-instead-of-`length` misreport that motivated it. The correction is kept
as a conservative safety net and documented as such rather than as a workaround
for a currently-observable defect. Streaming is deliberately not covered; see
the docstring.

Tests: 41 pass (was 32). openai_like + cometapi regressions: 126 passed,
8 skipped.

(cherry picked from commit c2dda4fd8b7a4e1874642b25d73656292cb52468)
2026-10-03 13:55:26 -04:00
Daniel JB Clark
2553cbd337
feat(clinepass): add ClinePass provider
ClinePass (the Cline API) is OpenAI-compatible apart from two quirks:

1. Non-streaming completions are nested under a `data` envelope --
   `{"data": {"choices": [...]}, "success": true}` -- rather than returning
   `choices` at the top level. Against the openai SDK this surfaces as
   `r.choices` being None, not as an error. Streaming responses are *not*
   enveloped, so SSE needs no special handling.
2. A bare model id is rejected with HTTP 400 "invalid model format. Expected
   format: modelType/model", but LiteLLM strips its own `clinepass/` routing
   prefix before the request is built, so it has to be restored.

Both are handled in ClinePassConfig, which inherits OpenAIGPTConfig and
overrides only transform_request/async_transform_request (prefix) and
transform_response (unwrap). Unwrapping rebuilds the httpx.Response around the
inner object, so the inherited OpenAI response transform -- including its
`reasoning` -> `reasoning_content` mapping, which Cline populates -- is reused
rather than duplicated.

A response transform is the reason this cannot be a declarative entry in
litellm/llms/openai_like/providers.json: JSON-configured providers are
dispatched to `_complete_custom_openai`, which builds the request via
provider_config but parses the response with convert_to_model_response_object
and never calls provider_config.transform_response. Under that dispatch the
unwrap is unreachable, so ClinePass gets an explicit branch in main.py routing
it to base_llm_http_handler.completion.

`clinepass` is still listed in openai_compatible_providers, which is what
routes an upstream 401 through _map_openai_exception to AuthenticationError;
the explicit dispatch branch precedes that list's catch-all, so membership does
not send it back to the SDK path.

Tests drive litellm.completion()/acompletion() against a mocked transport
rather than calling the transforms directly, so dispatch itself is covered --
a unit test that calls transform_response() proves the function is correct but
not that anything invokes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 86b3486e7cc8834ec9067e45f966b18b7466ec3b)
2026-10-03 13:55:26 -04:00
moe-berri
ad8babae33
fix(lens): paginate trace reads within ClickHouse limits (#44384) 2026-10-03 10:38:40 -07:00
devin-ai-integration[bot]
8b1990b4bc
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register typesafe as a provider so Jev deployments load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move provider endpoints under llms and validate proxy bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add Cloudflare Clef and Strands Decider backends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register decisions routes for managed agents and gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(decisions): use raw regex for cloudflare missing account match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): avoid cast in Cloudflare response unwrapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): default model, evaluation health probe, short Cloudflare names

The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.

* fix(decisions): let health_check_params override the evaluation probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit the decisions endpoint across providers, limits, health and chaos

Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).

The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.

* fix(decisions): send env API keys to a configured api_base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register the routes through the lazy feature registry

The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.

The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.

* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell

* fix(proxy): let a config pass-through beat a lazily registered route in eager mode

With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:38:38 +00:00
devin-ai-integration[bot]
b024950353
feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop (#44391)
* feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(harness): name per-tool spec FunctionTool

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-03 17:33:01 +00:00
berriai-litellm-provider-info-sync[bot]
231a46e40b
feat(azure): add azure_ai/kimi-k2-thinking from Azure Kimi pricing page (#44382)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 10:04:59 -07:00
devin-ai-integration[bot]
a5e9f275ea
fix(proxy): keep the database error when the log_db_metrics failure hook raises (#44383)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 16:38:15 +00:00
devin-ai-integration[bot]
4c6c84afb7
perf(proxy): stop prompt-cache eligibility from tokenizing the whole conversation (#44221)
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them

Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
797353f13a
fix(otel): name postgres service spans by operation and table (#44240)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
564d236985
fix(otel): nest cache spans under their operation and name service spans by purpose (#44150)
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.

A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
devin-ai-integration[bot]
53d2ab6b05
fix(proxy): emit postgres service spans only on real DB reads in auth cache helpers (#44148)
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.

A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.

Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
berriai-litellm-provider-info-sync[bot]
66a422ea50
fix(azure): add MAI-Image max_input_tokens from models sold directly page (#44375)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 08:49:09 -07:00
devin-ai-integration[bot]
4ece6c9fb8
build(deps): suppress unfixed braces GHSA-vfj7-8cjw-p6xm to clear osv-scan (#44347)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 07:55:14 -07:00
devin-ai-integration[bot]
e768ad55ce
refactor(types): replace Any with proven types in 4 files (#44370)
* refactor(types): replace Any with proven types in 8 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert unproven email logger protocol and iterator annotation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert unproven deepagents and http handler annotations

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 03:52:28 -07:00
devin-ai-integration[bot]
e340e546e2
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* wip

* wip

* wip

* chore(trace): checkpoint ongoing Rust migration

* refactor(trace): group Python bridge under trace package

* refactor(traces): read span conventions through a Convention trait

Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.

The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.

* feat(trace): export Rust-owned wire schemas and enforce contract bounds

* fix(trace): bound quoted counts in ClickHouse wire schemas

* feat(trace): generate Python wire contracts with datamodel-code-generator

* test(trace): validate migrated callers and generated contracts at the native boundary

* refactor(traces): rename normalization convention to format

* fix(traces): reconcile spend evidence and preserve unknown costs

* feat(traces): normalize additional telemetry formats

* test(traces): cover captured normalization fixtures

* refactor(traces): isolate SDK normalization rules

* feat(tracing): seed all trace exports for local dashboard

* fix(clickhouse): preserve custom LiteLLM request metadata

* docs(traces): define normalization module boundaries

* docs(traces): define resolution and OTLP boundaries

* fix(ui): normalize nullable trace message names

* refactor(traces): split resolver modules and cover resolution behavior

* test(traces): replace normalization snapshots with behavior assertions

* fix(ui): align dashboard API contracts with generated types

* refactor(traces): type normalization and storage boundaries

* fix(traces): seed captured SDK spend and preserve provider identities

* wip

* test(traces): verify guide discovery and content ordering

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:59:51 +00:00
devin-ai-integration[bot]
d260765652
refactor: clean up fresh tech debt from 2026-10-02 (#44362)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 02:12:03 -07:00
devin-ai-integration[bot]
e92bd50de1
refactor(ui): inject Lens backends as services instead of a fake HTTP client (#44369)
The interactive Lens demo used to build a fetch shim that encoded in-memory
fixtures as HTTP responses so the shared ApiClient could decode them again,
and every trace view branched on demo vs live to pick a URL. Lens and the
trace views now depend on two small service interfaces, LensApi and
TracesApi, with named operations. The live layer wraps the existing HTTP
calls, the demo layer reads fixtures directly, and a React context provides
whichever one the session runs on. Without a provider the hooks fall back to
the live implementation, so the live app and existing tests are unchanged.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 09:11:08 +00:00
devin-ai-integration[bot]
09313bcf1b
refactor(ui): reorganize Lens dashboard components (#44354)
* refactor(ui): reorganize Lens dashboard components

* refactor(ui): align Lens forms with react-hook-form

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep Lens submit errors out of form validity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): reduce Lens lint budget usage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): use a query key factory for Lens queries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 08:29:16 +00:00
moe-berri
cb00dbecd7
feat(ui): improve trace inspection and ROI estimation (#44351)
* feat(lens): simplify trace inspection in the gateway drawer

* feat(lens): add a conversation view for full traces

* test(lens): keep normalized message fixtures type safe

* refactor(lens): make conversation view read like a chat

* refactor(lens): use a quiet trace view menu

* refactor(lens): make trace view tabs explicit

* refactor(lens): restore compact trace view switch

* fix(lens): preserve complete conversation history and tool types

* feat(lens): add trace full-screen and close controls

* fix(lens): show forwarded answers and agent errors once

* fix(lens): reset full screen when closing a trace

* fix(roi): estimate linked authors and clarify model selection

* fix: preserve trace errors and ROI results across partial failures

* fix(roi): correct pagination variable typing

* fix(ui): place loaded root failures in conversation order

* fix(roi): read estimator recommendations from model catalog

* revert: remove catalog-driven ROI recommendations
2026-10-03 01:18:16 -07:00
devin-ai-integration[bot]
5724117116
fix(responses): merge bridged tool calls into the same choice as the text (#44346)
* revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295)

This reverts commit ca1994e403.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): merge bridged tool calls into the same choice as the text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 07:00:52 +00:00
tin-berri
8e32d4568c
feat(ui): move LiteAdmin into the header with a docked side panel (#44293)
The floating bottom-right LiteAdmin button covered page controls such as
the Logs pagination buttons, and Playground had to hide it entirely.
Render the trigger as a pill in the header tools ahead of Docs and open
LiteAdmin as a panel docked beside the content column, which narrows the
page instead of covering it. Add a Cmd/Ctrl+J toggle and drop the
Playground override.

The Logs and trace drawers treated Cmd+J as a plain J and advanced the
selection, so they now share RunDrawer's rule that letter shortcuts
yield to modified presses and typing.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 00:00:10 -07:00
devin-ai-integration[bot]
0c86d6bfc9
refactor(dashboard): migrate to zod 4 and openai 6 (#44345)
* refactor(dashboard): migrate to zod 4 and openai 6

Bump the dashboard to real zod 4.6.5 and openai 6.49.0 so every module
imports from bare "zod" instead of "zod/v4". Ports the ten files that
still used the zod 3 API (error params, record, passthrough, strict,
email/date validators, union discriminator codes) and adapts the
LiteAdmin tool schemas and form plumbing where openai 6 and zod 4
changed behaviour.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(dashboard): prettier format zod 4 schema files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(dashboard): restore system one missing state message under zod 4

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:53:41 +00:00
devin-ai-integration[bot]
d96e56c76f
revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295) (#44344)
This reverts commit ca1994e403.

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:47:11 +00:00
devin-ai-integration[bot]
7b432d78d2
fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model (#44341)
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): streamed alias matching a capability rule bills the deployment price

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): assert every streamed chunk carries the client alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): log the client alias on the priced streamed response, the same as non-streamed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:44:52 +00:00
devin-ai-integration[bot]
a76ba8c01e
revert(cost): revert "fix(cost): price rule-only model names at the deployment's rate" (#44144) (#44335)
This reverts commit f9a32ffcb5.

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:09:00 +00:00
devin-ai-integration[bot]
0fce5bccbc
fix(model-prices): mark azure us/eu responses-only models as mode responses (#44323)
* fix(model-prices): mark azure us/eu responses-only models as mode responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): keep azure us/eu o3-deep-research on chat, which Azure lists as chat capable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 23:05:55 -07:00
moe-berri
caee45fed4
feat(roi): add GitLab sources and branch cost attribution (#44324)
* feat(roi): support GitLab and tagged branch costs

* fix(roi): count tagged branches independently of estimation status

* test(roi): capture live GitHub and GitLab report validation

* fix(roi): open estimate details at the start

* fix(roi): clarify cost views and unify report layout

* feat(roi): showcase per-PR costs in the sample report

* fix(roi): separate report tabs and preserve branch cost attribution

* fix(roi): preserve demo previews and align progress spacing

* fix(roi): isolate demo loading and parallelize fork lookups

Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression

* fix(roi): separate demo and live loading states

Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers

* fix(roi): ignore refreshes from a previous source

* fix: trust gateway context for ROI estimator exclusion

* fix: preserve historical ROI estimator exclusion
2026-10-03 06:01:43 +00:00
yucheng-berri
12d4b75b7a
fix(search_tools): encrypt search tool litellm_params at rest (#43631)
* fix(search_tools): encrypt search tool litellm_params at rest

Encrypt every string value of a search tool's litellm_params on create and
update and decrypt on every DB read, so legacy plaintext rows load unchanged.
Include the table in master key rotation, LITELLM_MIGRATE_FROM_MASTER_KEY and
the migrate-encryption scan.

* fix(search_tools): keep edits made while the master key rotates

Write each rotated search tool row only if it still holds the litellm_params
that were read, and re-read and rotate it again if it was edited in between,
so a PUT that lands during /key/regenerate is not overwritten.

* fix(search_tools): retry rotation writes until the row stops changing

Rotate a search tool row again for as long as it keeps being edited instead of
giving up after five attempts, and stop with a warning only when the conditional
write fails on an unchanged row. Build the decrypted read result without
mutating it in place.

* refactor(search_tools): rotate edited rows in a loop, drop the step comment

Retry the conditional rotation write in a loop instead of recursion so sustained
edits cannot deepen the call stack, drop the step comment on the rotation call,
and stop mutating local state in the rotation tests.

* test(search_tools): drop the rotation test docstring

* Store search tool params as written when no encryption key is configured

* Rotate search tools under the salt key, keep non-ciphertext values and loaded tools that do not decrypt

* Treat a search tool as undecryptable only when its provider is ciphertext-length

* Drop suppressions the type discipline gate on main now reports as unused

* Show the loaded search tool in the admin list and info views when its DB params do not decrypt

* Keep the DB row's other fields when the admin views substitute loaded params
2026-10-02 22:53:18 -07:00
yucheng-berri
5d42cb7cfa
fix(guardrails): encrypt guardrail litellm_params secrets at rest (#43627)
* fix(guardrails): encrypt guardrail litellm_params secrets at rest

* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation

- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals

* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor

- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes

* test(guardrails): drive the real guardrail rotator from the master key rotation test

- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key

* Annotate guardrail param encryption collections for type-discipline gate

* Type guardrail param recursion through validated JSON containers

* Type guardrail registry test helpers and drop section comment

* Reject client-supplied encrypted values in guardrail litellm_params

* Allow depth-bounded contains_encrypted_marker in the recursion detector

* Keep a loaded guardrail when its DB params do not decrypt with the current key

* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models

* Keep the loaded guardrail when an undecryptable param has no loaded value

* Drop suppressions the type discipline gate on main now reports as unused

* Assert what the reinitialized guardrail holds after an edit to an undecryptable one

* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize

* Type the rotation test helpers and drop the new test docstrings

* fix(guardrails): refuse to approve a submission whose params do not decrypt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:49:41 -07:00
devin-ai-integration[bot]
6d8434f940
fix(proxy): return 4xx instead of 500 for missing required params, invalid pagination and unknown ids (#43787)
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run search_endpoints tests in proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(llm_http_handler): keep provider error text when re-raising mapped errors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): allow promptless image edits and default search models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): default missing image edit image to None

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): build image edit defaults without mutating request data

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): inject a fake router for the search default model test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on vector store and file lookup handlers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): keep the lookup handler raise block to a single statement

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover missing required body params and provider lookup status codes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bind spend-row request id with partial to satisfy B023

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): only reject non-positive page_size on vector store list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): remove unreachable fine-tuning body validation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover streaming anthropic messages reaching the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): count only provider calls when asserting missing params never reach the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve merge-base request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve interaction completion model defaults

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-02 22:48:53 -07:00
devin-ai-integration[bot]
8efb4a21f6
fix(vector_stores): return managed file ids from vector store file list (#43800)
* fix(vector_stores): return managed file ids from vector store file list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(vector_stores): cover managed file list route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vector_stores): only map round-trippable managed ids and index flat file ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy-extras): build managed file gin index concurrently

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(proxy-extras): move the managed file gin index migration after main's newest

* fix(vector_stores): satisfy lint gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(proxy): drop the stale no-index note on the raw file-id guard

* test(vector_stores): cover managed file ids on the vector store file list end to end

Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 21:28:50 -07:00
Ankit Jha
0d5ea45c86
feat(scaleway): add rerank support (#44160)
* feat(scaleway): add rerank support

Fixes #43856

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>

* fix(scaleway): caller headers cannot replace the provider key; inject the async client in tests

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>

* refactor(utils): fold Scaleway into the Jina rerank branch to stay inside the complexity budget

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>

---------

Signed-off-by: Ankit Jha <ankit.jha@tradomate.one>
2026-10-02 21:20:05 -07:00
michelligabriele
119179942e
fix(guardrails): run unified guardrails on /v1/images/edits (#44195) 2026-10-02 21:14:35 -07:00
Koh Jun Hao
5ce81631fe
fix(bedrock): send the anthropic-workspace-id header to Bedrock Mantle (#44173)
aws_bedrock_project_id reached Bedrock Mantle as anthropic-workspace, a
header AWS ignores, so requests ran under account defaults and models
that require a project data-retention mode failed with a 400. Send the
header AWS documents, anthropic-workspace-id, on both Mantle routes.

Re-lands #31994 by @gunjanjaswal, merged into the since-deleted
litellm_oss_staging branch on 2026-08-22, which never reached main.

Fixes #31947

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 21:12:33 -07:00
joeym82956
2394fc3682
fix(streaming): keep the finish_reason of the provider's last content chunk in the logged response (#44318)
* fix(streaming): keep the finish_reason of the provider's last content chunk in the logged response

When a provider sends the last content or tool-call delta and the finish_reason in one chunk (vLLM does this when speculative decoding finishes a reply in one step), the wrapper strips the finish_reason from that chunk and re-emits it on a terminal chunk from finish_reason_handler(). The terminal chunk reached the caller but was never added to self.chunks, so the response built for callbacks and SpendLogs defaulted to finish_reason "stop": tool calls and replies truncated by max_tokens were logged as plain stops. Append the terminal chunk in both the sync and async end-of-stream paths.

* fix(streaming): only log the terminal chunk when the provider sent a finish_reason

Review feedback: when a stream ends before any finish_reason arrives (e.g. an Anthropic stream cut after message_start), finish_reason_handler() synthesizes "stop"; adding that chunk to self.chunks made the response builder take the placeholder usage instead of estimating it. Append the terminal chunk only when received_finish_reason or intermittent_finish_reason is set, and test that case.

---------

Co-authored-by: joeym82956 <244252723+joeym82956@users.noreply.github.com>
2026-10-02 21:11:40 -07:00
devin-ai-integration[bot]
8596fe954d
feat(interactions): durable cross-pod settlement for background interaction billing (#41955)
* feat(interactions): durable cross-pod settlement for background interaction billing

Background interaction billing lived only in the creating replica's memory, so a
DELETE routed to another replica, or a restart of the creating one, never billed
the completed provider work and the budget reservation was refunded at the poll
timeout. The create now registers the billing context in a settlement store
before returning, the proxy installs a Prisma-backed store at boot
(LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica
claims the row once through a conditional update before billing or releasing,
startup resumes every unclaimed row with its remaining timeout, and a give-up
records an unsettled outcome instead of silently reconciling to zero. The SDK
keeps an in-memory store and behaves as before.

* fix(interactions): survive a settlement install failure at boot and stop carrying request headers

* fix(interactions): drop the stored request context once a settlement row is settled

* fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline

* fix(interactions): carry a missing model through the settlement context for agent-only background creates

An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it.

* fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere

The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials.

* fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill

* fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed

A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler.

A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere.

* fix(interactions): settle an unverified registration through the durable claim

A create whose settlement-store write raised no longer bills through a
private in-memory gate that a later boot's resume cannot see. The claim
asks the durable store first and falls back to the local gate only when
the store answers that no row exists, and a missing settlement table reads
as no rows so a replica without the migration still settles in process.

* test(proxy): keep the settlement test where the proxy-infra shard collects it

The merge of main moved test_background_interaction_settlement.py under
tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy
files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it
and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected
whole by the proxy-infra shard, which is where the test ran before the merge.

* fix(interactions): raise on a non-2xx Gemini interaction fetch

AsyncHTTPHandler.get never raises for status and the Gemini GET transform
only raised when the body was not JSON, so a 500 or 404 carrying Gemini's
JSON error body parsed as an interaction with no status. A delete on a
replica other than the creator then claimed the settlement as released and
forwarded the delete instead of failing closed, and the bill was lost. The
transform now raises GeminiError with the vendor's status, as the delete
transform already does; the in-process poll already retries a fetch that
raises

* test(integration): audit durable background interaction settlement across replicas

Twenty-six deterministic cells drive a one-worker creator and a two-worker
settler against an owned scripted Gemini upstream: cross-replica deletes
bill once, failed and cancelled interactions release, a later replica
resumes unclaimed rows, custom deployment pricing bills at the deployment
rate, a fetch the settler cannot make fails the delete closed, odd ids are
refused, a missing settlement table keeps in-process billing, polling
disabled registers nothing, the budget reservation is released by the
settler, an upstream outage mid-burst fails closed and recovers, killed
workers hand their polls to the respawned ones, and concurrent deletes on a
slow upstream settle exactly once. The support upstream gains a scripted
interaction store with per-id GET status and delay, and the process helper
gains an owned upstream a test can stop and restart

* test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate

* chore(ui): regenerate dashboard API types after merging main

* test(integration): accept the 422 budget refusal and a respawned worker's resume

The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged

* test(integration): delete the pinned key's interaction with a second key

A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 21:08:38 -07:00
Matthew Lapointe
ad566c90dd
fix(proxy): use OpenAI workload identity federation tokens on /openai_passthrough (#44140)
* fix(proxy): use OpenAI workload identity federation tokens on /openai_passthrough

The OpenAI passthrough routes (HTTP and websocket) only looked up a static
OpenAI API key, so proxies authenticating to OpenAI through workload identity
federation failed with "Required 'OPENAI_API_KEY'". When no non-empty static
key is configured, exchange the workload identity subject token for a bearer
token, scoped to the passthrough's own OPENAI_API_BASE so the token is never
sent to a non-OpenAI host

* fix(proxy): close the OpenAI websocket passthrough cleanly when the workload identity exchange fails

A rejected, unreachable or unreadable workload identity exchange raised out of
the websocket route before the handshake was accepted, so clients saw a bare
handshake failure. Log the cause and close with 1011 and a fixed reason instead

* refactor(openai): resolve workload identity bearer tokens for an api base in the OpenAI provider module

The passthrough route only owns the static key lookup now and asks the OpenAI
provider module for a workload identity bearer token scoped to its api base

---------

Co-authored-by: Krrish Dholakia <krrish+github@berri.ai>
2026-10-03 04:07:09 +00:00
ishaan-berri
80a2f4d8a8
feat(lens): add preset watch-for checks to investigation setup (#44313)
* feat(lens): add preset watch-for checks for common agent failures

* feat(lens): add keyboard-driven watch-for picker with lens dot animation

* feat(lens): use the watch-for picker in investigation setup

* feat(lens): show preset checks by name in the criteria tab

* test(lens): cover saving and editing watch-for presets

* feat(lens): shorten watch-for summaries and start with three presets on

* feat(lens): lay out watch-for presets as toggle tiles with a clear add-your-own button

* feat(lens): open a custom check from the watch-for picker

* test(lens): cover watch-for tiles and the add-your-own button

* fix(lens): draw the selected tile border inside the tile so the dialog edge cannot clip it
2026-10-02 21:06:04 -07:00