mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-07 02:59:05 +00:00
53646 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8364f88cbb
|
ci: run unit selections from GHA test-path and drop the CircleCI unit jobs (#44461)
* ci: run unit selections from GHA test-path and drop the CircleCI unit jobs * ci: keep existing shard token order and header note * ci: throwaway, drop tests/unit/repositories from the unit shard to show assert-ci-coverage fails * ci: revert throwaway assert-ci-coverage check * ci: stop crediting --ignore paths as invoked in assert_ci_coverage * ci: install the caching, extra_proxy and proxy-runtime extras in the GHA unit sync --------- Co-authored-by: yuneng <yuneng@berri.ai> |
||
|
|
9b6a6a0b71
|
refactor(tracing): generate existing HTTP request models from Rust schemas (#44591)
* test(tracing): pin HTTP request compatibility Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(tracing): generate existing HTTP request models from Rust schemas Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(tracing): bind trace query params to generated request models Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci(rust): raise the native wheel size gate to 48 MB Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * perf(tracing): read the trace list clock without a thread-pool dispatch Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(traces): share the trace page-size bounds between schema and reader Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(ui): type trace request queries against the generated OpenAPI schema Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs: encode the trace contract boundary in AGENTS.md Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(ui): format trace request aliases Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
d946706744
|
feat(guardrails): add llm shield pii redaction and rehydration guardrail (#42645)
* feat(guardrails): add llm shield pii redaction and rehydration guardrail
LLM Shield is a self-hosted PII gateway. This adds it as a guardrail so a
proxy operator can redact personal data out of outbound requests and have
the original values restored in the model's reply.
The substitution is reversible, which is the difference from a masking
guardrail. Outbound text is replaced with placeholders held in a session
vault inside the operator's own LLM Shield deployment, and the reply is
restored before it reaches the caller, so the end user still sees real
values while the provider never received them.
Streaming responses are restored incrementally. LLM Shield holds back only
the trailing characters that could still turn out to be part of a
placeholder, so tokens are forwarded as they arrive rather than the whole
response being collected first. A placeholder split across two chunks is
never emitted in fragments.
The integration talks to LLM Shield over HTTP and adds no dependency.
Notes for reviewers:
- The guardrail sets use_native_lifecycle_hooks, since redaction and
restoration need the native pre-call, post-call and streaming hooks
rather than the unified path.
- Per-request state lives on the request dict, never on the guardrail
instance, because the proxy registers a single instance process-wide.
The streaming carry-over is a local of the generator for the same reason.
- Every failure blocks the request. A redaction guardrail that fails open
would send the exact data it exists to protect to the provider.
* feat(ui): list llm shield in the guardrail garden
Adds the card, preset and logo so operators can pick LLM Shield from the
guardrails page the same way as the other partner guardrails.
* docs(guardrails): add llm shield example config
Shows both modes on one entry. Listing only pre_call redacts the request
and then hands the placeholders back to the end user, so the test asserts
both hooks are enabled.
* feat(ui): use the llm shield brand mark for the guardrail logo
* fix(guardrails): restore llm shield values in anthropic replies
The /v1/messages reply is a plain dict with a content block list and no
choices, so it fell through the restore path and went back to the caller
still carrying placeholders. The request was redacted correctly, which is
what made this easy to miss.
Found by running all three endpoints against a live provider; the mocked
tests all passed because they only built the OpenAI shape. Adds tests for
the message shape and for leaving non-text blocks alone.
* docs(guardrails): correct the llm shield start command
* fix(guardrails): redact every request shape and restore every reply shape
Three gaps, all of which let an enabled guardrail hand data to the provider
or hand placeholders to the caller.
Requests only walked `messages`. The Responses API `input` and tool call
`arguments` went out untouched. Measured against a live provider: a request
sent through `/v1/responses` reached the model with the real address in it
while the guardrail reported as enabled. Request traversal now covers chat
content (string and multimodal), tool call arguments, and `input` as a bare
string or a list of items.
Fixing that exposed the matching gap on the way back: the Responses API reply
carries `output` items rather than `choices`, so it returned to the caller
still holding placeholders. It now gets its own walk, handling text blocks as
dicts or objects.
The dashboard preset seeded only pre_call, so a guardrail created from the UI
would redact the request and return the placeholders to the user. Presets can
now seed both modes; the form already normalised either shape.
Adds tests for each request shape, for both Responses API reply forms, and
replaces a test that had asserted the `input` bypass as correct behaviour.
* fix(guardrails): narrow the stream delta before writing to it
basedpyright could not prove the delta was non-None on the write path, and
reportOptionalMemberAccess has a zero budget. The guard is also clearer than
relying on the text check to imply it.
* fix(guardrails): mint the vault id instead of trusting the caller's
The vault id was taken from caller-supplied session metadata, and every
caller shares one LLM Shield key. Someone who knew or guessed another
caller's session id could send a placeholder, have the model echo it back,
and get that caller's plaintext restored into their own reply.
Vault ids are now minted per request behind a per-process prefix, so a
caller cannot name a vault this process uses. Redaction mints, restoration
reads back, and a reply whose id does not match is left holding its
placeholders rather than resolved against some other vault.
Also covers two more request fields that were reaching the provider intact:
the Responses API `instructions`, and the legacy `function_call.arguments`
alongside `tool_calls`.
The collectors move to module level, which drops the traversal back under
the complexity limit and lets the code carry its own explanation instead of
the comments that were restating it.
* fix(guardrails): drop Final from a loop-assigned local
basedpyright rejects a Final assigned inside a loop, and
reportGeneralTypeIssues sits one over its budget ceiling.
* fix(guardrails): redact completion prompts and responses tool items
Two more provider-bound request shapes were reaching the model intact while
the guardrail reported as enabled.
/v1/completions carries its text in a top-level `prompt`, which the
traversal never looked at. It is handled as a string and as the array form,
where each entry is rewritten in place.
Responses input items hold tool data outside `content`: a function_call item
in `arguments`, a function_call_output item in `output`. Both are now
collected alongside the item's content.
Adds a test per shape.
* fix(guardrails): redact the anthropic system prompt and string-array input
Two more provider-bound shapes, found by walking the request types rather
than waiting for them to be reported.
/v1/messages carries its system prompt at the top level, as a string or a
list of text blocks. It is one of the endpoints this guardrail claims to
cover, and a system prompt is a natural place to put a customer's details.
`input` as an array of bare strings, the embeddings and moderations shape,
was skipped because the loop only handled item dicts.
Verified against a live provider: a system prompt holding an address now
reaches the model as a stand-in and is restored in the reply.
* fix(guardrails): narrow prompt and input to a list before iterating
Guarding with a conditional iterable left the value un-narrowed, so passing
it on was an argument-type error and the element checks read as unreachable.
An early return narrows it properly and reads better.
* fix(guardrails): restore every streaming choice, not just the first
Streaming rehydration read and rewrote choices[0] only, so with n>1 every
later choice went back to the caller still holding its placeholders.
Each choice is its own token stream, so the sliding window is now tracked
per choice index rather than once per stream. A single shared window would
have been worse than the bug: it would splice the characters held back for
one choice onto the next one's delta.
The final flush walks every choice the same way, and the two helpers that
only ever looked at choices[0] are gone.
Adds a test that both choices come back restored, and one that each choice
gets its own window handed back rather than its neighbour's.
* refactor(guardrails): name the guardrail llm_shield_proxy throughout
The integration was called llm_shield in code, llm-shield in the example
config, and LLM Shield in the dashboard, while the product and its PyPI
package are both llm-shield-proxy. An operator who saw the guardrail in
LiteLLM could not tell what to install.
One identifier now: llm_shield_proxy for the enum value, module, directory,
class, config model, logo and environment variables, with LLM Shield Proxy
as the display name. That matches `pip install llm-shield-proxy`.
Renames only; no behaviour change.
* feat(guardrails): redact the participant name on a message
`name` on a user or assistant turn identifies a person and was going to the
provider intact. The proxy this integrates with already redacts it, so the
integration was the weaker of the two.
On a tool or function turn the same field carries the function's name, which
has to arrive unchanged or the call stops routing. That case is skipped, and
a test asserts the value is never even sent to the shield.
* fix(guardrails): flush every held choice, and cover tool results and suffix
Three review findings.
The trailing flush walked the last chunk's choices, so a choice that finished
earlier and stopped appearing lost whatever text was still held for it and its
answer was truncated. It is now driven by the windows themselves and emits one
chunk per choice, synthesising the choice when the terminal chunk omits it.
That was data loss, not just under-redaction.
An Anthropic tool_result carries its own content, as a string or as further
blocks, and only each part's `text` was being collected. Handled recursively;
image and audio parts still fall through untouched.
The legacy completions `suffix` is forwarded to providers that support it and
was never collected. Note the placement: it has to be gathered before the
string-prompt early return, which is what the new test pins.
* fix(guardrails): walk nested tool results iteratively, with a depth bound
CI flagged _collect_content as recursive. It was, and worse, it was unbounded:
a tool_result nests its own content, the nesting is caller controlled, and the
descent had nothing to stop it. That is a JSON bomb, not a style issue.
Now an explicit queue with a depth bound of 8. Real payloads nest one or two
deep. The queue is walked in document order because the shield maps its replies
back by position, so collection order is part of the contract.
* fix(guardrails): redact Responses PromptObject variables
A Responses request can send `prompt` as a PromptObject rather than a string.
Its `variables` are substituted into the stored prompt on the provider side, so
they are caller text, and the dict shape was falling through untouched.
`id` and `version` pick which stored prompt to run and are left unchanged.
* test(guardrails): assert the depth bound instead of only reaching the end
The depth test asserted nothing, so it passed whether or not the bound held,
and the test-quality gate counted it as a zero-assert test. It now sends a
shallow value alongside a 200-deep chain and asserts the shallow one is
collected while the value past the bound is not.
* fix(guardrails): keep system-prompt values out of the restored reply
Redaction put every span of a request into one vault, and the reply was restored
against that same vault. System prompts are written by the application and the
caller never sees them, so a caller who got the model to echo a placeholder back
had its plaintext restored into their own reply -- a way to read a system prompt
they were never shown.
Server-authored spans now go into a vault of their own: system and developer
turns, Anthropic's top-level `system`, and the Responses API `instructions`.
Its id is deliberately never stored, so nothing restores against it. The reply
is restored against the caller's vault alone, and an echoed placeholder from a
system prompt comes back as the placeholder.
Values the caller also wrote themselves are unaffected -- they are in the
caller's vault too, and still restore. The extra round trip happens only when a
request actually carries server-authored text.
* style(guardrails): satisfy ruff format and annotate the new tests
`ruff format` wanted the widened `_collect_responses_fields` signature on one
line, and the three tests added with the split-vault fix needed return
annotations to keep ANN201 level with the base.
* fix(guardrails): restore tool calls in the LLM Shield guardrail
The request walk redacted a tool call's `arguments` -- plus the legacy `function_call`,
Anthropic `tool_use.input` leaves and the Responses API's `function_call` /
`function_call_output` fields -- while the response walk restored only `message.content`.
A placeholder therefore reached the caller inside a tool call, and nothing raised.
This is the same change as the out-of-tree example adapter this file is copied from, kept
body-identical on purpose: the response side now collects every restorable span in one
positional rehydrate batch, streaming keeps a window per (choice index, tool-call index)
and flushes each into the chunk carrying the finish_reason, and `apply_guardrail` restores
`inputs["tool_calls"]` on the response side. The declared limit on restoring values inside
a JSON string is documented in the module.
* fix(guardrails): import copy, keep the vault id off the provider, drop recursion
Three defects Greptile and veria-ai found on the reopened PR, all real:
- `copy.deepcopy` was called in `apply_guardrail` with no `import copy`, a
guaranteed NameError on every response carrying tool calls. It landed on
2026-09-13, ten days after the review that rated this branch safe, and no test
reached it: every tool-call test covered the request side. Adds the import and
a regression test on the response side.
- The vault session id was stored in `metadata`, which is forwarded to the
provider on /v1/responses. A provider holding the placeholders and the session
id can call the shield's rehydrate endpoint and read back the plaintext this
guardrail exists to withhold. Moves it to `litellm_metadata`, which is not
forwarded, and reads it back from there only.
- `_collect_json_leaves` recursed over model-controlled JSON; the repo's
recursive_detector gate rejects that. Rewritten with an explicit stack, same
depth bound.
52 tests pass. ruff format, ruff-strict and check_type_discipline all clean, with
LIT counts identical to the merge base.
* fix(guardrails): build llm_shield_proxy stream deltas without new mutable literals
The lint job's LIT002 budget gate failed on this PR: the file added 11
mutable-collection constructions and the tree sits at its limit. Build the
index-only tool-call continuation in one helper, keep read-only inputs as
tuples, and annotate the lists the delta and texts fields require.
Adds tests for the two tool-call flush paths the refactor touches, which
had no coverage: held arguments landing in the finish_reason chunk next to
that chunk's own fragment, and the trailing flush of a stream that ends
without a finish_reason.
* fix(guardrails): drop Final from loop-body locals in llm_shield_proxy
basedpyright rejects Final on a name assigned inside a loop, and the eleven
such locals put reportGeneralTypeIssues over its budget (112/101). The LIT010
Final rule already exempts loop-body assignments, so the annotations go.
* feat(guardrails): restore llm_shield_proxy placeholders on native streams
Anthropic /v1/messages and /v1/responses streams have no `choices`, so the
streaming hook passed them through with placeholders still in them. Both
are now restored incrementally, with the same per-stream windows as chat:
- /v1/messages arrives as raw SSE. Frames are cut at event boundaries,
text_delta and input_json_delta are restored per block index, and held
text is emitted as one more delta ahead of content_block_stop. Signed
thinking deltas, frames from other endpoints and non-SSE raw streams
pass through unchanged.
- /v1/responses events are restored per item and part. Held text goes out
as a copy of the stream's last delta before its .done event, and the
events that repeat the reply (.done, content_part.done, output_item.done,
response.completed) are restored in full.
The request side now also redacts Anthropic tool_use inputs and Responses
reasoning summaries, and sends tool and function descriptions (including
parameter schema descriptions) and the user / safety_identifier fields to
the non-restorable vault, like system prompts. Tool results stay
restorable: the model reads them to answer, so restoring them returns what
the caller would have seen without the guardrail.
* fix(guardrails): redact llm_shield_proxy predicted outputs and output schemas
`prediction.content` is the caller's own draft of the reply, so it is
redacted into the caller vault and restored with the reply. The
descriptions in a structured-output schema (Chat
response_format.json_schema, Responses text.format) are application
authored like tool schemas, so they go to the non-restorable vault.
* fix(guardrails): fail closed on deep llm_shield_proxy requests, widen coverage
- Request walks no longer skip what lies past their depth bound. Content
nested past it, and tool inputs or schemas past the new JSON bound, now
block the request instead of reaching the provider unredacted. The old
depth test asserted the skip; it now asserts the block.
- Tool and output schemas are walked by their JSON Schema structure, and
give up `title`, `examples` and `default` as well as `description`.
`enum` and `const` still go out as sent.
- Responses events are matched by shape: any `*.delta` with a string delta
is a token stream, and any `*.done` restores every non-identifier text
field plus the `part` or `item` it repeats. This covers
reasoning_summary_part.done and MCP arguments, and future families.
Audio deltas are left alone.
- An SSE stream whose first chunk ends partway through a field name
(`b"eve"`) is no longer taken for a non-SSE stream.
* fix(guardrails): scan llm_shield_proxy schemas by default
The schema walk collected an allowlist of keywords, so any keyword it did
not list -- draft-07 `dependencies`, `$comment`, vendor `x-` extensions --
went to the provider in clear. Invert it: every string is collected except
under keywords whose value must go out verbatim (types, formats, patterns,
references, required lists, enum, const). Name -> subschema maps still
treat their keys as property names, so a property called `type` is
walked, not skipped.
* fix(guardrails): redact llm_shield_proxy schema enum and const values
`enum` and `const` were skipped by the schema walk, so a value holding PII
went to the provider in clear. They now go to the caller's vault rather
than the non-restorable one: the model emits the stand-in in its tool
arguments or structured output, and restoring the reply turns it back into
the value the schema allows, so the call still routes.
* fix(guardrails): redact llm_shield_proxy web search user locations
Web search forwards the user's approximate location, and its free-text
`city` and `region` fields can hold an address. Collect them into the
non-restorable vault, from Chat `web_search_options.user_location` and
from the `user_location` of Responses and Anthropic web-search tools.
* fix(guardrails): drop unused llm_shield_proxy suppressions
Upstream added LIT013 (a *-ok marker that suppresses nothing) and LIT014
(at most one for and one if per comprehension). Remove the 34 markers
that no longer suppress anything and flatten the finished streams with
itertools.chain.from_iterable.
* fix(guardrails): type the llm_shield_proxy request and reply walks
Narrowing with isinstance(x, dict) leaves keys and values unknown, so
every call that passed a narrowed value counted against the
reportUnknownArgumentType budget. Parse into dict[str, object] and
list[object] once, in _as_object and _as_array, type the carry keys and
accumulators, and bind writers with functools.partial instead of lambdas.
The shield's batch reply is now also checked to hold only strings.
* fix(guardrails): keep restored llm_shield_proxy replies out of the cache, widen coverage
Addresses the open veria-ai and Cursor Bugbot findings on #42645.
- Restore a copy of the reply and of each stream chunk, never LiteLLM's own object.
LiteLLM caches and logs that object, and placeholders are numbered per request, so
two callers' redacted requests can share a cache key: restoring in place cached one
caller's plaintext for the next. The deployment hook no longer restores either,
since LiteLLM caches what it returns; the proxy's post-call hook restores
model-level guardrails after the cache write.
- Restore /v1/completions replies, streamed and not, which carry `choice.text`.
- Redact Responses replay fields the reply side already restores: tool output sent
as input_text parts, custom_tool_call `input`, code_interpreter_call `code`.
- Redact typed Responses prompt variables (`{"type": "input_text", "text": ...}`).
- Put Responses system and developer input items in the non-restorable vault, like
their Chat counterparts.
- Expose LLMShieldProxyGuardrailConfigModel through get_config_model, so the
dashboard can collect the Shield URL and key.
* fix(guardrails): redact llm_shield_proxy plain-text document blocks
An Anthropic document block carries text inline, in a text source's `data` or a
content source's `content`, and that text reached the provider unredacted. Collect
both, plus the block's `title` and `context`; base64, URL and file sources pass
untouched.
* fix(guardrails): redact llm_shield_proxy extra_body overrides
LiteLLM merges extra_body over the transformed request just before sending, so text
placed there (input, messages, system, ...) replaced the redacted field on the wire.
Walk extra_body with the same collectors as the request, keeping the caller /
application split.
* test(guardrails): import InMemoryCache directly in the llm_shield_proxy cache test
litellm keeps a deprecated module-level `caching` bool, so `litellm.caching.caching`
resolves to that bool once an earlier test in the same worker has set it, and the test
failed with AttributeError depending on test order.
* fix(guardrails): restore llm_shield_proxy replies for model-level use outside the proxy
|
||
|
|
2d82915084
|
fix(mcp): bind OAuth clients to their upstream issuer (#37777)
* feat(mcp): advertise the SDK's latest spec revision and validate the RFC 9207 iss MCPSpecVersion stopped at 2025-06-18 while the pinned SDK negotiates 2025-11-25, and the version LiteLLM puts on its own outbound initialize was a hardcoded historical member. Add the missing revision, name the highest revision we speak once, and pin it to the SDK's LATEST_PROTOCOL_VERSION with a test so the two cannot drift apart silently. /authorize now seals the issuer it sent the user to into the OAuth state, and /callback holds the authorization response's RFC 9207 iss against it, refusing to forward a code that came back from an authorization server we never sent the user to. An absent iss, an unanchored server row and a state minted before the seal all keep their current behavior. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep params, query and fragment significant in issuer comparison The shared canonicalizer drops all three, so two issuers differing only outside the path compared equal and a response from another tenant's authorization server would have continued through the flow. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): refresh generated API snapshots Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): cover OAuth client isolation and lint checks * fix(mcp): preserve registered clients in the existing save payload * test(mcp): cover optional OAuth registration metadata * fix(mcp): preserve compatible OAuth registrations across edits * fix(mcp): retain OAuth state through pending authorization * fix(mcp): guard pending OAuth at form submission * fix(mcp): discard canceled OAuth edit snapshots * test(mcp): preserve complete OAuth registration assertions * refactor(mcp): construct OAuth credential updates without mutation * fix(mcp): simplify issuer binding and reject unverifiable callbacks * fix(mcp): preserve replacement clients and pending redirect bindings * fix(mcp): retain clients with replacement authentication methods * fix(mcp): preserve cached clients and pin manual OAuth issuers --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
2e81db03b9
|
fix(auto-router): align preview and serving default resolution (#44436)
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com> |
||
|
|
3286782dea
|
fix(caching): key response cache by router model group in litellm_metadata (#44542)
* test: align integration fixtures with current behavior Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): wait for the spend flush before asserting its trace placement Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(caching): key response-cache entries by model group from litellm_metadata /v1/responses routes through the router with model_group in litellm_metadata, which the cache key ignored, so identical requests to different model groups sharing one underlying model hit each other's cached responses Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
eeb192d4a7
|
chore!: retire the integrated ROI calculator (#44477)
* chore!: retire the integrated ROI calculator * chore(ui): remove unused ROI demo notice |
||
|
|
902736bfe7
|
ci: move Postgres, MCP and Redis suites to CircleCI integration (#44453)
* ci: move Postgres, MCP and Redis suites to CircleCI integration * ci: throwaway, drop tests/proxy_behavior from its CircleCI job to show assert-ci-coverage fails * ci: revert throwaway assert-ci-coverage check * ci: keep the e2e helpers the gate tests still use * ci: move the roi-database Postgres shard to CircleCI integration * ci: run redis-compat without CircleCI's Azure and cassette env, cover postgres_suite test_path * ci: match the GitHub env for the moved Postgres and Redis jobs * ci: unset provider keys in the CircleCI MCP job and drop unused e2e-stack helpers --------- Co-authored-by: yuneng <yuneng@berri.ai> |
||
|
|
29b4f20572
|
fix(azure): set gpt-4o-transcribe retirement date from the retirement schedule (#44585)
Price-Sync: litellm-providers Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com> |
||
|
|
21881c5711
|
refactor(ui): extract shared timeline and time-range controls (#44584)
Move the timeline renderer to shared/timeline/Timeline taking buckets, a selected window, and callbacks as props, and TimeRangeControls to shared/timeline. Lens keeps bucketRuns as the adapter that converts loaded traces into buckets, preserving behavior without the histogram endpoint. Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
fdc0e7dc2b
|
refactor(rust): track ClickHouse migrations in a checksummed ledger (#44580)
* refactor(rust): track ClickHouse migrations in a checksummed ledger Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(clickhouse): harden migration execution and reuse retention SQL * refactor(rust): track ClickHouse migrations in a checksummed ledger Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs(rust): keep ClickHouse retention TTLs in the current policy list Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
adb59a7af8
|
fix(ui): remove Top models by task card from Model Leaderboard (#44502)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> |
||
|
|
9cc15e9320
|
feat(ui): add persistent columns and loading skeletons to Lens runs (#44579)
Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
02f61c9c42
|
chore(cost-map): add azure retirement dates for gpt-4o-realtime-preview-2024-10-01 and jamba-instruct (#44567)
Price-Sync: litellm-providers Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com> |
||
|
|
254c2f6b3b
|
refactor(mcp): drop Sequence/list return-type mismatch and collapse record_listed_tools wrapper (#44556)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
7a7d27c550
|
fix(guardrails): run the end-of-stream post_call scan when the client disconnects mid-stream (#43839)
* fix(guardrails): run end-of-stream post_call scan when the client disconnects mid-stream Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): close the guardrail stream chain in async_data_generator on client disconnect Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): leave the raw upstream response to the shielded finalizer on client disconnect Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): keep disconnect cleanup going when a streaming callback cleanup raises Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): assert the refund through a recorder instead of the mock Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): inspect tool calls released before a disconnect under incremental_diff and record a failed scan marker The incremental_diff transform stream now scans tool calls it already released when the client disconnects, and a disconnect scan whose translation raises after the guardrail recorded success also records guardrail_failed_to_respond Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): pin that a text-only disconnect scan is not handed a tool_calls finish Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): cover disconnect scans on every streaming endpoint and client, plus outage, worker-kill and cache-hit cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): prove the cache-hit twin is served from the cache Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): type the disconnect-close streams so basedpyright stops reporting unknown arguments * fix(guardrails): give the guardrail metadata cast a reason so the type discipline gate accepts it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): scan released Messages and Responses tool calls on disconnect Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(guardrails): format the disconnect scan unit tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(guardrails): drop mutable-ok markers that no longer suppress a rule Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): scan released Responses output after a finished item and end only in-flight Chat choices Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): type the disconnect scan test helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): type request_data in the disconnect scan helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): type request_data in the disconnect scan test doubles Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): pin that chat streams with no tool call in flight end as released Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): scan only released chunks on disconnect and skip it once a block owns the verdict The disconnect scan now uses the chunks actually yielded to the client, copies them before scanning, skips when a mid-stream block or HTTP error already settled the verdict, and the iterator wrapper only closes hooks that are async generators so plain async iterator hooks keep working Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): pin that a delivered guardrail error or final chunk settles the disconnect verdict Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): close any hook iterator that exposes aclose when the stream ends early Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): accept a synchronous aclose on custom streaming hook iterators Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): swallow callback aclose errors at end of stream A custom callback whose async_post_call_streaming_iterator_hook returns a non-generator async iterator with a raising aclose() failed the finished stream: content plus usage reached the client and then the stream surfaced an error SSE with no [DONE], or aborted a post_call pipeline's buffering loop into a 500 with an empty body. Wrap the aclose invocation in _wrap_streaming_iterator_with_enrichment in try/except and log a warning naming the callback and the cleanup error, matching close_guarded_stream and _close_guarded_layers. Iteration-time hook exceptions still propagate. * fix(proxy): log only the error type when a callback aclose raises Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
1238cfe90f
|
feat(proxy): add LITELLM_FIPS_MODE startup gate with provider assertion and loud password migration failure (#42700)
* feat(proxy): LITELLM_FIPS_MODE startup gate with provider assertion, TLS verify guard and loud password migration Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): match ssl_verify off detection to runtime str_to_bool semantics Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): drop tautological fips probe test and satisfy CodeQL return checks Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): inject the fake Prisma client through the module boundary instead of patching _setup_prisma_client Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2e1a98f521
|
fix(logging): deduplicate streaming failure callbacks (#44442)
* update logic that marks a logging callback as complete * test(logging): cover streaming failure dedupe in mark_logging_complete Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): cover streaming failure dedupe in S3 and DataDog Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(logging): keep has_run_logging as a deprecated alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): audit streaming failure dedupe across surfaces, fallbacks and sink outage Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): assert anthropic upstream path in streaming failure audit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): count only provider posts in streaming audit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): assert the sink outage rejects uploads in burst audit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): configure datadog retries with router override Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
8f6546df9f
|
refactor(ui): route dashboard URL state through nuqs parsers (#44537)
* docs(ui): add url-state agent skill Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(ui): migrate dashboard URL state to nuqs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(ui): cover chat URL id sync with the real chat shell provider Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
be4481779e
|
chore(cost-map): sync openrouter prices from the models API (#44533)
Some checks failed
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-infra-root (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
Postgres Tests / proxy-behavior (push) Has been cancelled
Postgres Tests / roi-database (push) Has been cancelled
Postgres Tests / proxy-security (push) Has been cancelled
Postgres Tests / schema-migration (push) Has been cancelled
Price-Sync: litellm-providers Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com> |
||
|
|
5a5fd92d39
|
fix(dashboard): route dev API calls on Accept and fail fast on non-JSON 2xx (#44528)
* fix(dashboard): route dev API calls on Accept and fail fast on non-JSON 2xx The next dev rewrite that sends API calls to the proxy keyed on Content-Type: application/json, which openapi-fetch rightly omits on a bodyless GET, so GET /lens fell through to the Lens page and returned HTML. Route on Accept: application/json instead, send it from both HTTP clients, turn a non-JSON 2xx into a non-retryable ApiError in the typed client, and stop react-query from retrying ApiError below 500. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(dashboard): accept a JSON body served without a JSON content type Test fakes and some servers hand back JSON as text/plain, so the typed client only rejects a 2xx whose body does not parse as JSON. The system_one request test expects the Accept header the legacy client now sends. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
1d52985d03
|
fix(ci): repair the security sweep and Lens billing integration tests for ROI, JWKS, and release identity changes (#44529)
* fix(ci): keep the security sweep off the observed ROI GitHub route and give Lens integration tests a release identity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): expect the credential JWKS export to 404 for non-federation credentials in the security sweep Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
635085ac14
|
fix(azure): set gpt-realtime-2.1 retirement dates from the retirement schedule (#44526)
Price-Sync: litellm-providers Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com> Co-authored-by: kerry-berri <kerry@berri.ai> |
||
|
|
97fc72e5bb
|
fix(ci): send the Lens preview body as selection in the tracing endpoint tests (#44525)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
461a58c40a
|
refactor(lens): storage-independent trace reads, shared keyset pager, typed read failures (#44422)
* feat(lens): own trace reads behind a cached TraceStore port Move storage-independent trace reads into litellm-traces-cache behind a TraceStore port that ClickHouse implements. One keyset pager drives the span, list span and spend reads, and a run list batch reads spend once. Trace opens, pages and list summaries share one resolved read per trace in an in-process cache with single-flight loading. Live traces and reads with unknown spend expire after 5s, quiet traces after 10 minutes, failed reads are never cached, and the accepted list page size is remembered per scope. Trace read failures map to their own status and code (400, 409, 413, 503 with Retry-After), and the trace drawer retries temporary failures while offering only a refresh for changed or oversized traces. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * perf(lens): seed large profiles with server-side copies and long sessions Replay one copy through the proxy, then copy it inside ClickHouse and PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so every copy keeps its own spend. Add three long single-trace sessions for drawer paging and the oversized read path Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chores * style(lens): float the investigation setup badge on the tab edge Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(lens): restructure the trace drawer and polish its layout Split the 547-line TraceDrawer into run/, tree/, span/, content/ and conversation/ modules. Step rows now sit on one line with colored span family tiles, and the per-row timing bar moved into an optional Waterfall layout with a time axis. The steps and details panes are separated by the shadcn Resizable handle, with the split remembered per orientation. Span payloads go through one pure classifier (payloadView) that picks messages, a tool result, a nested field tree or text. JSON-encoded field values unfold into a tree, prose renders as markdown, repr and tracebacks stay monospace, and every section offers a Raw view. LangChain's serialized messages now render as conversation cards. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * typesafety * wip * fmt --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
e1d16f51d1
|
refactor(ui): share CopyButton between Lens traces and logs (#44513)
Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
1459e00430
|
refactor(ui): rename view_logs to logs and split request, audit and detail (#44505)
Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
05f1c73a3c
|
refactor(ui): move TraceView into components/lens/traces (#44501)
Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
6532dcb73b
|
test(ci): pin the ROI estimator flag and the prompt-cache counter in two drifted tests (#44499)
Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
a2bf67a037
|
test(mcp): keep the SSO assertion round trip from matching its refresh token inside random ciphertext (#44494)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
b4fcc5c1bc
|
fix(otel): read the registered v2 logger without importing the proxy (#44485)
* fix(otel): read the registered v2 logger without importing the proxy phase_span, which the router enters on every deployment pick since #44150, looked up the proxy's OTel logger by importing litellm.proxy.proxy_server. In an SDK process that import loads the whole proxy synchronously on the caller's event loop during its first request, which stalled litellm_router_unit_testing past its 5s wait. Read the module from sys.modules instead: when the proxy was never imported it has no registered logger. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): pin the no-proxy-import guarantee in-process Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
bb4f7211d7
|
refactor: clean up fresh tech debt from 2026-10-03 (#44484)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0b74ae9c5c
|
feat(mcp): hand listed-tool description and input schema to pre-call hooks per caller (#41162)
* feat(mcp): hand listed-tool metadata to pre-call hooks with per-caller catalog identity Track the tools each MCP server listed per caller identity so pre_mcp_call and during_mcp_call hooks receive the tool description and input schema the client saw. Servers with no caller-dependent inputs share one slot; user identity, forwarded headers, stdio env, relayed bearers, and server-specific auth get their own. Local registry and OpenAPI paths pass the registered metadata and admin description overrides. The Agent 365 guardrail reads the new fields into its evaluate payload. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mcp): drop the listed-tools empty sentinel and routine test docstrings Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): mark the listed-tools cache digest as a non-security hash Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key the listed-tools cache by the OBO subject token token_exchange servers list upstream with the caller's own Entra bearer, so two callers on one LiteLLM key with different subjects were sharing a catalog slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): resolve the BYOK credential before keying the listed-tools slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): drop the OAuth discovery cache when a server definition changes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mcp): drop a diff-narrating comment Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): never validate a supplied header on the tools/list BYOK path The pre-listing resolver ran the tool-call byok_auth_required check even when the caller already supplied x-mcp-auth, and it ran outside the per-server error boundary, so a single deprecated-header caller dropped the server from the aggregate list. Listing now returns a supplied header unchanged and falls back to the stored credential without raising Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): assert the BYOK listing lands in the caller's listed-tool slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): cover the deprecated string x-mcp-auth header on a BYOK tools/list Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key the per-caller listed-tool slot by the hashed token Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key discovery cache by the hashed token Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates Discovery-list cache identity now uses the hashed token instead of the raw api_key and treats MCPJWTSigner-signed servers as per caller. Server definition changes also drop the cached upstream OAuth metadata. OpenAPI listings look tools up under the normalized registry prefix with the separator, so an overlapping sibling prefix no longer leaks into the list. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mcp): derive the listed-tool caller identity from the discovery cache key Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys An upstream metadata fetch that started before a server edit could store its stale reply after invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch only stores when the generation it captured before I/O is unchanged. The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so both go back to the merge-base behavior. Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep OAuth metadata generations only while a fetch is in flight Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): return one masked text per scanned string in the selected-guardrail REST test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): fold the signed caller into the discovery digest instead of a second key hash Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): satisfy type discipline gate on listed-tool identity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): hand tools/call hooks the exact catalog entry tools/list served get_listed_tool re-applied the admin description override on top of the cached listing, so a guardrail-masked description was restored to its original wording at call time, and the OpenAPI / local-registry call path built its metadata from the registry instead of the guarded caller catalog. Both paths now return the cached entry as served, falling back to the registry only when no listing was recorded Adds tests/integration/mcp/test_mcp_listed_tool_metadata.py (red on the prior head for the two regressions, red on the merge base for the feature, green on this head) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key OpenAPI listed-tool entries per caller so tools/call reads its own guarded listing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): align listed-tool slot tests with per-caller keying Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep oauth2 listing on the minted or signed credential, not the stored BYOK secret The listing helper that keys the per-caller catalog by the stored BYOK credential also handed that credential to the upstream client, which on an oauth2 server short-circuited the client_credentials mint and the MCPJWTSigner gate. Split the two: the catalog identity keeps the stored credential so tools/call finds the caller's slot, while an oauth2 server's tools/list sends only the per-request header, letting the M2M mint or signed JWT proceed as on main Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): oauth2 BYOK listing sends the minted token, not the stored secret, through the real proxy Integration cell for the listing fix: a client_credentials BYOK server with a stored user credential, one tools/list as that user, the peer must see a live minted bearer and one /token mint. Red at the pre-fix tip (zero mints, stored secret upstream), green at the fixed head and at the merge base Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep the stored BYOK credential for catalog identity only on tools/list Listing used the resolved stored credential both to key the caller's catalog slot and as the upstream transport header, so REST api_key and bearer_token listings sent the user's secret instead of the server's static token and the MCPJWTSigner gate went quiet. The upstream client and the signer gate now read the caller-supplied mcp_auth_header for every auth type, exactly as before the catalog existed, and the stored credential only names the slot tools/call reads Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): type the listed-tool metadata read from pre-call kwargs for the basedpyright gate Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): hand never-listed tools/call hooks name and arguments only The local-registry call path fell back to the registry entry with the admin description override when no tools/list had been recorded for the caller, so a pre_mcp_call guardrail scanned a description the caller was never served and blocked OpenAPI calls that passed before, and base's own selected-guardrail REST test failed on the two-text redaction. _registered_tool_metadata now returns the listed entry or None, so a tools/call with no prior listing sends name and arguments only as promised, and that REST test double goes back to its base shape * fix(mcp): keep during_mcp_call hooks on name and arguments only call_tool handed the caller's listed entry to the during-hook task as well, so during_mcp_call guardrails scanned the description line and schema leaves of any listed tool after the upstream call had already run, blocking calls that passed before whenever the policy matched the description, returned a fixed-length texts list, or hit the depth guard on a deep schema. The listed entry is only disclosed for pre_mcp_call, so the during task no longer receives it and its request object carries no description or schema, as before * fix(mcp): key the BYOK catalog slot by the client's header, not the stored credential tools/list resolved the stored BYOK credential to pick the caller's catalog slot, which read the credential store before the classified try block. With Postgres down and a cold per-worker cache that made every REST tools/list on an is_byok server fail with tools=[] and no upstream call, and the read seeded the per-worker cache (including a negative entry), so a tools/call on another worker after a store, rotate or revoke on this one kept using the stale value. The slot is now keyed by what the client supplied plus the caller's hashed key, on both sides. _get_tools_from_server and call_tool take a keyword-only catalog_auth_header that defaults to mcp_auth_header as received (the default is the builtin Ellipsis so it survives a module reload). The /mcp fan-out and execute_mcp_tool, which swap the resolved credential into mcp_auth_header, pass the client's value explicitly. What goes upstream is unchanged. _byok_catalog_auth_header is gone. * fix(mcp): drop a listed catalog recorded across a server save _record_listed_tools ran after the awaited upstream fetch, so a PUT /v1/mcp/server that landed mid-fetch had its invalidation undone when the fetch completed: hooks then saw the pre-save description next to the post-save definition until the next listing, instead of name and arguments only. The manager now keeps a per-server listed-tools generation, bumped by _invalidate_server_definition_caches. _get_tools_from_server reads it before the fetch and _record_listed_tools skips the write when it moved; the next listing records normally. * fix(mcp): drop the catalog again once a saved OpenAPI server's registry is rebuilt add_server and update_server publish the saved definition before the OpenAPI registry entries are rebuilt from the spec, so a listing recorded during that fetch held the pre-save entries under the new generation. The generation is bumped a second time after the registry refresh. The during-hook task no longer accepts a listed entry, the one-line wrapper over get_listed_tool is inlined at its two call sites, and the per-server generation map is a plain dict. * fix(mcp): keep discovery and OAuth metadata caches across an OpenAPI spec re-read add_server and update_server ran the full server-definition invalidation a second time after the awaited OpenAPI spec fetch, which also dropped the prompts/resources/templates discovery entries and the OAuth protected-resource metadata filled under the already-published definition, so the next request went upstream again. Only the listed-tool catalog recorded during the fetch holds pre-save entries, so the post-fetch pass now drops just that catalog and bumps its generation via the new _drop_listed_tools helper, which the full invalidation also calls. * fix(mcp): look a called tool up in the listed catalog by its bare name only get_listed_tool stripped the server prefix a second time when the exact name was absent from the caller's listing, so a never-listed upstream tool whose bare name starts with the server prefix resolved to the listed sibling and that sibling's description and input schema reached the pre-call hooks for a call to a different tool. Every caller already passes the once-stripped bare name, so the lookup is now exact. Tests that looked the catalog up by a prefixed name now use the bare name the callers pass; two new tests pin the never-listed sibling case at the manager and at the tools/call path. * fix(mcp): record a listed-tool catalog only for a listing the caller is served _get_tools_from_server now records the catalog into the caller's listed-tools slot only when asked (record_listing=True), which the served listings pass: the /mcp and Responses API tools/list handlers via _get_tools_from_mcp_servers, MCPServerManager.list_tools, and the REST listing via _list_server_tools. Four internal listings stop recording, so a later tools/call hands pre_mcp_call hooks name and arguments only, as on main: - _list_tools_before_first_call, the implicit listing inside tools/call when this worker does not yet expose the tool - fetch_pinnable_tool_catalog, the admin pin snapshot listed without the catalog guard and without description overrides - _initialize_tool_name_to_mcp_server_name_mapping, the startup fill - get_tools_for_server, used by the semantic tool filter _create_prefixed_tools returns to its tool-name mapping job only; the record follows it in _get_tools_from_server. * fix(mcp): opt every listing out of catalog recording unless it is served The aggregate listing and _list_mcp_tools now default to record_listing=False, so a catalog fetched inside a tools/call no longer fills the caller's listed-tools slot. The /mcp/proxy meta-tools (call_tool, search_tools, get_tool_schema) and the tool-search virtual tool stop recording: /mcp/proxy serves only the meta-tools and the search serves only its hits, so a later pre_mcp_call hook was reading a description the caller never listed. The tools/list handler, the Responses MCP handler and the /v1/mcp/tools management listing opt in with record_listing=True, since each serves the catalog to the caller. * fix(mcp): key the listed-tool slot by the caller's admission identity and forwarded bearer The slot a tools/list records for a later tools/call was keyed by (user_id, api_key) only, so every team-only JWT caller shared one slot and one JWT user acting in two teams shared a slot; a tools/call then handed pre_mcp_call hooks a description another caller was served. The slot is now keyed by the hashed key, user, team and organization, plus the admission credential of a caller admitted with neither a key nor a user. The caller bearer split the slot only on client-forwarded-token and token-exchange servers; a legacy delegated oauth2 server (delegate_auth_to_upstream without client credentials) also forwards it upstream and served a different catalog per bearer into one slot. The bearer now splits the slot on every server whose egress forwards it (_consumes_caller_authorization) or exchanges it. * fix(mcp): record only tools served by the bridge Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep bridge tool metadata request-local * refactor(mcp): centralize listed catalog recording guard * fix(mcp): preserve base TPM reservations for listed tool calls * fix(mcp): preserve project token reservations for listed calls * fix(mcp): record served catalogs and preserve call message bytes * test(mcp): align listing expectation with deferred recording * test(mcp): audit listed metadata across callers and bridge lifecycles * test(mcp): preserve guardrail fixture worker affinity * refactor(mcp): expose listed catalog recording API * feat(mcp): pass served_tools through the anthropic messages bridge /v1/messages auto-execution now hands this request's resolved tool definitions to _execute_tool_calls, matching the Responses and chat completions bridges: the pre_mcp_call hook receives the description and input schema the model was shown for that call. Request-local only; the shared listed-tools catalog is untouched. * test(mcp): pin served_tools handoff on the anthropic messages bridge Mirrors the credentials-forwarding test: the request's resolved tool definitions must reach _execute_tool_calls under served_tools so pre_mcp_call hooks judge the call on the description and input schema the model was shown. Fails without the previous commit's one-liner. * style(mcp): sort the local import block ruff flagged * fix(mcp): keep the admin include_disabled_tools view off the listed-tools catalog GET /mcp-rest/tools/list?include_disabled_tools=true is the admin-only configuration view: apply_tool_filters is False, so it serves the full server catalog. Recording that response into the caller's listed-tools slot warmed tools/call metadata no runtime listing ever served, breaking the only-a-served-listing-records invariant (Bugbot). The record is now gated on apply_tool_filters; disabled tools stay unreachable (the call-time allowlist 403 fires before hooks), so the observable fix is the slot no longer warming from a settings view. Verified live: the new test fails on the unfixed head and passes here, and the rest of the listed-tool-metadata suite is unchanged. --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
9a5e828310
|
feat(ui): inline Lens settings tab and investigation editor (#44479)
* feat(ui): move worker status into the Lens notch and New investigation into the list toolbar Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): make the Lens notch entry a settings gear that houses the worker section Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): replace the Lens worker modal with an inline Settings tab Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): section the Lens settings tab with tracing status and worker cards Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): let Lens settings sections span the full card width Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): replace the investigation setup modal with an inline side-by-side editor New, edit, and duplicate now take over the Investigations tab body: matching activity on the left, every setting on the right, with no wizard steps or modal Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): step the inline investigation setup vertically with traces alongside Setup now sits on the left as three progressive steps (activity, criteria, run) that collapse to a summary once done and reopen on click. Matching activity stays on the right for every step. The editor gets a back control and the Investigations notch shows a New, Editing, or Duplicate badge while the editor is open Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): page the matching activity preview with useInfiniteQuery as it scrolls Replace the Previous/Next offset buttons with the same infinite query and near-tail prefetch the traces list uses, so the preview keeps loaded runs and its title while the next page arrives Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): make the investigation step field map exhaustive over the form schema Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): always show the Settings tab label in the Lens mode switch Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): collapse Lens worker cards into compact status rows Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): keep the Lens Settings tab icon-only in every state Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * wip * fix(ui): tick the Lens worker health dot so an expired heartbeat goes stale Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): regroup Lens settings, model and api layers Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): dedupe Lens formatting helpers Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): poll Lens once, drive the interval from data, settle mutations before invalidating Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): let Lens leaves fetch their own data Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): bind drawer and trace shortcuts through react-hotkeys-hook One useShortcut hook replaces the three hand-rolled keydown listeners in SidePanel, the trace step tree and the log drawer, with a layer option deciding which keys a pane claims from the panel around it. The span tree footer now renders ShortcutHints from what is actually bound instead of hand-typed kbd text. * refactor(ui): extract Inspector from SidePanel Inspector.Root owns the open item, J/K stepping, Escape and full screen; Inspector.Row marks a list entry with aria-selected and data-state and toggles it on click or Enter/Space; Inspector.Panel is the resizable side panel with the exit animation and click-outside rules. The runs table and section compose these parts directly, so RunDrawer and the SidePanel prop bag go away. * refactor(ui): model the Lens worker screen as a tagged union and slot in its ready action workerScreen() decides between list, form and install from the worker rows, the registration result and the edit target, so WorkerSettings switches on one value and each card owns its own copy. The post-install CTA is now a ReactNode slot that LensWorkspace fills instead of an onReady callback threaded through LensSettings and WorkerSettings. Clipboard copy state lives in WorkerInstall as mutations, and the styled settings leaves export Props types, set data-slot and accept native element props. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): compose the Lens setup stepper from SetupStep children Each step's heading, summary and fields now live together in one SetupStep instead of four parallel structures keyed by index, and the last-step spacing comes from CSS rather than a passed index. The mode prop is now required since InvestigationsView always passes it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): keep the Lens Settings panel mounted so a pending worker install survives tab switches The settings TabsContent unmounted WorkerSettings whenever another tab was active, dropping the one-time worker token shown during install. The panel now uses keepMounted, and the workspace test registers a worker, switches tabs and back, then follows the connected worker into the first investigation through the slotted CTA. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ui): cover the Lens worker install waiting-to-connected transition Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): give the Lens activity preview a grouped contract and a structural debounce useMatchingActivity now owns its return types (scope options, preview status, page and optional manual selection) instead of borrowing them from the components it feeds, and the preview takes those groups plus the section attributes. The clear-selection action moves into the preview footer, RunList becomes a RunRow leaf, and ScopeFields drops the unused nameField and id props now that MetadataFilters calls useId itself. The preview scope settles through a hashKey-based useDebouncedValue instead of JSON round-tripping into state, and the loading title follows isPlaceholderData since the query keeps previous data. RunFields and AnalysisModelField take the analysis models and the model gate as two objects instead of seven flat props. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): select the Lens investigations screen with a pure tagged union investigationScreen maps the list query and the route to one of loading, failed, welcome, list, detail, setup or missing, so the view can switch instead of juggling mutually exclusive booleans. The status model gains activeJob, carries connected inside Readiness and folds the activity probe into one ActivityCheck value for the welcome page Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ui): cover the Lens preview footer clear action and the preview debounce Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): let Lens investigation leaves own their URL slice and express intent InvestigationsView renders the screen union and owns every write through useInvestigationActions, so leaves receive on* handlers instead of the API writer. useInvestigationResults becomes useRunSnapshot; FindingsTab, HistoryTab, RunPicker and RequestEvidenceSheet read their own nuqs slice and run their own queries. The finding sheet becomes an Inspector side panel (FindingDetails) keyed per finding, with the trace and request evidence sheets grouped in EvidenceSheets. WatchAllBanner owns its mutation, the welcome page takes the readiness and activity values, the progress sampler records on the wall clock outside render, and run history invalidates when the list reports a scheduler-started job Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ui): cover Lens investigation intents, pause, cancel, history refresh and the finding panel Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): keep the inspector open behind sheet overlays and mount HotkeysProvider from a client wrapper * feat(ui): open Lens runs and evidence in the inspector panel A quote's original trace step or logged request now stacks inside the finding panel behind a back link, keeping the finding and its feedback draft mounted. The detail Runs tab and the setup activity preview open runs in the same panel with J/K stepping, so the TraceSheet and RequestEvidenceSheet modals are gone. Picking a different finding or run clears any stacked evidence from the URL. * feat(ui): open Lens investigations in the inspector panel beside the list The investigations list stays on screen and a row opens its investigation in the side panel, so J/K walk investigations and their open findings in display order and the selected row carries the same highlight as runs. The panel body is the former detail page; findings and runs opened inside it nest their own inspector, which claims the keys from the one around it while open. Opening an investigation and peeking at a finding now replace each other in the URL. * fix(ui): run the Lens notch border along the tab pill and flag only a disconnected worker Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): show the shortcut hints in every inspector panel Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): extract the Lens dot field into composable DotFieldRoot and DotFieldCanvas Move the dot grid model and canvas painter out of TracesTimeline into components/lens/dotField so other Lens surfaces can reuse it. The root owns layout and context; overlays compose as children. agoLabel moves to lens/model/format and the unused columnTop helper is dropped. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): drop the manual refresh button from the runs time controls Live polls and range changes refetch, so the button only cleared the zoom, which Escape already does Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ui): extract the Lens run search into a composable SearchBox primitive Move the query parser, glob matcher and autocomplete out of runSearch into components/lens/search, generic over a QueryLanguage (field specs plus what free text searches). SearchBox.Root owns the ProseMirror state, menu and keyboard; SearchBox.Input and SearchBox.Suggestions compose under it. Clause highlighting becomes a ProseMirror plugin built from the language. RunSearch now only declares the run fields and composes the parts, so the investigations tab can define its own language and reuse the same box. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(ui): label the runs range by preset while Live and pin it once paused Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ui): split the Lens search language from where its data lives QueryLanguage is now pure vocabulary (keys, groups, icons). Reading fields off loaded items moves to a ClientIndex consumed by a separate evaluator, and value suggestions come from an injectable ValueSource, so a server-backed runs list can plug in a facet lookup while the investigations tab keeps filtering in memory. The parsed query serializes to a typed SearchQuery (text terms plus eq/neq/glob/nglob filters) that the client evaluator consumes today and a server can consume later. The suggestion menu shows a loading row while a source is still answering. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * style(ui): format the Lens SearchBox and its test with the project prettier config Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): mirror the Lens run filters as trace SQL with a copyable curl in the search footer The suggestions footer gains a slot, and SearchBox.ApiHint fills it with the API equivalent of the typed query: a dialect chip, a one-line preview and a Copy as curl button. The runs box translates each filter to a predicate over the agent_traces_by_key rollup, bounded to the range the list shows, and copies a POST to /v1/traces/query. Any other list can plug its own translate into the same part. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): show the Lens introduction as a first-visit dialog with a typed don't-show-again The guided setup no longer replaces the Lens tabs. It opens in a dialog on the first visit of a session or from ?setup=lens, with a close and a "Don't show this again" checkbox in its top-right corner. The header "Set up Lens" button is gone. Dismissal state lives in a new schema-validated web storage helper (src/lib/storage.ts) that reads through useSyncExternalStore, so server renders see the fallback and other tabs stay in sync; only Lens uses it for now. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): keep only Copy as curl in the Lens run search footer Drop the SQL chip and predicate preview; SearchBox.ApiHint becomes SearchBox.CopyCommand, which takes the command for the current query. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(ui): align the Lens investigations list with the traces list Use the shared query SearchBox with investigation fields (name, agent, status, schedule), match the traces toolbar, and drop the count footer and inner padding. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ui): let Lens settings bring back the introduction after don't show again Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ui): share one InspectorTable between the traces list and the Lens investigations tree Compose TanStack Table, react-virtual and the shadcn table cells into InspectorTable parts (Root, Grid, Header, Body, Row, Indent). Investigations get findings as real sub-rows with TanStack expansion instead of a hand-rolled flattener. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ui): break Lens import cycles and move shared pieces out of lens Search and the dot field go to components/shared, run search and the preview button go to view_logs where they are consumed. Lens api, services and demo live under data/, all URL state in route.ts, storage keys in storage.ts, and the session frame styles become cva variants. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ui): one Lens readiness source and one onboarding flow Readiness is computed once in model/readiness and read through useLensReadiness, replacing useLensSetup, status.readiness and the welcome screen's own checks. The Investigations welcome now renders the same onboarding steps as the introduction dialog, with permissions and actions coming from an OnboardingProvider instead of props passed down four levels. StepIndicator and StateMessage are shared lens components, and the step panels are labelled accordion regions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ui): read the Lens token from services and use semantic status colors Lens services carry the access token they were built for, so trace evidence, readiness and onboarding read it from context instead of a prop threaded through six components. List and history invalidation lives in one data hook. Status colors use the success, warning and destructive tokens, and template-literal class names go through cn. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ui): split Lens demo fixtures from the fake demo APIs Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ui): show Lens check history as a dot timeline Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chores * fix(ui): clear stale Lens evidence on run change and keep read-only users off Settings Also names inline option objects to bring local/no-large-inline-object-arg back under budget. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
984b4134be
|
fix(proxy-extras): log v1 migration failures at ERROR so LITELLM_LOG=ERROR shows them (#44202)
* fix(proxy-extras): log v1 migration failures at ERROR so LITELLM_LOG=ERROR shows them Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy-extras): reuse litellm secret redaction and mask configured DB passwords exactly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy-extras): name the password alternation in _redact_credentials Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(proxy-extras): wrap v1 migration ERROR lines at 120 columns Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): integration cells for v1 migration ERROR logging and password redaction Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): integration cells for component DB env vars, JSON logs, migration Job and v2 resolver Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): give every v1 migration integration cell the 900s timeout Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): bind the recovery forwarder before migrating so the retry cannot race it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): drop the slow P3005 integration cell and bound migration subprocesses Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy-extras): keep command repr and tolerate non-sequence cmd in migration error logs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): double the subprocess boundary in the cmd=None retry test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: jesus <jesus@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
cf22deb96a
|
test(ci): fix six CircleCI test regressions on main (#44429)
* test(ci): fix four CircleCI regressions on main Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(ci): stop reloading auth_checks in unit tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(constants): cover CLI JWT expiry env parsing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(gateway): restore proxy lifespan after importing gateway.main in launch tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2936307661
|
feat(lens): guide setup through the first investigation (#44475)
* feat(lens): guide setup through the first investigation * fix(lens): restore the onboarding reference visuals * fix(lens): compact onboarding and animate gateway flow * feat(lens): refine onboarding motion and linked examples * feat(lens): turn the LED swarm into organized dot groups * fix(lens): make the LED dot flow visibly animate * feat(lens): refine the swarm scale palette and motion * fix(lens): preserve investigations during activity refresh errors |
||
|
|
b61376a99e
|
feat(sdk): add run_tool_loop and arun_tool_loop helpers (#44381)
* feat(sdk): add run_tool_loop and arun_tool_loop helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(tests): allow-list bounded tool-loop recursion in recursive detector Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(sdk): harden run_tool_loop per review Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(deps): keep uv.lock at revision 3 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(tests): authorize typing-extensions PSF-2.0 license Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
5d513d8053
|
fix(lens): align source setup with available worker images (#44476)
* fix(lens): document working preview setup before coordinated releases * docs(lens): keep current setup guidance factual and preserve Helm options * fix(lens): support explicit worker images on Compose 2 |
||
|
|
5ddcc45b3a
|
feat(ui): share trace drawer as a closable SidePanel and polish Lens (#44473)
Some checks failed
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-infra-root (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* feat(ui): extract trace drawer into a shared SidePanel that closes on outside press
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): center Lens mode switch in a notch joined to the content card
Larger Traces/Investigations switch, a subtle dot when an investigation is running or queued, and a bigger Lens title with a docs link.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): stronger Lens frame border and header spacing
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): align Lens notch fillet with the notch border
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): centered Lens loading/error states and proxy JSON calls in dev
The dev server answered GET /lens with the Lens page HTML because the UI route shadowed the proxy fallback rewrite. JSON API requests now go to the proxy before page routes, and a non-JSON success body raises a readable ApiError.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(ui): theme-scale type and shape in trace views, aligned pane bars
Add local/no-arbitrary-design-value, scoped to TraceView and Lens, banning
arbitrary font size, tracking, leading, radius, border and CSS property
values. Map the Figma-export values onto the theme scale and replace hex
colors with info/destructive tokens.
Add PaneBar, a fixed-height bordered row, and build the step tree and span
detail headers from it so their borders line up across the split.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ui): restore AgentTracesSection emptied in
|
||
|
|
c80e6e2474
|
test(ui): wait for step search value to settle in TraceDrawer test (#44470)
* test(ui): wait for step search value to settle in TraceDrawer test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(ui): dispatch search typing and nav keys deterministically in TraceDrawer test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
62fb808d4b
|
feat: improve Lens dev seeding and live UI (#44468)
* feat: seed Lens dev with configurable load profiles * feat: run Lens UI live through the dev launcher * fix: verify Lens UI startup before seeding * fix(ui): render run timestamps on one compact line The agent runs table printed the long locale form with timezone, which wrapped to two lines per row. Use a fixed-width 24h form with milliseconds that matches the timeline axis, keep the long form in the hover title, and show the timezone once in the column header Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): use the root query client for the Lens demo Drop the demo's nested QueryClient. Cache keys are already partitioned by scope, and the root client now skips retries on 4xx ApiErrors, which covers the demo's not-in-demo and read-only rejections and live 4xx alike. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore: drop unused synthetic_spend reference from query help Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(traces): drop the deeplite fixtures and map every SDK to a logo The deeplite captures predate the example repository and carried synthetic spend rows, which leaked a fixture-only column into the query help SQL and pinned tests to its shape. Replace them with the SDK captures in the ClickHouse round trip and query API tests, and derive the seed tenant lookup from the capture metadata The runs table only knew the two Anthropic framework slugs. Register the slugs the normalizer emits for LangChain, LangGraph, Deep Agents, CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands, Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic agent glyph when a run has no known framework Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(ui): inject the Lens sample through data sources instead of demo checks Components no longer ask whether they are in the demo. The traces source carries live and handoff, navigation state comes from a URL or memory route, and the preview action comes from context instead of onDemo props Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ui): infinite scroll for the agent runs list Replace the Load more button with a sentinel that fetches the next cursor page as the list nears its end. Placeholder rows hold the tail while more runs exist, and a failed page stops auto-loading until Retry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ui): keep the whole Lens view in the URL Lens navigation now lives entirely in query params through nuqs: the sample session (demo=true, with a Demo data switch in the header), the open run (trace, trace_ref), the selected step, view and detail section (span, view, span_tab) and the list filters and range (q, agent, status, hours). Any Lens view is a shareable link and the back button walks runs RunView takes its selection injected: the drawer feeds it URL state and the investigations evidence sheet keeps a local one, so a finding's original run never writes step ids into the URL. Leaving the sample session clears every Lens key except the tab so sample ids never point at live data Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * style(traces): cargo fmt captures tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
8520626e7a
|
fix(lens): preserve approved worker digests and harden its image (#44467)
* fix(lens): pin worker dependencies and support approved image digests * fix(lens): include locked dependencies and release identity in build context * fix(lens): select the dev worker package for SHA-tagged charts |
||
|
|
a66adb4ff8
|
chore(cost-map): sync openrouter prices from the models API (#44466)
Price-Sync: litellm-providers Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com> |
||
|
|
9d16412341
|
fix(caching): count tool_call cache_control marks in the injection census (#43556)
* fix(caching): count tool_call cache_control marks in the injection census
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): remove cache census casts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): only skip injection on message or content marks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip injection on messages whose tool calls carry marks
Reverts
|
||
|
|
6370104c53
|
feat(enterprise): bundle LiteAdmin Slack with native gateway login (#44444)
* feat(enterprise): bundle LiteAdmin Slack worker with native gateway login * fix(enterprise): preserve gateway prefixes during native Slack linking * fix(enterprise): retain native Slack linking on the admin backend * fix(enterprise): reuse shared native Slack connection services |
||
|
|
4b67a2b845
|
feat(roi): default people and branch lists to matched accounts (#44465)
* feat(roi): show matched people by default in contributor lists * fix(roi): keep matched filter tabs readable on narrow screens * fix(roi): retain spend-only users and support older browsers |
||
|
|
f850b2c324
|
test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment (#44451)
* test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): compare every non-transport provider header and check for late provider requests at session end Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): name TranslationTestCase fields after litellm and provider sides and drop regressions Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): prefix checked TranslationTestCase fields with expected_ and name the fake reply mock_provider_response Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs(integration): name TranslationTestCase fields in the translation README Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): add a claude-opus-5-5 base case and deployment next to claude-sonnet-4-6 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): name translation cases <MODEL>_TEST_CASE and document the naming rule Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs(integration): move translation test rules into tests/integration/translation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0ed1c08f02
|
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one commit on top of litellm_internal_staging without the dashboard changes. Deployments on anthropic/ without a static api_key can exchange an OIDC workload assertion for a short-lived sk-ant-oat01 token through a shared RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file, an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment, per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation fields are server-owned: refused inline in request bodies and on POST /model/new, proxy-admin only on credentials, and the token exchange is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS adds a host. GET /credentials/{name}/jwks exports the public key set of a LiteLLM-signed credential for the Claude Console. The OpenAI federation trio from #39613 rides along on the backend side with the same server-owned handling. Fixes #28607 Resolves LIT-6107 Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com> * fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries The files handler enabled workload identity on batch-result downloads but never received the deployment's litellm_params, so a deployment authenticating through a named credential could only mint from process-wide env vars. It now threads litellm_params through to the auth header the way the batch retrieve path already does. LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme, so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the gateway. Entries are now parsed as network locations whether or not they carry a scheme. * fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle * test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries * fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly * fix(proxy): decrypt stored litellm_params before the WIF write gate * fix(proxy): hide WIF secret references from /health output * fix(proxy): keep the proxy error shape on credential endpoint refusals * fix(proxy): hide identity token file paths from /health output * fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working The Bedrock Claude Platform route already reads anthropic_workspace_id from optional_params, so banning that spelling as a server-owned federation parameter broke a pre-existing client capability. The federation field is now anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID), which restores the base branch's behavior for Bedrock callers, drops the Bedrock-specific hint from the refusal message, and deletes the unconditional ban constant that no longer had a reader * fix(auth): share one exchanged token across workers reading the same assertion Anthropic accepts each identity assertion exactly once, so two uvicorn workers reading the same token file both minting from it means the second exchange is denied with jti_reused. Minted tokens now land in a per-user 0700 cache directory guarded by a file lock, so workers on the same host reuse one exchange until the token expires or the assertion rotates. A 401 is only retried when the re-read assertion actually differs, and the denial hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache and an empty value disables it * fix: keep anthropic federation from being shadowed or leaked An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated deployment sent an empty x-api-key on every call instead of minting a token. Blank values now read as unset, and a real static key on a federated deployment logs once that it outranks federation and nothing is being federated. The exchange-host allowlist matched hostnames only, so a second process on another port of an allowed host was trusted with the workload's identity token. An entry that names a port now trusts that port alone, while a bare host still trusts every port. The shared token store exists so the workers reading one projected token file do not each spend its single-use jti. A source that mints its own assertion per exchange shares nothing with another worker, so it no longer writes a live token to disk for a lookup that can never hit. * fix: unlink a staged token file a failed write leaves behind The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads. * refactor: move anthropic jwks derivation behind a provider-owned tagged union * fix: unlink the staged token file when its write fails at close A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token. * fix(anthropic): close the staging descriptor before writing the shared token file * fix(wif): judge federation writes by what they set, not what is stored The admin gate read the stored deployment, so a team admin lost edit, delete and Test Connection on any deployment carrying federation params. It now returns early unless the submitted fields touch the federation surface, and a Test Connection probe that points the deployment at its own api_base is still refused, with the 403 no longer wrapped into a 500 The rest of the same review pass: POST /model/new refuses only a blocking value of `blocked`, so a client that always sends `blocked: false` is not turned away; a request body can no longer pick which federated identity to mint as by naming a stored credential; an advisory refresh the executor refuses disarms the entry instead of wedging the identity until the follower timeout; the static-key shadow warning resolves its env fallback inside the cache instead of once per request; credential writes drop nulls before storing them; the token exchange validates the endpoint URL before reading an assertion and keeps refusing redirects across a client heal; /health hides every server-owned federation field from non-admins; and the async create_file and create_batch paths say which setting is missing when the provider resolves no URL * fix(proxy): let a deployment write name a federated credential reject_federated_credential_reference runs from is_request_body_safe, which pre_db_read_auth_checks calls on every route, so it also fired on POST /model/new, /model/update, /model/{id}/update and /health/test_connection. A proxy admin could no longer attach a federated credential to a deployment over the API or the Admin UI, leaving a static config.yaml entry as the only way to configure the feature the rejection told the caller to go configure, and _reject_non_admin_wif_write never got to make the call it exists to make. is_request_body_safe now takes the route and skips only the credential-reference check on the routes that reach can_user_make_model_call. Federation fields typed inline into a body stay refused everywhere, and a call naming a federated credential still cannot pick the identity it mints as. * refactor(proxy): derive health display policy from the federation key sets The health check module hand-copied the five workload identity fields whose value is a credential, so a shared proxy surface named provider-specific parameters and a newly added secret-bearing field would have gone on being displayed until someone remembered both places WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of, types/utils derives secret_bearing_wif_litellm_params from it, and the health layer splats that tuple the same way it already splats the admin-only one * fix(anthropic_wif): treat blank identity-source fields as unset * test(proxy): classify the federation params in the credential slot registry main's registry test (#43298) now fails the build for any credential-named deployment param without a classification. The five federation fields that carry a token, a token file path, or a signing or client secret reference are Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak settings name a URL, a client id, an auth method, or a scope and are NotSecret * fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics Register the 18 Anthropic and 3 OpenAI federation params as frozen ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel and the request-body ban list read one declaration. Pass the deployment api_base through to the count-tokens handler instead of a pre-suffixed URL, which doubled the /count_tokens path on main's prompt-cache predictor. Add the five litellm_anthropic_wif_* families to the all-metrics Grafana dashboard. * fix(credentials): gate PATCH on WIF fields resolved from model_id The credential PATCH handler checked server-owned workload identity federation fields only on the values the caller sent, while a body that named a deployment through model_id had its credential values resolved after that check. A non-admin could therefore copy a federated deployment's WIF fields onto an ordinary credential. Resolve the incoming values first and run the non-admin gate on them, matching the POST path * fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header Count-tokens walked its own credential ladder: a static key, else skip minting when ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the same deployment authenticated with that token. The handler now takes the auth header that AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills use, and merges the oauth beta a minted or consumer token carries with the token-counting beta --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com> Co-authored-by: mateo-berri <happymvw@gmail.com> |
||
|
|
dfdd496db8
|
fix(tests): match the OS bind error in the owned-proxy port-race retry (#44462)
* fix(tests): match the OS bind error in the owned-proxy port-race retry * test(integration): keep the port-race predicate pure so its unit tests stay in-process --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |