mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
320 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
902736bfe7
|
ci: move Postgres, MCP and Redis suites to CircleCI integration (#44453)
* ci: move Postgres, MCP and Redis suites to CircleCI integration * ci: throwaway, drop tests/proxy_behavior from its CircleCI job to show assert-ci-coverage fails * ci: revert throwaway assert-ci-coverage check * ci: keep the e2e helpers the gate tests still use * ci: move the roi-database Postgres shard to CircleCI integration * ci: run redis-compat without CircleCI's Azure and cassette env, cover postgres_suite test_path * ci: match the GitHub env for the moved Postgres and Redis jobs * ci: unset provider keys in the CircleCI MCP job and drop unused e2e-stack helpers --------- Co-authored-by: yuneng <yuneng@berri.ai> |
||
|
|
7a7d27c550
|
fix(guardrails): run the end-of-stream post_call scan when the client disconnects mid-stream (#43839)
* fix(guardrails): run end-of-stream post_call scan when the client disconnects mid-stream Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): close the guardrail stream chain in async_data_generator on client disconnect Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): leave the raw upstream response to the shielded finalizer on client disconnect Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): keep disconnect cleanup going when a streaming callback cleanup raises Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): assert the refund through a recorder instead of the mock Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): inspect tool calls released before a disconnect under incremental_diff and record a failed scan marker The incremental_diff transform stream now scans tool calls it already released when the client disconnects, and a disconnect scan whose translation raises after the guardrail recorded success also records guardrail_failed_to_respond Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): pin that a text-only disconnect scan is not handed a tool_calls finish Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): cover disconnect scans on every streaming endpoint and client, plus outage, worker-kill and cache-hit cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): prove the cache-hit twin is served from the cache Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): type the disconnect-close streams so basedpyright stops reporting unknown arguments * fix(guardrails): give the guardrail metadata cast a reason so the type discipline gate accepts it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): scan released Messages and Responses tool calls on disconnect Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(guardrails): format the disconnect scan unit tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(guardrails): drop mutable-ok markers that no longer suppress a rule Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): scan released Responses output after a finished item and end only in-flight Chat choices Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): type the disconnect scan test helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): type request_data in the disconnect scan helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): type request_data in the disconnect scan test doubles Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): pin that chat streams with no tool call in flight end as released Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): scan only released chunks on disconnect and skip it once a block owns the verdict The disconnect scan now uses the chunks actually yielded to the client, copies them before scanning, skips when a mid-stream block or HTTP error already settled the verdict, and the iterator wrapper only closes hooks that are async generators so plain async iterator hooks keep working Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): pin that a delivered guardrail error or final chunk settles the disconnect verdict Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): close any hook iterator that exposes aclose when the stream ends early Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): accept a synchronous aclose on custom streaming hook iterators Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): swallow callback aclose errors at end of stream A custom callback whose async_post_call_streaming_iterator_hook returns a non-generator async iterator with a raising aclose() failed the finished stream: content plus usage reached the client and then the stream surfaced an error SSE with no [DONE], or aborted a post_call pipeline's buffering loop into a 500 with an empty body. Wrap the aclose invocation in _wrap_streaming_iterator_with_enrichment in try/except and log a warning naming the callback and the cleanup error, matching close_guarded_stream and _close_guarded_layers. Iteration-time hook exceptions still propagate. * fix(proxy): log only the error type when a callback aclose raises Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
1238cfe90f
|
feat(proxy): add LITELLM_FIPS_MODE startup gate with provider assertion and loud password migration failure (#42700)
* feat(proxy): LITELLM_FIPS_MODE startup gate with provider assertion, TLS verify guard and loud password migration Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): match ssl_verify off detection to runtime str_to_bool semantics Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): drop tautological fips probe test and satisfy CodeQL return checks Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): inject the fake Prisma client through the module boundary instead of patching _setup_prisma_client Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2e1a98f521
|
fix(logging): deduplicate streaming failure callbacks (#44442)
* update logic that marks a logging callback as complete * test(logging): cover streaming failure dedupe in mark_logging_complete Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): cover streaming failure dedupe in S3 and DataDog Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(logging): keep has_run_logging as a deprecated alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): audit streaming failure dedupe across surfaces, fallbacks and sink outage Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): assert anthropic upstream path in streaming failure audit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): count only provider posts in streaming audit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): assert the sink outage rejects uploads in burst audit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(logging): configure datadog retries with router override Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
1d52985d03
|
fix(ci): repair the security sweep and Lens billing integration tests for ROI, JWKS, and release identity changes (#44529)
* fix(ci): keep the security sweep off the observed ROI GitHub route and give Lens integration tests a release identity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ci): expect the credential JWKS export to 404 for non-federation credentials in the security sweep Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0b74ae9c5c
|
feat(mcp): hand listed-tool description and input schema to pre-call hooks per caller (#41162)
* feat(mcp): hand listed-tool metadata to pre-call hooks with per-caller catalog identity Track the tools each MCP server listed per caller identity so pre_mcp_call and during_mcp_call hooks receive the tool description and input schema the client saw. Servers with no caller-dependent inputs share one slot; user identity, forwarded headers, stdio env, relayed bearers, and server-specific auth get their own. Local registry and OpenAPI paths pass the registered metadata and admin description overrides. The Agent 365 guardrail reads the new fields into its evaluate payload. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mcp): drop the listed-tools empty sentinel and routine test docstrings Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): mark the listed-tools cache digest as a non-security hash Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key the listed-tools cache by the OBO subject token token_exchange servers list upstream with the caller's own Entra bearer, so two callers on one LiteLLM key with different subjects were sharing a catalog slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): resolve the BYOK credential before keying the listed-tools slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): drop the OAuth discovery cache when a server definition changes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mcp): drop a diff-narrating comment Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): never validate a supplied header on the tools/list BYOK path The pre-listing resolver ran the tool-call byok_auth_required check even when the caller already supplied x-mcp-auth, and it ran outside the per-server error boundary, so a single deprecated-header caller dropped the server from the aggregate list. Listing now returns a supplied header unchanged and falls back to the stored credential without raising Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): assert the BYOK listing lands in the caller's listed-tool slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): cover the deprecated string x-mcp-auth header on a BYOK tools/list Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key the per-caller listed-tool slot by the hashed token Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key discovery cache by the hashed token Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates Discovery-list cache identity now uses the hashed token instead of the raw api_key and treats MCPJWTSigner-signed servers as per caller. Server definition changes also drop the cached upstream OAuth metadata. OpenAPI listings look tools up under the normalized registry prefix with the separator, so an overlapping sibling prefix no longer leaks into the list. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mcp): derive the listed-tool caller identity from the discovery cache key Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys An upstream metadata fetch that started before a server edit could store its stale reply after invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch only stores when the generation it captured before I/O is unchanged. The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so both go back to the merge-base behavior. Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep OAuth metadata generations only while a fetch is in flight Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): return one masked text per scanned string in the selected-guardrail REST test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): fold the signed caller into the discovery digest instead of a second key hash Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): satisfy type discipline gate on listed-tool identity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): hand tools/call hooks the exact catalog entry tools/list served get_listed_tool re-applied the admin description override on top of the cached listing, so a guardrail-masked description was restored to its original wording at call time, and the OpenAPI / local-registry call path built its metadata from the registry instead of the guarded caller catalog. Both paths now return the cached entry as served, falling back to the registry only when no listing was recorded Adds tests/integration/mcp/test_mcp_listed_tool_metadata.py (red on the prior head for the two regressions, red on the merge base for the feature, green on this head) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): key OpenAPI listed-tool entries per caller so tools/call reads its own guarded listing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): align listed-tool slot tests with per-caller keying Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep oauth2 listing on the minted or signed credential, not the stored BYOK secret The listing helper that keys the per-caller catalog by the stored BYOK credential also handed that credential to the upstream client, which on an oauth2 server short-circuited the client_credentials mint and the MCPJWTSigner gate. Split the two: the catalog identity keeps the stored credential so tools/call finds the caller's slot, while an oauth2 server's tools/list sends only the per-request header, letting the M2M mint or signed JWT proceed as on main Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): oauth2 BYOK listing sends the minted token, not the stored secret, through the real proxy Integration cell for the listing fix: a client_credentials BYOK server with a stored user credential, one tools/list as that user, the peer must see a live minted bearer and one /token mint. Red at the pre-fix tip (zero mints, stored secret upstream), green at the fixed head and at the merge base Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep the stored BYOK credential for catalog identity only on tools/list Listing used the resolved stored credential both to key the caller's catalog slot and as the upstream transport header, so REST api_key and bearer_token listings sent the user's secret instead of the server's static token and the MCPJWTSigner gate went quiet. The upstream client and the signer gate now read the caller-supplied mcp_auth_header for every auth type, exactly as before the catalog existed, and the stored credential only names the slot tools/call reads Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): type the listed-tool metadata read from pre-call kwargs for the basedpyright gate Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): hand never-listed tools/call hooks name and arguments only The local-registry call path fell back to the registry entry with the admin description override when no tools/list had been recorded for the caller, so a pre_mcp_call guardrail scanned a description the caller was never served and blocked OpenAPI calls that passed before, and base's own selected-guardrail REST test failed on the two-text redaction. _registered_tool_metadata now returns the listed entry or None, so a tools/call with no prior listing sends name and arguments only as promised, and that REST test double goes back to its base shape * fix(mcp): keep during_mcp_call hooks on name and arguments only call_tool handed the caller's listed entry to the during-hook task as well, so during_mcp_call guardrails scanned the description line and schema leaves of any listed tool after the upstream call had already run, blocking calls that passed before whenever the policy matched the description, returned a fixed-length texts list, or hit the depth guard on a deep schema. The listed entry is only disclosed for pre_mcp_call, so the during task no longer receives it and its request object carries no description or schema, as before * fix(mcp): key the BYOK catalog slot by the client's header, not the stored credential tools/list resolved the stored BYOK credential to pick the caller's catalog slot, which read the credential store before the classified try block. With Postgres down and a cold per-worker cache that made every REST tools/list on an is_byok server fail with tools=[] and no upstream call, and the read seeded the per-worker cache (including a negative entry), so a tools/call on another worker after a store, rotate or revoke on this one kept using the stale value. The slot is now keyed by what the client supplied plus the caller's hashed key, on both sides. _get_tools_from_server and call_tool take a keyword-only catalog_auth_header that defaults to mcp_auth_header as received (the default is the builtin Ellipsis so it survives a module reload). The /mcp fan-out and execute_mcp_tool, which swap the resolved credential into mcp_auth_header, pass the client's value explicitly. What goes upstream is unchanged. _byok_catalog_auth_header is gone. * fix(mcp): drop a listed catalog recorded across a server save _record_listed_tools ran after the awaited upstream fetch, so a PUT /v1/mcp/server that landed mid-fetch had its invalidation undone when the fetch completed: hooks then saw the pre-save description next to the post-save definition until the next listing, instead of name and arguments only. The manager now keeps a per-server listed-tools generation, bumped by _invalidate_server_definition_caches. _get_tools_from_server reads it before the fetch and _record_listed_tools skips the write when it moved; the next listing records normally. * fix(mcp): drop the catalog again once a saved OpenAPI server's registry is rebuilt add_server and update_server publish the saved definition before the OpenAPI registry entries are rebuilt from the spec, so a listing recorded during that fetch held the pre-save entries under the new generation. The generation is bumped a second time after the registry refresh. The during-hook task no longer accepts a listed entry, the one-line wrapper over get_listed_tool is inlined at its two call sites, and the per-server generation map is a plain dict. * fix(mcp): keep discovery and OAuth metadata caches across an OpenAPI spec re-read add_server and update_server ran the full server-definition invalidation a second time after the awaited OpenAPI spec fetch, which also dropped the prompts/resources/templates discovery entries and the OAuth protected-resource metadata filled under the already-published definition, so the next request went upstream again. Only the listed-tool catalog recorded during the fetch holds pre-save entries, so the post-fetch pass now drops just that catalog and bumps its generation via the new _drop_listed_tools helper, which the full invalidation also calls. * fix(mcp): look a called tool up in the listed catalog by its bare name only get_listed_tool stripped the server prefix a second time when the exact name was absent from the caller's listing, so a never-listed upstream tool whose bare name starts with the server prefix resolved to the listed sibling and that sibling's description and input schema reached the pre-call hooks for a call to a different tool. Every caller already passes the once-stripped bare name, so the lookup is now exact. Tests that looked the catalog up by a prefixed name now use the bare name the callers pass; two new tests pin the never-listed sibling case at the manager and at the tools/call path. * fix(mcp): record a listed-tool catalog only for a listing the caller is served _get_tools_from_server now records the catalog into the caller's listed-tools slot only when asked (record_listing=True), which the served listings pass: the /mcp and Responses API tools/list handlers via _get_tools_from_mcp_servers, MCPServerManager.list_tools, and the REST listing via _list_server_tools. Four internal listings stop recording, so a later tools/call hands pre_mcp_call hooks name and arguments only, as on main: - _list_tools_before_first_call, the implicit listing inside tools/call when this worker does not yet expose the tool - fetch_pinnable_tool_catalog, the admin pin snapshot listed without the catalog guard and without description overrides - _initialize_tool_name_to_mcp_server_name_mapping, the startup fill - get_tools_for_server, used by the semantic tool filter _create_prefixed_tools returns to its tool-name mapping job only; the record follows it in _get_tools_from_server. * fix(mcp): opt every listing out of catalog recording unless it is served The aggregate listing and _list_mcp_tools now default to record_listing=False, so a catalog fetched inside a tools/call no longer fills the caller's listed-tools slot. The /mcp/proxy meta-tools (call_tool, search_tools, get_tool_schema) and the tool-search virtual tool stop recording: /mcp/proxy serves only the meta-tools and the search serves only its hits, so a later pre_mcp_call hook was reading a description the caller never listed. The tools/list handler, the Responses MCP handler and the /v1/mcp/tools management listing opt in with record_listing=True, since each serves the catalog to the caller. * fix(mcp): key the listed-tool slot by the caller's admission identity and forwarded bearer The slot a tools/list records for a later tools/call was keyed by (user_id, api_key) only, so every team-only JWT caller shared one slot and one JWT user acting in two teams shared a slot; a tools/call then handed pre_mcp_call hooks a description another caller was served. The slot is now keyed by the hashed key, user, team and organization, plus the admission credential of a caller admitted with neither a key nor a user. The caller bearer split the slot only on client-forwarded-token and token-exchange servers; a legacy delegated oauth2 server (delegate_auth_to_upstream without client credentials) also forwards it upstream and served a different catalog per bearer into one slot. The bearer now splits the slot on every server whose egress forwards it (_consumes_caller_authorization) or exchanges it. * fix(mcp): record only tools served by the bridge Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep bridge tool metadata request-local * refactor(mcp): centralize listed catalog recording guard * fix(mcp): preserve base TPM reservations for listed tool calls * fix(mcp): preserve project token reservations for listed calls * fix(mcp): record served catalogs and preserve call message bytes * test(mcp): align listing expectation with deferred recording * test(mcp): audit listed metadata across callers and bridge lifecycles * test(mcp): preserve guardrail fixture worker affinity * refactor(mcp): expose listed catalog recording API * feat(mcp): pass served_tools through the anthropic messages bridge /v1/messages auto-execution now hands this request's resolved tool definitions to _execute_tool_calls, matching the Responses and chat completions bridges: the pre_mcp_call hook receives the description and input schema the model was shown for that call. Request-local only; the shared listed-tools catalog is untouched. * test(mcp): pin served_tools handoff on the anthropic messages bridge Mirrors the credentials-forwarding test: the request's resolved tool definitions must reach _execute_tool_calls under served_tools so pre_mcp_call hooks judge the call on the description and input schema the model was shown. Fails without the previous commit's one-liner. * style(mcp): sort the local import block ruff flagged * fix(mcp): keep the admin include_disabled_tools view off the listed-tools catalog GET /mcp-rest/tools/list?include_disabled_tools=true is the admin-only configuration view: apply_tool_filters is False, so it serves the full server catalog. Recording that response into the caller's listed-tools slot warmed tools/call metadata no runtime listing ever served, breaking the only-a-served-listing-records invariant (Bugbot). The record is now gated on apply_tool_filters; disabled tools stay unreachable (the call-time allowlist 403 fires before hooks), so the observable fix is the slot no longer warming from a settings view. Verified live: the new test fails on the unfixed head and passes here, and the rest of the listed-tool-metadata suite is unchanged. --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
984b4134be
|
fix(proxy-extras): log v1 migration failures at ERROR so LITELLM_LOG=ERROR shows them (#44202)
* fix(proxy-extras): log v1 migration failures at ERROR so LITELLM_LOG=ERROR shows them Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy-extras): reuse litellm secret redaction and mask configured DB passwords exactly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy-extras): name the password alternation in _redact_credentials Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(proxy-extras): wrap v1 migration ERROR lines at 120 columns Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): integration cells for v1 migration ERROR logging and password redaction Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): integration cells for component DB env vars, JSON logs, migration Job and v2 resolver Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): give every v1 migration integration cell the 900s timeout Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): bind the recovery forwarder before migrating so the retry cannot race it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): drop the slow P3005 integration cell and bound migration subprocesses Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy-extras): keep command repr and tolerate non-sequence cmd in migration error logs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy-extras): double the subprocess boundary in the cmd=None retry test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: jesus <jesus@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
cf22deb96a
|
test(ci): fix six CircleCI test regressions on main (#44429)
* test(ci): fix four CircleCI regressions on main Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(ci): stop reloading auth_checks in unit tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(constants): cover CLI JWT expiry env parsing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(gateway): restore proxy lifespan after importing gateway.main in launch tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
9d16412341
|
fix(caching): count tool_call cache_control marks in the injection census (#43556)
* fix(caching): count tool_call cache_control marks in the injection census
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): remove cache census casts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): only skip injection on message or content marks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip injection on messages whose tool calls carry marks
Reverts
|
||
|
|
f850b2c324
|
test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment (#44451)
* test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): compare every non-transport provider header and check for late provider requests at session end Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): name TranslationTestCase fields after litellm and provider sides and drop regressions Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): prefix checked TranslationTestCase fields with expected_ and name the fake reply mock_provider_response Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs(integration): name TranslationTestCase fields in the translation README Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): add a claude-opus-5-5 base case and deployment next to claude-sonnet-4-6 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): name translation cases <MODEL>_TEST_CASE and document the naming rule Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs(integration): move translation test rules into tests/integration/translation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
dfdd496db8
|
fix(tests): match the OS bind error in the owned-proxy port-race retry (#44462)
* fix(tests): match the OS bind error in the owned-proxy port-race retry * test(integration): keep the port-race predicate pure so its unit tests stay in-process --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
6d73fa6b49
|
fix(vertex_ai): forward system and tools to partner model count_tokens (#43900)
* fix(vertex_ai): forward system and tools to partner model count_tokens Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(vertex_ai): avoid mutable token request construction Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(vertex_ai): return partner count_tokens provider errors as values so the proxy falls back locally * test(integration): cover Vertex AI partner count_tokens forwarding and local fallbacks Add wire-level cells for /v1/messages/count_tokens, /utils/token_counter, /v1/responses/input_tokens and the Gemini countTokens route on a Vertex AI Claude deployment: the system prompt and tools reach the partner count-tokens endpoint verbatim, null fields stay out of the body, malformed tools are rejected before any peer call, peer, token-endpoint and connection failures fall back to the local tokenizer unless disable_token_counter is set, generation on the same deployment keeps working, and concurrent bursts survive a peer outage, a slow peer and a worker SIGKILL. The sdk cells cover litellm.acount_tokens the same way. The _support/process.py and _support/client.py harness files are brought to main's content so the self-booting cells read INTEGRATION_PROXY_READY_SECONDS instead of a fixed 70 s boot budget. --------- Co-authored-by: jesus <jesus@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
1ef0fe9790
|
fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check (#43741)
* fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(auth): key the zero-cost cache by the resolved model group Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(auth): keep the zero-cost verdict per requested name and include hidden aliases * test(integration): audit the zero-cost bypass through hidden model_group_alias names * test(router): cover the extracted routing strategy switch * test(integration): record a pre-flip burst before the alias flip --------- Co-authored-by: jesus <jesus@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
1a7023366f
|
feat(roi): measure shipping velocity, quality, and recorded spend (#44426)
* feat(ui): prototype observed engineering ROI dashboard * feat(roi): replace effort estimates with measured repository metrics * fix(roi): finish connection recovery and generated API contracts * fix(roi): show merged changes before accounts are linked * fix(roi): preserve selected report tab across refreshes * fix(roi): recover app authorization and keep detail values readable * fix(roi): reuse the shared OAuth HTTP client * feat(roi): combine providers and compare equal reporting periods * docs: explain ROI metrics for first-time readers * fix(roi): preserve connections and scheduled reports during setup * ci(roi): assign database contracts to the active Postgres shard * fix(roi): preserve issue counts and normalized connections * fix(ui): compact ROI dashboard header and metrics * fix(ui): show ROI repository count with expandable list * fix(ui): wrap ROI controls within narrow panels * fix(roi): restore sample report preview and simplify setup |
||
|
|
01b4ffe16b
|
fix(bedrock_mantle): route Claude chat completions to the native Messages endpoint (#43646)
* fix(bedrock_mantle): route Claude chat completions to the native Messages endpoint Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(bedrock_mantle): price Claude chat on the Mantle row and route region-prefixed ids A Mantle request always carries a region, so a Claude id with no bedrock_mantle/<region>/ row fell through the model-info lookup to the bare Bedrock row, which the bedrock provider family also matches, and billed about 10 percent under the Mantle price. The lookup now tries the provider's region-free row before the bare model. The Claude route test asserts the Mantle row, a region-prefixed Claude id is covered end to end, and the provider config map references the Mantle config directly. * test(bedrock_mantle): cover supported params for Claude and open-weight Mantle ids * test(bedrock_mantle): audit the Claude chat bridge on the integration rig Adds the deterministic cells from the /audit of the Mantle Claude chat bridge: wire-level translation on every chat, responses, and messages route, SigV4 and bearer auth, region prefixes, api_base suffixes, unsupported params with and without drop_params, malformed model ids, upstream errors, the response-cache hit, the health check, pricing from the Mantle row for Claude and non-Claude ids with a bare Bedrock twin, and chaos cells for a mixed burst, an upstream outage, slow streams, and a worker kill on an owned two-worker proxy --------- Co-authored-by: jesus <jesus@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
e9bf2cfd01
|
refactor(types): replace Any with proven types in 9 files (#44389)
* refactor(types): replace Any with proven types in 13 files Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): keep provider error paths for malformed prefetch and poll JSON Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): keep main's login body parsing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): keep main's Copilot auth and budget alert typing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit cells for Any sweep 20261003_2 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): load audit video deployments at proxy start Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): keep main's video-edit prefetch handling Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): fix audit chaos Responses cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): probe every route after audit worker kill Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
fe683ea139
|
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models
GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.
* fix(proxy): offer a Codex service tier only when every deployment of the model lists it
* fix(codex-catalog): an invalid service_tiers value offers no tier for the model
* fix(codex-catalog): read service tiers off the deployments the key's team can route to
A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them
The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default
* test(codex-catalog): drop the redundant module docstring and sort the imports
* test(integration): add the Codex catalog audit cells and the multi-worker convergence note
* test(integration): clean up every catalog test model and answer the refresh GET
* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers
Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.
* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut
The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns
The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
|
||
|
|
1b98748528
|
fix(bedrock): honor the per-request timeout on Converse and Invoke streaming (internal copy of #38210) (#44134)
* fix(bedrock): propagate timeout to streaming requests * test(bedrock): prove streaming fails at the request timeout against a slow upstream * test(bedrock): simulate the slow upstream in process instead of over a local socket * test(bedrock): audit the Converse and Invoke stream timeout on the proxy Two integration files drive the per-request timeout on Bedrock streams through the real proxy against an owned wire peer: the wire file covers every surface (chat, messages, responses, invoke, pass-through), the sad, edge and precedence rows, and the chaos file covers bursts, a dropping upstream, a killed worker and a proxy stopped mid-burst. The wire peer gains Reply.drop_connection so a cell can close the socket before any response, and the harness's graceful stop grace is now INTEGRATION_PROXY_STOP_SECONDS (default unchanged at 30), since a two-worker supervisor's interpreter finalization takes longer than that on a loaded box. * test(bedrock): pin the fallback audit cell to one proxy worker The fallback cell created both deployments through /model/new on one worker and sent the chat request to the other, whose registry read-through loads only the requested model, so the fallback target was unknown there until the periodic DB poll. The cell now warms the fallback model and sends the request over one keep-alive client, so one TCP connection stays with one uvicorn worker, and it expects the fallback upstream to see both requests. --------- Co-authored-by: Sainyam Kapoor <hello@sainyam.me> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
a849093d89
|
test(cost): pin OpenAI reported web search count with mixed actions (#44414)
Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
6c32384d8c
|
fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3 (#44292)
* fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3
Bedrock prices Kimi K3 cache reads through implicit caching but Converse
rejects the explicit cachePoint marker ("This model doesn't support the
cachePoint field"), so any cache_control on the request answered 400.
Mark the three K3 rows supports_prompt_cache_breakpoint: false and have
bedrock_model_accepts_cache_points honor that flag before falling back to
supports_prompt_caching, keeping cached-token pricing intact.
* test(bedrock): assert cache points per request section
* fix(bedrock): honor a deployment's cache breakpoint flag for unmapped models
* fix(bedrock): read a converse-routed deployment's cache breakpoint flag
* refactor(bedrock): look up cache breakpoint flags by key
* test(bedrock): add the Kimi K3 cache point wire audit
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
|
||
|
|
8b1990b4bc
|
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): register typesafe as a provider so Jev deployments load Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(decisions): move provider endpoints under llms and validate proxy bodies Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(decisions): add Cloudflare Clef and Strands Decider backends Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): register decisions routes for managed agents and gateway Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(decisions): use raw regex for cloudflare missing account match Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): avoid cast in Cloudflare response unwrapping Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): default model, evaluation health probe, short Cloudflare names The proxy validates only state and questions, so a request without a model falls through to the configured default model like every other route. Health checks probe evaluation-mode deployments through the Decisions API instead of failing with an unsupported mode, and cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short names resolve a mode and a price. The registry no longer claims typed decisions for a provider with no backend. * fix(decisions): let health_check_params override the evaluation probe Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit the decisions endpoint across providers, limits, health and chaos Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port). The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load. * fix(decisions): send env API keys to a configured api_base Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(decisions): register the routes through the lazy feature registry The Decisions router was included at import, ahead of the config and DB pass-through endpoints, so a pass-through configured at /v1/decisions was skipped and answered 400 as an unknown Decisions provider. The routes now register through LAZY_FEATURES, which splices them in after every eager route, so a pass-through at /v1/decisions keeps its route while /decisions still serves natively. The lazy OpenAPI snapshot carries the two paths so the schema shows them before the first call. The audit cells add the env-key egress to a configured api_base, the client api_base opt-in shared with chat, the pass-through precedence on an owned proxy, and the Strands evaluation health check resolved from the cost map. The integration config exports the Perplexity env key the first cell needs. * fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell * fix(proxy): let a config pass-through beat a lazily registered route in eager mode With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at startup, so SafeRouteAdder treated a config pass-through at exactly /v1/decisions as already registered and dropped it. In lazy mode a pass-through created through the API after the first native call was skipped the same way. Routes a lazy feature owns no longer count as registered, and a route added at one of their paths is placed ahead of them, the precedence lazy mode gives a config pass-through when the feature has not loaded yet. --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
5724117116
|
fix(responses): merge bridged tool calls into the same choice as the text (#44346)
* revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295)
This reverts commit
|
||
|
|
d96e56c76f
|
revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295) (#44344)
This reverts commit
|
||
|
|
7b432d78d2
|
fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model (#44341)
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): streamed alias matching a capability rule bills the deployment price Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): assert every streamed chunk carries the client alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(logging): log the client alias on the priced streamed response, the same as non-streamed Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
a76ba8c01e
|
revert(cost): revert "fix(cost): price rule-only model names at the deployment's rate" (#44144) (#44335)
This reverts commit
|
||
|
|
caee45fed4
|
feat(roi): add GitLab sources and branch cost attribution (#44324)
* feat(roi): support GitLab and tagged branch costs * fix(roi): count tagged branches independently of estimation status * test(roi): capture live GitHub and GitLab report validation * fix(roi): open estimate details at the start * fix(roi): clarify cost views and unify report layout * feat(roi): showcase per-PR costs in the sample report * fix(roi): separate report tabs and preserve branch cost attribution * fix(roi): preserve demo previews and align progress spacing * fix(roi): isolate demo loading and parallelize fork lookups Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression * fix(roi): separate demo and live loading states Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers * fix(roi): ignore refreshes from a previous source * fix: trust gateway context for ROI estimator exclusion * fix: preserve historical ROI estimator exclusion |
||
|
|
6d8434f940
|
fix(proxy): return 4xx instead of 500 for missing required params, invalid pagination and unknown ids (#43787)
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: run search_endpoints tests in proxy-endpoints shard Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(llm_http_handler): keep provider error text when re-raising mapped errors Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): allow promptless image edits and default search models Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): default missing image edit image to None Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): build image edit defaults without mutating request data Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): inject a fake router for the search default model test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(llms): cover provider error status on vector store and file lookup handlers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(llms): keep the lookup handler raise block to a single statement Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover missing required body params and provider lookup status codes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): bind spend-row request id with partial to satisfy B023 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): only reject non-positive page_size on vector store list Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): remove unreachable fine-tuning body validation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover streaming anthropic messages reaching the upstream Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): count only provider calls when asserting missing params never reach the upstream Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): preserve merge-base request compatibility Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): preserve interaction completion model defaults Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): retry model read-through before rejecting params a DB-only deployment may default Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: shivam <shivam@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
8efb4a21f6
|
fix(vector_stores): return managed file ids from vector store file list (#43800)
* fix(vector_stores): return managed file ids from vector store file list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vector_stores): cover managed file list route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(vector_stores): only map round-trippable managed ids and index flat file ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): build managed file gin index concurrently
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy-extras): move the managed file gin index migration after main's newest
* fix(vector_stores): satisfy lint gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(proxy): drop the stale no-index note on the raw file-id guard
* test(vector_stores): cover managed file ids on the vector store file list end to end
Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
|
||
|
|
8596fe954d
|
feat(interactions): durable cross-pod settlement for background interaction billing (#41955)
* feat(interactions): durable cross-pod settlement for background interaction billing Background interaction billing lived only in the creating replica's memory, so a DELETE routed to another replica, or a restart of the creating one, never billed the completed provider work and the budget reservation was refunded at the poll timeout. The create now registers the billing context in a settlement store before returning, the proxy installs a Prisma-backed store at boot (LiteLLM_BackgroundInteractionSettlement, schema-only migration), any replica claims the row once through a conditional update before billing or releasing, startup resumes every unclaimed row with its remaining timeout, and a give-up records an unsettled outcome instead of silently reconciling to zero. The SDK keeps an in-memory store and behaves as before. * fix(interactions): survive a settlement install failure at boot and stop carrying request headers * fix(interactions): drop the stored request context once a settlement row is settled * fix(interactions): bill the completed response a poll already saw when its claim only answers at the deadline * fix(interactions): carry a missing model through the settlement context for agent-only background creates An interaction created with an agent and no model reaches the poll with no model name, exactly as on main. The settlement context now stores that None instead of rejecting the create, which answered the client with a 500 after the provider had already accepted it. * fix(interactions): leave an unfetchable background interaction to its creating poll when a delete lands elsewhere The remote pre-delete path fetches with only the delete's credentials, so a fetch it cannot make says nothing about the interaction. It used to claim the settlement row and release the reservation anyway, which stopped the creating replica's poll and lost the bill when the delete then failed the same way. It now returns without claiming; the in-process path keeps releasing on an unfetchable state, since its context carries the create's own credentials. * fix(interactions): fail a cross-replica delete when its pre-delete fetch fails so the creating poll keeps the bill * fix(interactions): keep the stored settlement gate when registration raises after landing, and fail resumed-poll deletes closed A registration that raised after its row committed moved the poll to a private in-memory gate, so the creating worker billed while the stored row stayed unclaimed for another replica's delete or the next boot to bill again. The row is now read back once and, when it landed, the poll claims through it like every other settler. A worker that resumed the poll after a restart is not the creator, so its delete on a failed pre-delete fetch now fails with the fetch's error instead of releasing and deleting. After a fleet restart every worker holds resumed polls, which left the fail-closed path applying nowhere. * fix(interactions): settle an unverified registration through the durable claim A create whose settlement-store write raised no longer bills through a private in-memory gate that a later boot's resume cannot see. The claim asks the durable store first and falls back to the local gate only when the store answers that no row exists, and a missing settlement table reads as no rows so a replica without the migration still settles in process. * test(proxy): keep the settlement test where the proxy-infra shard collects it The merge of main moved test_background_interaction_settlement.py under tests/unit/proxy/spend_tracking, but the proxy-db shards claim tests/unit/proxy files one by one in .circleci/scripts/unit_selection.sh, so no CI shard ran it and codecov/patch dropped. tests/test_litellm/proxy/spend_tracking is collected whole by the proxy-infra shard, which is where the test ran before the merge. * fix(interactions): raise on a non-2xx Gemini interaction fetch AsyncHTTPHandler.get never raises for status and the Gemini GET transform only raised when the body was not JSON, so a 500 or 404 carrying Gemini's JSON error body parsed as an interaction with no status. A delete on a replica other than the creator then claimed the settlement as released and forwarded the delete instead of failing closed, and the bill was lost. The transform now raises GeminiError with the vendor's status, as the delete transform already does; the in-process poll already retries a fetch that raises * test(integration): audit durable background interaction settlement across replicas Twenty-six deterministic cells drive a one-worker creator and a two-worker settler against an owned scripted Gemini upstream: cross-replica deletes bill once, failed and cancelled interactions release, a later replica resumes unclaimed rows, custom deployment pricing bills at the deployment rate, a fetch the settler cannot make fails the delete closed, odd ids are refused, a missing settlement table keeps in-process billing, polling disabled registers nothing, the budget reservation is released by the settler, an upstream outage mid-burst fails closed and recovers, killed workers hand their polls to the respawned ones, and concurrent deletes on a slow upstream settle exactly once. The support upstream gains a scripted interaction store with per-id GET status and delay, and the process helper gains an owned upstream a test can stop and restart * test(integration): refuse a repeated delete in the scripted upstream and pin the settlement budget below one estimate * chore(ui): regenerate dashboard API types after merging main * test(integration): accept the 422 budget refusal and a respawned worker's resume The budget cell pinned a 400 that the proxy stopped answering when budget refusals moved to 422, so it now asserts the status and the budget_exceeded error type the sibling budget tests pin. The later-booting replica cell accepts a claimer that is any worker started after the creates, since uvicorn's supervisor can respawn the creator's worker under load and the respawned worker's boot resume claims the rows by design; the single spend row check is unchanged * test(integration): delete the pinned key's interaction with a second key A key whose budget is filled by its own reservation is refused on every route, the DELETE included, so the cell now asserts that 422 and sends the delete with a second key, which is what the reservation release on another replica needs in order to be observable at all --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
6b6222578e
|
test(integration): let run.py select cells by pytest node id (#44319)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
f9a32ffcb5
|
fix(cost): price rule-only model names at the deployment's rate (#44144)
* fix(cost): price streamed aliases that only match a capability rule from the deployment model Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost): satisfy basedpyright delta after the cost-candidate sort Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * perf(cost): skip the capability-rule check for exact cost-map keys Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): price streamed alias rows from the deployment's own rates Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): isolate streamed alias deployments with per-run model names Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(streaming): keep the unpriceable-stamp case on a truly unmapped model Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): treat capability-rule matches as unmapped in every cost lookup Route every model-info lookup on cost paths through get_priced_model_info and _cached_get_priced_model_info_helper, which raise ModelNotMappedError when the only match is a pricing-free fallback-generalizations rule. An alias that matches a capability rule now falls through to the deployment's real model instead of billing 0, and the earlier candidate-sorting fix is reverted since the priced helper is the single choke point. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): assert the rule alias bills the same as the plain alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): treat router-registered rule-only model_cost entries as unmapped Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): count string rates and check pricing before the capability-rule match Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(cost): repoint cost-path patches at get_priced_model_info Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(cost): give the model_cost row cast a cast-ok reason Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): type the lazy get_priced_model_info export Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost): narrow the fix to ordering rule-only cost candidates last Drops get_priced_model_info and its call-site swaps, the lazy-import entry, and the cost-path code check. Only the candidate sort in completion_cost and pricing_entry_for_cost_calc stays, with its tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): cover rule-only base_model billing on every endpoint, client and failure mode Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
ca1994e403
|
fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice (#44295)
* fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): preserve deferred logging bridge coverage Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): cover gpt-6 family bridge routing and merged tool calls in integration Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
233db9f285
|
feat(bedrock): serve gpt-5.6+ chat completions natively by default, with chat_completions/ opt-in for gpt-oss and grok (#44307)
* feat(bedrock): send grok chat completions through runtime openai path Unspecified bedrock grok was rewritten to Converse. Chat completions now hit bedrock-runtime /openai/v1/chat/completions, and converse/ still uses Converse * feat(bedrock): serve gpt-oss and gpt-5.6 chat completions on runtime's native openai path * fix(bedrock): route gpt-oss response_format to Converse and decide the route once from the raw request * fix(bedrock): serve region-path and GovCloud gpt-oss ids on native Chat Completions The cost-map parity tests require every regional variant of a flagged id to carry the same supports_ flags, so the six us-gov gpt-oss entries now carry the native-route flags too. A region path in the model name (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model: the route is looked up on the id after the path, the path's region picks the endpoint and the SigV4 scope, an explicit aws_region_name still wins, and the body carries the bare id AWS expects * fix(bedrock): keep params AWS refuses natively off the chat completions route Drop the params each family 400s or 503s on runtime Chat Completions from the native config's supported list (GPT-5.6 penalties, stop, and logprobs, Grok penalties, gpt-oss logit_bias) so drop_params drops them as Converse did, gate legacy functions on GPT-5.6 the same way as tools, and send an Anthropic-style thinking block to Converse, the only route that forwards it * fix(bedrock): keep schema-less json_object on Converse for the chat completions models * fix(bedrock): keep every json_object response_format on Converse for the chat completions models * fix(rust): declare the bedrock runtime chat completions flags on ModelInfo * fix(bedrock): opt into the native chat completions route through supported_endpoints * docs(cost-map): describe the bedrock native chat completions capability flags * revert: docs(cost-map): describe the bedrock native chat completions capability flags This reverts commit |
||
|
|
5dbe4f95e8
|
feat(bedrock): drop lookaround regex patterns from tool schemas for Converse models that reject them (#44138)
* feat(bedrock): drop lookaround regex patterns from tool schemas for Converse models that reject them * fix(bedrock): rename the lookaround flag to supports_regex_lookaround and keep dropped patternProperties names allowed The cost-map flag becomes a generic supports_regex_lookaround capability, which the cost-map schema admits as a supports_* boolean, and the Converse transform now owns the drop decision instead of the shared tools factory. A patternProperties key dropped from an object closed by additionalProperties: false leaves its value schema as that object's additionalProperties, so the names it allowed stay allowed, on the OpenAI non-Python-regex drop too. tool_with_sanitized_parameters also cleans Anthropic-shape tools (input_schema). * fix(router): keep a deployment's supports_regex_lookaround off the shared cost-map entry A deployment's model_info.supports_regex_lookaround was written to the shared bedrock/<model> cost-map key, so every sibling deployment of that model id inherited one deployment's choice. The flag now stays under the deployment's own id, which is what the Converse lookaround check reads first, and the shared entry keeps the cost map's value * test(bedrock): audit the Converse lookaround drop on the wire across endpoints, SDKs, flags and chaos --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
fe910889f7
|
fix(responses): drop bridge-minted reasoning items from OpenAI replays (#44132)
* fix(responses): drop bridge-minted reasoning items from OpenAI replays * fix(responses): send id-less stored reasoning items without a made-up id The chat-to-Responses bridge gave a stored reasoning item with no id an rs_<n> id that OpenAI rejects (404 without encrypted content, 400 with it); an id-less item is accepted and verified by OpenAI itself. Decoding encrypted_content now keeps the verifiable thinking blocks of a mixed array instead of rejecting the whole array, and the verifiable-block rule lives in the shared module. * test(integration): audit minted reasoning item replay across responses, chat bridge and messages --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
dac31e3d5d
|
fix(anthropic): stop repeating streamed thinking text in the signature chunk (#44127)
* fix(anthropic): stop repeating streamed thinking text in the signature chunk The Anthropic stream handler emitted, on the signature_delta, a thinking block carrying every thinking delta seen so far plus the signature, after it had already streamed that text as per-delta blocks. Every additive consumer (the Agents SDK, stream_chunk_builder, the Responses bridge) stored the text twice under one signature and replayed the doubled block on the next turn The signature chunk now carries a signature-only block, matching the Bedrock emitter, so accumulators rebuild the text once and the saved history replays to Anthropic exactly as it was streamed * test(anthropic): audit the signature-only thinking chunk across wire providers * test(anthropic): pin the cached replay's thinking shape in the wire audit --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
677205b3f5
|
fix(proxy): share ownership permissions for spend logs and traces (#44239)
* refactor(proxy): extract shared spend log read policy * test(proxy): use named bindings for spend scope regression * test(proxy): reuse existing spend log query harness * test(proxy): cover spend log permission lookup adoption * chore(proxy): relocate existing spend query baseline * refactor(proxy): make scope query returns explicit * refactor(proxy): inject deferred log permission lookup * test(proxy): cover teamless management compatibility lookup * refactor(proxy): compose user and team log grants * refactor(proxy): share generic authorization composition * refactor(proxy): compose trace read permissions * refactor(proxy): centralize spend and trace authorization * refactor(proxy): strengthen spend and trace scope types * refactor(proxy): flatten log read scope into owned logs Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope, and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened the tree back into that shape. A caller with no user id now gets an empty scope instead of matching ownerless rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs is removed, and the management facets inject the log team lookup and reuse read_scope_sql instead of the list shim. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(proxy): run spend scope tests through one SQLite emulator Replace the string-matching payload emulator and the hand-rolled Prisma where interpreter with one SQLite helper that runs the real scope SQL. Session scope tests now go through the endpoint, including the no-user caller that must not match ownerless rows. Drop duplicated lookup-failure and trace mapping cases. load_permitted_log_team_ids returns no teams without a database instead of relying on the resolver's broad except. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(proxy): unify log and trace ownership permissions * test(tracing): align fixtures with ownership read scopes * refactor(tracing): align query scopes with row ownership * refactor(spend): make ownership SQL predicates explicit * test(spend): validate ownership SQL against PostgreSQL * docs(traces): drop key-row visibility from query help guide Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): reach the empty-memberships branch in team lookup test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(ui): regenerate dashboard API types Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
a292fd409f
|
test: fix three order-dependent and timing-flaky tests (#44271)
* test(integration): answer the model-info refresh GET in the mixed MCP responses wire peer The proxy's periodic model-info refresh sends GET /v1/models to the deployment api_base, which tripped the peer's /responses-only assertion when a tick landed mid-test * test(e2e): wait for the api-keys URL after clicking Virtual Keys in onboarding /ui already renders the Virtual Keys heading, so the helper returned before navigation finished. The late route change moved focus and closed the account menu popover in hideLiteAdmin * test(proxy): stop two unit modules leaking app.openapi_schema and a session-wide Router test_custom_openapi cached a stripped schema on app.openapi_schema and never cleared it, breaking later openapi route tests. test_proxy_reject_logging built a module-level Router that stayed in the live router registry all session and re-added cost-map keys during a reload. Reset the schema via monkeypatch and make the Router a function-scoped fixture |
||
|
|
b94eaa65c3
|
test(integration): opt the config pass-through spend-log case into auth (#44265)
The failure spend-log row from #42695 is written for pass-through routes that run as LLM API routes, which a config route only does with auth: true. The test omitted auth and passed only while config wins (#41779) registered config entries through the typed model, where auth defaults to true. #43962 restored the pre-config-wins registration, so the route lost that status and the row was never written. Set auth: true on the route so the test covers the logging it was written for without depending on that side effect |
||
|
|
8b4de39ad7
|
fix(bedrock): surface unrecognized converse-stream event frames instead of an empty turn (#44112)
* fix(bedrock): surface unrecognized converse-stream event frames instead of an empty turn * fix(bedrock): track distinct unknown stream event types and cover the stream error builder * test(bedrock): audit converse-stream event frames on the proxy and the sync decoder --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
04bc354525
|
refactor(mcp): extract upstream preparation and support modern clients (#44232)
* refactor(mcp): extract upstream preparation and support modern clients * fix(mcp): reject incompatible upstream transport before saving * fix(mcp): serialize protocol validation with server updates --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
9b5562f89b
|
test: repair stale and polluting tests red on scheduled main CI (#44229)
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context * test(secret-detection): give the hand-built redaction request an ASGI path Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path * test(integration): isolate litellm callback lists per sdk test usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards * test(integration): keep the owner-lookup fault proxy off the shared read replica The owned proxy points DATABASE_URL at a scratch database but inherited DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared database and rejected the freshly created key with token_not_found_in_db. Drop the replica variable like the other scratch-database owned proxies * test(integration): request every seeded key in the team owner breakdown The aggregated team activity endpoint now caps breakdown.api_keys at the top 100 keys by default (#43398), so the 300 seeded keys came back as 100 rows. The test guarantees each key is reported with its own owner, so ask for an api_key_limit that covers all seeded keys * test(integration): give every owned Redis its own port in the redis-cache container On CircleCI every owned Redis ran on the fixed port 16379 inside the shared redis-cache container. When an earlier server still held that port, the new one failed to bind, readiness pinged the old server, the pidfile read failed and cleanup then reported "Owned Redis still serves after shutdown" Reserve an ephemeral port for the docker-exec path the same way the local binary path already does, and refuse to start when something already serves the chosen port so the failure names the real cause * test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail * test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to * test(e2e): check only stored message content for a leaked card number The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there * test(e2e): assert the proxy decodes token-array embeddings for titan The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer * test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5 * test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed * test(e2e): cite the tokenizer and date behind the titan token-array fixture * test(e2e): let migration seed replicas finish their request-log indexes before cloning Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema * test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope #43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake * test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly |
||
|
|
0238ec9721
|
fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key (#43978)
* fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key A path-less add/replace op (RFC 7644 3.5.2, what Okta Push Groups sends on a rename) carries a partial Group resource. Each of its attributes now applies as if sent with that path, so displayName updates the team alias and externalId and members get their usual handling, and the pushed attributes merge into the scim_data snapshot the PUT path already writes. A path-less remove or a path-less op without an object value is rejected with a 400. Any group PATCH drops an empty metadata key an earlier push left behind, and the Admin UI metadata form skips an empty key so an affected team can save its settings. * fix(scim): let a later path op win over an earlier path-less value in the group snapshot * fix(scim): type the stored team metadata before the JSON object check * test(scim): run the real group transformation in the path-less replace test * test(scim): assert the renamed group comes back from the path-less replace * test(scim): audit the path-less group PATCH on the live proxy --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
f932e292c5
|
fix(proxy): carry key, team and project tags into pass-through spend logs (#42662)
* fix(proxy): carry key, team and project tags into pass-through spend logs Pass-through endpoints built their request metadata without the key, team and project controls that native routes apply, so spend rows for configured routes and provider pass-throughs like /anthropic dropped the key, team and project tags and the key and team spend_logs_metadata. The native team and project controls now live in a shared helper that both paths call, key spend_logs_metadata is copied instead of aliased from the cached key, client metadata cannot overwrite user_api_key_ fields, and header tags dedupe with the same merge used on native routes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): drop covers markers from pass-through tag tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): type the shared team and project control helper Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): validate pass-through endpoint list instead of suppressing pyright Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover body, streaming, hostile, forged and native cells for pass-through tags Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: mock pass-through request url as httpx.URL after rebase on main * refactor(proxy): merge spend_logs_metadata sources without a stacked comprehension * refactor(proxy): name the team and request spend_logs_metadata merge --------- Co-authored-by: ryan <ryan@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
f95446ea39
|
fix(otel): tolerate non-dict callback_settings.otel and ignore bare EXCLUDED_SERVICES env (#44086)
* test(otel): cover non-dict callback_settings.otel and bare EXCLUDED_SERVICES env Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): tolerate non-dict callback_settings.otel and ignore bare EXCLUDED_SERVICES env Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(otel): simplify settings_customise_sources signature Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): poll for present spans instead of waiting the full window Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): audit null otel block and bare EXCLUDED_SERVICES across the excluded-services matrix Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): assert per-trace datastore spans in the unconfigured burst cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): keep pydantic-settings runtime options on the OTel v2 env source Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): require post-auth datastore spans at the tenant on the cache-hit twin Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): cover pydantic-settings runtime options in callback_settings.otel through the proxy Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
b21e44cbf9
|
feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375)
* test(e2e): jwt auto_register map-existing-key repro Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(jwt): route existing-key lookup through VerificationTokenRepository Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key Skip Admin UI session keys and keys whose allowed_routes restrict them to anything other than llm_api_routes (management, read_only, password-reset sessions). Mapping a JWT to one of those left the user with 401s or 403s on every LLM call, since the mapping persists. * fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team Only reuse a key whose team_id matches the team auth_builder resolved for the JWT (no team matches no team), so a personal key can no longer bypass the resolved team's model and budget limits. With the flag on, the first JWT request now falls through to the same virtual-key checks later mapped requests get, instead of returning early, so a reused key's own limits apply from request one rather than 200 then 403. Flag off keeps the early return unchanged. * fix(jwt): keep the early return when no master key is set Without a master key the generic virtual-key path returns a bare INTERNAL_USER object, so falling through on the first auto-registered request dropped the key's team, models and budgets. Only fall through when a master key is configured. Tests now assert the reused key per team rather than the query shape, and cover the flag-off early return and the no-master-key case. * test(jwt): assert on race-loser's returned key, not only mocks (TQ002) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key) Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401 Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter * test(e2e): create the reused key in the team the JWT resolves to The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass * test(integration): match the held statement across TCP reads The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read * fix(jwt): gate key reuse on the claim field, not on the claim value Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path * fix(jwt): let an issuer's own user field replace the global one when gating key reuse An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key * test(jwt): make the flag-off test fail when the flag no longer gates key reuse The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings * chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops --------- Co-authored-by: yuneng <yuneng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mrinal <mrinal@berri.ai> Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
fe9b6fd603
|
refactor(types): replace Any with proven types in 7 files (#43844)
* refactor(types): replace Any with proven types in 9 files Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): pin xecguard-adjacent guardrail retry, mcp mixed tools, jwt routing and complexity router paths Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): keep only live-provable Any removals Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): type cache_hit as bool | None on the sync success path Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
7be2983f11
|
test(straiker): assert a saved api_version v1 with an sk_agt_ key routes to v3 (#44153)
* test(straiker): assert a saved api_version v1 with an sk_agt_ key routes to v3 (#44106) Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(straiker): ignore zombie workers when counting uvicorn children after a kill Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(straiker): describe the saved v1 cell by the v3 route it asserts Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
e32b25817f
|
fix(proxy): keep tool payloads and logprobs unmasked in stored spend logs (#44075)
* fix(proxy): keep tool payloads and logprobs unmasked in stored spend logs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): audit cells for stored spend-log tool payloads and logprobs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): assert cache-hit spend-log rows Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): handle model-list requests in spend-log audit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): assert unauthenticated requests never reach upstream Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
d131c43782
|
fix(guardrails): restore Azure guardrail get_user_prompt dispatch and allow logging (#44067)
* fix(guardrails): restore Azure guardrail get_user_prompt dispatch and allow logging Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): use transport-level doubles in Azure dispatch regression tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): add Azure dispatch audit matrix integration cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): restore global callback lists after Router reset in SDK cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): make proxy restart cell independent of shutdown timing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |