mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-02 02:11:58 +00:00
258 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
008fcb4fe3
|
feat(tool-policies): show the user who owns the key that discovered a tool (#43892)
* feat(tool-policies): show the user who owns the key that discovered a tool
GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tool-policies): bound the owner lookup with chunked membership queries
The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tool-policies): cover the owner column across the tool routes and the dashboard
Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
||
|
|
65a316fc92
|
ci(circleci): test Redis behavior against local Redis and print short tracebacks (#44062)
* ci(circleci): test Redis behavior against local Redis and print short tracebacks The redis_caching_unit_tests job ran three legacy files against the shared remote Redis. The DualCache and batch-read logic that never needed a server now lives in tests/unit/caching/test_dual_cache.py with a mocked RedisCache, and the behavior that does need one (the increment-with-floor Lua script, read-through, deletes, batch reads) moved to tests/integration, which starts a local Redis. test_returned_settings only read REDIS_PORT and is replaced by a unit test of Router.get_settings CircleCI pytest runs now use --tb=short so failure output stays readable in the test results tab * ci(integration): print short tracebacks from run.py and allow Redis in the sdk shard |
||
|
|
bba85f0b6c
|
chore(lint): remove the LIT002 mutable-construction rule (#43971)
* chore(lint): remove the LIT002 mutable-construction rule
Drop LIT002 from scripts/check_type_discipline.py along with its helpers,
its budget entry, its unit tests, and the AGENTS.md and gate docstring
mentions. `# mutable-ok` now only suppresses LIT001, so the markers that
only existed to silence LIT002 became LIT013 stale suppressions and are
removed. The files whose layout depended on those trailing comments are
reformatted with ruff format.
Every other LIT rule count is unchanged and the ASTs of all touched
litellm/ files match main apart from one docstring.
* chore(lint): keep the leftover mutable-ok markers for a follow-up
Restore the ~1.4k `# mutable-ok` markers stripped in the previous commit
so this PR only touches the checker, its tests, the budget, and docs.
Those markers no longer suppress anything, so `# mutable-ok` is exempt
from LIT013 until a follow-up strips them.
* Revert "chore(lint): keep the leftover mutable-ok markers for a follow-up"
This reverts commit
|
||
|
|
259c166ef6
|
refactor(lens)!: rename internal engine code and API (#44034)
* chore(lens): remove deployment screenshots * refactor(lens)!: rename internal engine package and API * fix(lens): pin worker image for renamed API * test(lens): cover fresh and populated rename migrations * fix(lens): protect db-push upgrades and restore routing and CI * fix(lens): resolve migration tables across schemas and include database driver |
||
|
|
bfd3f39dca
|
feat(s3_v2): add s3_partition_granularity option for hourly S3 folders (#43748)
* feat(s3_v2): add s3_partition_granularity option for hourly S3 folders Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover s3 v2 partition granularity across surfaces, settings and chaos Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover previous_response_id history rebuilt from an hourly cold storage object Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): reuse the cold storage key only when s3_v2 owns cold storage Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): cover hour rollover, postgres outage, in-flight switches, key/team vars and real S3 layout Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(liccheck): authorize libfaketime, the GPLv2 dev-only clock the s3 rollover integration test preloads Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): wait for the rejected-request cell's payloads by id, not by line count Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): declare the postgres outage cell's models in config and trip the relay on burst ids Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): drop the libfaketime hour rollover cell and its dev dependency Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): deselect the s3_v2 live e2e on the stage-mirror stack The stage-mirror config enables no s3_v2 callback, so every test in test_s3_log_e2e.py fails its readiness check there. The file keeps running in the Buildkite e2e lane, which configures s3_v2 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): declare the sink outage burst models in config Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(s3_v2): read cold storage metadata without an empty dict default Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): wait for the proxy to reconnect before the postgres outage recovery request Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mrinal <mrinal@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
6ca90b927c
|
test(ci): repair stale tests and flaky CI infrastructure (#43983)
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden #43063 stamps used_client_oauth_token into spend-log metadata, so test_async_gcs_pub_sub_v1 failed on main with an extra metadata key * test(ui): give the auto-router threshold save wait room for the availability debounce #42625 keeps Save disabled while a 300ms-debounced availability check runs. This test waits for Save right after the change, so the whole debounce lands inside waitFor's 1s default and it times out under CI load. It is the recurring UI Unit Tests failure on main since #42625 landed * test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default #42870 added both the rule that a served default or standard tier bills at base pricing and records no service_tier, and streamed tests expecting the row to record 'default'. They have failed on every scheduled litellm-e2e run since. The tests now map the served tier to the pricing basis the bill must record and check input is billed at that basis's rate; the messages case registers custom rates so the rate check has something to compare against * test(e2e-ui): wait for the call-id search before hovering the logs row The row the spec hovers is already on the unfiltered first page, so it was found before the search request returned. The search response then re-rendered the table under the mouse, and the Base UI tooltip never opened. Reproduced with Playwright against a local proxy: hovering right after the fill never shows the tooltip, hovering after the search response shows the call id every time * test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off The case picked the cheapest Together row flagged supports_response_schema. DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model the reasoning_effort=none case already exercises, and Together lists it with structured output support * test(integration): read the agent 365 guardrail status by its own name in spend logs The MCP shard runs under xdist against one database, and a sibling file creates a default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so that filter's 'success' entry could land first in guardrail_information and the test read it instead of the agent 365 verdict * test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check gc.collect() inside the caplog window can collect a pending task an earlier test left on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this test's records. The check still counts every LiteLLM logger, and unretrieved task exceptions on this loop still go through the asserted exception handler * test(e2e-ui): fill the create-tag fields inside the dialog #42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description') match two elements and Playwright's strict mode fails the create step * test(integration): run integration proxies with the CI license Multi-worker proxies start each uvicorn worker in a fresh process, so every worker reads the license from its environment. Forward LITELLM_LICENSE into the proxy and test runner environments * ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5 Every pull request saved its own uv, maturin, Rust and Prisma caches, about 4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries within minutes. Pull request jobs then missed every cache, downloaded all dependencies from PyPI and hit the install step timeouts. Pull requests now restore only, and main keeps the caches warm for them. test-linting and check-ui-api-types run only on pull requests and keep saving codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity keybase account, so every upload failed signature verification. 5.5.5 reads it from codecovsecops; the key ID matches the one signing the current CLI * test(unit): join the session-minting thread before collecting the handler asyncio.to_thread resumes the test as soon as the worker sets its result, while the pool thread can still hold the work item and through it the handler. gc.collect() then cannot finalize the handler and the session stays open. A pool that shuts down before the test continues drops that reference * test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early owned_proxy_process released its reserved port and the proxy bound it only after full startup, so another xdist worker or an outgoing connection could take it first and the proxy exited with 'address already in use'. The launch now retries on a fresh port when that happens and stops every failed attempt. uvicorn closes idle keep-alive connections after 5 seconds and httpx expired them at the same 5 seconds, so a request sent right at that mark could reuse a socket the server was closing and get 'Connection reset by peer'. Gateway clients now drop idle connections after 2 seconds * ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build The release profile builds with fat LTO and one codegen unit, so the final link of litellm-cache-s3 runs silently for minutes. Successful builds take 711 to 749 seconds, right at the default 10 minute no-output limit, and about 30% of recent runs were killed there * test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections The proxy retries the database about every 30 seconds and each retry opens roughly one connection, so a 5-connection outage took 3 to 4 retries to clear and recovery landed between 60 and 90 seconds, straddling the test's 80 second reset window. A fixed 10 second outage still refuses the immediate reconnect and recovers on the next retry * ci: move the unit-test uv cache split into a composite action check_workflow_startup_safety sums every setup step's timeout, so the save and restore variants each counted 5 minutes although only one runs. One composite step keeps the setup ceiling at 35 minutes * test(unit): point tiktoken at the bundled cache for every unit test The rust_bridge tokenizer tests loaded o200k_base before any test in their xdist worker had imported default_encoding, so tiktoken fell back to the temp cache and tried to download under pytest-socket. Move the session fixture from litellm_core_utils/conftest.py to the root unit conftest. * test(integration): answer model discovery probes in the hosted_vllm wire tests The router's periodic upstream model info refresh sends GET /v1/models to hosted_vllm deployments, so a wire server that is live during a refresh sees an extra request. Answer the probe with an empty model list and leave it out of the provider-call assertions, matching the responses bridge tests. |
||
|
|
d9f73245be
|
feat(lens): track worker spend through virtual keys (#43989)
* feat(lens): bill worker analysis through virtual keys * fix(lens): pin the verified worker image and add setup proof * fix(lens): preserve network checks and redact billed analysis logs * test(lens): preserve legacy worker result submission during upgrade * fix(lens): enforce trusted worker IPs and restore coverage uploads * docs(lens): explain trusted proxy requirements for worker allowlists * fix(lens): yield to worker disconnects after the synthetic body |
||
|
|
6f123b7083
|
test(proxy): migrate DB and Redis backed proxy tests into tests/integration (#43996)
* test(proxy): migrate DB and Redis backed proxy tests into tests/integration Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): drop a type suppression comment from the key metadata integration test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): scope integration test cleanup to owned rows and wait for backend stats flush Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): seed NULL cache_hit and bound recovery reads from below Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yuneng <yuneng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
4b06d04334
|
test(anthropic): native /v1/messages reasoning integration tests built on a captured Claude Code request (#43361)
* test(integration): group /v1/messages contracts under tests/integration/messages Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): make ci coverage census collect nested test dirs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): nest /v1/messages contracts under messages_endpoint/providers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): replay a real Claude Code /v1/messages request through the native wire Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): drop legacy covers marker from claude code wire test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): inline the Claude Code request instead of a json fixture Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): drop the legacy covers marker from the new contract Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge) (#43386) * test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): strengthen bot-flagged assertions in the Claude Code matrix Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): type the usage mapping parameter in the shared builders Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): drop low-priority Claude Code error and count_tokens tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): send the full 24-tool Claude Code request and pin upstream headers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): drop responses bridge Claude Code tests to keep this PR Anthropic direct only Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): fix duplicate WebSearch tool, drop mutation in stream builders, ignore pings Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): assert upstream request order in multi-turn Claude Code tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): name Claude Code wire tests by behavior and move provider-agnostic ones to routing and streaming Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): sort imports after moving the Claude Code fixture into _support Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): group anthropic messages tests into feature subfolders Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): drop the pre-move anthropic test paths Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): cover native reasoning translation, response and pricing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): narrow PR to new reasoning tests, restore moved files and drop non-reasoning tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): scope reasoning tests to reasoning and cover betas, thinking usage and streamed pricing Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): write reasoning cases as literal sent and received fields Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): require stopped stream blocks and check upstream model on switch Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0980f756bd
|
fix(guardrails): scan Responses API input in Azure Text Moderation (#43965)
* fix(guardrails): scan Responses API input in Azure Text Moderation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): log Azure Text Moderation prompts at debug and cover streamed Responses blocking Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
19842da059
|
fix(guardrails): scan Responses API input in Azure Prompt Shield (#43786)
* fix(guardrails): scan Responses API input in Azure Prompt Shield and Text Moderation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): tolerate unmodeled Responses input items in Azure prompt extraction Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): pick Azure prompt source by call type so a messages stub cannot hide Responses input Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): tighten Azure Content Safety endpoint test types Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): suppress Azure cast lint violations with cast-ok reasons Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): audit Azure content safety across endpoints Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): isolate worker-kill audit rig and cover during_call on chat Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(guardrails): shorten Azure cast-ok reasons to fit the line limit Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): inline spend row count in the concurrency audit cell Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): assert caller-observed outcomes in Azure call type unit tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): reuse the existing text moderation response helper Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): assert no duplicate rows instead of exact row count after worker kill Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): poll worker-kill spend rows to settle before the duplicate check Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): keep Azure Text Moderation on messages only so this PR stays Prompt Shield scoped Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: shivam <shivam@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
a308a8e579
|
feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (#43134)
* feat(guardrails): honor litellm_params.timeout in every HTTP guardrail Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): accept timeout kwarg in presidio and responses-handler post stubs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): bound hiddenlayer startup jwt call by configured timeout, drop akto from timeout coverage Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): narrow hiddenlayer startup auth timeout without cast Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): bound hiddenlayer jwt refresh by configured timeout Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): keep provider timeout defaults when unset and bound only rubrik moderation calls Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): cover model_armor and run timeout probes concurrently Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): match sink calls to the exact guardrail name Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
c168199e33
|
test(ci): repair stale tests and move retired OpenAI text-completion fixtures (#43958)
* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups Request fakes now carry the scope a real Starlette request has, the GCS pub/sub spend-log golden gains the agent identity keys from #43722, the auto-router session tests follow the baseline_models contract from #43348, and the Interactions spec checks resolve the create body and resource paths from the live spec instead of hardcoded names * test(ci): move retired OpenAI text-completion fixtures to live vehicles OpenAI still serves native /v1/completions on the gpt-5.4 family, so the single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt batches and echo with logprobs now 500 on every OpenAI model, so those cases keep the same text-completion-openai transport pointed at Fireworks, which documents both. The optional-params test asserts the request body actually sent instead of a success callback whose assertions were swallowed * test(ci): use a serverless Fireworks model for the text-completion batch and echo cases gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not deployed; glm-5p3-flash is listed as serverless * test(ci): skip the ROI calculator repository listing in the security route sweep GET /roi-calculator/repositories (#43669) lists repositories from the configured GitHub API, api.github.com by default, so the S2 sweep's GET of every route made the owned proxy reach an external host and failed the egress check in 31 integration-security tests. It joins /get/latest_release_info in the deny list |
||
|
|
54ae4c5bbf
|
fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (#43082)
Some checks are pending
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: rerun integrations shard after unrelated gitlab prompt manager timeout Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit azure_storage client reuse against a local Data Lake sink Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(azure_storage): cover the exact TTL expiry boundary Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(azure_storage): wait for a rejected write before flipping the sink back Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): add azure_storage log delivery cells behind an opt-in lane Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(azure_storage): restart the proxy mid burst and bound the loss to the unflushed queue Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(azure_storage): drop the redundant stop after the owned proxy exits Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): read azure_storage objects at the auth-mode-dependent layout Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): install the datalake sdk in the e2e lint environment Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(azure_storage): drop the opt-in real Azure e2e cells and their e2e-dev dependency Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
a3a7650569
|
fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (#43956)
* fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): patch the shared proxy logger directly in the straiker api_version test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): keep the straiker stray-version block marker separate from the shared block marker Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): drop redundant comments on the straiker api_version integration tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2c3866ebb4
|
fix(azure_storage): name Data Lake objects without base64 padding or slashes (#43914)
Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
6997223068
|
fix(grayswan): send request conversation and tool calls to post-call monitor (#43770)
* fix(grayswan): send request conversation and tool calls to post-call monitor Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(grayswan): tighten post-call context typing and wire test helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(grayswan): resolve post-call surface from request route before call_type Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(grayswan): omit tools from post-call monitor when request context is empty Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(grayswan): apply ruff format to post-call context changes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(grayswan): merge response text and tool calls into one assistant monitor message Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(grayswan): only merge tool calls into the response text for single-choice responses Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): audit post-call context across endpoints, modes and outages Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): share the upstream model probe reply across audit responders Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): assert the full generic guardrail body and kill a real serving worker Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): normalize the client user agent in the generic body assert Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): normalize accept-encoding in generic body assertion Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): keep volatile header placeholders only when the header is present Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): capture monitor calls immutably in the unit test client Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(grayswan): type the test helper parameters Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
c42d06fb80
|
fix(router): bill service tiers at catalog rates for custom-priced deployments (#43890)
* fix(router): inherit catalog service-tier rates for custom-priced deployments Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): apply tier-suffixed long-context rates when only tier thresholds are set Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(router): cover canonical cost-map backend model resolution Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(router): use descriptive names for service-tier pricing fixtures Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
6fd9334751
|
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at
|
||
|
|
6b9766fa0c
|
feat(proxy): add native ROI calculator for gateway spend vs merged PRs (#43669)
* feat(proxy): add native ROI calculator for gateway spend vs merged PRs Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * fix(proxy): serialize ROI Prisma inputs with builtin containers Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * style(proxy): format ROI calculator backend files Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * fix: parse fenced ROI estimates and retain completed reports Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * fix(roi-calculator): correct estimator and dashboard behavior Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * chore(ui): drop next dev generated AGENTS.md block Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * fix(proxy): chunk ROI spend user lookup Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * fix(ui): show reused ROI estimates after sync Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * fix(security): address ROI CodeQL alerts Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * test(proxy): make ROI calculator unit tests discoverable Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * test(ci): run ROI calculator tests in proxy infra shard Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * fix(roi): page repository search and recover polling errors Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> * feat(roi): bring scheduled analysis and guided setup into the gateway * fix(roi): recover interrupted syncs and resolve review findings * fix(roi): preserve cached estimates across report scope changes * fix(roi): normalize scheduler timestamps to UTC * fix(roi): fence cancelled syncs and read reports from writer * fix(roi): preserve reports during metadata outages * refactor(roi): isolate outage validation and verify uncached retry * fix(roi): make scheduled job registration repeatable * style(roi): format scheduler import * fix(roi): continue syncing accessible repositories * fix(roi): preserve reports and identity during upstream outages * fix(roi): persist refreshed identities for reused estimates * perf(roi): skip writes for unchanged cached identities --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com> Co-authored-by: moe-berri <moe@berri.ai> |
||
|
|
f285229b51
|
fix(proxy): delete large teams without per-member transaction fan-out (#42998)
* fix(proxy): delete large teams without per-member transaction fan-out
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): evict email-only member caches and reset team members metric on delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep new delete-team literals within the LIT002 ceiling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve deleted-team member ids before the locked delete
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup
`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.
* fix(repositories): slice find_by_emails into bounded IN statements
The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.
* fix(proxy): delete a team once when /team/delete repeats its id
The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.
* test(integration): audit cells for /team/delete on large, legacy and concurrent teams
Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.
Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.
Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.
* test(integration): pin each chaos outage to a live /team/delete
The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
|
||
|
|
657bb777fa
|
fix(hosted_vllm): keep reasoning_content on replayed assistant messages (#43599)
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm chat templates consume it, so popping it made reasoning models lose earlier reasoning across tool loops. thinking_blocks is still removed for vLLM compatibility. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(hosted_vllm): forward replayed reasoning_content only when it is a string * test(integration): cover hosted_vllm reasoning_content replay across endpoints * test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: yassin <yassin@berri.ai> |
||
|
|
50f5cc9bbb
|
feat(otel v2): excluded_services opt-out for datastore spans on tenant destinations (#43278)
* feat(otel): excluded_services opt-out for datastore spans on tenant destinations Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): keep upstream support unchanged Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): excluded_services resolves from the otel callback config only Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): name the otel callback logger so excluded_services owner lookup matches Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): assert no aux datastore traces reach the tenant sink Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): read bogus-start proxy log from the results dir Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): read only this invocation's bogus-start proxy log Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): assert operator kept db spans over the whole recorded window Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): split operator db-span asserts by trace scope Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): build the otel logger after preset callbacks and validate the exclusion env at boot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): tolerate a bogus exclusion env when callback config wins Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): keep bogus exclusion env fatal when a preset parses it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): hoist the preset check out of the callback loop Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): use a rule-scoped pyright suppression Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): log and drop unknown excluded_services instead of failing boot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): wait for the operator spend-writer span before checking the tenant for postgres Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel v2): leave callback init and boot untouched when excluded_services is unset Read callback_settings.otel.excluded_services directly instead of making the otel callback build its own logger, and drop the new boot-time parse of callback_settings.otel, so a proxy without the setting behaves exactly as on main Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel v2): normalize callback_settings excluded_services without rereading OTel env vars Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel v2): log and ignore malformed excluded_services instead of failing startup Lowercase and trim names, drop non-string items, and add an integration matrix over endpoints, clients, cache hits, destination outages and setting shapes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel v2): cover failed upstream calls in the excluded_services matrix Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel v2): pin operator Langfuse credentials in preset-only excluded_services tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mrinal <mrinal@berri.ai> |
||
|
|
264b09ac8d
|
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present. Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment. Resolves LIT-8931 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): reject empty guardrail rewrites instead of forwarding raw input An explicit texts=[] answer from a guardrail now fails the count check and raises UnappliableRequestRewrite like any other misaligned rewrite; only a missing texts key means no rewrite. Types the out-param as dict[str, object] and adds integration coverage for instructions blocking, masking, empty instructions, tool loops, latest-only and concurrent workers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): type the texts-replacing guardrail helper explicitly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): honor skip_system_message_in_guardrail for instructions and system input items Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): cover skip_system_message_in_guardrail on the live proxy Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system Trust a guardrail's structured_messages_cover_full_request claim only when it returns as many rows as the full normalized request, otherwise merge the scoped rows back so skipped instructions and system items survive the write-back. Make PANW's Responses reasoning alignment skip-aware so latest-only still picks the latest user turn when system content is excluded from texts. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): annotate new guardrail tests with return types Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): type the guardrail test doubles explicitly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
82d8b3797c
|
fix(proxy): attribute completed batch cost rows to /batches in daily activity (#43870)
* fix(proxy): attribute completed batch cost rows to /batches in daily activity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): wait for priced batch tokens before asserting team endpoint activity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
9dda4d895f
|
fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (#43764) | ||
|
|
6bc17f98d7
|
refactor: clean up fresh tech debt from 2026-09-29 (#43830)
* refactor: clean up fresh tech debt from 2026-09-29 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(routing): pin usage-based routing Redis reads through the proxy and SDK Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
ba6d6d1a95
|
test(ci): repair MCP Responses and budget fixtures (#43788)
* test(ci): repair MCP Responses and budget fixtures * test(auth): verify delegated budget changes persist |
||
|
|
61a73c59b0
|
fix(proxy): look up hashed key names with two spend log rows per key (#43656)
* fix(proxy): look up hashed key names with two spend log rows per key The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged. * fix(proxy): cap each spend log name probe at 100 rows per key * fix(proxy): bound the newest-row probe at where the oldest probe stopped The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale. * test(integration): add spend log alias probe cells for the daily activity routes --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
c129ea4fc9
|
fix(mcp): scope OpenAPI listings to the exact server prefix and drop upstream OAuth metadata when a server is saved (#43608)
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates Discovery-list cache identity now uses the hashed token instead of the raw api_key and treats MCPJWTSigner-signed servers as per caller. Server definition changes also drop the cached upstream OAuth metadata. OpenAPI listings look tools up under the normalized registry prefix with the separator, so an overlapping sibling prefix no longer leaks into the list. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys An upstream metadata fetch that started before a server edit could store its stale reply after invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch only stores when the generation it captured before I/O is unchanged. The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so both go back to the merge-base behavior. Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep OAuth metadata generations only while a fetch is in flight Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
f5a1c9f1f1
|
fix(proxy): recover session key owners from daily spend for usage attribution (#43642)
* fix(proxy): recover daily spend key owners Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): simplify daily spend owner recovery Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(proxy): format daily activity metadata Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): cover recovered owner metadata merge Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): bound the daily spend owner lookup with the statement timeout * test(integration): audit the daily activity key owner fallback on every usage route Thirty five integration cells under tests/integration/spend cover the daily spend owner fallback on all nine daily activity routes and /usage/ai/chat: the happy path per route, the unanimity rules (two users, blank and null rows, an owner the user table lacks, live and deleted keys with and without their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a second user landing between reads, a concurrent burst across the unified endpoints, a killed worker, and a proxy restart The traffic cells ignore the GET /v1/models call the proxy's five minute token limit refresh makes to every registered OpenAI compatible deployment, since it lands on a test's provider wire whenever the refresh instant falls inside the test --------- Co-authored-by: jesus <jesus@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
e814532033
|
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(streaming): satisfy type-discipline and strict ruff budgets Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(streaming): stamp the served service_tier on every Responses bridge chunk Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(service-tier): cover anthropic and responses served-tier billing paths Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(service-tier): bill disconnects through the router's anthropic stream wrapper Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style: apply ruff format to the anthropic stream changes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(coverage): ignore delegating properties the ast scan cannot see Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style: keep the cast-ok reasons on the cast call line Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(streaming): keep service_tier on OpenAI-compatible parsed chunks Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(streaming): parameterize delegated chunks and messages types Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(tests): follow the anthropic pass_through rename after merging main Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): drain the logging worker between response cache tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(databricks): keep the served service_tier on streamed chunks and bill it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(databricks): type the served service_tier chunk without a loose kwargs dict Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): bill the served service_tier over the requested one Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost): drop explanatory comment from the tier resolution Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: kerry <kerry@berri.ai> |
||
|
|
abc85c2651
|
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins Move Bedrock commitment rows to cost_per_second so they bill once Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost_calculator): drop legacy per-second fields from chat paths Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost_calculator): recognize output-only per-second rates Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(pricing): cover cost_per_second and legacy per-second aliases through the proxy Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
336c7c0849
|
test(integration): sweep proxy logs, metrics, a Datadog intake and the Logs drawer for credential canaries (#43306)
* test(integration): credential canary suite harness Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix. * test(integration): widen canary route sweep and harden the rig Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy. * test(integration): descend into any decoded value that can still hold an encoded canary * test(integration): bound canary decoding by depth and decoded bytes * test(integration): scope log-table and spend-log reads to the scenario window * test(integration): sweep spend-log rows in the scenario date window * test(integration): keep spend-log date window summarized * test(integration): sweep proxy logs, metrics, a gzip Datadog intake and the Logs drawer for credential canaries * test(e2e): treat an unset prompt-storage setting as unset and restore it * test(integration): name the Datadog sink slot G1d * test(e2e): search the Logs page for base64 forms of the deployment key * test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot * test(integration): pass the resolved deployment id to the Datadog route sweep * test(integration): expect 404 from the caller-scoped team membership route * test(integration): use the rig's own master key and expect 404 from submission lookups * test(integration): check the overridden rig key without assuming the default key is unknown |
||
|
|
cede93e826
|
test(integration): request-path credential canary slots D1-D4 (#43307)
* test(integration): credential canary suite harness Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix. * test(integration): widen canary route sweep and harden the rig Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy. * test(integration): descend into any decoded value that can still hold an encoded canary * test(integration): bound canary decoding by depth and decoded bytes * test(integration): scope log-table and spend-log reads to the scenario window * test(integration): sweep spend-log rows in the scenario date window * test(integration): keep spend-log date window summarized * test(integration): request-path credential canary slots D1-D4 * test(integration): read the Logs drawer and spend-log filter for failed request rows * test(integration): check the marker in each failed row's spend-log filter; run header slots on chat-family routes * test(integration): run D2 on embeddings again; only the client-header slot runs on chat-family routes * test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot * test(integration): pass the slot deployment's model_info id to the route sweep * test(integration): expect 404 from the caller-scoped team membership route * test(integration): use the rig's own master key and expect 404 from submission lookups * test(integration): check the overridden rig key without assuming the default key is unknown |
||
|
|
b3dcf8208d
|
test(integration): callback credential canary slots C1-C3 and D5 (#43630)
* test(integration): credential canary suite harness Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix. * test(integration): widen canary route sweep and harden the rig Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy. * test(integration): descend into any decoded value that can still hold an encoded canary * test(integration): bound canary decoding by depth and decoded bytes * test(integration): scope log-table and spend-log reads to the scenario window * test(integration): sweep spend-log rows in the scenario date window * test(integration): keep spend-log date window summarized * test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot * test(integration): expect 404 from the caller-scoped team membership route * test(integration): use the rig's own master key and expect 404 from submission lookups * test(integration): check the overridden rig key without assuming the default key is unknown * test(integration): callback credential canary slots C1-C3 and D5 Team callback, team callback_settings, config default_team_settings and key metadata.logging Langfuse secrets, a team Datadog dd_api_key, and request-body Langfuse keys (allow_client_side_credentials) must reach only their sink. Each scenario checks its sink received the canary as auth and that the marker is visible at the stored body, the Logs drawer route and the sink. Adds a unit test that the stored request body snapshot carries no callback parameter. * test(integration): give the callback sink waits a wider bound * test(integration): sweep provider requests for callback credentials |
||
|
|
d46304900f
|
fix(router): stream /v1/messages lifecycle frames live when no fallback can take over (#43600)
* fix(router): stream anthropic messages lifecycle frames live when no fallback can take over The /v1/messages streaming wrapper buffered message_start and content_block_start until the first content_block_delta and dropped pings behind buffered frames unconditionally, even for requests no fallback could ever recover. With adaptive thinking on Bedrock or Vertex the client saw no bytes for the whole thinking pass and hit read timeouts. Buffering now applies only while a fallback can still take over (generic or refusal chain resolving), and a ping is always forwarded live since it carries no lifecycle and keeps the connection alive. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(router): mirror every dispatcher fallback path in the anthropic stream gate Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(router): skip already-tried order levels in the anthropic stream gate Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(router): keep a transport-split ping behind buffered lifecycle frames instead of forwarding its head live Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(router): credit the #39566 branch this fix supersedes Co-authored-by: Radu Swigler <radu.porumba@gmail.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Radu Swigler <radu.porumba@gmail.com> |
||
|
|
9dfa42dcde
|
refactor(types): replace Any with proven types in 7 files (#43704)
* refactor(types): replace Any with proven types in 11 files Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): revert Any changes that broke existing callers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(types): drop prompt factory helper wrappers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover typing sweep surfaces Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): tighten sweep audit tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
39d14bd855
|
feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers (#43641)
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(fireworks_ai): drive the router request test through an httpx MockTransport Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(fireworks_ai): assert tool definitions reach the router upstream Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
3572d359a1
|
test(integration): stored-config credential canary slots (#43309)
* test(integration): credential canary suite harness Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix. * test(integration): widen canary route sweep and harden the rig Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy. * test(integration): descend into any decoded value that can still hold an encoded canary * test(integration): bound canary decoding by depth and decoded bytes * test(integration): scope log-table and spend-log reads to the scenario window * test(integration): sweep spend-log rows in the scenario date window * test(integration): keep spend-log date window summarized * test(integration): stored-config credential canary slots Add canary slots for credentials the proxy holds in its env, config or database: virtual key raw value, master key, deployment api_key via /model/new, named credentials, AWS secret key, Vertex service-account JSON and its minted token, team model_config credential overrides, config guardrail api_key, and sink credentials from env. * test(integration): resolve the config guardrail id and require detail routes to return 200 * test(integration): drop repeated timeout comments * test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot * test(integration): expect 404 from the caller-scoped team membership route * test(integration): use the rig's own master key and expect 404 from submission lookups * test(integration): check the overridden rig key without assuming the default key is unknown * test(integration): sweep config-deployment routes with the real model id and use the rig admin for the master-key slot * test(integration): mark the configure-hook config edits as intended * test(integration): drop suppression markers that suppress nothing * Use claude-haiku-4-5 for the Bedrock stored-credential test model |
||
|
|
db9c307e1a
|
test(integration): credential canary slots for MCP and pass-through credentials (#43308)
* test(integration): credential canary suite harness Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix. * test(integration): widen canary route sweep and harden the rig Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy. * test(integration): descend into any decoded value that can still hold an encoded canary * test(integration): bound canary decoding by depth and decoded bytes * test(integration): scope log-table and spend-log reads to the scenario window * test(integration): sweep spend-log rows in the scenario date window * test(integration): keep spend-log date window summarized * test(integration): credential canary slots for MCP and pass-through credentials Adds slots F1 (MCP static auth), F2 (per-user MCP OAuth token), F2E (per-user MCP env var), F3 (x-mcp client auth header), H1 (pass-through credential header), H2 (vector store api_key) and H2S (search tool api_key) to the credential canary suite. The OAuth double gains an optional mint hook so a test can choose the issued access token. * test(integration): wait for MCP spend rows by call type * test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot * test(integration): canary MCP and pass-through slots pass resolved ids * test(integration): expect 404 from the caller-scoped team membership route * test(integration): use the rig's own master key and expect 404 from submission lookups * test(integration): check the overridden rig key without assuming the default key is unknown |
||
|
|
dab2deb5ed
|
test(integration): credential canary suite harness (#43300)
* test(integration): credential canary suite harness Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix. * test(integration): widen canary route sweep and harden the rig Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy. * test(integration): descend into any decoded value that can still hold an encoded canary * test(integration): bound canary decoding by depth and decoded bytes * test(integration): scope log-table and spend-log reads to the scenario window * test(integration): sweep spend-log rows in the scenario date window * test(integration): keep spend-log date window summarized * test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot * test(integration): expect 404 from the caller-scoped team membership route * test(integration): use the rig's own master key and expect 404 from submission lookups * test(integration): check the overridden rig key without assuming the default key is unknown |
||
|
|
7e383c9f6a
|
revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377)
* Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" * revert: "feat(usage): search keys beyond the top-N usage subset (#42827)" (#43378) * Revert "feat(usage): search keys beyond the top-N usage subset (#42827)" * revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595) * Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table" * Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596) Co-authored-by: yassin <yassin@berri.ai> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
317430db4e
|
fix(panw_prisma_airs): honor experimental_use_latest_role_message_only on every request shape (#42447)
* fix(panw_prisma_airs): apply experimental_use_latest_role_message_only to every request shape Explicit true/false now applies to chat completions, Anthropic /v1/messages and /v1/responses alike; unset keeps latest-only for Anthropic and full history otherwise. Text indices are mapped back to their source message by value instead of by count, so Responses instructions, function_call_output and reasoning items no longer derail the alignment and silently rescan the whole history Co-authored-by: scthornton <scthornton@gmail.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(panw_prisma_airs): type latest-message helpers against AllMessageValues Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): require forward and reverse text attribution to agree A Responses function_call_output whose text equals the latest user turn could claim that turn's slot in a forward-only walk and demote the latest-only scan to an earlier message. Walk both directions and fall back to the full role-filter scan when they disagree Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): pick the latest human turn from messages, not from aligned texts An image-only latest user turn no longer promotes an earlier user turn into the latest-only scan; it scans nothing on the request side, as the Anthropic path did before. A latest user/developer message whose text never reached texts (a trailing Responses reasoning item) falls back to the role-filter scan instead of narrowing. Types the test helpers, drops the narrating docstrings and adds regressions for both shapes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): log when latest-only selection leaves nothing to scan An image-only latest user turn with experimental_use_latest_role_message_only=true intentionally yields zero scanner calls. Emit a debug line naming the call_id so operators can tell this apart from the guardrail not firing. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): keep Responses reasoning items out of latest-turn selection The Responses translation handler gives reasoning input items the default user role, so a reasoning item with text content after the latest prompt was picked as the latest human turn and the real prompt went unscanned under experimental_use_latest_role_message_only. Map reasoning items back to their texts positions from the raw input and exclude them; fall back to the role-filter scan when the raw items do not account for every text Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: scthornton <scthornton@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
fe76c2473d
|
Revert "feat(proxy): server-side Team Usage export beyond the top-N key cap (#42996)" (#43376)
This reverts commit
|
||
|
|
3726ce2cfc
|
refactor(guardrails): fix agent 365 to the production endpoint and log the opt-in fail_open at error level (#43189)
* feat(guardrails): fail open by default when Agent 365 cannot evaluate and count it in Prometheus Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(tests): ruff format the Prometheus fail-open registry test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(guardrails): add Agent 365 authority host override, fail-open integration test and per-guardrail YAML default Add `authority_host` to the Agent 365 config (also read from AGENT365_AUTHORITY_HOST, then AZURE_AUTHORITY_HOST) so sovereign clouds and the integration test can point the OBO exchange at a different Entra host. Add tests/integration/mcp/test_mcp_agent_365_guardrail.py, a real proxy test with Postgres, Redis, a scripted MCP upstream and local Entra and Agent 365 doubles covering the default fail-open, explicit fail-closed and fail-open, Defender Skipped, policy denial, persisted status and Prometheus counter. Use PrometheusLogger.get_instance for the fail-open metric lookup instead of a hand-rolled callback scan. Clarify the config description: gateway credential failures fail open, caller token failures block. Extract the dashboard YAML preview into teamGuardrailConfigYaml.ts so the effective per-guardrail default is unit tested and the "default" hint only shows when nothing was set explicitly. Regenerate the lazy OpenAPI snapshot and schema.d.ts for the new field. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): default a scheme-less Agent 365 authority host to https and treat a null fallback as unset in the YAML preview Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit cells for the Agent 365 fail-open default across entry points, Entra faults, throttling and two workers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): prove both Agent 365 workers serve and that a killed worker is replaced Each fresh connection reports its worker pid from /debug/memory/summary and its MCP catalog on the same connection, so the two-worker readiness wait covers both workers by identity. The kill test now kills a pid the proxy reported as a worker and waits for a replacement pid, instead of the first psutil child Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): drop the prometheus fail-open counter from the agent 365 guardrail Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): keep agent 365 fail closed by default and make fail_open an explicit opt-in Restores the shared unreachable_fallback default and the sibling guardrail initializers, drops the Admin UI YAML preview that only existed for the per-guardrail default, and reworks the unit and integration tests so the default blocks with HTTP 503 while unreachable_fallback: fail_open lets availability failures through as Unscanned Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): append authority_host after the existing Agent365Guardrail parameters Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(guardrails): agent 365 fails open by default and hides the production overrides from the UI form Agent 365 sits in the runtime path of every MCP tool call, so an Entra or Agent 365 outage now lets the call through unscanned (logged at error level, recorded as Unscanned with guardrail_failed_to_respond) instead of blocking it. unreachable_fallback: fail_closed stays as the opt-in strict mode. Policy blocks, throttling, 4xx rejections and a rejected caller token still block The shared unreachable_fallback field becomes nullable so each guardrail owns its default; every sibling still resolves None to fail_closed and typesafe keeps failing open api_base, resource_app_id and agent_id have production defaults and leave the dashboard form (ui_hidden); they stay available in config.yaml and env. The authority_host override and its env keys are gone, the OBO exchange always uses login.microsoftonline.com. The integration suite keeps only the cells that need no Entra double, the evaluation paths live in unit tests with an injected handler Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(ui): regenerate openapi snapshot and schema.d.ts for the nullable unreachable_fallback Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): fix agent 365 to the production endpoint and keep fail_closed as the default Remove api_base, resource_app_id and agent_id from the Agent 365 config model, their AGENT365_* env fallbacks and the _is_ui_hidden helper: the evaluation URL and the Agent Tools app id are fixed production constants and the agent identity is always the caller's key alias. Revert the fail_open default; unreachable_fallback: fail_open stays an explicit opt-in. Restore the shared unreachable_fallback field, the sibling guardrail initializers and typesafe to main. Move the Entra dependent cells from the subprocess integration suite to unit tests with an injected HTTP handler. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): warn when agent 365 yaml still carries the removed override keys Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): inject the http handler into the agent 365 initializer instead of assigning it after construction Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
6c34c6e7ee
|
test(integration): pin org-admin status codes in the team-admin matrix (#43592)
Run every route in the team-admin matrix again on a team that belongs to an organization, as an admin of that organization and as an admin of another organization. 90 new cases pin what org admins get today, so collapsing the team-admin helpers into one gate can prove parity for org admins too Scenario gains organization() and org_member() helpers that clean up the organization, its budget row and the member users |
||
|
|
7aba77197d
|
feat(otel): add SigNoz preset for OpenTelemetry v2 (#43296)
Some checks are pending
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
* feat(otel): add SigNoz preset for OpenTelemetry v2 Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite Absorbs the work from https://github.com/BerriAI/litellm/pull/38206 Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(otel): drop explanatory comments from the SigNoz preset Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(types): keep signoz dynamic param lines within ruff format width Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(signoz): assert the missing-endpoint boot path directly instead of in an except block Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(ui): regenerate schema.d.ts for the signoz health service Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
eea1d0f269
|
fix(responses): stream guardrail pre-call block as SSE with a typed output item (#42507)
* fix(responses): stream guardrail pre-call block as SSE with a typed output item Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): import blocked usage helper from the guardrail utils module Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): drop narrating docstrings and poll without rebinding Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover pre-call guardrail block on /v1/responses stream and json Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit cells for responses guardrail block contract Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): observe upstream on the recorded chat route for responses denial cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): tidy responses denial audit cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): wait for worker count to recover after SIGKILL Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): require a replacement worker after SIGKILL Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): type the blocked response test helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
849f3037b4
|
fix(langtrace): deliver spans to app.langtrace.ai/api/trace with x-api-key (#43322)
* test(langtrace): integration test for the built-in callback wire (path, x-api-key) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(langtrace): deliver spans to app.langtrace.ai/api/trace with x-api-key The built-in langtrace callback posted to the dead host langtrace.ai, sent the key as api_key instead of x-api-key, and let the OTLP endpoint normalizer append /v1/traces to the complete /api/trace path, so every export returned 404. Default the host to https://app.langtrace.ai, honor LANGTRACE_API_HOST for self-hosted servers, pass the key as an exporter header instead of a process-wide env var, and keep the /api/trace path unchanged for traces on the langtrace callback only Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(langtrace): keep LANGTRACE_API_HOST that already ends in /api/trace Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(langtrace): deterministic audit inventory for the built-in callback (surfaces, failures, endpoints, chaos) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(langtrace): build the repeated-request body once Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(langtrace): outage cell asserts at-most-once delivery, not span loss The OTLP HTTP exporter reposts once on ConnectionError and the batch processor may still be flushing the previous burst when the sink closes, so whether the outage burst is lost or delivered after revival depends on timing. The invariant is no duplicate and recovery on the same port Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(langtrace): delivered spans must carry the prompt in the gen_ai.content.prompt event The upstream echoes the marker into the completion, so a whole-span match alone would still pass if the prompt event disappeared Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(langtrace): disable model info refresh so the scripted upstream only sees completion requests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |