Commit graph

312 commits

Author SHA1 Message Date
devin-ai-integration[bot]
b61376a99e
feat(sdk): add run_tool_loop and arun_tool_loop helpers (#44381)
* feat(sdk): add run_tool_loop and arun_tool_loop helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(tests): allow-list bounded tool-loop recursion in recursive detector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(sdk): harden run_tool_loop per review

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(deps): keep uv.lock at revision 3

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(tests): authorize typing-extensions PSF-2.0 license

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 20:10:30 -07:00
devin-ai-integration[bot]
fe683ea139
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models

GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.

* fix(proxy): offer a Codex service tier only when every deployment of the model lists it

* fix(codex-catalog): an invalid service_tiers value offers no tier for the model

* fix(codex-catalog): read service tiers off the deployments the key's team can route to

A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them

The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default

* test(codex-catalog): drop the redundant module docstring and sort the imports

* test(integration): add the Codex catalog audit cells and the multi-worker convergence note

* test(integration): clean up every catalog test model and answer the refresh GET

* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers

Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.

* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut

The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns

The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 20:42:55 +00:00
devin-ai-integration[bot]
b024950353
feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop (#44391)
* feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(harness): name per-tool spec FunctionTool

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-03 17:33:01 +00:00
devin-ai-integration[bot]
564d236985
fix(otel): nest cache spans under their operation and name service spans by purpose (#44150)
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.

A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
yucheng-berri
5d42cb7cfa
fix(guardrails): encrypt guardrail litellm_params secrets at rest (#43627)
* fix(guardrails): encrypt guardrail litellm_params secrets at rest

* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation

- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals

* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor

- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes

* test(guardrails): drive the real guardrail rotator from the master key rotation test

- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key

* Annotate guardrail param encryption collections for type-discipline gate

* Type guardrail param recursion through validated JSON containers

* Type guardrail registry test helpers and drop section comment

* Reject client-supplied encrypted values in guardrail litellm_params

* Allow depth-bounded contains_encrypted_marker in the recursion detector

* Keep a loaded guardrail when its DB params do not decrypt with the current key

* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models

* Keep the loaded guardrail when an undecryptable param has no loaded value

* Drop suppressions the type discipline gate on main now reports as unused

* Assert what the reinitialized guardrail holds after an edit to an undecryptable one

* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize

* Type the rotation test helpers and drop the new test docstrings

* fix(guardrails): refuse to approve a submission whose params do not decrypt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:49:41 -07:00
yujonglee
677205b3f5
fix(proxy): share ownership permissions for spend logs and traces (#44239)
* refactor(proxy): extract shared spend log read policy

* test(proxy): use named bindings for spend scope regression

* test(proxy): reuse existing spend log query harness

* test(proxy): cover spend log permission lookup adoption

* chore(proxy): relocate existing spend query baseline

* refactor(proxy): make scope query returns explicit

* refactor(proxy): inject deferred log permission lookup

* test(proxy): cover teamless management compatibility lookup

* refactor(proxy): compose user and team log grants

* refactor(proxy): share generic authorization composition

* refactor(proxy): compose trace read permissions

* refactor(proxy): centralize spend and trace authorization

* refactor(proxy): strengthen spend and trace scope types

* refactor(proxy): flatten log read scope into owned logs

Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.

A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(proxy): run spend scope tests through one SQLite emulator

Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.

load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(proxy): unify log and trace ownership permissions

* test(tracing): align fixtures with ownership read scopes

* refactor(tracing): align query scopes with row ownership

* refactor(spend): make ownership SQL predicates explicit

* test(spend): validate ownership SQL against PostgreSQL

* docs(traces): drop key-row visibility from query help guide

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): reach the empty-memberships branch in team lookup test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate dashboard API types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 01:38:54 +00:00
devin-ai-integration[bot]
6c2ede00ac
test: remove 130 legacy tests owned by stronger unit proofs (#44157)
* test: remove 130 legacy tests owned by stronger unit proofs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: list router _embedding and _aembedding as covered via public embedding calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 10:18:50 -07:00
devin-ai-integration[bot]
c208306f60
fix(daily_activity): keep NULL entity ids when excluding entity ids (#44139)
The shared exclusion predicate negated a PostgreSQL ANY comparison without handling NULL, so NULL entity ids evaluated to UNKNOWN and dropped out of the Unassigned bucket whenever exclude_*_ids was set. The paginated daily rows Prisma filter had the same NOT IN shape and gets the same IS NULL OR NOT IN treatment

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 21:19:48 -07:00
devin-ai-integration[bot]
54a51c80df
feat(proxy): bounded daily activity routes for all usage entities (#43408)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 16:57:23 -07:00
devin-ai-integration[bot]
aa601ce4e8
refactor(repositories): daily activity repository with centralized bounded usage queries (#43398)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 16:02:01 -07:00
ishaan-berri
0f6a06a6b9
feat: add litellm.agent() to run claude code, codex, opencode and deep agents through the ai gateway (#43885)
* feat(harness): add litellm/__init__.py

* feat(harness): add litellm/constants.py

* feat(harness): add litellm/harness/__init__.py

* feat(harness): add litellm/harness/adapters/__init__.py

* feat(harness): add litellm/harness/adapters/base.py

* feat(harness): add litellm/harness/adapters/claude_code.py

* feat(harness): add litellm/harness/adapters/codex.py

* feat(harness): add litellm/harness/adapters/opencode.py

* feat(harness): add litellm/harness/endpoint.py

* feat(harness): add litellm/harness/errors.py

* feat(harness): add litellm/harness/options.py

* feat(harness): add litellm/harness/runtime.py

* feat(harness): add litellm/harness/sandbox/__init__.py

* feat(harness): add litellm/harness/sandbox/base.py

* feat(harness): add litellm/harness/sandbox/docker.py

* feat(harness): add litellm/harness/sandbox/local.py

* feat(harness): add litellm/harness/sandbox/snapshot.py

* feat(harness): add litellm/harness/sync.py

* feat(harness): add litellm/harness/types.py

* feat(harness): add litellm/sandbox/__init__.py

* feat(harness): add README.md

* feat(harness): add tests/harness_e2e/__init__.py

* feat(harness): add tests/harness_e2e/conftest.py

* feat(harness): add tests/harness_e2e/test_harness_e2e.py

* feat(harness): add tests/test_litellm/harness/__init__.py

* feat(harness): add tests/test_litellm/harness/adapters/__init__.py

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/test_claude_code.py

* feat(harness): add tests/test_litellm/harness/adapters/test_codex.py

* feat(harness): add tests/test_litellm/harness/adapters/test_opencode.py

* feat(harness): add tests/test_litellm/harness/core_fakes.py

* feat(harness): add tests/test_litellm/harness/sandbox/__init__.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_docker.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_local.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_snapshot.py

* feat(harness): add tests/test_litellm/harness/test_endpoint.py

* feat(harness): add tests/test_litellm/harness/test_init.py

* feat(harness): add tests/test_litellm/harness/test_runtime.py

* feat(harness): add tests/test_litellm/harness/test_sync.py

* feat(harness): add tests/test_litellm/harness/test_types.py

* test(harness): use word recall in stream e2e test

* refactor(harness): update litellm/__init__.py

* refactor(harness): update litellm/constants.py

* refactor(harness): update litellm/harness/__init__.py

* refactor(harness): remove litellm/harness/adapters/__init__.py

* refactor(harness): update litellm/harness/context.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/__init__.py

* refactor(harness): update litellm/harness/handlers/base.py

* refactor(harness): update litellm/harness/handlers/cli_handler.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/docker.py

* refactor(harness): update litellm/harness/sandbox/local.py

* refactor(harness): update litellm/harness/sync.py

* refactor(harness): update litellm/harness/types.py

* refactor(harness): update litellm/llms/base_llm/harness/__init__.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/__init__.py

* refactor(harness): update litellm/llms/claude_code/harness/__init__.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/__init__.py

* refactor(harness): update litellm/llms/codex/harness/__init__.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/__init__.py

* refactor(harness): update litellm/llms/deepagents/harness/__init__.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/__init__.py

* refactor(harness): update litellm/llms/opencode/harness/__init__.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/utils.py

* refactor(harness): update README.md

* refactor(harness): update tests/harness_e2e/conftest.py

* refactor(harness): update tests/harness_e2e/test_harness_e2e.py

* refactor(harness): remove tests/test_litellm/harness/__init__.py

* refactor(harness): remove tests/test_litellm/harness/adapters/__init__.py

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/test_claude_code.py

* refactor(harness): remove tests/test_litellm/harness/adapters/test_codex.py

* refactor(harness): remove tests/test_litellm/harness/adapters/test_opencode.py

* refactor(harness): remove tests/test_litellm/harness/core_fakes.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/__init__.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_docker.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_local.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_snapshot.py

* refactor(harness): remove tests/test_litellm/harness/test_endpoint.py

* refactor(harness): remove tests/test_litellm/harness/test_init.py

* refactor(harness): remove tests/test_litellm/harness/test_runtime.py

* refactor(harness): remove tests/test_litellm/harness/test_sync.py

* refactor(harness): remove tests/test_litellm/harness/test_types.py

* refactor(harness): update tests/unit/harness/__init__.py

* refactor(harness): update tests/unit/harness/core_fakes.py

* refactor(harness): update tests/unit/harness/handlers/__init__.py

* refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py

* refactor(harness): update tests/unit/harness/sandbox/__init__.py

* refactor(harness): update tests/unit/harness/sandbox/test_docker.py

* refactor(harness): update tests/unit/harness/sandbox/test_local.py

* refactor(harness): update tests/unit/harness/sandbox/test_snapshot.py

* refactor(harness): update tests/unit/harness/test_endpoint.py

* refactor(harness): update tests/unit/harness/test_init.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/harness/test_sync.py

* refactor(harness): update tests/unit/harness/test_types.py

* refactor(harness): update tests/unit/llms/claude_code/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/api_error.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/max_turns.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/resume_turn.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/structured_output.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/success_tools.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/codex/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/reasoning.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/structured_output.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn1_bash.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn2_resume_apply_patch.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn_failed.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/deepagents/__init__.py

* refactor(harness): update tests/unit/llms/deepagents/harness/__init__.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/opencode/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/api_error.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/endpoint_requests.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/readonly_denied_bash.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn1_write_read.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn2_session_skill.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* ci: allowlist tests/harness_e2e, which needs live runtimes and a gateway

* fix(harness): bound the turn event queue

* fix(harness): bound the turn event queue with backpressure

* style: sort imports in utils

* ci: exclude agent-harness config folders from provider docs check

* refactor(harness): update tests/harness_e2e/conftest.py

* refactor(harness): update tests/harness_e2e/test_harness_e2e.py

* refactor(harness): update tests/unit/harness/core_fakes.py

* refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/harness/test_sync.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/options.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/snapshot.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/sandbox/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update tests/code_coverage_tests/recursive_detector.py

* refactor(harness): update tests/unit/llms/base_llm/harness/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/__init__.py

* refactor(harness): update litellm/harness/__init__.py

* refactor(harness): update litellm/harness/context.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/__init__.py

* refactor(harness): update litellm/harness/handlers/base.py

* refactor(harness): update litellm/harness/handlers/cli_handler.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/__init__.py

* refactor(harness): update litellm/harness/sandbox/base.py

* refactor(harness): update litellm/harness/sandbox/docker.py

* refactor(harness): update litellm/harness/sandbox/local.py

* refactor(harness): update litellm/harness/sandbox/snapshot.py

* refactor(harness): update litellm/harness/sync.py

* refactor(harness): update litellm/harness/types.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/types/llms/custom_http.py

* refactor(harness): update tests/unit/harness/test_endpoint.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update tests/unit/harness/test_init.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py
2026-10-01 22:27:49 +00:00
devin-ai-integration[bot]
c66c8288c3
fix(proxy-extras): build the SpendLogs indexes in the migration job instead of in migrations (#43948)
The two SpendLogs index migrations shipped in v1.103.0 each break one table shape: the plain CREATE INDEX holds a SHARE lock on a large unpartitioned table and the CONCURRENTLY one fails with 0A000 on a partitioned parent. Both files are now inert and the indexes are built by a table-driven, shape-aware step after migrate deploy: CONCURRENTLY on a plain table, ON ONLY the parent plus per-partition CONCURRENTLY and ATTACH PARTITION on a partitioned one. The migration job builds them synchronously and exits non-zero on failure; a serving proxy that ran migrate deploy itself builds them in the background off the readiness path. A valid index of the same definition under another name is renamed and reused, an invalid one is rebuilt, and extra copies are reported with their DROP INDEX statement instead of being dropped. The migration checker rejects any CREATE INDEX on LiteLLM_SpendLogs or LiteLLM_ErrorLogs in future migrations

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 14:12:22 -07:00
yuneng-jiang
e725832cd8
chore(deps): drop unused pytest-postgresql dev dependency (#44056)
The pytest-postgresql based proxy tests moved to tests/integration on the real
Postgres harness in #43996, so nothing loads the plugin anymore. uv.lock is
edited by hand to drop the package and its now orphaned mirakuru and port-for
deps; uv lock --check passes and a full relock resolves the same package set
2026-10-01 18:53:54 +00:00
moe-berri
259c166ef6
refactor(lens)!: rename internal engine code and API (#44034)
* chore(lens): remove deployment screenshots

* refactor(lens)!: rename internal engine package and API

* fix(lens): pin worker image for renamed API

* test(lens): cover fresh and populated rename migrations

* fix(lens): protect db-push upgrades and restore routing and CI

* fix(lens): resolve migration tables across schemas and include database driver
2026-10-01 11:02:14 -07:00
ryan-crabbe-berri
ae05f7d2c1
test(e2e): typed per-test metadata for the e2e suite (#42044)
* feat(e2e): give e2e tests typed metadata for what they drive

@meta(Subject(domain, route, providers, models, capabilities, mode)) declares
what a test is about with closed enums, and each field lands in the JUnit report
as a property. The quota_management suites are the first to declare it.

* docs(e2e): say e2e_metadata avoids litellm, not that it is stdlib-only

It already imports pydantic and pytest, both of which the suite needs to collect. The rule that matters is no litellm import

* test(e2e): declare models through the constant each test drives

43 @meta declarations in quota_management typed the model name out again, so changing the call would leave the coverage report naming the old model. Each file now has one constant used by both, and a guard fails on any model written as a string literal in @meta

* refactor(e2e): set route only when the endpoint is what the test checks

A budget or rate-limit test whose chat call only triggers the block now leaves route unset, since its steps already name the call. Tests of an endpoint keep it: budget CRUD, key creation, spend reporting reads, and the per-endpoint spend tests for chat, messages, embeddings, batches and health. The two /spend/logs tests tagged chat_completions are now spend_reporting

* refactor(e2e): build the declared properties without mutating a list

subject_properties seeded a list and grew it with append and extend. It now flattens one tuple per field, and the plural-name table is a read-only mapping

* fix(e2e): tag each spend-route probe with the endpoint it checks

The breadth test gave all 33 probes spend_reporting, so /key/list, /user/list, /team/list, /organization/list and /customer/list counted as spend reporting. Each case now carries its own route, with organization and customer management added to Route
2026-09-30 21:03:21 -07:00
ryan-crabbe-berri
424bfd8758
feat(e2e): record each e2e test's steps, starting with ProxyClient (#42393)
* feat(e2e): record each e2e test's steps, starting with ProxyClient

@step on a harness method records a plain-English line for every call, in
order, as repeated JUnit step properties. Labels are templates filled from the
call's parameters, like "Generate a virtual key with models: claude-haiku-4-5
and rpm limit: 3", and secret request fields are marked Field(repr=False) so
they never print. ProxyClient and the rate-limit QuotaClient carry steps first;
the other harnesses follow one area at a time. The recorder and JUnit tests run
in the Code Quality workflow's test_e2e_metadata step.

* docs(e2e): rewrite the recorded test steps guide in plain language

* fix(e2e): keep logging callback credentials out of recorded steps

* fix(e2e): mask the run's credentials in every recorded step

* fix(e2e): attach steps before the oauth failure snapshot

The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach

* fix(e2e): name the saved credential in its recorded step

The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
2026-09-30 19:33:53 -07:00
moe-berri
6fd9334751
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main

Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.

* feat(lens): add trace analysis and standalone worker

* fix(lens): clarify review limits and finalize main integration

* fix(lens): simplify worker setup and show the next check

* fix(lens): simplify analyzer setup and resolve integration failures

* fix(lens): preserve durations and evidence from later trace reads

* fix(lens): trust server context for internal analysis exclusion

* fix(lens): pin reviewed analyzer image and verify request inclusion

* test(lens): select time units before entering custom duration

* test(lens): allow the standalone analyzer lifetime HTTP client

* test(lens): run analyzer tests in active proxy coverage shard
2026-09-30 22:42:09 +00:00
devin-ai-integration[bot]
ffb15f946f
perf(proxy): one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads (#43407)
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.

The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)

Resolves LIT-8882

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 16:42:13 -07:00
devin-ai-integration[bot]
e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
Yassin Kortam
7e383c9f6a
revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377)
* Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)"

* revert: "feat(usage): search keys beyond the top-N usage subset (#42827)" (#43378)

* Revert "feat(usage): search keys beyond the top-N usage subset (#42827)"

* revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595)

* Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table"

* Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596)

Co-authored-by: yassin <yassin@berri.ai>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:47:46 +00:00
ryan-crabbe-berri
96c008f420
ci: fail on new unbounded SQL IN lists and add a Prisma chunking helper (#42629)
* ci: warn on SQL IN lists with no written bound

Postgres caps a prepared statement at 32,767 bind parameters and a
membership filter binds one per value, so an IN list built from table
data breaks once the table outgrows the cap. That is how the budget reset
job froze every due budget (LIT-7535, #40564).

check_unbounded_in_lists.py reports every Prisma "in" / "not_in" filter
whose value has no fixed size and every raw SQL literal that splices a
list in after "IN (", unless the line carries "# bounded-ok: <reason>".
It only warns for now: the output is the inventory for RCA action item
AI-1, and it exits 0.

* ci: decide a constant IN list by its module binding, not its casing

An ALL_CAPS name imported or filled at runtime is as unbounded as any
other, so a name now passes only when the module binds it once to a
value of fixed size. Adds Final to the locals a loop does not forbid.

* ci: only a frozen module value makes an IN list constant

A module list bound once could still grow through append or extend, so
a name now counts as fixed only when it is bound to a tuple, frozenset
or constant. Trims the module docstring to what a reader needs.

* ci: chunk Prisma IN lists with a shared helper and fail on new unbounded ones

Add litellm.repositories.bounded_in: find_many_in, count_in, update_many_in
and delete_many_in split a deduplicated value list into 5,000-value chunks,
AND each chunk with the caller's where, run them in order (a transaction
handle works) and combine the results. Writes take a required atomicity
argument, and a where that already filters the chunked field is refused.

check_unbounded_in_lists.py now fails CI on any finding missing from
unbounded_in_baseline.txt and on any stale baseline entry, so the baseline
only shrinks. Entries are keyed by path, enclosing scope, kind, field and
occurrence, not line numbers. The helper module is exempt, a constant
spread into a frozen tuple counts as fixed, and messages point at the
helper for "in" and at an array parameter for "not_in" and raw SQL.

A real-Postgres integration test shows a raw 40,000-value filter rejected
for too many bind variables while the helpers handle it.

* refactor: rename bounded_in to chunked_in and let callers pick a chunk size

The helper module is litellm.repositories.chunked_in, and its unit and
integration tests, the checker's exemption path and its finding messages
follow the new name. The `# bounded-ok` marker is unchanged.

find_many_in, count_in, update_many_in and delete_many_in take a
keyword-only chunk_size, defaulting to IN_LIST_CHUNK_SIZE (5,000). A value
below 1 or above MAX_IN_LIST_CHUNK_SIZE (30,000) raises ValueError before
any query, which leaves the rest of the filter headroom under Postgres's
32,767 bind-parameter cap.

* refactor: flatten chunked_in's stacked comprehensions with chain.from_iterable

LIT014 (#42650) caps a comprehension at one for and one if clause. The four nested walks in the helper now chain their iterables instead, with the same order and results.

* refactor: recover user details with find_many_in, sending chunks as lists

_details_for_user_ids reads users through find_many_in instead of a raw
"in" filter, so its lookup stays under the bind-parameter cap for any
number of recovered keys. Up to 5,000 ids it still sends one find_many
with the same where dict, and a PrismaError from any chunk is still
logged and treated as no details.

The helper now sends each chunk as a list, so a chunked filter equals
the dict a hand-written call would send and a migrated call site's
existing assertions keep passing.

The site's baseline entry is gone.

* ci: skip functional TypedDict field maps in the unbounded IN list check

The dict passed as the field map of TypedDict("Name", {...}), or as its fields= keyword, names fields: an "in" or "notIn" key there is a type, not a filter. Only that dict is skipped, for TypedDict, typing.TypedDict and typing_extensions.TypedDict; a filter nested in a field value or passed to any other call is still reported. The two types/proxy/management_endpoints/team_endpoints.py entries leave the baseline, which is now 156.

* fix: refuse an update_many_in whose data writes the chunked field

Chunks run one after another, so an update that sets the chunked field can move a row into a later chunk, which updates it again and counts it twice: values ["old", "new"] with chunk_size=1 and data={"id": "new"} does exactly that. update_many_in now raises ChunkedFieldWriteError before any query when data has the chunked field as a top-level key, in any form, including Prisma operators such as {"set": ...}.

* docs: cut the unbounded IN list checker's docstring to what it flags and how to clear it

It now says what is reported, the three ways to clear a finding, and how the baseline and --update-baseline work, in 11 lines. The per-shape detail lives in the tests.

* ci: key an unbounded IN list finding by its filtered expression too

A baseline key of path, scope, kind, field and occurrence let a PR delete
a baselined filter and add a different unbounded one on the same field in
the same function, and the new one took over the old key. The key now
also carries the filtered expression's source, whitespace-normalized
(the Prisma value, or a raw-SQL `IN (...)` slot), so that swap reads as
one new and one stale entry and fails the run. The same expression
re-added in the same function is still the same finding.

Every baseline entry is rewritten in the new form; the 156 findings are
unchanged, and only occurrence indexes renumber where one field had
several different expressions.
2026-09-26 13:40:44 -07:00
devin-ai-integration[bot]
a09f8b84a4
fix(sentry): scrub PII and secrets inside object reprs and nested locals, add SENTRY_SEND_DEFAULT_PII opt-in (#43123)
* fix(sentry): scrub PII and secrets inside object reprs and nested locals, add SENTRY_SEND_DEFAULT_PII opt-in

* fix(sentry): keep the SDK denylist and filter the request headers a virtual key arrives in

* fix(sentry): leave source context lines unscrubbed

* fix(sentry): filter bracketed secret values and cap the JSON walk depth

* ci(deps): install sentry-sdk in the proxy-dev group so the unit shards import it

* fix(sentry): scrub source-context names outside real stack frames

* fix(sentry): tie the key pattern floor to the custom key minimum

* fix(sentry): keep the key pattern floor at or below a generated key's length

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-25 14:52:35 -07:00
yuneng-jiang
f6882246d4
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 11:30:43 -07:00
devin-ai-integration[bot]
1987133b4e
fix(router): match provider-prefixed fallback keys for bare model groups served by wildcard deployments (#43062)
* fix(router): match provider-prefixed fallback keys for bare model groups

* fix(router): infer the fallback key's provider the way routing does for bare model groups

A bare model group served by a wildcard deployment (claude-sonnet-4-6 routed to anthropic/*) now finds a fallback keyed <provider>/<group>. The provider is inferred through one shared helper, inferred_provider, which the pattern router already used inline, so the fallback lookup and routing agree on the prefix. The lookup only infers a provider when some fallback key ends in /<group>, so alias-style groups never hit the resolver

* fix(router): resolve context window and content policy fallback keys through the shared lookup

---------

Co-authored-by: Jason Dougherty <jasondoc3@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 18:08:56 -07:00
devin-ai-integration[bot]
248f0eb159
ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests (#42903)
* ci: fix the litellm-tests unit job with sysmon coverage, an env allowlist and coverage upload on failure

* test: replace key-dependent proxy, enterprise and mcp unit tests with synthetic values and integration and e2e coverage

* test: drop key reads at the legacy proxy, enterprise and mcp paths and wire the gemini pass-through split

* ci: move caching, proxy-extras, gateway and enterprise tests into tests/unit and run them from litellm-tests under their legacy flags

* ci: move caching, proxy-extras, gateway and enterprise tests into tests/unit and run them from litellm-tests under their legacy flags

* ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests

* ci: fail the unit shard when circleci tests split errors

* test: drop restating comments from the gemini pass-through split

* build: point the local proxy unit targets at the nested tests/unit/proxy tree

* ci: exit the unit shard cleanly when circleci tests split assigns it no files

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-09-24 22:59:11 +00:00
devin-ai-integration[bot]
b0407ad33e
ci: add merge smoke checks workflow with loopback-only harness and 11 curated cases (#42709)
* ci: add dashboard and core smoke checks across supported Python versions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: tighten merge smoke harness and keep mapped test diffs additive

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: terminate proxy on readiness timeout and use contextlib.suppress in teardown

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-23 11:01:08 -07:00
devin-ai-integration[bot]
b0ac23d385
feat(logger): dispatch Python logging through the Rust diagnostics processor (#42616)
* feat(logger): add shared Rust diagnostics and Python logging bridge

* feat(logger): dispatch diagnostic processing through Rust

* chore: regenerate Cargo.lock after rebase

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: allowlist bounded logging tree walkers in recursive detector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(logger): skip decoding plain access arguments

* test(logger): skip embedded-python logger test when litellm deps are absent

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: cargo fmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: expect NativeDiagnosticProcessor in the native public surface

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(stub): export NativeDiagnosticProcessor via __new__ in _native.pyi

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): rename logger crate and document host sink contract

* test(logger): cover exc, stack, and nested extras in the diagnostic filter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logger): keep rendered redacted line when template scan flags a key pattern

The blanket REDACTED for a changed msg/color template discarded lines
whose rendered form was already redacted by the same pipeline, e.g.
'password=%s' became 'REDACTED' instead of 'password=REDACTED'. Only
fall back to REDACTED when the rendered form did not change either,
which is where interpolation can mangle the key pattern the scrub
would otherwise see.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): install python deps so the logger bridge test runs

The end-to-end bridge test skipped silently when litellm's Python deps
were absent. uv sync --no-install-project installs them without a
maturin build, and PYTHONPATH makes them visible to the embedded
interpreter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 18:44:15 -07:00
devin-ai-integration[bot]
b682278aa9
ci(code-quality): allowlist _render_json in the recursive detector (#42442)
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 01:05:58 -07:00
devin-ai-integration[bot]
77d656a8ba
fix(e2e-stack): print add-mask lines only under GitHub Actions (#42423)
* fix(e2e-stack): print add-mask lines only under GitHub Actions

* refactor(e2e-stack): inline the add-mask lines into main

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-21 22:42:13 -07:00
devin-ai-integration[bot]
13b374873d
fix(otel v2): map completions, images, speech, transcription and moderation output onto the Langfuse generation output (#42394)
* fix(otel v2): map completions, images, speech, transcription and moderation output onto the generation output

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(redaction): redact text completion choices in the standard logging payload

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): compare decoded generation output text and follow the live moderation verdict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): compare logged byte counts with the received media and move e2e schemas into models.py

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(e2e): keep the otel_v2 Langfuse output e2e file out of the stage-mirror gate it cannot run in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:57:44 -07:00
Mateo Wang
4b2e96a5f5
Merge pull request #42143 from BerriAI/litellm_e2e_changed_keep_pytest_log
ci(e2e): fix the stage-mirror batch reds and keep a redacted pytest log
2026-09-21 11:39:25 -07:00
Yuneng Jiang
e34fd201b3
Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-19 23:59:26 -07:00
mateo-berri
468f74c628 ci(e2e): fix the stage-mirror batch reds and keep a redacted pytest log
The changed-test gate booted its stage-mirror stack without files_settings
or finetune_settings, so every raw upload with a custom_llm_provider hit a
500, and it exported the whole provider env into the gateways, so the
AWS_ROLE_NAME the assume-role test needs made the GovCloud deployment run
an AssumeRole with its static keys. The gate also deleted its pytest output,
so a red run left nothing to read. The mirror config now carries the
openai, azure, and vertex_ai file settings, gateways start without
AWS_ROLE_NAME, and the workflow uploads the pass logs and junit files with
every secret value, every field of a JSON-valued secret, and their
XML-escaped forms replaced before the raw files are removed.
2026-09-19 22:37:07 -07:00
joshua-berri
8df260a13d
Merge pull request #42051 from BerriAI/litellm_mcp_oauth_e2e_3467_rework
test(e2e): restore MCP OAuth happy-path coverage (LIT-3467)
2026-09-20 02:50:14 +00:00
ryan-crabbe-berri
a6c51ba3de fix(proxy): never treat plaintext that base64-decodes to nothing as a ciphertext during the master key migration
A string such as "*" or "..." has no base64 characters, so it decoded to no bytes and read as an empty plaintext under any key. The migration would have counted it and overwritten it with a ciphertext of the empty string. Also read from the writer database instead of a read replica, report a database error during the migration instead of crashing the boot, skip columns the connected schema lacks across every schema on the search path, cap the JSON walk depth for the recursion detector, and move the boot wiring into one tested function.
2026-09-19 17:53:17 -07:00
Joshua Valluru
368401e85c test(e2e): complete OAuth triggers and preserve failure diagnostics 2026-09-19 17:03:38 -07:00
Joshua Valluru
64452f76c2 test(e2e): restore LIT-3467 implementation for rework 2026-09-19 16:21:53 -07:00
Mateo Wang
af6a1798e2
Revert "test(e2e): cover MCP OAuth SSO and cold restart persistence" 2026-09-19 16:15:43 -07:00
Joshua Valluru
d5ac850feb test(e2e): isolate diagnostic reporter subprocess 2026-09-19 15:13:27 -07:00
Joshua Valluru
b7bab56d4d test(e2e): report safe OAuth failure locations 2026-09-19 12:52:28 -07:00
Joshua Valluru
fb56a14cd4 chore(mcp): merge main with unit test timeout safeguards 2026-09-19 09:42:08 -07:00
mateo-berri
aceae8e566 test: drop the recursive detector allowlist entry for the removed _walk_payload 2026-09-19 04:07:07 -07:00
Joshua Valluru
f5ab563499 fix(mcp): preserve session expiry signals and scope dependency CI 2026-09-18 22:52:10 -07:00
Yuneng Jiang
62f6ee9a16
Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-18 20:55:05 -07:00
Yujong Lee
bf7d1c0733 chore: consolidate CLAUDE.md into AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 02:30:35 +00:00
joshua
4bc3f1d0fc build(deps): migrate MCP integration to MCP SDK 2.2.0
Replace the bespoke dependency-install CI gate with a real migration:
require mcp>=2.2.0,<3 alongside httpx2>=2.5.0,<3 and pydantic>=2.12.0,<3
in the proxy and mcp extras, drop langchain-mcp-adapters (pins mcp<2)
from the dev group, and remove the dependency-install workflow and
tests/mcp_dependency_tests that only exercised the old pins.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 22:12:40 +00:00
Yuneng Jiang
a0a006f248
fix(e2e): own a shared fixture's deployment by the fixture's node, not the first test
A deployment registered while a module- or class-scoped fixture is being set up
was bound to whichever test asked for the fixture first, so every later test in
the module shared that partition. A session-scoped fixture is set up by every
xdist worker, so its deployment could never have one owner at all.

The e2e conftest now wraps pytest_fixture_setup and records the node the fixture
is scoped to: registrations made during a module or class fixture's setup carry
that node's slug, and a session- or package-scoped one has no owner and stays
live. The registration seam test moves from tests/e2e to the cache harness tests
beside the rest of the attribution coverage.
2026-09-16 17:35:05 -07:00
Yuneng Jiang
c63d0e6922
fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process
The cache edge keyed every recording on its own process's PYTEST_CURRENT_TEST.
Under xdist that names whatever test the serving worker is in, which is
unrelated to the caller: the proxy is a separate pod, and the Claude Code compat
matrix registered its shared aliases from every worker, each pointing at that
worker's edge, so the router spread one worker's calls across all eight edges.
Builds 234 and 235 of litellm-e2e, same commit, credited the same Bedrock
request to unrelated tests 92% of the time, and Bedrock never converged past a
~20% hit rate while OpenAI, whose deployments are per test, sat at 90%.

A deployment registered from inside a test now carries its test's slug in the
edge URL it is pointed at, `{edge}/{mount}/t/{slug}`, and the edge reads that
segment off every request before forwarding. A request without one is forwarded
live and never cached, and the edge no longer falls back to process state. The
compat aliases are registered with provider_live=True and stay on their real
provider path: no single test owns them, and the matrix exists to prove the real
CLI against real providers.
2026-09-16 17:08:17 -07:00
Yuneng Jiang
99545b5f26
Merge remote-tracking branch 'origin/main' into litellm_/buildkite-litellm-e2e-setup-ff714d 2026-09-16 15:15:22 -07:00
yassin
4513f78df5 fix(migrations-check): read the table name past comments, ignore referential SET DEFAULT
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:14:36 +00:00