* fix(caching): keep tool calls and tool results in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): keep semantic tool prompt helpers within lint budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): keep structured function_call_output text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split Responses text-field collection to stay within complexity budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): tag each tool result with the position of the call it answers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): encode tool result position and output together so tool text cannot forge result tags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): expect encoded tool result record in qdrant semantic prompt parity case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): cover tool result arrangements, SDK clients, concurrency and qdrant outage for semantic cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(caching): embed every semantic cache prompt field except volatile ones
Replace the per-shape allowlist in the Python and Rust semantic cache prompt
walkers with one include-by-default walker. Plain text keeps its old
concatenation; any other block or message is embedded as compact JSON with
call ids mapped to ordinals, cache_control dropped, and signatures, encrypted
content and base64 data replaced with a short sha256 digest.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts"
This reverts commit 86c82b949b.
* Revert "test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check"
This reverts commit 39efb5d9da.
* Revert "feat(caching): embed every semantic cache prompt field except volatile ones"
This reverts commit 5aed3ab3de.
* refactor(caching): rename get_str_from_messages_with_tools to get_semantic_cache_prompt_from_messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split semantic cache prompt extraction by API format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): drop TypeIs guard and register Responses prompt walker with the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): pick the Responses text field without a Final inside a loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): walk semantic cache prompts as plain dicts, dumping pydantic items once up front
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): write the semantic cache prompt builders as plain loops
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): drop formatting-only churn from the redis semantic cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(integration): drop the caching group wiring that main already carries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): recurse into tool_result content in the semantic cache prompt helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): check the max_messages cap on the shared exact-cache proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): read list-form function_call_output text in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): extract nested Responses input lookup to keep walker under complexity limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "refactor(caching): extract nested Responses input lookup to keep walker under complexity limit"
This reverts commit 0665296bf1.
* style(caching): suppress C901 on the Responses input walker instead of splitting it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag guardrails and logging tests with Subject metadata and record client steps
* test(e2e): leave the guardrails and logging harness unit tests untagged
* test(e2e): let the inner create_model step name the guardrail backend deployment
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): declare the default guardrail backend model on the tests that drive it
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag management tests with Subject metadata and record management client steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): keep the prompt out of the chat_status step so polled retries collapse
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag claude_code tests with Subject metadata and record CLI driver steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): decorate run_claude directly so the label gate discovers its step
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* test(e2e): tag llm_translation tests with Subject metadata and record harness steps
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* test(e2e): declare the realtime param tuples Final
Add litellm_settings.force_redis_hash_tag_grouping so Redis endpoints that enforce cluster slot rules behind a standalone protocol (Redis Enterprise clustering policy) group multi-key scripts by slot like RedisClusterCache, and return grouped batch values in the caller's key order.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Chenglun Hu <chenglunhu@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ollama): turn streamed prompt-based JSON tool calls into real tool calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ollama): separate replayed tool calls from text and tighten parser typing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(health): attribute background health check results to their own deployment
Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(health): annotate locals with Final and split nested comprehension
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Dennis Pfisterer <302635+pfisterer@users.noreply.github.com>
Co-authored-by: Suhas Hanamannavar <hanamannavarsuhas17@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(auth): clear the recent-miss user memo when /user/new creates the user
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(auth): pin the miss memo window and drop the class patch in the new_user regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): pin a first-time SSO user's first message and second sign-in on one worker
* test(integration): cover /user/new clearing the user-miss memo on the creating worker
The fake IdP now signs in the subject a login_hint names, so cells on the shared one-worker proxy can each use a fresh user. New cells: a plain-key miss followed by /user/new is budgeted at once (same on both legs, the auth prefetch loads the row) and the admin-created user's first SSO sign-in inside the window completes (500 at the callback before the fix). The first-sign-in cell moved onto the shared one-worker proxy fixture
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* refactor(types): replace Any with proven types in 16 files
* fix(types): import TypedDict from typing_extensions for pydantic on 3.10
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject the cleanup job stagger offset so the runtime sync test no longer depends on the runner's host and pid
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(policy_engine): count resolver calls instead of timing them so the linear dedup guard is deterministic under CI load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(policy_engine): annotate the line-count guard's locals as Final
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: give installing_litellm jobs their own Postgres
* ci: give the entrypoint jobs their own Postgres sidecar
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ui): render team_metadata_schema keys as fixed labels
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ui): update team metadata schema tests for fixed labels
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): derive team metadata schema labels from live key values
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ci): stop deferred pydantic builds leaking caller locals and add missing Lens FK migration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(types): drop narrating comment from caller-locals regression test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): guard the Lens review FK migration with DO blocks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(types): run the caller-locals regression in-process
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy-extras): scope the Lens review FK guards to their table
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): skip guardrail VAD session.update injection for transcription sessions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(realtime): flag transcription sessions from the route intent and backend events only
A client session.update declaring session.type transcription on a voice
session no longer sets the transcription flag, so it cannot switch off the
guardrail's create_response gate or skip the transcript guardrail
* fix(realtime): flag transcription sessions from provider-transformed session events
* test(integration): cover transcription sessions skipping the VAD auto-response injection
Adds the realtime transcript guardrail audit cells: transcription sessions on all three
realtime routes, the OpenAI SDK, the beta protocol, the whisper default deployment, the Azure
GA path over a TLS scripted upstream, and Meta Muse push-to-talk sessions keep the client's
session.update verbatim and get their transcript, while voice sessions keep the injected
create_response gate and a client-declared transcription type no longer bypasses it. Sad,
edge, and chaos cells cover malformed session fields, duplicate and older backend session
events, unauthenticated upgrades, repeated sessions, an upstream outage under open sessions,
and a worker kill with a proxy restart.
The scripted upstream now answers session.update the way the vendor does (session.updated,
or the missing turn_detection.type and session-type errors), records every websocket frame,
serves the Azure and Muse realtime paths, and can run over TLS from an owned upstream.
---------
Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(mcp): run the static root path issuer discovery test in-process
The test spawned a fresh interpreter with a 60 s deadline to import litellm
and the MCP discovery router cold, so on a loaded box it died with
subprocess.TimeoutExpired before any assertion ran. The discovery routes
bake SERVER_ROOT_PATH into their paths when the module executes, so the
test now reloads that one module under the gateway env, restores its
namespace afterwards, and asserts on the same four discovery documents.
PROXY_BASE_URL now names an origin distinct from the test client's, so the
assertions fail when it stops being honored.
* test(mcp): type the gateway discovery fixture and its test
* test(mcp): mark the registry fill the fixture hands the test
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test(bedrock): live e2e asserting nova sonic realtime delivers each assistant sentence once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): drop ticket reference from nova sonic e2e docstring
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): forward each Nova Sonic assistant sentence once over the realtime API
Nova 2 Sonic sends every assistant text block twice, a SPECULATIVE preview
next to the audio and a FINAL transcript once the audio turn has ended. The
realtime bridge forwarded both, so voice clients rendered each sentence twice
and the FINAL copies opened extra responses after response.done, the last of
which never closed. FINAL assistant text blocks are now dropped whole, so a
turn carries each sentence once inside the one response with its audio
* test(bedrock): tag the live Nova Sonic test and type its helpers
* test(bedrock): pin that the Nova Sonic barge-in marker is dropped with its FINAL block
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* test: delete unconditionally skipped legacy tests
* test: move whole-unit legacy test files into tests/unit
* test: keep moved legacy tests free of import-time global state
* test: keep the module-level invocation scan pointed at tests/local_testing
* ci: drop the agent_testing CircleCI job emptied by the move
* test: fix moved-test isolation and router coverage
* test: add Tinyfish search package marker
* test: isolate moved tests from logger state leaks
* test: isolate Helicone logging fixture state
* test: isolate Vertex pass-through credentials between moved tests
* test: cancel S3 periodic flush tasks started by moved tests
* ci: restore CircleCI assistant test selection after move
* test: prevent Datadog datetime import shadowing
* ci: drop the litellm_assistants_api_testing CircleCI job emptied by the move
* test: deduplicate imports in rebased unit tests
* test: remove duplicate passthrough router patch import
* test: remove moved legacy source files after rebase
* test: align moved tests with rebased main
* test: carry main's legacy-file edits into moved destinations
* test: make the moved cost map fallback tests assert the fetch and the backup
The four fallback cases only checked the result was non-empty, so they still
passed with integrity validation disabled. They now inject a mock client, assert
one fetch happened, and assert the result is exactly the local backup with the
fallback reason recorded.
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* feat(claude_code_gateway): issue rotating refresh tokens and a revocation endpoint
The Claude Code gateway's device-code grant now returns a refresh token
alongside the session JWT, so a Claude Code session renews itself before
the JWT expires instead of forcing the user back through the browser sign-in.
grant_type=refresh_token re-mints the JWT from the live user row, rotates the
refresh token, and refuses a replayed, foreign, or identity-only token with
invalid_grant. The discovery document now advertises an RFC 7009 revocation
endpoint, which Claude Code's /logout calls with both tokens, so sign-out
burns the refresh token. The refresh token is the MCP gateway's sealed
session refresh token bound to the fixed client id claude_code, so both
front doors share one single-use record.
* fix(claude_code_gateway): mint the whole credential before claiming the device code
A refresh token that failed to mint answered 500 after the device code was
already claimed and the login deleted, so the client could not redeem the
completed sign-in again. The response is now built first, and the code is
claimed only when it is a 200.
* fix(claude-code-gateway): keep device sign-in working when session signing is unusable
* feat(claude-code-gateway): end the whole refresh chain on a replay or a revocation
* fix(mcp-gateway): fail the single-use peek closed on a Redis fault
* fix(mcp-gateway): read the single-use marker under the cache namespace
* fix(mcp-gateway): refuse a replayed refresh token without ending its chain
A refresh token presented a second time is refused as already used and
nothing else happens to the chain it was rotated from. Claude Code renews
from its in-memory copy of the credential, so a second terminal on the
same machine presents the token the first terminal already rotated and
then recovers from the shared credential file; ending the chain there
would sign both terminals out at every expiry. Only a revocation ends a
chain, so the 503-on-unrecorded-chain-ending path and its tests go away.
* fix(mcp-gateway): end the refresh chain before burning the revoked token
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(exceptions): keep upstream 402 status and cool down 402 deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(exceptions): map 402 to PaymentRequiredError subclass of BadRequestError
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(exceptions): annotate PaymentRequiredError methods
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): skip 402 cooldown on single-deployment model groups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(exceptions): single prefix and 402 fallback response for PaymentRequiredError
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic): map billing_error to PaymentRequiredError regardless of status
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(router): honor explicit allowed-fails policy for single-deployment 402s
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(anthropic): drop billing_error body mapping to PaymentRequiredError
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(exceptions): type the PaymentRequiredError constructor parameters
* test(integration): cover 402 PaymentRequiredError mapping and cooldown
---------
Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
The migration_startup_tests job runs pytest over tests/e2e/migrations with
PYTHONPATH=tests/e2e, so it loads tests/e2e/conftest.py, which has required
LITELLM_MASTER_KEY at import since #44718. That PR gave every other e2e job a
"Generate LiteLLM master key" step but not this one, so all five scheduled
migration jobs died before collection with KeyError: 'LITELLM_MASTER_KEY'.
Give the job the same step the other jobs have.
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(mcp): refresh server catalog for each gateway operation
* fix(mcp): reject configuration changes during scoped dispatch
* fix(mcp): refresh shared catalog state for each operation
* fix(mcp): refresh catalog before native alias routing
* docs(mcp): clarify native route resolution order
* test(mcp): align streaming fixture and generated API documentation
* fix(mcp): preserve discovery published during catalog refresh
* fix(mcp): coalesce queued catalog reads without a stale window
* fix(mcp): reconcile concurrent route changes when publishing catalog
* test(mcp): provide catalog scope in post-call hook fixtures
* fix(mcp): retain valid live routes and handlers during refresh
* fix(mcp): keep refreshed OpenAPI operation membership authoritative
* test(mcp): preserve logging fixtures after catalog integration
* fix(mcp): isolate transport test admission state and exhaust OAuth outcomes
* refactor(mcp): narrow catalog consistency change to ticket scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(mcp): gate catalog refresh on a database revision marker
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): refresh waiters that observed a newer catalog revision
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(mcp): awaitable catalog revision doubles in prisma mocks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(mcp): satisfy catalog refresh lint and type gates
* test(mcp): provide catalog scope in toolset fixtures
* test(mcp): exercise temporary OAuth through catalog operations
* fix(mcp): share active catalog snapshots across discovery tasks
* fix(mcp): retain discovered routes across cached catalog operations
* fix(mcp): preserve local tool ownership across discovery
* refactor(mcp): satisfy tightened immutability lint ceiling
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(mcp): authorize local handlers by registered server ownership
* fix(mcp): check registered handler freshness and isolate fixtures
* fix(mcp): read catalog revisions and snapshots from the writer
* fix(mcp): resolve access groups from the operation catalog
* fix(mcp): preserve empty access group restrictions
* fix(mcp): retain rediscovered routes across concurrent updates
* fix(mcp): enforce registered ownership for local tool dispatch
* fix(mcp): preserve issuer discovery across catalog snapshots
---------
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): isolate Codex catalog and provider discovery fixtures
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): keep model discovery constant import-safe
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test: align test keys and CI env with generated master keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata
* docs(e2e): name every markerless harness test file that carries no Subject
* test(e2e): keep the step discovery comprehensions to one for clause
* feat(terraform): expose key type on virtual keys
* docs(terraform): remove in-tree key type docs
* fix(terraform): preserve server-derived key routes
* fix(terraform): keep unconfigured key routes plan-known and unsent
Two regressions from exposing key_type on litellm_key:
1. Marking allowed_routes Computed makes an omitted attribute unknown at
plan time ("known only after apply"), so any plan that consumes it
before the key exists fails, e.g.
for_each = toset(coalesce(litellm_key.x.allowed_routes, [])).
Computed is dropped again; server-derived routes still land in state
through reads, and a DiffSuppressFunc keyed on the raw config keeps a
config that never declares the attribute from showing a perpetual
removal diff against those routes (a config that shrinks the list or
sets it still diffs).
2. mapResourceDataToKey copies allowed_routes unconditionally and
UpdateKey sends it when non-empty, so once reads materialize the
server's routes into state, every update re-asserts them: an
alias-only rename POSTs allowed_routes (the pre-key_type provider
sent none), and with stale state (-refresh=false) it silently
overwrites routes managed outside Terraform. Updates now omit the
field whenever the raw config does not declare it.
The key_type flow is unchanged: create still sends key_type, the proxy
presets the routes, reads materialize them into state, and plans stay
drift-free.
* fix(terraform): reject allowed_routes alongside a presetting key_type
The proxy derives allowed_routes from the key_type preset and overwrites
whatever the request declared, so a config combining the two could never
match what gets stored: the key came back with the preset routes and
drifted against the declared list on every plan. A CustomizeDiff now
fails the plan with an actionable message when a presetting key_type
(llm_api, management, read_only) is combined with allowed_routes.
key_type "default" presets nothing and keeps declared routes.
* fix(terraform): scope key_type route rejection to create-shaped plans
/key/update stores an explicit allowed_routes verbatim and never reapplies
the key_type preset, so an existing or imported typed key can manage its
routes in place. Only plans that create a key (fresh, or a replacement
that changes key_type) still reject the combination, because there the
preset always overwrites the declared list. A replacement forced by
another ForceNew attribute converges on the next apply, which re-sends
the declared routes.
* fix(terraform): restore declared routes on typed key creation
/key/generate replaces a declared allowed_routes with the key_type
preset while /key/update stores the list verbatim, so any create that
carries both (a fresh key, or a replacement forced by key_type or
another ForceNew attribute) used to leave the key holding the preset
instead of the declared routes until a second apply. When the generate
response does not match the declared list, create now follows up with an
update that re-sends the full create payload against the new key hash,
so the first apply already stores the declared routes. This also
replaces the plan-time rejection of the combination: every config shape
now converges, and existing typed keys keep managing routes in place as
before.
* fix(terraform): delete the key when a route restore fails at create
If /key/generate succeeds but the restore update is rejected, the key
exists server-side while terraform holds no state for it: an active key
with the type preset would be orphaned and a retried apply would mint
another one. The restore failure path now deletes the created key, and a
delete that also fails names the key hash in the error so an operator
can remove it manually.
* fix(terraform): make the route restore surgical and keep supplied keys
Two sharp edges on the create-time route restore:
- Re-sending the full create payload rewrote fields the config never
declared: /key/update is a merge patch, so the empty metadata and
model_rpm_limit/model_tpm_limit maps the restored struct carried would
clear server-applied values such as team-inherited rate limits. The
restore now sends only the routes plus the two fields /key/update
requires non-null (permissions, model_max_budget); every other stored
value is kept.
- /key/generate upserts a config-supplied key value, so a restore
failure on such a key must not delete it: it may be an existing
credential that predates this apply. The compensating delete now runs
only for proxy-minted keys, and the error names the hash either way.
* fix(terraform): echo stored permissions and budgets in route restore
The surgical restore body carried empty permissions and model_max_budget
objects, and /key/update writes fields that are present: a key created
with declared permissions or model budgets next to a presetting key_type
and allowed_routes lost them on the first apply. The restore now echoes
the values /key/generate just stored (falling back to the configured
values when the response omits them), so the only field the restore ever
changes is allowed_routes.
* test(terraform): pin echoed budgets in the route restore
Adds the nonempty model_max_budget case Greptile asked for (the restore
must echo the stored map, never clear it) and drops a comment that
restated its own line.
* test(terraform): assert the declared budget reaches key generation
The budget echo case fed the raw config a malformed JSON string (a
template leftover), so nothing verified the declared budget actually
reached /key/generate. The config now carries the valid JSON and the
generate payload is asserted to match it.
* chore(terraform): trim the restore test preface to the proxy facts
---------
Co-authored-by: Roman Soletskyi <roman@mistral.ai>
* fix(traces): use the first user message for the run input preview
* test(traces): cover first user message as the input preview
* test(traces): check fixture previews against the first user message
* fix(lens): show the whole input preview on one line in the runs table
* test(lens): cover multi-line input previews in the runs table