* fix(proxy): scan batch records with the content hooks that are not guardrails
Guardrails were made to run on batch uploads by scanning each record through the pre-call hook
with the walk limited to guardrails. That limit exists because the same branch carries the rate
limiters and budget accounting, which must count an upload once rather than once per line. It
also excluded every enforcement hook written as a plain CustomLogger, so prompt-injection
detection, Azure content safety, banned keywords and the blocked-user check never saw a batch
record at all. Content that is a hard 400 online reached the provider verbatim through batch.
A CustomLogger now declares whether its pre-call hook judges the payload or merely counts the
request. The four that judge it opt in, the walk admits them, and both short-circuits learn
about them, including the one that decides whether the file is streamed off disk in the first
place: a proxy configured only with one of these hooks was skipping the scan entirely. Nothing
that counts a request is marked, so an upload still costs one slot and one budget check.
* refactor(proxy): drop the per-hook comment the attribute contract already states
* test(proxy): make the classification a ledger, and pin the wiring with a real hook
The classification test listed the two non-enterprise hooks by hand, so unmarking either
enterprise one changed nothing and the mutation matrix passed with both surviving. It now walks
the hook registries and fails on any pre-call CustomLogger that is on neither side, which also
gives the flag the forcing function it lacked: an enforcement hook added later would otherwise
default to off and silently skip batch records, which is the bug being fixed here.
Nothing exercised the path the bug actually lived on either, since every test raised its own
exception rather than a real hook's. One test now drives the shipped prompt-injection hook
through the scan, which pins the part no synthetic exception reaches: a chained exception reads
as a failure to judge, so refactoring any of these hooks to `raise ... from` would turn every
per-record drop into an aborted upload.
Also records why a hook that rewrites the payload for routing stays unmarked, and that only the
leaf class is consulted.
* test(proxy): set the callback list through monkeypatch rather than writing the global
* fix(proxy): read batch records the same way the upload validation does
The upload validation parses each JSONL line as bytes, where the json module sniffs the
encoding itself and accepts a leading byte order mark or a lone surrogate. The guardrail scan
that runs immediately after decoded each line to text first, which is stricter, so a file the
validation had just accepted could fail the scan. A `.jsonl` written by any of the editors that
emit a BOM, which includes PowerShell's Out-File and classic Notepad, uploaded fine until a
pre_call guardrail was configured and then returned 500 with a decode error and no indication
of which line or why. The scan now parses the same bytes the validation did, and an untouched
record is copied through as the bytes it arrived as rather than re-encoded.
A numeric custom_id was reported as null. The spec asks for a string, but callers do send
numbers, and null leaves the one field a caller reconciles on empty for exactly the records
that need it.
* fix(proxy): read the load-balancing record the same way, so a byte order mark keeps its routing
The first record is parsed to pick a deployment when batch load balancing is on, and it was
decoded to text before parsing, which rejects a leading byte order mark. The lookup returns None
on any parse failure, so such a file silently lost its routing and went to the default provider
rather than the configured one. That was already reachable for an upload no guardrail changed,
since the original bytes are passed straight through, and preserving the mark through a rewrite
widens it. Parsed as bytes now, like the validation and the scan.
* fix(proxy): find the routing record past a blank first line
The upload validation and the guardrail scan both skip blank lines, but deployment selection
read only the first physical line, so a file starting with a blank line lost its routing model
and went to the default provider rather than the configured one. It now skips blanks the way
the other two readers do, reading lazily so a large file is not read past its first record.
* fix(proxy): do not crash deployment selection on a record whose body is not an object
The upload validation checks that a record has a `body`, not that it is an object, so a record
can carry a string or a list there. Deployment selection called `.get` on it unconditionally and
raised, returning 500. That was already reachable for a plain file, and reading past a byte order
mark or a blank first line widened it to files that previously fell through to the default
provider instead. A record whose body names no readable model now resolves to no model, which is
the same answer the default-provider branch already handled.
* fix(proxy): keep a custom_id that cannot be encoded from failing the whole upload
A record identifier is echoed back in the create response. JSON parses a lone surrogate happily
but it cannot be encoded again, so a file the upload validation accepts returned 500 from the
response renderer rather than a report. Unencodable characters are replaced, which leaves every
ordinary identifier untouched and keeps a pathological one reconcilable.
This predates the reader change; reading past a byte order mark only altered which error the
same file produced first.
* fix(proxy): treat a url the parser rejects as one we do not recognize
Resolving a record's call type from its url runs the url through urlsplit, which raises on a few
malformed authorities such as an unclosed bracket. That happens before the try that wraps the
guardrail call, so it escaped the scan and returned 500 on a file the upload validation had just
accepted. An unreadable url is simply one we cannot recognize, which the body-shape fallback
already handles, so the record is still scanned rather than lost.
Reachable on staging today for a proxy running any guardrail. Enabling the scan for a proxy that
runs only a content-enforcing CustomLogger widens it to that configuration too, which is why it
is fixed here rather than left.
The auto-routed model group was only reachable through the
x-litellm-model-id response header. SDK and framework callers that do not
expose response headers had no way to read it, and under streaming there
was no body surface at all.
The response body `model` field is deliberately restamped back to the
client-requested alias on both paths, which is correct OpenAI semantics,
so this adds a separate namespaced `router_model_name` key instead of
redefining `model`. The key is written on non-streaming bodies and on
every SSE chunk, including the streaming fast path, and is emitted only
when an auto-routing strategy actually selected the deployment.
After a mid-stream fallback moves the request off the group the router
picked, the key is omitted rather than continuing to claim the original
tier. The router marker already supports per-chunk fallback signals via
`x-litellm-attempted-fallbacks` headers; this wires that signal into
the gate so no stale tier is claimed after a fallback fires.
Also removes a redundant function-local import in the streaming
generator that shadowed the module-level one for the whole function.
The Prisma query engine is a separate process whose resident memory grows with
what it is asked to hold and glibc never returns it, so a pod's memory floor
ratchets up to its worst statement and stays there for the life of the worker.
#34956 bounded a spend-log flush by payload bytes, which caps that floor when
prompts are stored and does nothing when they are not: rows carrying only
attribution metadata run about 1.2 KB, so a 1000-row statement is roughly
1.2 MB, the 2 MB byte budget never binds, and every statement stays at 1000
rows forever.
The engine charges per row as well as per byte. Measured on a container running
the same engine build (5.4.2) against real Postgres, with rows shaped like a
store_prompts_in_spend_logs=false deployment, writing the same 200,000 rows:
rows/statement engine RSS still resident after the flush
1000 179 MB
500 91 MB
250 41 MB
100 19 MB
None of those statements came near the byte budget, so the whole difference is
row count. The floor is a plateau rather than a leak: 1,000,000 rows written at
1000 per statement settles around 229 MB and stops climbing.
Adds SPEND_LOG_WRITE_BATCH_MAX_ROWS, default 100, applied alongside the
existing byte budget so whichever binds first splits the statement. Both are
needed, since bytes are what track a prompt-carrying row and rows are what
track the engine's per-row bookkeeping.
One consequence worth naming: a flush now issues more statements, and a
statement that fails under a poison flood costs one insert before any
isolation runs, so the irreducible floor rises by the statement count. The
isolation budget still caps the amplification on top of that, and the tests
assert the bound derived from the configured row cap rather than a constant.
Per-model budgets were three separate things pretending to be one. The
enforcement check, the post-call increment and the info endpoints each derived
their own cache key, so a budget could refuse traffic at 429 while /key/info
reported zero usage, and a Bedrock model id never matched a budget keyed on the
bare family name. /user/new echoed a model_max_budget back and stored an empty
dict, and nothing enforced a user-scoped per-model budget at all.
One owner now builds the counter key from the configured budget model, and
enforcement, the increment and the info endpoints all read it. Bedrock ids
resolve through the model-cost map. Auth carries the user's budget onto the
token on every branch that reaches the spend hook, including JWT and
auto-registration. Native passthrough attaches the three budget metadata keys
its StandardLoggingUserAPIKeyMetadata does not carry, so /anthropic/... and
/bedrock/... traffic is counted and capped like /v1/chat/completions.
The dashboard gains the per-model budget editor it never had, on the key create,
key edit and internal-user edit forms. It is read-only without an enterprise
license, matching the write gate the proxy already enforces, and an untouched
budget is left out of an update so an unrelated edit cannot trip that gate.
The editor hydrates from either BudgetConfig spelling, since model_max_budget is
a plain dict that the proxy stores exactly as the client sent it, and it carries
through the fields it does not model. Without both, editing one model would drop
another model row entirely and silently discard its tpm_limit and rpm_limit.
/user/info refreshes its local copy of the user field by field after a save, so
model_max_budget joins that list. Left out, a saved cap read back as the old one
when the form was reopened, and clearing the row to recover would then wipe the
value that had actually persisted.
A zero-dollar cap is the strictest limit expressible, not the absence of one,
so it is enforced rather than skipped on falsiness, spend exactly at the cap is
refused the way every sibling budget check already refuses it, and a counter
that was never written reads as zero spend rather than as unknown. The usage
endpoints read every counter in one batched lookup, so a large model_max_budget
cannot fan out into one concurrent cache call per configured model.
Every auth path honours the same zero-cost skip flag, so none of them can refuse
a free request that another serves. The custom-auth helper gains the flag it
never had, which also changes its pre-existing key and end-user checks.
The compaction summary gate checks the user scope alongside the key and end-user
ones. This file propagates all three budgets into the summary subrequest, so
enforcing only two let compaction increment a counter it could not be refused by.
Custom auth attaches the user's budget to the token unconditionally, since the
post-call spend hook reads it there: gating the attach on the same condition as
enforcement left the counter uncharged whenever the request was not itself
enforceable. An entry that will not validate is skipped rather than raised on,
so one malformed scope cannot abort every other scope's increment or turn a
config typo into a 500.
The edit forms re-seed the budget editor when a different key or user is loaded.
Its rows are seeded once and cannot re-read their own value prop, so without
this a save wrote the previously loaded record's budgets onto the current one.
Only the built-in provider pass-through routes carry the budget metadata.
get_model_from_request deliberately resolves no model for a user-defined
pass-through, since its body is forwarded verbatim and names an upstream model,
so attaching there would charge a counter nothing on that route can refuse.
The rollup read litellm.proxy.proxy_server.llm_router out of sys.modules, so a run
priced and swept whatever deployments anything else in the process had left on that
module. Under xdist the shard's module-to-worker assignment varies per run, which made
three rollup tests fail or pass on the same commit depending on ordering.
Callers now hand the router in, and the proxy's scheduled job passes its own.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
tests/litellm/ was a second mirror beside tests/test_litellm/ that no workflow,
Makefile target, or CircleCI job ever named. Its other 33 files were reconciled
during August 2026; this one stayed behind under a ci-coverage-allowlist entry
asking a later pass to decide which of its five orphan behaviours still hold.
They no longer hold as written: 25 of its 32 cases fail against today's code,
because the file froze on the day it stopped being collected and the endpoints
kept moving. Three of the five are already covered by the live twin, and better.
test_get_request_base_url_xff_trust_gate parametrizes the trust gate in both
directions, including the exact untrusted-caller case the orphan asserted, and
the standard and legacy protected-resource shapes are both exercised through
use_standard_pattern.
The other two were the only tests anywhere for validate_trusted_redirect_uri
under that same gate, so they are ported rather than dropped, rebuilt on the
live file's request-mock conventions. Both directions are load-bearing: forcing
is_request_from_trusted_proxy to True fails the untrusted case, forcing it to
False fails the trusted one.
313 tests pass in the live file, up from 311. Dropping the dead file clears one
zero-assert TQ001 violation, so its ceiling ratchets down with it.
Master-key auth stamps the stable alias litellm_proxy_master_key instead of the
raw key, so spend logs carry a readable, non-secret identifier for those rows.
The new redaction path only recognized sha256 and hashed-jwt shapes, so it
hashed that alias and broke continuity with every master-key row written
before this change. The alias joins the recognized non-secret values, still
behind the same provenance gate, so a caller who sends the alias string as
their own bearer token still gets it hashed.
The proxy conftest already snapshots master_key and prisma_client around every
test, because a value left behind on litellm.proxy.proxy_server poisons the rest
of the xdist worker. llm_router has the same problem. The PTU rollup reads the
running router out of sys.modules, so a router a sibling test left behind lands
in its deployment scan and three test_ptu_flat_cost_rollup tests fail or pass
depending on how xdist happens to split the shard.
The hashed-jwt branch trusted the value's shape alone, so a caller-supplied key in that shape was stored unhashed. Both pass-throughs now require the value to match the auth-time user_api_key_hash, and the shape check is a full match.
The spend-log helper no longer treats a 64-hex shape as proof a value was already hashed, so this case has to say where the hash came from. Reconciles the test that came in with #31799 against that change.
`pytest.raises(Exception)` with no `match=` passes on any error that broad. A
TypeError from a refactor, a botched fixture, an import that moved: all of them
read as the rejection the test claims to police, so the test goes green for the
wrong reason and stays green after the behaviour it guards is gone.
PT011 closes that gap for the 317 sites B017 could not reach, because B017 only
fires on a single-statement body with no `as e` binding. Each pattern here is the
message the code actually raised, recorded by running the sites under a plugin
that logged the concrete type and text per call site, so the assertions describe
observed behaviour rather than a guess. Where a site raises more than one message
across its parametrize cases, the pattern is an alternation of what was seen;
where the exception carries an empty `str()` and puts the text on `.message`, the
site keeps a narrow `noqa` with the reason.
PT014 removes four parametrize cases that were listed twice. The duplicate re-runs
an assertion that already passed, and it usually marks a case someone meant to
vary and forgot to edit.
`members_with_roles` is a denormalized JSON snapshot written at add-time.
`_update_team_members_list` backfilled `user_id` from `user_email` but never
the reverse, so a member added by `user_id` alone was stored with
`user_email=None` permanently - and `/team/info` returns that blob verbatim
with no join to `LiteLLM_UserTable`, so the Admin UI's member table renders
"-" for a user that plainly has an email.
Fix both ends:
- write path: `_resolve_member_identity` resolves identity both ways off the
user rows the add just touched, so new roster entries stop being born blank.
- read path: `/team/info` fills blank emails from `LiteLLM_UserTable` in one
indexed `user_id IN (...)` query, repairing rows already in the database.
Members that already carry an email are passed through untouched and cost
no query, so this only ever turns a null into the right value.
* test: enforce PT012 so a pytest.raises block cannot hide dead assertions
`with pytest.raises(...)` stops at the first statement that raises. Anything
sequenced after it inside the block never runs, so an assertion written there is
never checked and the test still reports green.
Two sites were doing exactly that, and both assertions turned out to be wrong
once they started running. tests/llm_translation/test_prompt_factory.py asserted
the bedrock rejection names "requires at least one non-system message", which
holds. tests/proxy_unit_tests/test_proxy_server.py asserted the prisma startup
failure mentions "httpx.ConnectError", which never appears: the failure is an
httpx.ConnectError whose message is "All connection attempts failed", so that
test now asserts the type. Its DATABASE_URL override moves to monkeypatch, since
the old restore sat below the assertion and leaked the invalid URL into every
later DB test the moment the assertion started being able to fail.
The remaining 72 sites are rewritten without changing what they exercise: setup
that cannot raise moves above the block, a nested `patch` moves outside it, and
bodies with real control flow (a stream drain, an if/else on sync_mode, a
retry loop) move into a local closure the block calls.
Fixing PT012 unmasked two B017s, since ruff only reports a blind
pytest.raises(Exception) once the block holds a single statement.
tests/proxy_unit_tests/test_auth_checks.py narrows to the ProxyException
can_key_call_model actually raises. tests/local_testing/test_completion_cost.py
was asserting vertex_ai/medlm-medium has no cost entry, which stopped being true
at some point; that dead first half is gone and the rest of the test, which
checks medlm pricing resolves above zero, now runs instead of being skipped.
* chore(ci): ratchet TQ004 to 768 after the prisma test moved to monkeypatch
SCIM group members were matched against litellm user ids only. An identity
provider that lists people by email or by the OIDC subject therefore matched
nothing, and the member fell through to placeholder creation.
Since #37688 made a failed member creation fail the group sync rather than drop
the member, that fallthrough is no longer quiet: the placeholder is created with
user_email set to the member value, the duplicate-email check rejects it, and the
whole group push answers 500. So on current staging a group listing anyone by
their email fails outright, every other member in the payload included.
An unmatched member id is now looked up across sso_user_id and user_email in one
query. Searching either field first would hide a value that names one account by
its SSO identity and another by its email, and hand the group to whichever was
searched first. The two are not compared alike: an email is matched the way
new_user matches one before accepting a new account, case-insensitively, because
matching more strictly than the layer that would reject the placeholder is what
turned an id whose casing differed from the stored email into that same 500. An
SSO identity is matched exactly, since OIDC defines sub as case-sensitive and
nothing folds its case on the way in.
An exact user_id hit is checked the same way rather than trusted outright, since a
value can be one account's id and another's SSO identity or email. That is not a
corner case: the placeholders this bug provisioned are keyed by the very id the
provider keeps pushing, so on a tenant that already has them the placeholder wins
the id lookup and the real account can never be matched. Refusing names the
problem instead of silently landing on the placeholder again. Those rows still
have to be deleted before the real account resolves; making the sync heal itself
needs a trustworthy way to tell a placeholder from an account someone created, and
created_via lives in caller-writable metadata, so it is left to a follow-up.
A value that names more than one account is refused with a 400 naming the id
rather than attributed to one of them.
Removals resolve too, since the roster holds canonical user ids and a directory
removes people by the id it added them with. A removal counts the members one
value names: the id as written when the roster holds it verbatim, which is how an
earlier release recorded a member it could not match, together with the members it
resolves to. Counting only the accounts on the roster keeps someone removable
after a second account takes their email, which resolving table-wide would not,
and counting both ways of naming a member together stops one value revoking two
people when it is one member's canonical id and another's email. A value naming
two of the group's own members is undecidable and fails rather than guessing or
reporting a removal it did not perform.
Resolves LIT-5383
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Pre-call processing rewrites request_data["model"] for aliasing and routing, so
matching either key let a routed model count as the client's own name and put the
wrapper model back on an Azure Model Router row.
A chunk carrying usage is stored as a pre-restamp copy, so an alias-restamped
stream reaches disconnect billing with its first chunk still on the deployment
model and every later chunk on the client's name. That is the same shape Azure
Model Router produces, and the previous guard read it as a routed model and
left the alias on the row, which is the unpriced name this PR set out to stop.
Compare the assembled model against the name the proxy stamps chunks with, so
the alias goes back to the deployment's model and the routed model stays.
Pricing a frame through the request's own logging object is what makes custom
deployment pricing work, but _response_cost_calculator does not only return a
number. It also stamps cost_breakdown onto the live logging object, and on a
pricing failure it writes response_cost_failure_debug_information into
model_call_details.
On an ordinary proxy stream that is harmless, because the success handler
recomputes cost_breakdown at end of stream and overwrites whatever the frames
left behind. The pass-through handlers are the problem: they compute their final
cost with a bare completion_cost call and never touch cost_breakdown again, so a
breakdown derived from one mid-stream frame would survive to the end and land in
the spend log's metadata. response_cost itself is unaffected either way, so this
was a reporting surface bug rather than a billing one, but the spend row would
have gone from null to a populated breakdown for a partial frame.
Snapshot both writes and put them back once the cost is read, so pricing a frame
stays a read as far as the rest of the request is concerned. The returned cost is
unchanged, so nothing about the injected usage.cost moves.
team_allowed_routes and admin_allowed_routes only matched exact strings or named route groups, so a whole prefix of pass-through endpoints had to be listed route by route in config. Match trailing-wildcard patterns with the same helper the key-level allowed_routes check uses, so "/prefix/*" covers endpoints registered later.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The logging-object pricing applies to streamed /v1/chat/completions too, not
just Anthropic message_delta, so a deployment with negotiated per-token prices
now gets that price in the streamed usage.cost there as well. Nothing asserted
that half. Adds the discounted and the sticker-fallback case for the OpenAI
chunk shape, plus the branch where the pricer raises and the frame falls back
to model-name pricing instead of breaking the stream.
The disconnect billing path was stamping the wrapper's model over whatever
stream_chunk_builder assembled. For Azure Model Router that throws away the
routed model: the proxy deliberately leaves those chunks unrestamped so the
builder can pick the real model off a later chunk, and overwriting it prices
the row at the router alias instead.
Only apply the wrapper's model when the builder did not find a model beyond
the first chunk's, which is every case except Model Router.
* test(lint): ban blind pytest.raises(Exception) with ruff B017
A bare pytest.raises(Exception) accepts whatever the body throws. The TypeError
a refactor introduces satisfies it exactly as well as the rejection the test was
written for, so the crash reads as a pass and the test never goes red.
All 111 existing sites are narrowed here. A runtime probe recorded the concrete
exception each one actually catches, and each site now names that type. Where
the code under test genuinely raises a bare Exception, the site pins a stable
slice of the message with match= instead.
Two sites tell on themselves. The shared responses-API cancel test raises
"custom_llm_provider is required but passed as None" rather than talking to a
provider at all, because cancel_responses takes a provider, not a model. And
test_bedrock_guardrails_with_streaming was the only test in its file still
passing without AWS credentials, because the NoCredentialsError boto3 raised
long before the guardrail ran satisfied the blind raises.
* fix(test): widen the openai batch-dispatch assertion to OpenAIError
The narrowed NotFoundError only holds where OPENAI_API_KEY is set. Without one
the SDK raises OpenAIError while building the client, long before any 404, so CI
went red. OpenAIError covers both and still rejects a TypeError from a refactor.
cache_read_input_tokens and cache_creation_input_tokens are pydantic extras
on Usage, not declared fields, so filling them in created keys that were not
there before rather than replacing a None. Readers that test for presence
then took the new zero as authoritative: the spend log writer skipped its
own copy from prompt_tokens_details, turning a real cache read of 500 into
0, and the prometheus provider cache counters stopped incrementing.
Carry the prompt_tokens_details counts up before defaulting to zero, so a
partial row reports the same cache numbers a complete one does. Renamed the
helper to say what it now does.
Every pod schedules the budget reset job, so a fleet re-read the whole due
population and wrote it back against one Postgres at the same calendar
boundary, multiplying a single sweep by its replica count. The job now takes
the shared PodLockManager lease, so one pod sweeps per tick. A deployment with
no Redis keeps its previous behavior, and a Redis that cannot answer sweeps
unguarded rather than stranding every expired budget at its cap.
The per-window scan read every row carrying budget_limits in one statement, so
its cost grew with the deployment's key count. It is now keyset-paginated and
walks to the end of the table on every sweep. A per-run cap would need a resume
position, and no pod can hold one because the lease rotates between ticks, so
the strictly advancing cursor is what terminates the walk.
Found and updated rows were also JSON-serialized into the service hook's
metadata and into debug lines on every chunk, on the event loop, whether or not
anything consumed them. The hooks now carry counts, and the debug payload is
deferred until a record is actually emitted.
Resolves LIT-4793
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
SCIM roster writes were swallowed, so a group or user push returned 200 while the
team roster never received the membership. Surfacing the failure fixes that, but
aborting on the first failed write leaves the rest of the batch unattempted on top
of unrolled-back, which is worse than what it replaces.
Every roster write in a reconciliation is now attempted, and the ones that did not
land are reported together, naming each failed add and remove. Rollback would be the
other option and it is not safe here: the compensating write can fail too, and it can
strip a membership that pre-dated the push. SCIM reconciliation is idempotent, so a
named partial failure is what the IdP's next push needs to close the gap.
The reported status still follows the failures, so a unanimous 404 stays a 404 and
only a batch whose failures disagree falls back to 500.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The read replica never received the operator's DB pool settings, so its
Prisma pool fell back to `num_physical_cpus * 2 + 1` and the configured cap
was not enforced. Both startup paths now pass the same params to the reader:
the CLI, and the componentized entrypoints that go through
`DatabaseURLSettings.apply_to_env`.
Only pool and timeout params are inherited, through a single allowlist both
paths share. Anything that decides which tables a query resolves against
stays on the writer, including entries smuggled in through
`database_extra_connection_params`, so a writer `search_path` cannot repoint
reader queries. Params the operator pinned on the replica URL still win.
Co-authored-by: Yassin Kortam <yassin.kortam@gmail.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Cognition serves an OpenAI-compatible /v1/chat/completions endpoint, so it has been onboarded as
custom_llm_provider: openai. That books its traffic as OpenAI, which means OpenAI-specific cost
discounts and provider-level reporting apply to it.
Registers cognition through the JSON provider registry: a providers.json entry with
COGNITION_API_KEY and COGNITION_API_BASE, LlmProviders.COGNITION, the constants.py provider lists,
cost map entries for swe-1.6 and swe-1.7, the provider endpoints matrix, the dashboard provider
fields, and tests. JSON providers can now also be resolved from their base url alone, so an
api_base pointing at a known provider no longer falls through to an unresolved provider.
A JWKS fetch had no retry, so a single connect timeout to the identity provider
failed authentication outright, and once the cached copy expired there was
nothing to fall back on. How that surfaced depended on the outage shape:
httpx.ConnectTimeout was missing from DB_CONNECTION_ERROR_TYPES so it fell
through to the generic auth handler as a 401 with an empty detail, while a read
timeout took the database path and reported a healthy database as unreachable.
Transport failures are now retried three times with a short backoff, and the
last-known-good JWKS stays usable for a bounded window past public_key_ttl.
That window is public_key_stale_ttl, a new config field defaulting to 3600s and
settable to 0 to fail closed. It is checked on every read against the current
setting rather than baked into the cache entry when it is written, so lowering
it binds immediately instead of waiting for entries written under the old value
to age out, which matters because a shared cache survives the restart an
operator performs to make the change take effect. A copy whose write time
cannot be established is not servable. Only httpx.TransportError unlocks the
stale copy, so an identity provider that answers at all, including with a
narrowed key set, revokes on the next refresh. Every stale serve logs the kid
it authenticated, how long ago that copy was refreshed, and how long until it
stops being trusted.
A sustained outage is remembered for 30s per key url, so it costs one fetch per
window instead of three timeouts per request serialised behind the refresh lock.
Non-200 JWKS responses now raise instead of being cached as the key set, which
previously let an error body overwrite the last-known-good copy. An unreachable
identity provider with no cached copy left returns 503 auth_provider_unavailable.
Resolves LIT-5524
Co-authored-by: Yassin Kortam <yassin@berri.ai>
A request can name more than one model, through a comma-separated model or target_model_names on
the batch and fine-tuning routes, and the gate only looked at the string case, so an unpriced model
riding alongside a priced one went through and billed. Check every candidate and name the unpriced
ones in the 403
Aliases had the same problem on the other side: a group that prices itself through its model_info
block lands in the cost map under its deployment id, and the explicit-cost check walked the raw
model list by group name, so an alias pointing at that group read as unpriced. Resolve the group
through the router the way the pricing check already does
Also correct the 403 copy. Providers that return their own usage cost still bill for these models,
so the accurate claim is that litellm has no pricing of its own for them
Some IdPs, ADFS among them, return only `sub` from UserInfo and put the real
identity claims in the ID token or the access token. Those users land in the
Admin UI with no username, email, groups or teams.
Adds an opt-in `GENERIC_INCLUDE_TOKEN_CLAIMS` that merges token claims into the
UserInfo response before the existing `GENERIC_USER_*_ATTRIBUTE` mappings run.
Precedence is UserInfo, then id_token, then access token, and it applies to both
the PKCE and non-PKCE login flows. With the flag unset, behavior is unchanged.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Adds a chat_completions route module to litellm-core, mirroring the messages
route, plus Anthropic Messages and Bedrock Converse provider configs. The
per-model `rust: true` opt-in now covers /chat/completions for both providers.
The core accepts an allowlisted subset (text conversations, non-streaming) and
returns CoreError::Unsupported for anything else, so tool calls, multimodal
content and streaming fall back to the Python path transparently.
Resolves LIT-5698
tiktoken's BPE merge loop is quadratic in the length of a single regex piece, so a long
run of one repeated character turns a multi-MB payload into minutes of CPU. Encoding in
bounded chunks makes that linear, at a drift of at most ~1 token per chunk boundary.
Chunking alone only makes the stall shorter, so the async paths now count in a worker
thread: tiktoken releases the GIL for its Rust encode, so the loop keeps serving other
requests while a count is in flight. The /utils/token_counter endpoint awaits the new
atoken_counter, and the router's async deployment selection counts off-loop and hands
the result to _pre_call_checks instead of making it count inline.
The chunk size knob is bounded to [1, 4096]: a non-positive value used to raise or
silently report zero tokens, and an arbitrarily large one restored the quadratic cost
this exists to remove. Out-of-range and unparseable values warn and fall back to 1024.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Budget reservation tokenized every request twice, once for the max-cost
estimate and once for the input-cost estimate, and again per pricing
candidate. Tokenizing is O(prompt) and ran inline, so admitting one large
request stalled every other request the worker was serving.
Count the input tokens once per request and reuse the counts for both
estimates. Prompts above 30K characters of input text are counted in a
worker thread so the event loop stays free. The size heuristic renders the
body rather than walking its values, so tool-schema property names count
toward the threshold, and it sizes every field the counter tokenizes,
tool_choice included.
Co-authored-by: Yassin Kortam <yassin@berri.ai>