Applies the review suggestions. The cost map now carries the rates published on
https://scx.ai/pricing, GLM-5.2 at 0.61 in, 0.22 cached, 1.98 out and
Qwen3.8-Max at 1.65 in, 0.21 cached, 4.99 out per million tokens, and cites that
page as the source rather than a third party gateway. The provider link is
corrected to https://docs.litellm.ai/docs/providers/scx_ai to match the page
that shipped as scx_ai.md. Both the primary files and their backup mirrors are
updated.
Resolves the three conflicts against the JSON provider registry refactor. The
hardcoded api.scx.ai base-url branch in get_llm_provider_logic.py is dropped in
favour of the generic JSONProviderRegistry.get_by_base_url lookup, which reads
the same base_url and api_key_env from providers.json and additionally honours
an explicitly passed api_key. constants.py and types/utils.py keep both the
cognition and scx-ai entries added on either side.
Master-key auth stamps the stable alias litellm_proxy_master_key instead of the
raw key, so spend logs carry a readable, non-secret identifier for those rows.
The new redaction path only recognized sha256 and hashed-jwt shapes, so it
hashed that alias and broke continuity with every master-key row written
before this change. The alias joins the recognized non-secret values, still
behind the same provenance gate, so a caller who sends the alias string as
their own bearer token still gets it hashed.
The proxy conftest already snapshots master_key and prisma_client around every
test, because a value left behind on litellm.proxy.proxy_server poisons the rest
of the xdist worker. llm_router has the same problem. The PTU rollup reads the
running router out of sys.modules, so a router a sibling test left behind lands
in its deployment scan and three test_ptu_flat_cost_rollup tests fail or pass
depending on how xdist happens to split the shard.
The hashed-jwt branch trusted the value's shape alone, so a caller-supplied key in that shape was stored unhashed. Both pass-throughs now require the value to match the auth-time user_api_key_hash, and the shape check is a full match.
The spend-log helper no longer treats a 64-hex shape as proof a value was already hashed, so this case has to say where the hash came from. Reconciles the test that came in with #31799 against that change.
`pytest.raises(Exception)` with no `match=` passes on any error that broad. A
TypeError from a refactor, a botched fixture, an import that moved: all of them
read as the rejection the test claims to police, so the test goes green for the
wrong reason and stays green after the behaviour it guards is gone.
PT011 closes that gap for the 317 sites B017 could not reach, because B017 only
fires on a single-statement body with no `as e` binding. Each pattern here is the
message the code actually raised, recorded by running the sites under a plugin
that logged the concrete type and text per call site, so the assertions describe
observed behaviour rather than a guess. Where a site raises more than one message
across its parametrize cases, the pattern is an alternation of what was seen;
where the exception carries an empty `str()` and puts the text on `.message`, the
site keeps a narrow `noqa` with the reason.
PT014 removes four parametrize cases that were listed twice. The duplicate re-runs
an assertion that already passed, and it usually marks a case someone meant to
vary and forgot to edit.
sagemaker_chat never put X-Amzn-SageMaker-Inference-Component on the request, so any endpoint
backed by inference components answered 400 INFERENCE_COMPONENT_NAME_MISSING and the call never
reached the container. The legacy sagemaker provider has built that header from model_id since
#8889, and this brings the chat provider in line. It goes on in validate_environment, which runs
before the request is SigV4-signed, so the signature covers it
The request body also always named the endpoint rather than the served model, which containers
that validate the body's model answer with a 404. hf_model_name now becomes the body's model
when it is set, and endpoints that do not set it keep sending exactly what they send today
Add sync and async regression tests for the BaseLLMHTTPHandler streaming
path, which forwards provider response headers for the ~30 providers that
ride the generic handler and had no coverage. Also drop redundant setup
prose from the moonshot invoke test docstring.
`members_with_roles` is a denormalized JSON snapshot written at add-time.
`_update_team_members_list` backfilled `user_id` from `user_email` but never
the reverse, so a member added by `user_id` alone was stored with
`user_email=None` permanently - and `/team/info` returns that blob verbatim
with no join to `LiteLLM_UserTable`, so the Admin UI's member table renders
"-" for a user that plainly has an email.
Fix both ends:
- write path: `_resolve_member_identity` resolves identity both ways off the
user rows the add just touched, so new roster entries stop being born blank.
- read path: `/team/info` fills blank emails from `LiteLLM_UserTable` in one
indexed `user_id IN (...)` query, repairing rows already in the database.
Members that already carry an email are passed through untouched and cost
no query, so this only ever turns a null into the right value.
The cost map shipped cognition/swe-1.7 at $2.50 in / $12.50 out per million with
$1.00 cache reads. Those are the Lightning numbers. Cognition's own model list at
https://docs.devin.ai/desktop/models has uid swe-1-7 at $0.50 / $2.50 with $0.20
cache reads, and uid swe-1-7-lightning at $2.50 / $12.50 with $1.00 cache reads,
so every swe-1.7 call has been costed at 5x since the entry landed.
swe-1.7 now carries the standard rates and the Lightning tier gets its own entry,
in both cost map copies. The source field on both moves to the desktop models page,
which is the one that lists both tiers.
* test: enforce PT012 so a pytest.raises block cannot hide dead assertions
`with pytest.raises(...)` stops at the first statement that raises. Anything
sequenced after it inside the block never runs, so an assertion written there is
never checked and the test still reports green.
Two sites were doing exactly that, and both assertions turned out to be wrong
once they started running. tests/llm_translation/test_prompt_factory.py asserted
the bedrock rejection names "requires at least one non-system message", which
holds. tests/proxy_unit_tests/test_proxy_server.py asserted the prisma startup
failure mentions "httpx.ConnectError", which never appears: the failure is an
httpx.ConnectError whose message is "All connection attempts failed", so that
test now asserts the type. Its DATABASE_URL override moves to monkeypatch, since
the old restore sat below the assertion and leaked the invalid URL into every
later DB test the moment the assertion started being able to fail.
The remaining 72 sites are rewritten without changing what they exercise: setup
that cannot raise moves above the block, a nested `patch` moves outside it, and
bodies with real control flow (a stream drain, an if/else on sync_mode, a
retry loop) move into a local closure the block calls.
Fixing PT012 unmasked two B017s, since ruff only reports a blind
pytest.raises(Exception) once the block holds a single statement.
tests/proxy_unit_tests/test_auth_checks.py narrows to the ProxyException
can_key_call_model actually raises. tests/local_testing/test_completion_cost.py
was asserting vertex_ai/medlm-medium has no cost entry, which stopped being true
at some point; that dead first half is gone and the rest of the test, which
checks medlm pricing resolves above zero, now runs instead of being skipped.
* chore(ci): ratchet TQ004 to 768 after the prisma test moved to monkeypatch
SCIM group members were matched against litellm user ids only. An identity
provider that lists people by email or by the OIDC subject therefore matched
nothing, and the member fell through to placeholder creation.
Since #37688 made a failed member creation fail the group sync rather than drop
the member, that fallthrough is no longer quiet: the placeholder is created with
user_email set to the member value, the duplicate-email check rejects it, and the
whole group push answers 500. So on current staging a group listing anyone by
their email fails outright, every other member in the payload included.
An unmatched member id is now looked up across sso_user_id and user_email in one
query. Searching either field first would hide a value that names one account by
its SSO identity and another by its email, and hand the group to whichever was
searched first. The two are not compared alike: an email is matched the way
new_user matches one before accepting a new account, case-insensitively, because
matching more strictly than the layer that would reject the placeholder is what
turned an id whose casing differed from the stored email into that same 500. An
SSO identity is matched exactly, since OIDC defines sub as case-sensitive and
nothing folds its case on the way in.
An exact user_id hit is checked the same way rather than trusted outright, since a
value can be one account's id and another's SSO identity or email. That is not a
corner case: the placeholders this bug provisioned are keyed by the very id the
provider keeps pushing, so on a tenant that already has them the placeholder wins
the id lookup and the real account can never be matched. Refusing names the
problem instead of silently landing on the placeholder again. Those rows still
have to be deleted before the real account resolves; making the sync heal itself
needs a trustworthy way to tell a placeholder from an account someone created, and
created_via lives in caller-writable metadata, so it is left to a follow-up.
A value that names more than one account is refused with a 400 naming the id
rather than attributed to one of them.
Removals resolve too, since the roster holds canonical user ids and a directory
removes people by the id it added them with. A removal counts the members one
value names: the id as written when the roster holds it verbatim, which is how an
earlier release recorded a member it could not match, together with the members it
resolves to. Counting only the accounts on the roster keeps someone removable
after a second account takes their email, which resolving table-wide would not,
and counting both ways of naming a member together stops one value revoking two
people when it is one member's canonical id and another's email. A value naming
two of the group's own members is undecidable and fails rather than guessing or
reporting a removal it did not perform.
Resolves LIT-5383
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Pre-call processing rewrites request_data["model"] for aliasing and routing, so
matching either key let a routed model count as the client's own name and put the
wrapper model back on an Azure Model Router row.
A chunk carrying usage is stored as a pre-restamp copy, so an alias-restamped
stream reaches disconnect billing with its first chunk still on the deployment
model and every later chunk on the client's name. That is the same shape Azure
Model Router produces, and the previous guard read it as a routed model and
left the alias on the row, which is the unpriced name this PR set out to stop.
Compare the assembled model against the name the proxy stamps chunks with, so
the alias goes back to the deployment's model and the routed model stays.
Pricing a frame through the request's own logging object is what makes custom
deployment pricing work, but _response_cost_calculator does not only return a
number. It also stamps cost_breakdown onto the live logging object, and on a
pricing failure it writes response_cost_failure_debug_information into
model_call_details.
On an ordinary proxy stream that is harmless, because the success handler
recomputes cost_breakdown at end of stream and overwrites whatever the frames
left behind. The pass-through handlers are the problem: they compute their final
cost with a bare completion_cost call and never touch cost_breakdown again, so a
breakdown derived from one mid-stream frame would survive to the end and land in
the spend log's metadata. response_cost itself is unaffected either way, so this
was a reporting surface bug rather than a billing one, but the spend row would
have gone from null to a populated breakdown for a partial frame.
Snapshot both writes and put them back once the cost is read, so pricing a frame
stays a read as far as the rest of the request is concerned. The returned cost is
unchanged, so nothing about the injected usage.cost moves.
Every live together_ai call in CI has answered 503 Service unavailable since
2026-08-20, across two runs 2.5 hours apart, while Together's status page
reported no incident in either window. These are real calls, not replayed
cassettes: the VCR layer runs filter_non_2xx_response, so a 503 is never
written to a cassette and cannot be replayed back.
Qwen/Qwen2.5-7B-Instruct-Turbo does not appear anywhere on Together's monitored
component list, whose Qwen entries are all Qwen3.x, so a model-level outage
there would never surface as an incident. The same 503 already forced
test_basic_rerank_together_ai to be skipped on a different together_ai model,
so per-model 503s are an established failure mode here rather than a platform
outage.
openai/gpt-oss-20b is the cheapest together_ai entry that carries real pricing
and the capabilities these suites exercise, at $0.05/$0.20 per 1M tokens with
function calling, response schema and tool choice. Together monitors it as a
served component. The retired model also carries null pricing in the cost map,
which is its own liability now that unpriced models are blocked.
test_multiple_deployments.py keeps the old id: it is a router fallback list
that is green today, and busting its cassette to prove a point would trade a
passing test for a live call this change cannot vouch for.
a369cb0da7 made completion() hand the prefixed model back to responses(), so
that responses() running get_llm_provider() a second time becomes a no-op
instead of stripping a prefix the model id owns. That was deliberate, and it
shipped with its own unit test, but it left two older assertions behind still
expecting the bare id.
#37744 corrected the openai one in test_openai.py. This is its azure sibling,
which llm_translation_testing has been failing on ever since.
Only the expected value moves. The neighbouring custom_llm_provider assertion
already passes and stays as it is.
The logging-object pricing applies to streamed /v1/chat/completions too, not
just Anthropic message_delta, so a deployment with negotiated per-token prices
now gets that price in the streamed usage.cost there as well. Nothing asserted
that half. Adds the discounted and the sticker-fallback case for the OpenAI
chunk shape, plus the branch where the pricer raises and the frame falls back
to model-name pricing instead of breaking the stream.
The disconnect billing path was stamping the wrapper's model over whatever
stream_chunk_builder assembled. For Azure Model Router that throws away the
routed model: the proxy deliberately leaves those chunks unrestamped so the
builder can pick the real model off a later chunk, and overwriting it prices
the row at the router alias instead.
Only apply the wrapper's model when the builder did not find a model beyond
the first chunk's, which is every case except Model Router.
* test(lint): ban blind pytest.raises(Exception) with ruff B017
A bare pytest.raises(Exception) accepts whatever the body throws. The TypeError
a refactor introduces satisfies it exactly as well as the rejection the test was
written for, so the crash reads as a pass and the test never goes red.
All 111 existing sites are narrowed here. A runtime probe recorded the concrete
exception each one actually catches, and each site now names that type. Where
the code under test genuinely raises a bare Exception, the site pins a stable
slice of the message with match= instead.
Two sites tell on themselves. The shared responses-API cancel test raises
"custom_llm_provider is required but passed as None" rather than talking to a
provider at all, because cancel_responses takes a provider, not a model. And
test_bedrock_guardrails_with_streaming was the only test in its file still
passing without AWS credentials, because the NoCredentialsError boto3 raised
long before the guardrail ran satisfied the blind raises.
* fix(test): widen the openai batch-dispatch assertion to OpenAIError
The narrowed NotFoundError only holds where OPENAI_API_KEY is set. Without one
the SDK raises OpenAIError while building the client, long before any 404, so CI
went red. OpenAIError covers both and still rejects a TypeError from a refactor.
The support matrix says cognition serves /v1/responses, but the README row
left that column blank, so the two disagreed. Every other provider row tracks
the matrix, so bring this one in line.
The Model-Specific Limits rows now carry Input TPM and Output TPM, and a
limit the operator removes is sent as an explicitly empty map so
/project/update actually drops it instead of leaving the stored quota
enforced behind a UI that shows it gone.
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
cache_read_input_tokens and cache_creation_input_tokens are pydantic extras
on Usage, not declared fields, so filling them in created keys that were not
there before rather than replacing a None. Readers that test for presence
then took the new zero as authoritative: the spend log writer skipped its
own copy from prompt_tokens_details, turning a real cache read of 500 into
0, and the prometheus provider cache counters stopped incrementing.
Carry the prompt_tokens_details counts up before defaulting to zero, so a
partial row reports the same cache numbers a complete one does. Renamed the
helper to say what it now does.
The swe-1.7 rates were briefly lowered to the standard tier. The docs page
records the API-served swe-1.7 as the Cerebras-served Lightning tier, so put
the matching rates back rather than have the cost map and the docs disagree.
Cognition also answers /v1/responses through the chat-completions bridge, the
same as every other provider in the JSON registry, so the endpoints support
matrix should say so instead of under-declaring it.
The swe-1.7 rates were carried over from the closed prior attempt and
match SWE-1.7 Lightning, 5x the SWE-1.7 Max and Medium rates the vendor
publishes. swe-1.6 was already on the standard tier, so the two entries
disagreed with each other. Both now read 0.5 in, 2.5 out, 0.2 cached per
million tokens.
Also drops the redundant registry comment in constants.py.