litellm/.github
Soham Chatterjee 38bdda9ca7
feat(scaledown): add native text provider support (#44168)
* Add ScaleDown model pricing

Registers the five ScaleDown models in the pricing map. ScaleDown bills on
input tokens only, so output_cost_per_token is zero for every entry. The
classify and decisions models bill at $0.04 per million input tokens, and
extract, summarize, and compress bill at $0.05 per million.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Price all ScaleDown models at $0.04 per million and drop unverified context limits

Every ScaleDown model now bills $0.04 per million input tokens. The
32000-token max_input_tokens/max_output_tokens values were not confirmed
against the service, so they are removed rather than registered.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Price all ScaleDown models at $0.05 per million input tokens

This is the standard rate in ScaleDown's billing and matches the usage dashboard.
The Decisions usage.cost field is not used for pricing.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add regression tests for the ScaleDown pricing entries

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Cite the source of the ScaleDown rate in the pricing tests and type them

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add ScaleDown as a chat provider

ScaleDown serves five models over two wire protocols. extract, summarize, and
compress speak OpenAI chat on /v1/chat/completions, while classify and
decisions implement the Jev Decisions API on /v1/scaledown, taking a state
object plus a questions map and returning one typed answer per question.

The provider authenticates with x-api-key rather than a bearer token, and the
decisions models need a non-chat request body, so it routes through the HTTP
handler instead of the OpenAI passthrough. Malformed decisions requests are
rejected locally with a message naming the problem, so callers do not have to
decode an opaque 422 from upstream. Decisions bills one model call per
question, so the cost upstream reports is passed through rather than
re-estimated from token counts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Route ScaleDown domain models to the native API and stop trusting usage.cost

The /v1/chat/completions and /v1/models routes return 403 on api.scaledown.xyz,
so extract, summarize and compress now call the native /extract,
/summarization/abstractive and /compress/raw/ endpoints. The default api_base is
the bare host; a trailing /v1 is stripped and re-added only for decisions.

The Decisions usage.cost field was about 600x the dashboard amount, so it is no
longer passed through; cost is computed from input tokens at the registered rate.
Native responses are returned unchanged as JSON, including the _value/_span_anchor
wrappers on nested extraction. They carry no output token count, so
completion_tokens is reported as 0 and documented as unmeasured.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Move ScaleDown translation into llms/ and address review findings

- Dispatch moves to llms/scaledown/chat/handler.py; main.py keeps a thin call
  and get_llm_provider_logic.py no longer has a ScaleDown branch, since the
  config resolves key and base URL itself.
- questions passed via extra_body are merged and validated after the shared
  merge (sign_request), so the documented extra_body path works.
- extra_body can no longer change the upstream model, and decisions state may
  carry a document but not text, so guardrails always see the text.
- A custom api_base requires its own api_key; the environment key is only sent
  to the default host or SCALEDOWN_API_BASE.
- Image content parts become the decisions document; max_completion_tokens maps
  to max_tokens; stream is served as one chunk for every model.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Refuse prompt-bearing extra_body keys and empty-key fallback to the env credential

extra_body is merged after guardrails inspect the messages, so each model now
accepts only option keys there (extract threshold/top_n, compress
compression_rate, decisions questions/state); text, instructions, prompt,
context and model are rejected. An empty api_key no longer lets the
SCALEDOWN_API_KEY from the environment go to a custom api_base.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read compress token usage from the real response shape

/compress/raw/ nests the per-prompt counts under results and reports the total
as a top-level input_tokens; the provider read a top-level original_prompt_tokens
that is not there, so compress reported zero tokens and zero cost.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Accept the router's max_retries and word the extra_body error clearly

The proxy adds max_retries to every call, which map_openai_params rejected, so
every proxied ScaleDown request failed with a 400. It is now accepted and not
forwarded.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Return extracted fields from extract, follow schema refs, and meet the type-discipline gate

- extract content is now the extracted fields (span anchors dropped, wrappers
  unwrapped), so it validates against the caller's response_format; the raw
  payload stays on _hidden_params['scaledown_response'].
- Local $ref/$defs in the response_format schema are resolved.
- The adapter is rewritten with Final locals, read-only annotations, one raising
  helper and a typed handler; check_type_discipline reports no violations.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Fix stream_options, schema-guided extract cleaning, document input and multi-image handling

- stream_options is accepted and not sent upstream, so include_usage streams work.
- Extract output is cleaned against the requested entities, so a field named
  _value or *_span_anchor is no longer dropped.
- Extract and summarize send an image or PDF as the native document fields.
- More than one image is rejected instead of silently using the first.
- The cost test derives its rate from the cost map.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Cut recursive schema refs at the cycle and cap the expanded entity count

A small response_format schema with a self-referencing $def expanded to 4^16
leaves before any request was sent. Refs that point back into themselves are now
cut off at that point, and the total expanded entities are capped at 1000 so
repeated acyclic refs are rejected with a clear error.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Bound ref chains and schema nesting so long schemas fail cleanly

Resolving a chain of thousands of distinct $refs recursed until RecursionError.
Ref resolution is now an iterative walk capped at 32 hops, and nested properties
are capped at 32 levels; both raise a clear 400 instead.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* fix(scaledown): support native classify and preserve extraction results

* test(scaledown): include provider regressions in CI shard

* fix(scaledown): pass registry typing and bounded recursion checks

* docs(scaledown): capture provider setup before and after

* fix(scaledown): preserve nullable schemas and reject malformed requests

* fix(scaledown): isolate provider dispatch from existing providers

* fix(scaledown): preserve the shared dispatch return contract

* fix(scaledown): preserve failure logging and reject lossy schemas

* fix(scaledown): keep timeout typing annotation on its expression

* fix(scaledown): resolve escaped extraction schema references

* fix(scaledown): annotate the bounded reference walk

* chore(scaledown): remove unused loop suppression

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: moe-berri <moe@berri.ai>
2026-10-10 14:01:17 -07:00
..
actions ci: split slow unit shards and build the Rust bridge once per run (#44622) 2026-10-05 23:43:01 +00:00
assets/scaledown feat(scaledown): add native text provider support (#44168) 2026-10-10 14:01:17 -07:00
codeql ci(codeql): run the default suite on full scans and security-extended on PRs (#45149) 2026-10-07 13:24:28 -07:00
e2e-stack ci: move Postgres, MCP and Redis suites to CircleCI integration (#44453) 2026-10-05 09:33:14 -07:00
ISSUE_TEMPLATE ci: classify new issues into domain, provider, kind, priority and lift labels 2026-09-18 11:37:55 -07:00
observatory Add observatory test workflow for RC/stable releases 2026-03-01 15:30:09 -03:00
prompts ci: classify new issues into domain, provider, kind, priority and lift labels 2026-09-18 11:37:55 -07:00
PULL_REQUEST_TEMPLATE fix(security): remove the publicly known master key from the repo (#44718) 2026-10-06 10:55:24 -07:00
screenshots fix(team_endpoints): auto-add SSO team members to org on move (proxy admin only) (#26377) 2026-04-24 08:36:25 -07:00
scripts perf(ci): isolate and shard integration suites (#44146) 2026-10-10 11:15:19 -07:00
workflows feat(scaledown): add native text provider support (#44168) 2026-10-10 14:01:17 -07:00
ci-coverage-allowlist.yml test: move routing, caching and callback tests out of local_testing (#45761) 2026-10-10 09:33:09 -07:00
CODEOWNERS chore(codeowners): replace kerry-berri with kerrylu-berri (#45173) 2026-10-07 17:02:37 -07:00
dependabot.yaml chore: fixes 2026-04-05 01:30:57 -07:00
deploy-on-aws.png feat: add LiteLLM Rust workspace with Mistral OCR bridge (#31033) 2026-06-23 13:16:47 -07:00
deploy-on-gcp.png feat: add LiteLLM Rust workspace with Mistral OCR bridge (#31033) 2026-06-23 13:16:47 -07:00
deploy-to-aws.png Add files via upload 2023-10-25 16:33:53 -07:00
FUNDING.yml Update FUNDING.yml 2023-09-22 09:51:35 -07:00
issue-labels.json ci(issue-classifier): name every workflow, script and job after the issue it works on 2026-09-18 11:37:55 -07:00
mutmut-coverage.rc fix(ci): let the mutation workflow find covered lines so it generates mutants 2026-08-25 23:16:40 -07:00
template.yaml fix(proxy): drop legacy telemetry key from persisted WORKER_CONFIG before initialize 2026-09-20 03:14:11 +00:00