mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
* Add ScaleDown model pricing Registers the five ScaleDown models in the pricing map. ScaleDown bills on input tokens only, so output_cost_per_token is zero for every entry. The classify and decisions models bill at $0.04 per million input tokens, and extract, summarize, and compress bill at $0.05 per million. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Price all ScaleDown models at $0.04 per million and drop unverified context limits Every ScaleDown model now bills $0.04 per million input tokens. The 32000-token max_input_tokens/max_output_tokens values were not confirmed against the service, so they are removed rather than registered. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Price all ScaleDown models at $0.05 per million input tokens This is the standard rate in ScaleDown's billing and matches the usage dashboard. The Decisions usage.cost field is not used for pricing. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add regression tests for the ScaleDown pricing entries Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Cite the source of the ScaleDown rate in the pricing tests and type them Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add ScaleDown as a chat provider ScaleDown serves five models over two wire protocols. extract, summarize, and compress speak OpenAI chat on /v1/chat/completions, while classify and decisions implement the Jev Decisions API on /v1/scaledown, taking a state object plus a questions map and returning one typed answer per question. The provider authenticates with x-api-key rather than a bearer token, and the decisions models need a non-chat request body, so it routes through the HTTP handler instead of the OpenAI passthrough. Malformed decisions requests are rejected locally with a message naming the problem, so callers do not have to decode an opaque 422 from upstream. Decisions bills one model call per question, so the cost upstream reports is passed through rather than re-estimated from token counts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Route ScaleDown domain models to the native API and stop trusting usage.cost The /v1/chat/completions and /v1/models routes return 403 on api.scaledown.xyz, so extract, summarize and compress now call the native /extract, /summarization/abstractive and /compress/raw/ endpoints. The default api_base is the bare host; a trailing /v1 is stripped and re-added only for decisions. The Decisions usage.cost field was about 600x the dashboard amount, so it is no longer passed through; cost is computed from input tokens at the registered rate. Native responses are returned unchanged as JSON, including the _value/_span_anchor wrappers on nested extraction. They carry no output token count, so completion_tokens is reported as 0 and documented as unmeasured. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Move ScaleDown translation into llms/ and address review findings - Dispatch moves to llms/scaledown/chat/handler.py; main.py keeps a thin call and get_llm_provider_logic.py no longer has a ScaleDown branch, since the config resolves key and base URL itself. - questions passed via extra_body are merged and validated after the shared merge (sign_request), so the documented extra_body path works. - extra_body can no longer change the upstream model, and decisions state may carry a document but not text, so guardrails always see the text. - A custom api_base requires its own api_key; the environment key is only sent to the default host or SCALEDOWN_API_BASE. - Image content parts become the decisions document; max_completion_tokens maps to max_tokens; stream is served as one chunk for every model. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Refuse prompt-bearing extra_body keys and empty-key fallback to the env credential extra_body is merged after guardrails inspect the messages, so each model now accepts only option keys there (extract threshold/top_n, compress compression_rate, decisions questions/state); text, instructions, prompt, context and model are rejected. An empty api_key no longer lets the SCALEDOWN_API_KEY from the environment go to a custom api_base. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Read compress token usage from the real response shape /compress/raw/ nests the per-prompt counts under results and reports the total as a top-level input_tokens; the provider read a top-level original_prompt_tokens that is not there, so compress reported zero tokens and zero cost. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Accept the router's max_retries and word the extra_body error clearly The proxy adds max_retries to every call, which map_openai_params rejected, so every proxied ScaleDown request failed with a 400. It is now accepted and not forwarded. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Return extracted fields from extract, follow schema refs, and meet the type-discipline gate - extract content is now the extracted fields (span anchors dropped, wrappers unwrapped), so it validates against the caller's response_format; the raw payload stays on _hidden_params['scaledown_response']. - Local $ref/$defs in the response_format schema are resolved. - The adapter is rewritten with Final locals, read-only annotations, one raising helper and a typed handler; check_type_discipline reports no violations. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Fix stream_options, schema-guided extract cleaning, document input and multi-image handling - stream_options is accepted and not sent upstream, so include_usage streams work. - Extract output is cleaned against the requested entities, so a field named _value or *_span_anchor is no longer dropped. - Extract and summarize send an image or PDF as the native document fields. - More than one image is rejected instead of silently using the first. - The cost test derives its rate from the cost map. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Cut recursive schema refs at the cycle and cap the expanded entity count A small response_format schema with a self-referencing $def expanded to 4^16 leaves before any request was sent. Refs that point back into themselves are now cut off at that point, and the total expanded entities are capped at 1000 so repeated acyclic refs are rejected with a clear error. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Bound ref chains and schema nesting so long schemas fail cleanly Resolving a chain of thousands of distinct $refs recursed until RecursionError. Ref resolution is now an iterative walk capped at 32 hops, and nested properties are capped at 32 levels; both raise a clear 400 instead. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * fix(scaledown): support native classify and preserve extraction results * test(scaledown): include provider regressions in CI shard * fix(scaledown): pass registry typing and bounded recursion checks * docs(scaledown): capture provider setup before and after * fix(scaledown): preserve nullable schemas and reject malformed requests * fix(scaledown): isolate provider dispatch from existing providers * fix(scaledown): preserve the shared dispatch return contract * fix(scaledown): preserve failure logging and reject lossy schemas * fix(scaledown): keep timeout typing annotation on its expression * fix(scaledown): resolve escaped extraction schema references * fix(scaledown): annotate the bounded reference walk * chore(scaledown): remove unused loop suppression --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: moe-berri <moe@berri.ai> |
||
|---|---|---|
| .. | ||
| actions | ||
| assets/scaledown | ||
| codeql | ||
| e2e-stack | ||
| ISSUE_TEMPLATE | ||
| observatory | ||
| prompts | ||
| PULL_REQUEST_TEMPLATE | ||
| screenshots | ||
| scripts | ||
| workflows | ||
| ci-coverage-allowlist.yml | ||
| CODEOWNERS | ||
| dependabot.yaml | ||
| deploy-on-aws.png | ||
| deploy-on-gcp.png | ||
| deploy-to-aws.png | ||
| FUNDING.yml | ||
| issue-labels.json | ||
| mutmut-coverage.rc | ||
| template.yaml | ||