litellm/tests/code_coverage_tests/recursive_detector.py
Soham Chatterjee 38bdda9ca7
feat(scaledown): add native text provider support (#44168)
* Add ScaleDown model pricing

Registers the five ScaleDown models in the pricing map. ScaleDown bills on
input tokens only, so output_cost_per_token is zero for every entry. The
classify and decisions models bill at $0.04 per million input tokens, and
extract, summarize, and compress bill at $0.05 per million.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Price all ScaleDown models at $0.04 per million and drop unverified context limits

Every ScaleDown model now bills $0.04 per million input tokens. The
32000-token max_input_tokens/max_output_tokens values were not confirmed
against the service, so they are removed rather than registered.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Price all ScaleDown models at $0.05 per million input tokens

This is the standard rate in ScaleDown's billing and matches the usage dashboard.
The Decisions usage.cost field is not used for pricing.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add regression tests for the ScaleDown pricing entries

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Cite the source of the ScaleDown rate in the pricing tests and type them

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add ScaleDown as a chat provider

ScaleDown serves five models over two wire protocols. extract, summarize, and
compress speak OpenAI chat on /v1/chat/completions, while classify and
decisions implement the Jev Decisions API on /v1/scaledown, taking a state
object plus a questions map and returning one typed answer per question.

The provider authenticates with x-api-key rather than a bearer token, and the
decisions models need a non-chat request body, so it routes through the HTTP
handler instead of the OpenAI passthrough. Malformed decisions requests are
rejected locally with a message naming the problem, so callers do not have to
decode an opaque 422 from upstream. Decisions bills one model call per
question, so the cost upstream reports is passed through rather than
re-estimated from token counts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Route ScaleDown domain models to the native API and stop trusting usage.cost

The /v1/chat/completions and /v1/models routes return 403 on api.scaledown.xyz,
so extract, summarize and compress now call the native /extract,
/summarization/abstractive and /compress/raw/ endpoints. The default api_base is
the bare host; a trailing /v1 is stripped and re-added only for decisions.

The Decisions usage.cost field was about 600x the dashboard amount, so it is no
longer passed through; cost is computed from input tokens at the registered rate.
Native responses are returned unchanged as JSON, including the _value/_span_anchor
wrappers on nested extraction. They carry no output token count, so
completion_tokens is reported as 0 and documented as unmeasured.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Move ScaleDown translation into llms/ and address review findings

- Dispatch moves to llms/scaledown/chat/handler.py; main.py keeps a thin call
  and get_llm_provider_logic.py no longer has a ScaleDown branch, since the
  config resolves key and base URL itself.
- questions passed via extra_body are merged and validated after the shared
  merge (sign_request), so the documented extra_body path works.
- extra_body can no longer change the upstream model, and decisions state may
  carry a document but not text, so guardrails always see the text.
- A custom api_base requires its own api_key; the environment key is only sent
  to the default host or SCALEDOWN_API_BASE.
- Image content parts become the decisions document; max_completion_tokens maps
  to max_tokens; stream is served as one chunk for every model.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Refuse prompt-bearing extra_body keys and empty-key fallback to the env credential

extra_body is merged after guardrails inspect the messages, so each model now
accepts only option keys there (extract threshold/top_n, compress
compression_rate, decisions questions/state); text, instructions, prompt,
context and model are rejected. An empty api_key no longer lets the
SCALEDOWN_API_KEY from the environment go to a custom api_base.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read compress token usage from the real response shape

/compress/raw/ nests the per-prompt counts under results and reports the total
as a top-level input_tokens; the provider read a top-level original_prompt_tokens
that is not there, so compress reported zero tokens and zero cost.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Accept the router's max_retries and word the extra_body error clearly

The proxy adds max_retries to every call, which map_openai_params rejected, so
every proxied ScaleDown request failed with a 400. It is now accepted and not
forwarded.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Return extracted fields from extract, follow schema refs, and meet the type-discipline gate

- extract content is now the extracted fields (span anchors dropped, wrappers
  unwrapped), so it validates against the caller's response_format; the raw
  payload stays on _hidden_params['scaledown_response'].
- Local $ref/$defs in the response_format schema are resolved.
- The adapter is rewritten with Final locals, read-only annotations, one raising
  helper and a typed handler; check_type_discipline reports no violations.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Fix stream_options, schema-guided extract cleaning, document input and multi-image handling

- stream_options is accepted and not sent upstream, so include_usage streams work.
- Extract output is cleaned against the requested entities, so a field named
  _value or *_span_anchor is no longer dropped.
- Extract and summarize send an image or PDF as the native document fields.
- More than one image is rejected instead of silently using the first.
- The cost test derives its rate from the cost map.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Cut recursive schema refs at the cycle and cap the expanded entity count

A small response_format schema with a self-referencing $def expanded to 4^16
leaves before any request was sent. Refs that point back into themselves are now
cut off at that point, and the total expanded entities are capped at 1000 so
repeated acyclic refs are rejected with a clear error.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Bound ref chains and schema nesting so long schemas fail cleanly

Resolving a chain of thousands of distinct $refs recursed until RecursionError.
Ref resolution is now an iterative walk capped at 32 hops, and nested properties
are capped at 32 levels; both raise a clear 400 instead.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* fix(scaledown): support native classify and preserve extraction results

* test(scaledown): include provider regressions in CI shard

* fix(scaledown): pass registry typing and bounded recursion checks

* docs(scaledown): capture provider setup before and after

* fix(scaledown): preserve nullable schemas and reject malformed requests

* fix(scaledown): isolate provider dispatch from existing providers

* fix(scaledown): preserve the shared dispatch return contract

* fix(scaledown): preserve failure logging and reject lossy schemas

* fix(scaledown): keep timeout typing annotation on its expression

* fix(scaledown): resolve escaped extraction schema references

* fix(scaledown): annotate the bounded reference walk

* chore(scaledown): remove unused loop suppression

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: moe-berri <moe@berri.ai>
2026-10-10 14:01:17 -07:00

167 lines
9.8 KiB
Python

import ast
import os
IGNORE_FUNCTIONS = [
"_json_cost", # bounded at depth 32 and consumed under byte/node limits.
"_format_type",
"remove_additional_properties",
"remove_strict_from_schema",
"filter_schema_fields",
"text_completion",
"_check_for_os_environ_vars",
"clean_message",
"unpack_defs",
"convert_anyof_null_to_nullable", # has a set max depth
"add_object_type",
"strip_field",
"transform_prompt",
"mask_dict",
"_serialize", # we now set a max depth for this
"_sanitize_request_body_for_spend_logs_payload", # testing added for circular reference
"_sanitize_value", # testing added for circular reference
"set_schema_property_ordering", # testing added for infinite recursion
"process_items", # testing added for infinite recursion + max depth set.
"_can_object_call_model", # max depth set.
"encode_unserializable_types", # max depth set.
"filter_value_from_dict", # max depth set.
"normalize_json_schema_types", # max depth set.
"_extract_fields_recursive", # max depth set.
"remove_json_schema_refs", # max depth set.,
"_convert_schema_types", # max depth set.,
"_fix_enum_empty_strings", # max depth set.,
"get_access_token", # max depth set.,
"_redact_base64", # max depth set.
"_contains_vision_content", # max depth set.
"_read_all_bytes", # max depth set.
"_fix_enum_types", # max depth set.
"_collect_argument_paths", # max depth set.
"_split_text", # max depth set.
"_mask_sequence", # max depth set.
"_encrypted_param", # max depth set.
"_decrypted_param", # max depth set.
"contains_encrypted_marker", # max depth set.
"_rotate_guardrail_row", # bounded by attempts_left.
"_delete_nested_value_custom", # max depth set (bounded by number of path segments).
"filter_exceptions_from_params", # max depth set (default 20) to prevent infinite recursion.
"__getattr__", # lazy loading pattern in litellm/__init__.py with proper caching to prevent infinite recursion.
"_validate_inheritance_chain", # max depth set (default 100) to prevent infinite recursion in policy inheritance validation.
"_basic_json_schema_validate", # max depth set.
"extract_text_from_a2a_message", # max depth set (default 10) to prevent infinite recursion in A2A message parsing.
"_convert_to_json_serializable_dict", # max depth set (default 20) and circular reference protection to prevent infinite recursion.
"dict", # max depth set. _LiteLLMParamsDictView.dict() calls builtin dict(), not itself.
"_read_image_bytes", # max depth set.
"get_masked_values", # max depth set (default 20) to prevent infinite recursion while masking nested sensitive config dicts.
"_redact_sensitive_litellm_params", # max depth set (default 10).
"_redact_secret_values_in_obj", # max depth set (default 10, _REDACT_SECRET_MAX_DEPTH); fails closed by returning "REDACTED" at the cap.
"_resolve", # OCI: $ref resolver bounded by `resolving_stack` cycle guard.
"resolve_oci_schema_anyof", # OCI: bounded by JSON-schema tree depth (no cycles possible in well-formed input).
"sanitize_oci_schema", # OCI: bounded by JSON-schema tree depth.
"_freeze_for_dedupe", # OTEL: max depth set (default 16, _FREEZE_MAX_DEPTH); fails closed by returning repr(value) at the cap.
"apply_json_merge_patch", # max depth set (_MAX_MERGE_DEPTH=64); fails closed by raising ValueError at the cap.
"_filter_argument_value", # max depth set (DEFAULT_MAX_RECURSE_DEPTH); fails closed by blocking the tool call at the cap.
"_redact_scanned_content", # max depth set (DEFAULT_MAX_RECURSE_DEPTH); fails closed by returning "[REDACTED]" at the cap.
"replace_ciphertexts", # max depth set (DEFAULT_MAX_RECURSE_DEPTH); walks stored JSON, which has no cycles, and leaves values below the cap untouched.
"_iter_fallback_targets", # max depth set (2 * ROUTER_MAX_FALLBACKS); fails closed by raising ValueError at the cap.
"_mergeable_branch", # max depth set (_MAX_SCHEMA_FLATTEN_DEPTH=32) plus a seen_refs cycle guard; passes the schema through untouched at the cap.
"json_string_leaves", # max depth set (MAX_STRUCTURED_CONTENT_SCAN_DEPTH); fails closed by raising at the cap so nothing goes unscanned.
"strict_json_schema", # harness: max depth set (DEFAULT_MAX_RECURSE_DEPTH); fails closed by raising ValueError at the cap.
"toml_value", # harness/codex: max depth set (DEFAULT_MAX_RECURSE_DEPTH); fails closed by raising OptionsMismatch at the cap.
"with_json_string_leaves", # transitively bounded: only runs on a tree json_string_leaves already walked under the cap.
"_clean_extraction", # ScaleDown: traverses the entity map built under MAX_SCHEMA_DEPTH=32 and MAX_ENTITIES=1000.
"json_unrewritable_labels", # max depth set (MAX_STRUCTURED_CONTENT_SCAN_DEPTH); returns the None sentinel at the cap so the caller blocks.
"_flatten_form_field", # bounded by the nesting depth of the already-parsed request body (a finite JSON tree, no cycles possible).
"_flatten_form_data_field", # bounded by the nesting depth of the already-parsed request body (a finite JSON tree, no cycles possible).
"_json_safe", # max depth set (_MAX_DEPTH) plus a seen-ids cycle guard for self-referential input.
"_redact_agent_params_tree", # max depth set (default 10), same shape as _redact_sensitive_litellm_params.
"_restore_redacted_nested_value", # max depth set (default 10), mirrors _redact_agent_params_tree on the write side.
"_unqualified", # bounded by the qualifier depth of a static TypedDict annotation (Annotated, Required/NotRequired, ReadOnly around one type, no cycles possible).
"_render_json", # bounded by the nesting depth of a pydantic-validated JsonValue from the operator's config (a finite JSON tree, no cycles possible).
"completion_cost", # max depth 1: recursion only fires for mixed-tier Responses WS logging objects, and each split part carries a single service_tier so _split_responses_ws_logging_object_by_service_tier returns None.
"_string_leaves", # bounded by the nesting depth of a safe_json_structure output (a finite JSON tree, no cycles possible).
"_replace_string_leaves", # bounded by the nesting depth of a safe_json_structure output (a finite JSON tree, no cycles possible).
"_sort_processed_sets", # bounded by the nesting depth of the log-record extra it walks (a finite JSON tree, no cycles possible).
"scrub_json_strings", # max depth set (MAX_SCRUB_DEPTH); fails closed by returning "[Filtered]" for anything nested past the cap.
]
class RecursiveFunctionFinder(ast.NodeVisitor):
def __init__(self):
self.recursive_functions = []
self.ignored_recursive_functions = []
def visit_FunctionDef(self, node):
# Check if the function calls itself
if any(self._is_recursive_call(node, call) for call in ast.walk(node)):
if node.name in IGNORE_FUNCTIONS:
self.ignored_recursive_functions.append(node.name)
else:
self.recursive_functions.append(node.name)
self.generic_visit(node)
def _is_recursive_call(self, func_node, call_node):
# Check if the call node is a function call
if not isinstance(call_node, ast.Call):
return False
# Case 1: Direct function call (e.g., my_func())
if isinstance(call_node.func, ast.Name) and call_node.func.id == func_node.name:
return True
# Case 2: Method call with self (e.g., self.my_func())
if isinstance(call_node.func, ast.Attribute) and isinstance(
call_node.func.value, ast.Name
):
return (
call_node.func.value.id == "self"
and call_node.func.attr == func_node.name
)
return False
def find_recursive_functions_in_file(file_path):
with open(file_path, "r") as file:
tree = ast.parse(file.read(), filename=file_path)
finder = RecursiveFunctionFinder()
finder.visit(tree)
return finder.recursive_functions, finder.ignored_recursive_functions
def find_recursive_functions_in_directory(directory):
recursive_functions = {}
ignored_recursive_functions = {}
for root, _, files in os.walk(directory):
for file in files:
print("file: ", file)
if file.endswith(".py"):
file_path = os.path.join(root, file)
functions, ignored = find_recursive_functions_in_file(file_path)
if functions:
recursive_functions[file_path] = functions
if ignored:
ignored_recursive_functions[file_path] = ignored
return recursive_functions, ignored_recursive_functions
if __name__ == "__main__":
# Example usage
# raise exception if any recursive functions are found, except for the ignored ones
# this is used in the CI/CD pipeline to prevent recursive functions from being merged
directory_path = "./litellm"
recursive_functions, ignored_recursive_functions = (
find_recursive_functions_in_directory(directory_path)
)
print("UNIGNORED RECURSIVE FUNCTIONS: ", recursive_functions)
print("IGNORED RECURSIVE FUNCTIONS: ", ignored_recursive_functions)
if len(recursive_functions) > 0:
# raise exception if any recursive functions are found
for file, functions in recursive_functions.items():
print(
f"🚨 Unignored recursive functions found in {file}: {functions}. THIS IS REALLY BAD, it has caused CPU Usage spikes in the past. Only keep this if it's ABSOLUTELY necessary."
)
file, functions = list(recursive_functions.items())[0]
raise Exception(
f"🚨 Unignored recursive functions found include {file}: {functions}. THIS IS REALLY BAD, it has caused CPU Usage spikes in the past. Only keep this if it's ABSOLUTELY necessary."
)