mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-05 02:41:56 +00:00
18 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0f6a06a6b9
|
feat: add litellm.agent() to run claude code, codex, opencode and deep agents through the ai gateway (#43885)
* feat(harness): add litellm/__init__.py * feat(harness): add litellm/constants.py * feat(harness): add litellm/harness/__init__.py * feat(harness): add litellm/harness/adapters/__init__.py * feat(harness): add litellm/harness/adapters/base.py * feat(harness): add litellm/harness/adapters/claude_code.py * feat(harness): add litellm/harness/adapters/codex.py * feat(harness): add litellm/harness/adapters/opencode.py * feat(harness): add litellm/harness/endpoint.py * feat(harness): add litellm/harness/errors.py * feat(harness): add litellm/harness/options.py * feat(harness): add litellm/harness/runtime.py * feat(harness): add litellm/harness/sandbox/__init__.py * feat(harness): add litellm/harness/sandbox/base.py * feat(harness): add litellm/harness/sandbox/docker.py * feat(harness): add litellm/harness/sandbox/local.py * feat(harness): add litellm/harness/sandbox/snapshot.py * feat(harness): add litellm/harness/sync.py * feat(harness): add litellm/harness/types.py * feat(harness): add litellm/sandbox/__init__.py * feat(harness): add README.md * feat(harness): add tests/harness_e2e/__init__.py * feat(harness): add tests/harness_e2e/conftest.py * feat(harness): add tests/harness_e2e/test_harness_e2e.py * feat(harness): add tests/test_litellm/harness/__init__.py * feat(harness): add tests/test_litellm/harness/adapters/__init__.py * feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl * feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl * feat(harness): add tests/test_litellm/harness/adapters/test_claude_code.py * feat(harness): add tests/test_litellm/harness/adapters/test_codex.py * feat(harness): add tests/test_litellm/harness/adapters/test_opencode.py * feat(harness): add tests/test_litellm/harness/core_fakes.py * feat(harness): add tests/test_litellm/harness/sandbox/__init__.py * feat(harness): add tests/test_litellm/harness/sandbox/test_docker.py * feat(harness): add tests/test_litellm/harness/sandbox/test_local.py * feat(harness): add tests/test_litellm/harness/sandbox/test_snapshot.py * feat(harness): add tests/test_litellm/harness/test_endpoint.py * feat(harness): add tests/test_litellm/harness/test_init.py * feat(harness): add tests/test_litellm/harness/test_runtime.py * feat(harness): add tests/test_litellm/harness/test_sync.py * feat(harness): add tests/test_litellm/harness/test_types.py * test(harness): use word recall in stream e2e test * refactor(harness): update litellm/__init__.py * refactor(harness): update litellm/constants.py * refactor(harness): update litellm/harness/__init__.py * refactor(harness): remove litellm/harness/adapters/__init__.py * refactor(harness): update litellm/harness/context.py * refactor(harness): update litellm/harness/endpoint.py * refactor(harness): update litellm/harness/handlers/__init__.py * refactor(harness): update litellm/harness/handlers/base.py * refactor(harness): update litellm/harness/handlers/cli_handler.py * refactor(harness): update litellm/harness/handlers/deepagents_handler.py * refactor(harness): update litellm/harness/runtime.py * refactor(harness): update litellm/harness/sandbox/docker.py * refactor(harness): update litellm/harness/sandbox/local.py * refactor(harness): update litellm/harness/sync.py * refactor(harness): update litellm/harness/types.py * refactor(harness): update litellm/llms/base_llm/harness/__init__.py * refactor(harness): update litellm/llms/base_llm/harness/transformation.py * refactor(harness): update litellm/llms/base_llm/harness/utils.py * refactor(harness): update litellm/llms/claude_code/__init__.py * refactor(harness): update litellm/llms/claude_code/harness/__init__.py * refactor(harness): update litellm/llms/claude_code/harness/transformation.py * refactor(harness): update litellm/llms/codex/__init__.py * refactor(harness): update litellm/llms/codex/harness/__init__.py * refactor(harness): update litellm/llms/codex/harness/transformation.py * refactor(harness): update litellm/llms/deepagents/__init__.py * refactor(harness): update litellm/llms/deepagents/harness/__init__.py * refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py * refactor(harness): update litellm/llms/deepagents/harness/transformation.py * refactor(harness): update litellm/llms/opencode/__init__.py * refactor(harness): update litellm/llms/opencode/harness/__init__.py * refactor(harness): update litellm/llms/opencode/harness/transformation.py * refactor(harness): update litellm/utils.py * refactor(harness): update README.md * refactor(harness): update tests/harness_e2e/conftest.py * refactor(harness): update tests/harness_e2e/test_harness_e2e.py * refactor(harness): remove tests/test_litellm/harness/__init__.py * refactor(harness): remove tests/test_litellm/harness/adapters/__init__.py * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl * refactor(harness): remove tests/test_litellm/harness/adapters/test_claude_code.py * refactor(harness): remove tests/test_litellm/harness/adapters/test_codex.py * refactor(harness): remove tests/test_litellm/harness/adapters/test_opencode.py * refactor(harness): remove tests/test_litellm/harness/core_fakes.py * refactor(harness): remove tests/test_litellm/harness/sandbox/__init__.py * refactor(harness): remove tests/test_litellm/harness/sandbox/test_docker.py * refactor(harness): remove tests/test_litellm/harness/sandbox/test_local.py * refactor(harness): remove tests/test_litellm/harness/sandbox/test_snapshot.py * refactor(harness): remove tests/test_litellm/harness/test_endpoint.py * refactor(harness): remove tests/test_litellm/harness/test_init.py * refactor(harness): remove tests/test_litellm/harness/test_runtime.py * refactor(harness): remove tests/test_litellm/harness/test_sync.py * refactor(harness): remove tests/test_litellm/harness/test_types.py * refactor(harness): update tests/unit/harness/__init__.py * refactor(harness): update tests/unit/harness/core_fakes.py * refactor(harness): update tests/unit/harness/handlers/__init__.py * refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py * refactor(harness): update tests/unit/harness/sandbox/__init__.py * refactor(harness): update tests/unit/harness/sandbox/test_docker.py * refactor(harness): update tests/unit/harness/sandbox/test_local.py * refactor(harness): update tests/unit/harness/sandbox/test_snapshot.py * refactor(harness): update tests/unit/harness/test_endpoint.py * refactor(harness): update tests/unit/harness/test_init.py * refactor(harness): update tests/unit/harness/test_runtime.py * refactor(harness): update tests/unit/harness/test_sync.py * refactor(harness): update tests/unit/harness/test_types.py * refactor(harness): update tests/unit/llms/claude_code/__init__.py * refactor(harness): update tests/unit/llms/claude_code/harness/__init__.py * refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/api_error.jsonl * refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/max_turns.jsonl * refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/resume_turn.jsonl * refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/structured_output.jsonl * refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/success_tools.jsonl * refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py * refactor(harness): update tests/unit/llms/codex/__init__.py * refactor(harness): update tests/unit/llms/codex/harness/__init__.py * refactor(harness): update tests/unit/llms/codex/harness/fixtures/reasoning.jsonl * refactor(harness): update tests/unit/llms/codex/harness/fixtures/structured_output.jsonl * refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn1_bash.jsonl * refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn2_resume_apply_patch.jsonl * refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn_failed.jsonl * refactor(harness): update tests/unit/llms/codex/harness/test_transformation.py * refactor(harness): update tests/unit/llms/deepagents/__init__.py * refactor(harness): update tests/unit/llms/deepagents/harness/__init__.py * refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py * refactor(harness): update tests/unit/llms/opencode/__init__.py * refactor(harness): update tests/unit/llms/opencode/harness/__init__.py * refactor(harness): update tests/unit/llms/opencode/harness/fixtures/api_error.jsonl * refactor(harness): update tests/unit/llms/opencode/harness/fixtures/endpoint_requests.jsonl * refactor(harness): update tests/unit/llms/opencode/harness/fixtures/readonly_denied_bash.jsonl * refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn1_write_read.jsonl * refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn2_session_skill.jsonl * refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py * ci: allowlist tests/harness_e2e, which needs live runtimes and a gateway * fix(harness): bound the turn event queue * fix(harness): bound the turn event queue with backpressure * style: sort imports in utils * ci: exclude agent-harness config folders from provider docs check * refactor(harness): update tests/harness_e2e/conftest.py * refactor(harness): update tests/harness_e2e/test_harness_e2e.py * refactor(harness): update tests/unit/harness/core_fakes.py * refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py * refactor(harness): update tests/unit/harness/test_runtime.py * refactor(harness): update tests/unit/harness/test_sync.py * refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py * refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py * refactor(harness): update litellm/harness/endpoint.py * refactor(harness): update litellm/harness/handlers/deepagents_handler.py * refactor(harness): update litellm/harness/options.py * refactor(harness): update litellm/harness/runtime.py * refactor(harness): update litellm/harness/sandbox/snapshot.py * refactor(harness): update litellm/llms/base_llm/harness/transformation.py * refactor(harness): update litellm/llms/base_llm/harness/utils.py * refactor(harness): update litellm/llms/claude_code/harness/transformation.py * refactor(harness): update litellm/llms/codex/harness/transformation.py * refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py * refactor(harness): update litellm/llms/deepagents/harness/transformation.py * refactor(harness): update litellm/llms/opencode/harness/transformation.py * refactor(harness): update litellm/sandbox/__init__.py * refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py * refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py * refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py * refactor(harness): update litellm/llms/base_llm/harness/utils.py * refactor(harness): update litellm/llms/codex/harness/transformation.py * refactor(harness): update tests/code_coverage_tests/recursive_detector.py * refactor(harness): update tests/unit/llms/base_llm/harness/__init__.py * refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/__init__.py * refactor(harness): update tests/unit/llms/codex/harness/fixtures/__init__.py * refactor(harness): update tests/unit/llms/opencode/harness/fixtures/__init__.py * refactor(harness): update litellm/harness/__init__.py * refactor(harness): update litellm/harness/context.py * refactor(harness): update litellm/harness/endpoint.py * refactor(harness): update litellm/harness/handlers/__init__.py * refactor(harness): update litellm/harness/handlers/base.py * refactor(harness): update litellm/harness/handlers/cli_handler.py * refactor(harness): update litellm/harness/handlers/deepagents_handler.py * refactor(harness): update litellm/harness/runtime.py * refactor(harness): update litellm/harness/sandbox/__init__.py * refactor(harness): update litellm/harness/sandbox/base.py * refactor(harness): update litellm/harness/sandbox/docker.py * refactor(harness): update litellm/harness/sandbox/local.py * refactor(harness): update litellm/harness/sandbox/snapshot.py * refactor(harness): update litellm/harness/sync.py * refactor(harness): update litellm/harness/types.py * refactor(harness): update litellm/llms/base_llm/harness/transformation.py * refactor(harness): update litellm/llms/base_llm/harness/utils.py * refactor(harness): update litellm/llms/claude_code/harness/transformation.py * refactor(harness): update litellm/llms/codex/harness/transformation.py * refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py * refactor(harness): update litellm/llms/deepagents/harness/transformation.py * refactor(harness): update litellm/llms/opencode/harness/transformation.py * refactor(harness): update litellm/types/llms/custom_http.py * refactor(harness): update tests/unit/harness/test_endpoint.py * refactor(harness): update litellm/harness/runtime.py * refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py * refactor(harness): update tests/unit/harness/test_init.py * refactor(harness): update tests/unit/harness/test_runtime.py * refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py |
||
|
|
438bffc26e
|
build(rust): package the gateway container (#43471)
* build(rust): package the gateway container * ci: exempt the gateway Dockerfile from the CI coverage gate Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Yujong Lee <yujong@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2b87b3c873
|
fix(redis): authenticate sync clusters with IAM credential providers (#40204)
* fix(redis): authenticate sync clusters with IAM credential providers Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com> * test(redis): exercise IAM cluster authentication over TCP Run Azure and GCP regressions against a real local cluster with only cloud token issuance stubbed. Build a checksum-verified Redis server in the compatibility workflow and report its coverage. Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com> * test(redis): separate unit and cluster integration coverage Keep the mapped test tree mock-only. Run the live cluster cases from the existing local caching integration file, selected by explicit node IDs in the Redis compatibility workflow. Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com> * test(ci): isolate workflow coverage audit fixtures Replace the stale unrun caching-file assumption with isolated workflow fixtures for file and node-ID selectors. Keep the unnamed-file negative check and clarify which live caching cases remain outside CI. Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com> --------- Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com> Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
754bdc0a25 | Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_ai_gateway_image_build | ||
|
|
aec083cdac
|
feat(proxy): per-worker admission control that rejects excess requests with 503 (#39352)
* feat(proxy): reject excess per-worker requests with 503 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): drop redundant suppressions in admission middleware Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): exempt the /metrics/ redirect target from admission control Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: allowlist live Granian saturation benchmark Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): regenerate dashboard API types Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): normalize root_path for admission exemptions, validate settings, inject state Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): cover prometheus metric factory, lifespan scope, and prefix lookalike paths Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): queue behind pending waiters, cache admission settings parsing, log invalid limits Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): simplify invalid admission settings handling Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2c30fe16b0
|
Merge pull request #38765 from BerriAI/litellm_ocr_sdk_parity_tests
test(harness): add OCR parity with migration strategy runners |
||
|
|
fbb799e240 | ci: drop the ai-gateway image from the coverage allowlist now that a job builds it | ||
|
|
ae0e8a20db
|
fix(ci): run the migration DDL guard, and stop it reading comments as SQL (#37791)
* fix(ci): make the migration DDL guard run, and stop it reading comments as SQL TestMigrationSQLIdempotency requires guarded DDL across litellm-proxy-extras and has never run in any job, so the convention eroded quietly. Four of its assertions fail today, and it was allowlisted rather than wired up because fixing the migrations is not an option: Prisma checksums an applied migration, so editing one breaks `migrate deploy` for every existing install. Two things were wrong with the guard itself. It scanned raw lines, so Prisma's own `-- CREATE INDEX CONCURRENTLY ...` explanations counted as the statements they describe, which is two of the reported migrations. And it had no way to say "these predate the rule", so the only options were editing immutable files or leaving the whole file unrun. Comments are now stripped before matching, on the drop-column rule too, and the migrations that already violate are named once in _PRE_GUARD_MIGRATIONS. The rules bind everything after them, so a new migration with bare CREATE TABLE, ADD COLUMN, CREATE INDEX or an unguarded ADD CONSTRAINT now fails a check instead of landing unnoticed. That set is 14 migrations, not the 13 previously recorded, measured after comment-stripping. It can only shrink: a test fails if an entry names no migration on disk, and another fails if an entry no longer violates anything. The file now runs as a proxy-extras shard and comes off the coverage allowlist. * fix(ci): strip block comments in the migration guard too Prisma opens a destructive migration with a /* Warnings: You are about to drop the column ... */ header. Nothing in the tree trips a rule on that text today, but it is prose about a statement rather than the statement, and the line-comment fix left the class open. Bodies are blanked rather than removed so the reported line number still points at the real statement. |
||
|
|
b31484ed19
|
ci: run the keyless caching tests that ran in no job (#37790)
The allowlist recorded eight files in tests/local_testing, 118 tests, that every job globbing that directory then deselects: local_testing_part1 and part2 carry `-k "... and not caching and not cache"`, and the other three keep one unrelated keyword each. They counted as covered while running nowhere. Five of the eight need nothing. Measured with no provider credentials and no Redis: test_cache_preset_key, test_caching_handler, test_prompt_caching, test_responses_stream_cache_keys and test_unit_test_caching pass, 45 tests together, and they now run as a caching-local shard. The other three stay allowlisted with what they actually need recorded rather than a question: test_caching wants Redis and a provider key for 37 of its 65, disk-cache wants OPENAI_API_KEY for 2 of 4, gcs-cache wants GCS credentials for all 4. Taking them off the allowlist exposed a gap in the slice guard itself: it reasoned only about CircleCI `-k` expressions, so a file every slice drops read as unrun even when a workflow names it outright. It now credits workflow test-paths the way the census already does, and only workflows, so a tree only CircleCI globs is still reported. |
||
|
|
13d4074492
|
test(mcp): retire the last file of the dead tests/litellm mirror
tests/litellm/ was a second mirror beside tests/test_litellm/ that no workflow, Makefile target, or CircleCI job ever named. Its other 33 files were reconciled during August 2026; this one stayed behind under a ci-coverage-allowlist entry asking a later pass to decide which of its five orphan behaviours still hold. They no longer hold as written: 25 of its 32 cases fail against today's code, because the file froze on the day it stopped being collected and the endpoints kept moving. Three of the five are already covered by the live twin, and better. test_get_request_base_url_xff_trust_gate parametrizes the trust gate in both directions, including the exact untrusted-caller case the orphan asserted, and the standard and legacy protected-resource shapes are both exercised through use_standard_pattern. The other two were the only tests anywhere for validate_trusted_redirect_uri under that same gate, so they are ported rather than dropped, rebuilt on the live file's request-mock conventions. Both directions are load-bearing: forcing is_request_from_trusted_proxy to True fails the untrusted case, forcing it to False fails the trusted one. 313 tests pass in the live file, up from 311. Dropping the dead file clears one zero-assert TQ001 violation, so its ceiling ratchets down with it. |
||
|
|
8a18e24faa
|
test: merge three stranded twins into the files that shadow them (#37600)
* test: merge three stranded twins into the files that shadow them
The second mirror's last four files each share a filename with a live test, so
the previous commit could not move them. Three of the four turn out to be plain
additions: their classes collide with nothing in the live file, so the tests are
extra coverage that has sat unrun rather than a competing version of anything.
Appending them takes the three files from 156 collected tests to 196, and all
196 pass. The 40 recovered are 13 OCI cases covering key normalization,
credential validation, complete-URL building and image-url transformation, 15
management-endpoint cases covering empty-value handling and the premium check,
and 12 DeepSeek thinking-parameter cases.
One assertion had to change. test_map_reasoning_effort_none_does_not_enable_thinking
asserted that reasoning_effort='none' leaves no thinking key, while the handler
maps it to {'type': 'disabled'} on purpose, documented in map_openai_params as
the OpenAI-style way to ask for thinking off. The test's stated intent holds,
since disabled does not enable anything, so it now asserts the disabled mapping
instead of the key's absence. Two imports moved to module scope for the
appended code, and no live test was touched.
test_discoverable_endpoints.py is the one left. Its twin grew from 1268 lines
to 9434, 25 of its assertions fail against today's code, and only 5 of its 19
tests have no counterpart, so deciding what survives that rewrite is a
judgement about the endpoints rather than a merge. The allowlist now holds
exactly that file and that reasoning.
* test(oci): stop the OCI suite reading credentials from the environment
validate_environment falls back to os.environ for every OCI credential and only
defaults the region when OCI_REGION is unset, so on a machine with OCI
configured the missing-credential test finds credentials it never passed and the
default-region test builds a URL for the ambient region. The suite then passes
or fails depending on who runs it.
A fixture drops the seven OCI variables for the four classes this branch added
and for TestOCIChatConfig, which had the same dependency before any of this and
fails the same way: with OCI_USER and friends exported, two of its cases fail on
origin/litellm_internal_staging today.
clean env: 83 passed
ambient OCI env: 83 passed
Same numbers either way, where the pre-existing file gave 68 passed / 2 failed
under the second.
|
||
|
|
3357ec8d34
|
test: run the 30 test files stranded in the second mirror (#37595)
* test: run the 30 test files stranded in the second mirror
tests/litellm sat beside tests/test_litellm, which is the mirror the repo
convention names, and no job collected it. The allowlist called the directory
unresolved and assumed it was a duplicate. It is not: 30 of its 34 files have no
counterpart in the real mirror, so they are tests nobody has run since they were
written, not copies of tests that run elsewhere.
Moving them in is byte-identical, and it is what makes them run. Every one is
now claimed by a shard's test-path rather than by an allowlist entry, and the
216 tests they hold pass. Directories that needed to become packages did, since
several files are named test_transformation.py and pytest cannot import two of
those from non-package directories in one session.
Never running is why three assertions had drifted away from the code:
* nvidia.nemotron-super-3-120b max_output_tokens, 32000 -> 32768
* sambanova/MiniMax-M2.7 max_input_tokens, 204800 -> 196608
* the Vertex text-to-speech handler moved from data= to json=, so the test
reads the decoded body off the json kwarg instead of parsing the data one
The first two follow model_prices_and_context_window.json, which the catalog
sync keeps current; the third follows the handler. In all three the test was the
stale side.
The lint workflow ran test_no_hardcoded_secrets.py by path and now points at the
new one.
Four files stay behind. Each shares a filename with a live test whose contents
are disjoint from it, so landing those means merging test bodies, which is a
content review rather than a move. The allowlist entry now names those four and
records how many tests each would bring, in place of calling the whole
directory unresolved.
* fix(ci): keep the secret scan out of the mirror's conftest
The secret-scan job runs pytest under uv run --no-project, so its environment
holds pytest and nothing else. That worked while the file sat in tests/litellm,
which has no conftest, and broke the moment it moved into tests/test_litellm,
whose conftest imports litellm on collection: ModuleNotFoundError: No module
named 'dotenv', before a single test ran.
The file is a repo-wide static scan that imports only base64, os, re and pytest,
so it belongs with the other repo-wide checks in tests/code_coverage_tests,
which has no conftest, rather than in the package mirror. Installing the full
dependency set into a 15-second job to satisfy a conftest it does not use would
be the wrong trade.
Verified with the job's exact command:
uv run --no-project --with 'pytest==9.0.2' pytest \
tests/code_coverage_tests/test_no_hardcoded_secrets.py -q
1 passed in 0.47s
|
||
|
|
a48baefc95
|
feat(ci): catch files a -k expression deselects from every job (#37601)
* feat(ci): catch files a -k expression deselects from every job The coverage census asks whether some job names a file. It cannot ask what that job's -k then does with it, and the gap is not hypothetical: tests/local_testing is globbed by five jobs, two of which carry -k "... and not router and not assistants and not langfuse and not caching and not cache" while the other three keep one keyword each. Any file whose path holds an excluded term is dropped by the first two and matched by none of the rest, so it runs nowhere while the census counts it as covered. 118 tests across eight caching files sit in exactly that hole today. The new mode reads the same CircleCI jobs the census already parses and asks whether each globbed file survives its job's selector. Two facts about -k make that decidable without running pytest: it matches an item's own name and its parents', so a term appearing in the module path deselects the whole file; and the names it can match are otherwise the classes and functions in the file, which ast reads. A positive term is therefore satisfied by the path or by a name inside, which is what keeps a langfuse-named test inside test_logging.py from being reported. Where the parser is unsure it stays quiet. An expression with or, parentheses, or a negated group is left unmodelled and its job is treated as claiming everything it globs, so an unparsed selector can never raise a false alarm. Glob translation learned character classes, without which tests/local_testing/**/test_[a-mA-M]*.py matches nothing and the guard would report that whole directory. The census and shard counts are unchanged by it, 2423 files and 327 shard children before and after. Validated against the real thing: collecting tests/local_testing under each job's own selector leaves 175 of 1577 tests unselected, in exactly the ten files this check derives statically, no more and no fewer. Two of the ten are named outright by other jobs, which the check credits, leaving the eight now recorded in the allowlist as a decision rather than an accident. Verified red-first: dropping one of those eight from the allowlist reports it, and adding 'and not embedding' to the two part jobs reports test_embedding.py and test_get_optional_params_embeddings.py. * fix(ci): keep the slice guard from pairing one command's -k with another's glob Two accuracy notes from review, both about the parser's model rather than its current verdicts. A job that runs several pytest commands offers no way to tell which glob a -k belongs to, since both are read out of the same flattened job text. Combining them could pair one command's exclusion with another command's glob and report a file that in fact runs. Such a job is now left unmodelled, which means it claims everything it globs, matching how the parser already treats an expression it cannot read. Only one job in the config has two globs today and it carries no -k at all, so no verdict changes. The second is a deliberate limit, now stated where it lives: an excluded term is only honoured when it sits in the module path, because that is the case that takes the whole file with it. A term matching one function inside drops that test and leaves the file running, and reporting it would be a false alarm. Answering per-test instead would need a baseline of test ids that churns on every rename, for a smaller failure than a file going dark. Both are pinned by tests. |
||
|
|
9b00fd9dd9
|
test: settle three allowlist entries that were open questions (#37598)
* test: settle three allowlist entries that were open questions The allowlist is meant to hold decisions, not deferrals, so an entry reading 'needs moving' or 'referenced by no job' is a gap wearing an exemption. These three each get an answer. The two prompt-factory tests move into the mirror, which is what their own entry said they needed. Both were passing the whole time, so the 23 tests they hold start running and the entry goes away rather than getting reworded. test_aio_http_image_conversion.py is not a test. It fetches live image URLs, times aiohttp against httpx, prints the ratio, and asserts nothing, and pytest cannot collect it because its functions take arguments rather than fixtures. Running it beside its siblings would buy CI a network dependency and a number nothing reads, so it stays exempt with that written down. test_litellm_proxy_extras_utils.py stays exempt with a measured reason. 24 of its 28 tests pass; the 4 in TestMigrationSQLIdempotency fail because nine migrations from 2026-04 onward use bare CREATE TABLE, ADD COLUMN and CREATE INDEX where that file requires guarded forms. The convention eroded quietly precisely because the test enforcing it has never run. Wiring it up is blocked on what to do about those migrations, and editing them is not the answer, since Prisma checksums an applied migration and a changed one breaks migrate deploy for existing installs. Allowlist entries 10 -> 9, paths 88 -> 86. * docs(ci): correct the migration count in the proxy-extras allowlist reason |
||
|
|
7de8f458e0
|
test: retire tests/old_proxy_tests, which holds no tests (#37605)
* test: retire tests/old_proxy_tests, which holds no tests Twenty files named test_*.py, and pytest collects nothing from any of them: uv run pytest tests/old_proxy_tests --collect-only -q no tests collected, 16 errors in 114.17s They are manual snippets against a running proxy, written at module level with no test function, no assertion and no entry point, so the only thing the name buys them is a place on the coverage allowlist. Sixteen of the twenty cannot even be imported in this environment, wanting langchain, llama_index or google.api_core, and ten still point at 0.0.0.0:8000, which stopped being the proxy's default port some time ago. Nothing outside the directory refers to it apart from the allowlist entry, which goes with it. The other loose contents go too: five load_test_*.py scripts, a bursty variant, two committed log files, an essay fixture and a stray .js snippet. Allowlist paths 88 -> 68, test files 2422 -> 2402, and no job loses anything it was running. Recoverable from history if a snippet turns out to be someone's habit. * test: drop the retired old_proxy_tests paths from the coverage allowlist |
||
|
|
5b1a9563d6
|
chore(ci): close the test-census blind spots and move scripts out of workflows/ (#37586)
The agent job's CircleCI glob collected `tests/agent_tests/**/test_*.py` and then piped it through `grep -v` to drop `local_only_agent_tests/`. `assert_ci_coverage.py` reads the glob but not the pipeline, so those two files looked covered and were invisible to the census. The glob now excludes them structurally and they carry an allowlist entry instead, which is a decision on the record rather than a hidden filter. The collected file set is unchanged: `tests/agent_tests/` holds exactly one CI-runnable test at the top level. `tests/scim_tests/` held a single JSON fixture and no tests, referenced from nowhere. `.github/workflows/` is for workflows. Both stray scripts move to `.github/scripts/` with their callers updated: the price-file updater is invoked by `auto_update_price_and_context_window.yml`, and the translation-report runner by `make test-llm-translation`. The audit listed the latter as orphaned, but Makefile line 317 still runs it, so it moves rather than being deleted. The rollout heads-up workflow was a deliberate one-shot for the agent-shin rollout. That rollout is done, the triage and auto-close workflows have been running daily since June, so the pre-flip warning window is long past. Its script and dedicated test go with it, and the sibling workflow-invariant test drops its entry. |
||
|
|
6ba744b340
|
test(docker): gate the componentized gateway and backend images on an arbitrary-uid offline boot (#36136) | ||
|
|
7a5d6a0548
|
ci: fail the build when a test file or Dockerfile is invoked by no job (#35991)
The absence of a CI signal is indistinguishable from a passing one, and that shape has now produced several independent holes: whole test directories no job runs, and shipped images no job builds. Nothing was watching for either, so each was found by accident. assert_ci_coverage.py enumerates every test_*.py under tests/ and every Dockerfile in the repo, then credits only what the workflows and the CircleCI config actually invoke. Paths filters and lint steps that merely name a directory do not count as coverage, because crediting a mention is the same mistake one level up. Anything neither invoked nor listed in .github/ci-coverage-allowlist.yml with a written reason fails the job. The guard found 275 uncovered test files and 6 unbuilt Dockerfiles. Fixed here: the root Dockerfile, the primary published image, now gets an image-scan leg that builds it and runs the offline migration check against it, and 8 tests/test_litellm subdirectories join the shards that already enumerate their siblings. Everything else is allowlisted per file, so a new file cannot inherit an exemption, and the remaining decisions are tracked rather than invisible. The job reports but is not in branch protection, so it does not block merges; promoting it is a separate change once the allowlist has survived contact with a few pull requests. |