litellm/tests
Cursor Agent 5ad351bbd2
fix(cron_vm): veria — isolate $HOME and hide credential dotdirs from claude
The cron systemd unit's `ProtectHome=read-only` blocks writes to
/home/mateo but still allows reads. With `HOME=/home/mateo` forwarded
to the `claude` subprocess, a compromised @anthropic-ai/claude-code
release (running during the `claude --version` probe) — or a
model-directed `Read` tool call during a PDF cell (which passes
`--allowed-tools Read`) — could read host credential files like
~/.config/gh/hosts.yml (gh-host token), ~/.ssh/, or ~/.bash_history
and exfiltrate them.

Two complementary mitigations, addressing veria's exact recommendation:

1. Per-invocation isolated HOME for every `claude` subprocess:
   * cli_driver.py: drop HOME from _CLI_ENV_ALLOWLIST; create a
     fresh empty tmpdir under tempfile.gettempdir() (`PrivateTmp=true`
     keeps it on a service-private tmpfs) and pass it as HOME to
     each `claude` invocation. Cleaned up in a `finally` so
     timeouts and CLI-not-found don't leak tmpdirs.
   * run_daily.sh: the up-front `claude --version` probe also runs
     under $CLAUDE_PROBE_HOME (a per-run dir under ${WORKDIR}) so
     the probe can never reach the runtime user's real home; the
     existing `cleanup` trap removes ${WORKDIR}.
   * Closes the `os.path.expanduser('~/.config/gh/hosts.yml')`-style
     attack from a compromised CLI / model.

2. Filesystem-level hiding of credential dotdirs in the systemd unit:
   * Add `InaccessiblePaths=-/home/mateo/.config/gh -/home/mateo/.ssh
     -/home/mateo/.aws -/home/mateo/.docker -/home/mateo/.kube
     -/home/mateo/.gnupg`. The kernel hides these paths from every
     process in the unit's mount namespace, defeating the absolute-path
     attack (`Read('/home/mateo/.config/gh/...')`) that the per-
     invocation HOME override alone cannot block.
   * Drop `/home/mateo/.config/gh` from `ReadWritePaths=` (it's
     now hidden, and we pass GH_TOKEN inline to every `gh` call).
   * Pass GH_TOKEN inline to `gh repo clone` in run_daily.sh
     (was relying on host gh-cli config); the docs repo is public
     so this is a no-op functionally, but it lets us drop the
     ~/.config/gh dependency entirely.

Tests:
  * test_run_claude_uses_isolated_per_invocation_home: pin that the
    CLI subprocess never sees the parent's $HOME, and that the
    isolated HOME is a fresh tmpdir prefixed claude-cli-home-.
  * test_run_claude_isolated_home_is_distinct_per_invocation: pin that
    each call gets its own dir (no cross-call planting).
  * test_run_claude_isolated_home_cleaned_up_after_run / on_subprocess
    _failure: pin that the tmpdir is rm-rf'd on both the happy path
    and the timeout/CLI-error path.
  * test_version_probe_uses_isolated_home_not_runtime_user_home: pin
    that run_daily.sh's probe forwards $CLAUDE_PROBE_HOME, not
    ${HOME}, into its `env -i` block.
  * test_systemd_unit_credential_isolation.py (new): pin that
    InaccessiblePaths covers all credential dotdirs, that
    .config/gh is not under ReadWritePaths, and that ProtectHome
    stays at least read-only.

All 349 existing claude_code unit tests still pass.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-05-19 05:49:18 +00:00
..
agent_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
audio_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
basic_proxy_startup_tests build: migrate packaging, CI, and Docker from Poetry to uv (#25007) 2026-04-09 11:46:23 -07:00
batches_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
benchmarks Add CodSpeed performance benchmarks (#23676) 2026-03-14 18:44:36 -07:00
claude_code fix(cron_vm): veria — isolate $HOME and hide credential dotdirs from claude 2026-05-19 05:49:18 +00:00
code_coverage_tests feat: add componentized proxy deployment with gateway, backend, ui, and migrations (#27557) 2026-05-16 09:25:17 -07:00
documentation_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
enterprise chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
guardrails_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
image_gen_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
litellm Add new chat model metadata (#27313) 2026-05-06 15:15:21 -07:00
litellm-proxy-extras style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
litellm_core_utils Merge branch 'litellm_internal_staging' into litellm_staging_03_22_2026 2026-04-20 19:56:00 +05:30
litellm_utils_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
llm_responses_api_testing fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
llm_translation fix(deepseek): use native /anthropic/v1/messages endpoint and sanitize tools (#28200) 2026-05-18 18:14:13 -07:00
load_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
local_testing fix(caching): replay openai/responses bridge cache hits as chat streams (#28158) 2026-05-18 16:27:06 -07:00
logging_callback_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
mcp_tests feat: litellm shin agent oss staging 05 10 2026 (#27631) 2026-05-11 20:31:43 -07:00
multi_instance_e2e_tests
ocr_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
old_proxy_tests/tests fix: cleanup tests 2026-03-30 16:24:35 -07:00
openai_endpoints_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
otel_tests feat(prometheus): add user_email and user_alias to user budget metrics (#28155) 2026-05-18 16:28:14 -07:00
pass_through_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
pass_through_unit_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
proxy_admin_ui_tests [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
proxy_e2e_anthropic_messages_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
proxy_security_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
proxy_unit_tests fix(managed_batches): convert raw output_file_id to managed ID in CheckBatchCost poller (#27984) 2026-05-15 04:41:38 -07:00
router_unit_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
scim_tests
search_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
spend_tracking_tests chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
store_model_in_db_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_litellm fix(deepseek): use native /anthropic/v1/messages endpoint and sanitize tools (#28200) 2026-05-18 18:14:13 -07:00
unified_google_tests fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
vector_store_tests fix: drop milvus dbName and partitionNames from MILVUS_OPTIONAL_PARAMS 2026-04-30 11:51:32 -07:00
windows_tests style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
__init__.py
_flush_vcr_cache.py [Infra] Promote internal staging to main (#27245) 2026-05-05 16:15:03 -07:00
_vcr_conftest_common.py fix(tests): stabilize image-edit VCR cassettes to stop live gpt-image-1 spend (#28110) 2026-05-18 09:15:39 -07:00
_vcr_redis_persister.py fix(caching): replay openai/responses bridge cache hits as chat streams (#28158) 2026-05-18 16:27:06 -07:00
eval_swe_bench.py Prompt Compression - add it to the proxy (#25729) 2026-04-20 15:08:00 -07:00
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
README.MD
test_budget_management.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_callbacks_on_proxy.py test(callbacks): harden flaky proxy callback-leak detector (#28195) 2026-05-18 16:39:02 -07:00
test_config.py
test_debug_warning.py
test_default_encoding_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_end_users.py chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
test_entrypoint.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_health.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_keys.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_litellm_proxy_responses_config.py chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
test_logging.conf
test_models.py test: replace test_add_and_delete_models integration test with mock 2026-03-30 21:30:57 -07:00
test_new_vector_store_endpoints.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_openai_endpoints.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_organizations.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_otel_thread_leak.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_passthrough_endpoints.py
test_presidio_latency.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_proxy_server_non_root.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_ratelimit.py chore(ci): modernize model references in tests and configs (#27856) 2026-05-15 15:44:28 -07:00
test_resource_cleanup.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_service_logger_otel.py
test_spend_logs.py
test_team.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_team_logging.py test: cleanup dead tests 2026-03-28 20:49:02 -07:00
test_team_members.py
test_users.py Fix: tag budget reset must drop stale management-cache entry (#27568) 2026-05-10 00:18:55 +00:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.