litellm/tests
moe-berri b69d744993
feat(lens): analyze trace workspaces with confined Python and compaction (#44640)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* feat(lens): integrate confined workspace analysis with live reviews

* fix(lens): synchronize confined Python process monitoring

* Update review.md

* fix(lens): allow mixed context capacities and correct review assertions

* fix(lens): retrieve evidence on demand and isolate failed reviews

* fix(lens): isolate incomplete evidence reads from peer reviews

* test(lens): await trace status filter option

* test(lens): wait for reclaimed review state to settle

* fix(lens): recover from incomplete cross-session evidence

---------

Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
2026-10-05 22:06:48 +00:00
..
_support fix(params): validate stream_chunk_size once, before any provider call (#43222) 2026-09-26 23:01:20 +00:00
agent_tests
audio_tests test: remove 130 legacy tests owned by stronger unit proofs (#44157) 2026-10-02 10:18:50 -07:00
base_sdk_tests fix(mcp): explain missing public client dependencies 2026-09-19 19:52:13 -07:00
basic_proxy_startup_tests
batches_tests test(e2e): move live-provider legacy tests into tests/e2e (#44120) 2026-10-02 00:02:18 -07:00
benchmarks feat(tokenizer): preserve Python defaults with opt-in Rust dispatch (#42174) 2026-09-22 04:41:11 +00:00
code_coverage_tests ci: move Postgres, MCP and Redis suites to CircleCI integration (#44453) 2026-10-05 09:33:14 -07:00
documentation_tests fix(s3_v2): upload fresh events first, drop terminal failures and hour-old retries by default, opt-in adaptive concurrency (#43022) 2026-09-26 14:58:28 -07:00
e2e fix(mcp): bind OAuth clients to their upstream issuer (#37777) 2026-10-05 10:25:47 -07:00
guardrails_tests test(e2e): move live-provider legacy tests into tests/e2e (#44120) 2026-10-02 00:02:18 -07:00
harness_e2e feat: add litellm.agent() to run claude code, codex, opencode and deep agents through the ai gateway (#43885) 2026-10-01 22:27:49 +00:00
image_gen_tests test: remove 130 legacy tests owned by stronger unit proofs (#44157) 2026-10-02 10:18:50 -07:00
integration test(integration): basic translation cases for the openai_responses route (#44663) 2026-10-05 15:06:10 -07:00
litellm_utils_tests test(ci): pin the ROI estimator flag and the prompt-cache counter in two drifted tests (#44499) 2026-10-04 08:40:23 -07:00
llm_responses_api_testing test: remove 130 legacy tests owned by stronger unit proofs (#44157) 2026-10-02 10:18:50 -07:00
llm_translation test: remove 130 legacy tests owned by stronger unit proofs (#44157) 2026-10-02 10:18:50 -07:00
load_tests
local_testing test: repair stale and polluting tests red on scheduled main CI (#44229) 2026-10-02 21:32:17 +00:00
logging_callback_tests test(ci): pin the ROI estimator flag and the prompt-cache counter in two drifted tests (#44499) 2026-10-04 08:40:23 -07:00
mcp_tests fix(mcp): preserve upstream tool schemas and parameter headers (#44425) 2026-10-03 15:00:43 -07:00
multi_instance_e2e_tests
ocr_tests refactor(ocr): remove the Python OCR execution path and require the Rust route (#43081) 2026-09-24 18:18:50 -07:00
openai_endpoints_tests test: remove 130 legacy tests owned by stronger unit proofs (#44157) 2026-10-02 10:18:50 -07:00
otel_tests test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
pass_through_tests fix(mcp): preserve legacy behavior on SDK2 and streamline verification 2026-09-18 22:28:31 -07:00
pass_through_unit_tests feat: add Bespoke Nimble gateway and OSS classifier support (#44246) 2026-10-02 19:37:08 -07:00
proxy_admin_ui_tests
proxy_behavior feat(lens): analyze trace workspaces with confined Python and compaction (#44640) 2026-10-05 22:06:48 +00:00
proxy_e2e_anthropic_messages_tests fix(test): run the all-beta-headers bedrock cases on Claude Fable 5.1 2026-09-19 23:33:29 +00:00
proxy_migration_tests fix(proxy-extras): bound the lock waits of the partitioned SpendLogs index build (#44109) 2026-10-01 18:28:50 -07:00
proxy_security_tests refactor(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key so the name says exactly what it permits 2026-09-19 18:53:14 -07:00
proxy_unit_tests ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests (#42903) 2026-09-24 22:59:11 +00:00
router_unit_tests fix(ci): stop stale CI reds, keep unit tests off the host env, retry CyberArk policy conflicts (#43294) 2026-09-26 09:25:13 -07:00
rust-python-harness refactor(ocr): remove the Python OCR execution path and require the Rust route (#43081) 2026-09-24 18:18:50 -07:00
search_tests fix(cost-map): retirement dates, chatgpt reasoning flags, bing pricing, bedrock mantle and mythos, azure gpt-5.6 alias, anthropic batch rates, new nebius, openrouter and xai rows (#42951) 2026-09-25 19:12:36 -07:00
spend_tracking_tests test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
store_model_in_db_tests test(ci): fix six CircleCI test regressions on main (#44429) 2026-10-03 20:55:01 -07:00
test_litellm fix(tracing): preserve spend identity and gateway correlation (#44421) 2026-10-03 16:00:49 -07:00
test_litellm_rust feat(tracing)!: return only data from SQL queries (#44609) 2026-10-05 12:11:59 -07:00
unified_google_tests test: remove 130 legacy tests owned by stronger unit proofs (#44157) 2026-10-02 10:18:50 -07:00
unit feat(lens): analyze trace workspaces with confined Python and compaction (#44640) 2026-10-05 22:06:48 +00:00
vector_store_tests
windows_tests fix(packaging): keep wheel paths under Windows MAX_PATH for Store Python (#43903) 2026-09-30 22:20:36 +00:00
__init__.py
_fake_openai_endpoint_server.py
_flush_vcr_cache.py
_live_test_helpers.py
_openai_record_replay_proxy.py
_process_helpers.py test: count a zombie grandchild as gone in the migrate deploy timeout test (#42570) 2026-09-22 14:59:26 -07:00
_vcr_conftest_common.py fix(tests): stop VCR recording and replaying a test's own localhost upstream (#43346) 2026-09-26 15:49:39 -07:00
_vcr_redis_persister.py
_wait_helpers.py
_ws_vcr.py
AGENTS.md ci(tests): wire tests/unit into CircleCI and drain legacy unit shards green 2026-09-20 07:05:42 +00:00
capturing_transport.py test(vcr): guard leaked cassette patches and make injected-transport embedding tests immune (#42542) 2026-09-22 14:20:18 -07:00
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD test: finish the non-proxy half of tests/test_litellm (#43281) 2026-09-25 22:43:41 -07:00
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
test_fallbacks.py test(e2e): move live-provider legacy tests into tests/e2e (#44120) 2026-10-02 00:02:18 -07:00
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
test_new_vector_store_endpoints.py
test_openai_endpoints.py test: remove 130 legacy tests owned by stronger unit proofs (#44157) 2026-10-02 10:18:50 -07:00
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py test(proxy): delete the legacy proxy test tree and shard tests/unit/proxy by glob (#44018) 2026-10-01 11:40:45 -07:00
test_resource_cleanup.py
test_rust_python_harness.py test: fix stale and state-leaking tests red on scheduled CircleCI (#43266) 2026-09-25 19:10:20 -07:00
test_service_logger_otel.py
test_spend_logs.py test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
test_team.py test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
test_team_logging.py
test_team_members.py test(proxy): move management_endpoints, management_helpers and guardrails tests into tests/unit/proxy (#44003) 2026-10-01 10:52:03 -07:00
test_users.py test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128) 2026-10-01 23:00:54 -07:00
white_100x100.png test: stop CI tests from downloading tokenizer files and images (#43257) 2026-09-25 19:27:48 -07:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/unit

This folder can only run mock tests.