GitNexus/eval/tests
2026-10-01 19:16:12 +03:00
..
fixtures test(eval): run the benchmark offline against a scripted provider (#3235) 2026-09-09 12:34:00 +01:00
__init__.py docs: agent development framework, GitHub templates, eval refactor (#479) 2026-03-25 06:48:41 +00:00
bench_fixtures.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
conftest.py docs: agent development framework, GitHub templates, eval refactor (#479) 2026-03-25 06:48:41 +00:00
test_ce_plugin_runtime.py feat(skills): GitNexus Engineering Tool Kits (#2566) 2026-07-19 15:07:24 +01:00
test_comparator_reuse.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_errors.py docs: agent development framework, GitHub templates, eval refactor (#479) 2026-03-25 06:48:41 +00:00
test_evolve.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_mcp_bridge.py feat(eval): evolve review skills against historical PRs 2026-09-04 05:32:31 +00:00
test_measure_evolution_cost.py feat(eval): Add bounded packed-scheduler primitives and offline replay benchmarks (#3206) 2026-09-08 08:22:12 +01:00
test_mock_provider.py test(eval): run the benchmark offline against a scripted provider (#3235) 2026-09-09 12:34:00 +01:00
test_model_gateway.py fix(eval): require finite gateway startup budgets 2026-09-05 11:03:23 +00:00
test_offline_session_integration.py test(eval): run the benchmark offline against a scripted provider (#3235) 2026-09-09 12:34:00 +01:00
test_offline_sweep_integration.py test(eval): run the benchmark offline against a scripted provider (#3235) 2026-09-09 12:34:00 +01:00
test_oracle_assets.py fix(eval): close CI and remaining review gaps 2026-09-05 10:29:58 +00:00
test_parse_run_id.py docs: agent development framework, GitHub templates, eval refactor (#479) 2026-03-25 06:48:41 +00:00
test_process_control.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_promotion_apply.py fix(eval): make evolution evidence valid and bounded 2026-09-05 10:08:35 +00:00
test_property_based.py docs: agent development framework, GitHub templates, eval refactor (#479) 2026-03-25 06:48:41 +00:00
test_proposer_sandbox.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_provider_usage.py feat(eval): record provider-native usage at the gateway instead of inferring it after translation (#3220) 2026-09-08 18:22:04 +01:00
test_provider_usage_capture.py test(eval): run the benchmark offline against a scripted provider (#3235) 2026-09-09 12:34:00 +01:00
test_reuse_round_trip.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_review_corpus.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_review_scoring.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_runner_hardening.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_sanitized_graph.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_session_progress.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_sweep_finalization.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00
test_task_assets.py Address PR review feedback (#2785) 2026-09-04 18:59:32 +00:00
test_tool_scripts.py docs: agent development framework, GitHub templates, eval refactor (#479) 2026-03-25 06:48:41 +00:00
test_workflow_bench.py chore(deps): consolidate pending dependency upgrades (#3441) 2026-10-01 19:16:12 +03:00
test_workflow_bench_evolution.py fix(eval): make evolution evidence valid and bounded 2026-09-05 10:08:35 +00:00
test_workflow_bench_sessions.py fix(eval): sweep evidence handling and measurement health, with guarded comparator reuse (#3207) 2026-09-08 09:25:45 +00:00