Extracted from the combined skill-evolution branch. This is the runtime change set: everything that alters how a sweep executes and what it records. The packed scheduler and its measurement harness were separated onto perf/skill-evolution-packed-scheduler, which is purely additive. Correctness. The review artifact was mounted as a writable FILE inside a read-only workspace while the agent's Write tool writes atomically - temp file beside the target, then rename - so the temp create failed EROFS and the artifact was never written. Four layers gate that path and three named the old location: the CLI's own allowWrite/allowRead policy, the task corpus verify command run in its own sandbox invocation, and host_text/host_path for the host-unsafe backend. Fixing the translator exposed a second defect, since "/review-output" appears twice in "/review-output/review-output.json"; matching is now anchored to a path boundary. parse_review_output folded OSError, UnicodeError and JSONDecodeError into one message, so an artifact that was never written looked like an encoding fault; each cause now names itself. Health classification. broken_incumbent_arms inferred a broken environment from an arm resolving zero tasks, which a reviewer facing a hard corpus falsifies - Actions run 33962002890 is exactly that shape. Arms are now classified from fresh execution and evidence outcomes as UNKNOWN, OBSERVED_OK, DEGRADED or UNUSABLE, and only UNUSABLE aborts. Resolution count is not consulted. The guard names no cause: an empty artifact establishes unusable evidence, not that a mount rejected the write. Comparator reuse. Reuse accepted evidence it should have rejected: a row without a runtime_digest passed the drift lock, recorded_at was overwritten with the copy time so a row could outlive its own max_age, and the binding ignored task-asset and dependency digests although this change alters sandbox_dependencies in the review corpus. Closing the last one required preparing asset snapshots before the reuse decision, which also removes the concurrent TaskAssetCache.prepare the prefetch thread could race. These three concerns share aggregate() and _run_sweep, which is why they ship together: separating them further would mean hunk-level surgery on a function all three modify, and the risk of a silent omission outweighs the reviewability gain. 592 eval tests pass at this base. The two test_model_gateway.py failures, test_locked_litellm_translates_messages_to_offline_responses and test_openai_gateway_never_leaves_proxy_output_on_an_undrained_pipe, fail identically on origin/main in this environment. Known gap, and the reason this is not ready to merge: no test drives _run_sweep end to end. enforce_measurement_health is unit-tested including the below-breaker unusable case, and the caller wiring is pinned structurally by reading _run_sweep's compiled code object, but interruption semantics, exit precedence and persisted artifacts are not exercised through the real path. |
||
|---|---|---|
| .. | ||
| agents | ||
| analysis | ||
| bridge | ||
| configs | ||
| environments | ||
| prompts | ||
| tests | ||
| utils | ||
| workflow_bench | ||
| .env.example | ||
| .gitignore | ||
| __init__.py | ||
| constants.py | ||
| pyproject.toml | ||
| README.md | ||
| run_eval.py | ||
| tool_registry.py | ||
| uv.lock | ||
GitNexus SWE-bench Evaluation Harness
Evaluate whether GitNexus code intelligence improves AI agent performance on real software engineering tasks. Runs SWE-bench instances across multiple models and compares baseline (no graph) vs GitNexus-enhanced configurations.
What This Tests
Hypothesis: Giving AI agents structural code intelligence (call graphs, execution flows, blast radius analysis) improves their ability to resolve real GitHub issues — measured by resolve rate, cost, and efficiency.
Evaluation modes:
| Mode | What the agent gets |
|---|---|
baseline |
Standard bash tools (grep, find, cat, sed) — control group |
native |
Baseline + explicit GitNexus tools via eval-server (~100ms) |
native_augment |
Native tools + grep results automatically enriched with graph context (recommended) |
Recommended: Use
native_augmentmode. It mirrors the Claude Code model — the agent gets both explicit GitNexus tools (fast bash commands) AND automatic enrichment of grep results with callers, callees, and execution flows. The agent decides when to use explicit tools vs rely on enriched search output.
Models supported (see configs/models/ for the current list):
- Claude Haiku 4.5, Claude Sonnet 4, Claude Opus 4
- MiniMax M1 2.5, MiniMax M2.5
- GLM 4.7, GLM 5
- DeepSeek
- Any model supported by litellm (add a YAML config)
Prerequisites
- Python 3.11+
- Docker (for SWE-bench containers)
- Node.js 22+ (for GitNexus)
- API keys for your chosen models
Setup
cd eval
# Install dependencies
pip install -e .
# Set up API keys — copy the template and fill in your keys
cp .env.example .env
# Then edit .env and paste your key(s)
All models are routed through OpenRouter by default, so a single OPENROUTER_API_KEY is all you need. To use provider APIs directly (Anthropic, ZhipuAI, etc.), edit the model YAML in configs/models/ and set the corresponding key in .env.
# Pull SWE-bench Docker images (pulled on-demand, but you can pre-pull)
docker pull swebench/sweb.eval.x86_64.django_1776_django-16527:latest
Debug logging
Set GITNEXUS_EVAL_DEBUG=1 to include full Python tracebacks in run summaries and logs. By default, errors are sanitized to avoid leaking host paths or stack traces.
Quick Start
Debug a single instance
# Fastest way to verify everything works
python run_eval.py debug -m claude-haiku -i django__django-16527 --subset lite
Run a single configuration
# 5 instances, Claude Sonnet, native_augment mode (default)
python run_eval.py single -m claude-sonnet --subset lite --slice 0:5
# Baseline comparison (no GitNexus)
python run_eval.py single -m claude-sonnet --mode baseline --subset lite --slice 0:5
# Full Lite benchmark, 4 parallel workers
python run_eval.py single -m claude-sonnet --subset lite -w 4
Run the full matrix
# All models x all modes
python run_eval.py matrix --subset lite -w 4
# Key comparison: baseline vs native_augment
python run_eval.py matrix -m claude-sonnet -m claude-haiku --modes baseline --modes native_augment --subset lite --slice 0:50
Analyze results
# Summary table
python -m analysis.analyze_results results/
# Compare modes for a specific model
python -m analysis.analyze_results compare-modes results/ -m claude-sonnet
# GitNexus tool usage analysis
python -m analysis.analyze_results gitnexus-usage results/
# Export as CSV for further analysis
python -m analysis.analyze_results summary results/ --format csv > results.csv
# Run official SWE-bench test evaluation
python -m analysis.analyze_results summary results/ --swebench-eval
List available configurations
python run_eval.py list-configs
Architecture
eval/
run_eval.py # Main entry point (single, matrix, debug commands)
agents/
gitnexus_agent.py # GitNexusAgent: extends DefaultAgent with augmentation + metrics
environments/
gitnexus_docker.py # Docker env with GitNexus + eval-server + standalone tool scripts
bridge/
gitnexus_tools.sh # Bash wrappers (legacy — now standalone scripts are installed directly)
mcp_bridge.py # Legacy MCP bridge (kept for reference)
prompts/
system_baseline.jinja # System: persona + format rules
instance_baseline.jinja # Instance: task + workflow
system_native.jinja # System: + GitNexus tool reference
instance_native.jinja # Instance: + GitNexus debugging workflow
system_native_augment.jinja # System: + GitNexus tools + grep enrichment docs
instance_native_augment.jinja # Instance: + GitNexus workflow + risk assessment
configs/
models/ # Per-model YAML configs
modes/ # Per-mode YAML configs (baseline, native, native_augment)
analysis/
analyze_results.py # Post-run comparative analysis
results/ # Output directory (gitignored)
How It Works
Template structure
mini-swe-agent requires two Jinja templates:
- system_template → system message: persona, format rules, tool reference (static)
- instance_template → first user message: task, workflow, rules, examples (contains
{{task}})
Each mode has a system_{mode}.jinja + instance_{mode}.jinja pair. The agent loads both automatically based on the configured mode.
Per-instance flow
- Docker container starts with SWE-bench instance (repo at specific commit)
- GitNexus setup: Node.js + gitnexus installed,
gitnexus analyzeruns (or restores from cache) - Eval-server starts:
gitnexus eval-serverdaemon (persistent HTTP server, keeps LadybugDB warm) - Standalone tool scripts installed in
/usr/local/bin/— works withsubprocess.run(no.bashrcneeded) - Agent runs with the configured model + system prompt + GitNexus tools
- Agent's patch is extracted as a git diff
- Metrics collected: cost, tokens, tool calls, GitNexus usage, augmentation stats
Tool architecture
Agent → bash command → /usr/local/bin/gitnexus-query
→ curl http://127.0.0.1:4848/tool/query (fast path: eval-server, ~100ms)
→ npx gitnexus query (fallback: cold CLI, ~5-10s)
Each tool script in /usr/local/bin/ is standalone — no sourcing, no env inheritance needed. This is critical because mini-swe-agent runs every command via subprocess.run in a fresh subshell.
Eval-server
The eval-server is a lightweight HTTP daemon that:
- Keeps LadybugDB warm in memory (no cold start per tool call)
- Returns LLM-friendly text (not raw JSON — saves tokens)
- Includes next-step hints to guide tool chaining (query → context → impact → fix)
- Auto-shuts down after idle timeout
CLI flags:
| Flag | Default | Purpose |
|---|---|---|
--port <port> |
4848 |
Port to listen on |
--host <host> |
127.0.0.1 |
Bind address — use 0.0.0.0 for cross-container access |
--idle-timeout <seconds> |
0 (disabled) |
Auto-shutdown after N seconds of inactivity |
READY signal:
When the server is ready, it writes to stdout:
# IPv4
GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848
# IPv6 (bracketed to avoid colon ambiguity)
GITNEXUS_EVAL_SERVER_READY:[::1]:4848
Parse the port as the last colon-segment (split(':').pop()) — not split(':')[1], which breaks for IPv6 and for non-loopback IPv4 hosts added in this release.
Custom port and host
run_eval.py does not expose --port or --host as CLI flags. Configure them in your mode YAML under the environment: key:
# configs/modes/native_augment.yaml (or whichever mode you're running)
environment:
eval_server_port: 4849 # change if 4848 is already in use on the host
eval_server_host: "0.0.0.0" # bind all interfaces — needed for cross-container setups
Defaults are port: 4848 and host: 127.0.0.1 (loopback only). Use 0.0.0.0 only when the agent container needs to reach the eval-server from a separate network namespace. The health probe and tool scripts connect via the configured bind host (defaulting to 127.0.0.1), which is reachable for both loopback and all-interface binds.
"localhost" is also a valid eval_server_host value. The OS resolves it at bind time — typically 127.0.0.1 on dual-stack or IPv4-only systems, and ::1 on IPv6-only systems. The exact result depends on your /etc/hosts and gai.conf. The READY signal will reflect the actual bound address (e.g. GITNEXUS_EVAL_SERVER_READY:127.0.0.1:4848 or GITNEXUS_EVAL_SERVER_READY:[::1]:4848), not the literal string localhost. Use this when you want the server to bind to whichever loopback address the OS prefers rather than forcing IPv4.
Running eval-server directly in Docker / Docker Compose:
# Bind to all interfaces so sibling containers can reach it
gitnexus eval-server --host 0.0.0.0 --port 4848
# Then probe from a sibling container via its service hostname
curl http://eval-container:4848/health
If you need a non-default port (e.g. to avoid conflicts), pass --port <port> alongside --host. The READY signal will reflect both:
GITNEXUS_EVAL_SERVER_READY:0.0.0.0:5000
Parse the port as the last colon-segment (split(':').pop()) — safe for both IPv4 and bracketed IPv6 forms.
Index caching
SWE-bench repos repeat (Django has 200+ instances at different commits). The harness caches GitNexus indexes per (repo, commit) hash in ~/.gitnexus-eval-cache/ to avoid redundant re-indexing.
Grep augmentation (native_augment mode)
When the agent runs grep or rg, the observation is post-processed: the agent class calls gitnexus-augment on the search pattern and appends [GitNexus] annotations showing callers, callees, and execution flows for matched symbols. This mirrors the Claude Code / Cursor hook integration.
Adding Models
Create a YAML file in configs/models/:
# configs/models/my-model.yaml
model:
model_name: "openrouter/provider/model-name"
cost_tracking: "ignore_errors" # if not in litellm's cost DB
model_kwargs:
max_tokens: 8192
temperature: 0
The model name follows litellm conventions.
Metrics Collected
| Metric | Description |
|---|---|
| Patch Rate | % of instances where agent produced a patch |
| Resolve Rate | % of instances where patch passes tests (requires --swebench-eval) |
| Total Cost | API cost across all instances |
| Avg Cost/Instance | Cost efficiency |
| API Calls | Number of LLM calls |
| GN Tool Calls | How many GitNexus tools the agent used |
| Augment Hits | How many grep/find results got enriched |
| Augment Hit Rate | % of search commands that got useful enrichment |