mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-09-30 01:51:20 +00:00
Merge upstream main into codex/atlascloud-wiki-provider
Resolve provider-list conflicts while preserving Atlas Cloud and Grok support. Isolate Atlas Cloud fallback key tests for both present and absent GITNEXUS_API_KEY cases.
This commit is contained in:
commit
eb94b41ea1
840 changed files with 302328 additions and 6874 deletions
|
|
@ -6,7 +6,7 @@
|
|||
"plugins": [
|
||||
{
|
||||
"name": "gitnexus",
|
||||
"version": "1.6.9",
|
||||
"version": "1.6.11",
|
||||
"source": {
|
||||
"source": "local",
|
||||
"path": "./gitnexus-claude-plugin"
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@
|
|||
"plugins": [
|
||||
{
|
||||
"name": "gitnexus",
|
||||
"version": "1.6.9",
|
||||
"version": "1.6.11",
|
||||
"source": "./gitnexus-claude-plugin",
|
||||
"description": "Code intelligence powered by a knowledge graph. Provides execution flow tracing, blast radius analysis, and augmented search across your codebase."
|
||||
}
|
||||
|
|
|
|||
|
|
@ -21,13 +21,21 @@ Run from the project root. This parses all source files, builds the knowledge gr
|
|||
|
||||
| Flag | Effect |
|
||||
| -------------- | ---------------------------------------------------------------- |
|
||||
| `--watch` | Keep a Git repository index current with serialized refreshes |
|
||||
| `--debounce <ms>` | Watch quiet period before refresh (default: 300 ms) |
|
||||
| `--force` | Force full re-index even if up to date |
|
||||
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
|
||||
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
|
||||
| `--pdg` | Build the program-dependence layers used by `explain` and `pdg_query` (taint, CDG, and REACHING_DEF). |
|
||||
| `--spring-actuator <path>` | Import opt-in Spring Boot Actuator mappings, beans, conditions, configprops, and env snapshots. Forces a full rebuild; unsupported with `--watch`. |
|
||||
| `--asyncapi-spec <path>` | Read opt-in AsyncAPI 3.x documents (directory or single file) and mint `Destination` nodes from their operations. 2.x is refused, not mapped. Unsupported with `--watch`. |
|
||||
|
||||
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
|
||||
|
||||
For Spring runtime enrichment, pass a JSON bundle, one endpoint JSON file, or a directory containing endpoint files. Route evidence is authoritative only when `runtimeConfirmed === true`; `runtimeSource` records provenance and may also accompany `handler-conflict`. Env/configprops values are never persisted.
|
||||
|
||||
Use `node .gitnexus/run.cjs analyze --watch` for a long-lived local Git repository. It performs an initial analysis, queues scanner-admitted file changes, and retries intact failed batches with bounded backoff. Watch refreshes update only the graph: they skip AGENTS.md / CLAUDE.md injection and standard skill installation, so run a one-shot `analyze` when those generated files need updating. Watch rejects one-shot or context-output flags including `--force`, embedding flags, `--skills`, `--default-branch`, `--skip-agents-md`, `--skip-skills`, `--no-stats`, `--self-commit`, `--index-only`, and `--skip-git`. It never pulls remotes. Scheduled remote clone/pull is a different command: `gitnexus auto-sync`. Bare `gitnexus watch` is reserved and does not start either job. Running MCP and `serve` processes periodically check for a published replacement and reopen it without a restart. MCP checks are throttled to once every five seconds, so a tool call before the next check can briefly use the previous index.
|
||||
|
||||
### status — Check index freshness
|
||||
|
||||
```bash
|
||||
|
|
@ -55,15 +63,19 @@ Deletes the `.gitnexus/` directory and unregisters the repo from the global regi
|
|||
node .gitnexus/run.cjs wiki
|
||||
```
|
||||
|
||||
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
|
||||
Generates repository documentation from the knowledge graph using an LLM. HTTP providers require an API key (saved to `~/.gitnexus/config.json` on first use). Local CLI providers (`--provider cursor|claude|codex|opencode|grok`) use your existing CLI login.
|
||||
|
||||
| Flag | Effect |
|
||||
| ------------------- | ----------------------------------------- |
|
||||
| `--force` | Force full regeneration |
|
||||
| `--force` | Force full regeneration, also required to re-generate an existing wiki in a different language |
|
||||
| `--provider <name>` | LLM provider: minimax, openai, openrouter, azure, custom, cursor, claude, codex, opencode, or grok (default: minimax). Local CLIs (`cursor`, `claude`, `codex`, `opencode`, `grok`) use your existing CLI login and skip `--api-key`. |
|
||||
| `--model <model>` | LLM model (default: MiniMax-M3) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
| `--timeout <seconds>` | LLM request timeout in seconds (default: disabled) |
|
||||
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
|
||||
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese) |
|
||||
| `--gist` | Publish wiki as a public GitHub Gist |
|
||||
|
||||
### list — Show all indexed repos
|
||||
|
|
@ -82,5 +94,5 @@ Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_
|
|||
## Troubleshooting
|
||||
|
||||
- **"Not inside a git repository"**: Run from a directory inside a git repo
|
||||
- **Index is stale after re-analyzing**: Restart Claude Code to reload the MCP server
|
||||
- **Index is stale after re-analyzing**: Wait for the next MCP tool call to reopen the published index; this normally takes no more than five seconds
|
||||
- **Embeddings slow**: Omit `--embeddings` (it's off by default) or set `OPENAI_API_KEY` for faster API-based embedding
|
||||
|
|
|
|||
|
|
@ -13,9 +13,28 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
|||
- "This endpoint returns 500"
|
||||
- Investigating bugs, errors, or unexpected behavior
|
||||
|
||||
## Bind the repository first
|
||||
|
||||
A root cause traced in the wrong repository is a wrong root cause.
|
||||
|
||||
Call `list_repos {}` before the first tool call. With one indexed repository,
|
||||
use the examples below as written. With more than one, pass `repo` on every
|
||||
call: an omitted `repo` normally errors, but under an MCP policy with a
|
||||
configured default it resolves to that default silently. If you cannot tell
|
||||
which repository is meant, stop and ask. This matters most for `cypher`, whose
|
||||
statement carries no in-band hint of which database it ran against.
|
||||
|
||||
`list_repos` is paginated, so page with `offset: pagination.nextOffset` until
|
||||
`hasMore` is false before concluding a repository is absent.
|
||||
|
||||
A stale index describes the code from before your bug, so refresh before
|
||||
trusting a trace, and state the repository and index freshness with the
|
||||
diagnosis.
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
0. list_repos {} → Bind repo
|
||||
1. query({search_query: "<error or symptom>"}) → Find related execution flows
|
||||
2. context({name: "<suspect>"}) → See callers/callees/processes
|
||||
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
|
||||
|
|
@ -27,6 +46,7 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
|||
## Checklist
|
||||
|
||||
```
|
||||
- [ ] list_repos {} — bind repo; explicit repo when >1 indexed, ask if ambiguous
|
||||
- [ ] Understand the symptom (error message, unexpected behavior)
|
||||
- [ ] query for error text or related code
|
||||
- [ ] Identify the suspect function from returned processes
|
||||
|
|
@ -34,6 +54,7 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
|||
- [ ] Trace execution flow via process resource if applicable
|
||||
- [ ] cypher for custom call chain traces if needed
|
||||
- [ ] Read source files to confirm root cause
|
||||
- [ ] State the repository and index freshness with the diagnosis
|
||||
```
|
||||
|
||||
## Debugging Patterns
|
||||
|
|
@ -44,7 +65,7 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
|||
| Wrong return value | `context` on the function → trace callees for data flow |
|
||||
| Intermittent failure | `context` → look for external calls, async deps |
|
||||
| Performance issue | `context` → find symbols with many callers (hot paths) |
|
||||
| Recent regression | `detect_changes` to see what your changes affect |
|
||||
| Recent regression | `detect_changes` to see what your changes affect — pass `worktree` for a linked worktree |
|
||||
| "How does A reach B?" | `trace` between the two symbols — shortest call chain in one call |
|
||||
|
||||
## Tools
|
||||
|
|
@ -52,7 +73,7 @@ description: "Use when the user is debugging a bug, tracing an error, or asking
|
|||
**query** — find code related to error:
|
||||
|
||||
```
|
||||
query({search_query: "payment validation error"})
|
||||
query({search_query: "payment validation error", repo: "my-app"})
|
||||
→ Processes: CheckoutFlow, ErrorHandling
|
||||
→ Symbols: validatePayment, handlePaymentError, PaymentException
|
||||
```
|
||||
|
|
@ -60,13 +81,15 @@ query({search_query: "payment validation error"})
|
|||
**context** — full context for a suspect:
|
||||
|
||||
```
|
||||
context({name: "validatePayment"})
|
||||
context({name: "validatePayment", repo: "my-app"})
|
||||
→ Incoming calls: processCheckout, webhookHandler
|
||||
→ Outgoing calls: verifyCard, fetchRates (external API!)
|
||||
→ Processes: CheckoutFlow (step 3/7)
|
||||
```
|
||||
|
||||
**cypher** — custom call chain traces:
|
||||
**cypher** — custom call chain traces. Pass `repo` alongside the statement; the
|
||||
Cypher text itself names no repository, so the result is unattributable without
|
||||
it:
|
||||
|
||||
```cypher
|
||||
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
|
||||
|
|
@ -76,7 +99,7 @@ RETURN [n IN nodes(path) | n.name] AS chain
|
|||
**trace** — shortest call chain between two symbols ("how does A reach B?"), one call instead of chaining `context` hops:
|
||||
|
||||
```
|
||||
trace({ from: "processCheckout", to: "fetchRates" })
|
||||
trace({ from: "processCheckout", to: "fetchRates", repo: "my-app" })
|
||||
→ status: ok, hopCount: 3
|
||||
→ hops: processCheckout → validatePayment → verifyCard → fetchRates
|
||||
→ edges: CALLS (1.0), CALLS (0.95), CALLS (1.0)
|
||||
|
|
@ -87,15 +110,22 @@ When no path exists, `trace` reports the furthest reachable node — exactly whe
|
|||
## Example: "Payment endpoint returns 500 intermittently"
|
||||
|
||||
```
|
||||
1. query({search_query: "payment error handling"})
|
||||
0. list_repos {}
|
||||
→ total: 2 (my-app, billing-api) — bind my-app explicitly on every call
|
||||
|
||||
1. query({search_query: "payment error handling", repo: "my-app"})
|
||||
→ Processes: CheckoutFlow, ErrorHandling
|
||||
→ Symbols: validatePayment, handlePaymentError
|
||||
|
||||
2. context({name: "validatePayment"})
|
||||
2. context({name: "validatePayment", repo: "my-app"})
|
||||
→ Outgoing calls: verifyCard, fetchRates (external API!)
|
||||
|
||||
3. READ gitnexus://repo/my-app/process/CheckoutFlow
|
||||
→ Step 3: validatePayment → calls fetchRates (external)
|
||||
|
||||
4. Root cause: fetchRates calls external API without proper timeout
|
||||
Repository: my-app Index: current
|
||||
```
|
||||
|
||||
With a single indexed repository, step 0 returns `total: 1` and the `repo`
|
||||
argument drops out of every call above.
|
||||
|
|
|
|||
|
|
@ -13,10 +13,22 @@ description: "Use when the user asks how code works, wants to understand archite
|
|||
- "Where is the database logic?"
|
||||
- Understanding code you haven't seen before
|
||||
|
||||
## Bind the repository first
|
||||
|
||||
Step 1 discovers what is indexed; every call after it must say which of those
|
||||
it means. With one indexed repository, use the examples below as written. With
|
||||
more than one, pass `repo` on every call: an omitted `repo` normally errors,
|
||||
but under an MCP policy with a configured default it resolves to that default
|
||||
silently. If you cannot tell which repository is meant, stop and ask. Report
|
||||
the bound repository and index freshness alongside your explanation.
|
||||
|
||||
`list_repos` is paginated, so page with `offset: pagination.nextOffset` until
|
||||
`hasMore` is false before concluding a repository is absent.
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
1. READ gitnexus://repos → Discover indexed repos
|
||||
1. list_repos {} or READ gitnexus://repos → Discover indexed repos
|
||||
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
|
||||
3. query({search_query: "<what you want to understand>"}) → Find related execution flows
|
||||
4. context({name: "<symbol>"}) → Deep dive on specific symbol
|
||||
|
|
@ -28,12 +40,14 @@ description: "Use when the user asks how code works, wants to understand archite
|
|||
## Checklist
|
||||
|
||||
```
|
||||
- [ ] list_repos {} — bind repo; explicit repo when >1 indexed, ask if ambiguous
|
||||
- [ ] READ gitnexus://repo/{name}/context
|
||||
- [ ] query for the concept you want to understand
|
||||
- [ ] Review returned processes (execution flows)
|
||||
- [ ] context on key symbols for callers/callees
|
||||
- [ ] READ process resource for full execution traces
|
||||
- [ ] Read source files for implementation details
|
||||
- [ ] State the repository and index freshness with the explanation
|
||||
```
|
||||
|
||||
## Resources
|
||||
|
|
@ -50,7 +64,7 @@ description: "Use when the user asks how code works, wants to understand archite
|
|||
**query** — find execution flows related to a concept:
|
||||
|
||||
```
|
||||
query({search_query: "payment processing"})
|
||||
query({search_query: "payment processing", repo: "my-app"})
|
||||
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
|
||||
→ Symbols grouped by flow with file locations
|
||||
```
|
||||
|
|
@ -58,16 +72,20 @@ query({search_query: "payment processing"})
|
|||
**context** — 360-degree view of a symbol:
|
||||
|
||||
```
|
||||
context({name: "validateUser"})
|
||||
context({name: "validateUser", repo: "my-app"})
|
||||
→ Incoming calls: loginHandler, apiMiddleware
|
||||
→ Outgoing calls: checkToken, getUserById
|
||||
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
|
||||
```
|
||||
|
||||
`repo` is required once more than one repository is indexed, and may be omitted
|
||||
with a single one.
|
||||
|
||||
## Example: "How does payment processing work?"
|
||||
|
||||
```
|
||||
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
|
||||
1. list_repos {} → total: 1 (my-app) — bind it
|
||||
READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
|
||||
2. query({search_query: "payment processing"})
|
||||
→ CheckoutFlow: processPayment → validateCard → chargeStripe
|
||||
→ RefundFlow: initiateRefund → calculateRefund → processRefund
|
||||
|
|
@ -75,4 +93,8 @@ context({name: "validateUser"})
|
|||
→ Incoming: checkoutHandler, webhookHandler
|
||||
→ Outgoing: validateCard, chargeStripe, saveTransaction
|
||||
4. Read src/payments/processor.ts for implementation details
|
||||
5. Answer, noting: Repository my-app, index current
|
||||
```
|
||||
|
||||
Had step 1 returned two repositories, every call above would carry
|
||||
`repo: "my-app"`.
|
||||
|
|
|
|||
|
|
@ -14,13 +14,42 @@ description: "Use when the user wants to know what will break if they change som
|
|||
- Before making non-trivial code changes
|
||||
- Before committing — to understand what your changes affect
|
||||
|
||||
## Bind the repository first
|
||||
|
||||
Impact analysis is the gate that authorizes an edit, so it must answer for the
|
||||
repository you are about to edit.
|
||||
|
||||
Call `list_repos {}` before the first tool call. With one indexed repository,
|
||||
use the examples below as written. With more than one, pass `repo` on every
|
||||
call: an omitted `repo` normally errors, but under an MCP policy with a
|
||||
configured default it resolves to that default silently. If you cannot tell
|
||||
which repository is meant, stop and ask — every result below an ambiguous
|
||||
identity inherits the ambiguity. `list_repos` is paginated, so page with
|
||||
`offset: pagination.nextOffset` until `hasMore` is false before concluding a
|
||||
repository is absent.
|
||||
|
||||
`detect_changes` takes `worktree` when your changes are in a linked worktree
|
||||
the MCP server was not launched from. The server auto-detects a worktree only
|
||||
when it was launched from inside one; otherwise `git diff` runs in the wrong
|
||||
checkout and reports zero changed symbols — a false clean check that carries
|
||||
none of the degradation flags described below. In the CLI fallbacks, `--repo .`
|
||||
means the current checkout; pass the intended repository path instead when you
|
||||
are not standing in it.
|
||||
|
||||
State the bound identity with your risk report:
|
||||
|
||||
```
|
||||
Repository: <name> (<path>) Worktree: <path> Index: <commit>, <n> behind HEAD
|
||||
```
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
0. list_repos {} → Bind repo (and worktree)
|
||||
1. impact({target: "X", direction: "upstream"}) or `node .gitnexus/run.cjs impact "X" --direction upstream --repo .`
|
||||
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
|
||||
3. detect_changes({scope: "all"}) or `node .gitnexus/run.cjs detect-changes --scope all --repo .`
|
||||
4. Assess risk and report to user
|
||||
4. Assess risk and report to user, echoing repo/worktree/index identity
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `node .gitnexus/run.cjs analyze` in terminal.
|
||||
|
|
@ -29,12 +58,14 @@ description: "Use when the user wants to know what will break if they change som
|
|||
## Checklist
|
||||
|
||||
```
|
||||
- [ ] list_repos {} — bind repo; explicit repo when >1 indexed, ask if ambiguous
|
||||
- [ ] impact({target, direction: "upstream"}) or CLI fallback to find dependents
|
||||
- [ ] Review d=1 items first (these WILL BREAK)
|
||||
- [ ] Check high-confidence (>0.8) dependencies
|
||||
- [ ] READ processes to check affected execution flows
|
||||
- [ ] detect_changes({scope: "all"}) or CLI fallback for pre-commit check
|
||||
- [ ] Assess risk level and report to user
|
||||
- [ ] Confirm the checkout you edited is the checkout that was diffed
|
||||
- [ ] Assess risk level and report, stating repo/worktree/index identity
|
||||
```
|
||||
|
||||
## Understanding Output
|
||||
|
|
@ -62,6 +93,15 @@ dispatch, cross-language calls), so few-callers ⇒ LOW does **not** apply. The
|
|||
result carries a `riskNote` saying so. Confirm with a text search before
|
||||
treating the symbol as safe to change or delete.
|
||||
|
||||
`risk` is the edit gate: warn on HIGH/CRITICAL and stop on UNKNOWN until the
|
||||
uncertainty is resolved. Within single-repo mode, compare File and symbol
|
||||
targets with local `riskSharedAxes` (direct/total only). Within group mode,
|
||||
compare only group results: their `riskSharedAxes` overlays resolved
|
||||
cross-repo crossings on that local value. Never use either field to waive the
|
||||
edit gate. Check `riskScale.unusedAxes` before comparing kinds: MCP File walks
|
||||
omit process/module axes, while web Graph-RAG expands File targets to in-file
|
||||
symbols before enrichment.
|
||||
|
||||
## Tools
|
||||
|
||||
**impact** — the primary tool for symbol blast radius. If MCP is unavailable, use `node .gitnexus/run.cjs impact <symbol> --direction upstream --repo .` instead:
|
||||
|
|
@ -69,6 +109,7 @@ treating the symbol as safe to change or delete.
|
|||
```
|
||||
impact({
|
||||
target: "validateUser",
|
||||
repo: "my-app", // required once >1 repository is indexed
|
||||
direction: "upstream",
|
||||
minConfidence: 0.8,
|
||||
maxDepth: 3
|
||||
|
|
@ -92,15 +133,26 @@ detect_changes({scope: "all"})
|
|||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
Add `repo` once more than one repository is indexed, and `worktree: "<abs
|
||||
path>"` when your changes are in a linked worktree the server was not launched
|
||||
from.
|
||||
|
||||
`partial: true` (a graph query failed) or `truncated: true` (the changed-symbol
|
||||
listing was capped) means the result is short of the truth, and reads like
|
||||
`UNKNOWN` above: a zero there means unseen, not unaffected. Re-run it rather
|
||||
than tick the pre-commit check.
|
||||
|
||||
A wrong-worktree zero carries neither flag and is shape-identical to a genuine
|
||||
clean result, so confirm the checkout you edited is the one that was diffed
|
||||
before treating an empty change set as a passed check.
|
||||
|
||||
## Example: "What breaks if I change validateUser?"
|
||||
|
||||
```
|
||||
1. impact({target: "validateUser", direction: "upstream"}) or `node .gitnexus/run.cjs impact "validateUser" --direction upstream --repo .`
|
||||
0. list_repos {}
|
||||
→ total: 2 (my-app, billing-api) — both define validateUser, so bind explicitly
|
||||
|
||||
1. impact({target: "validateUser", repo: "my-app", direction: "upstream"}) or `node .gitnexus/run.cjs impact "validateUser" --direction upstream --repo .`
|
||||
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
|
||||
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
|
||||
|
||||
|
|
@ -108,4 +160,8 @@ than tick the pre-commit check.
|
|||
→ LoginFlow and TokenRefresh touch validateUser
|
||||
|
||||
3. Risk: 2 direct callers, 2 processes = MEDIUM
|
||||
Repository: my-app (/abs/path/my-app) Worktree: same Index: current
|
||||
```
|
||||
|
||||
With a single indexed repository, step 0 returns `total: 1` and the `repo`
|
||||
argument drops out of every call above.
|
||||
|
|
|
|||
|
|
@ -13,9 +13,32 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
|||
- "Move this to a new file"
|
||||
- Any task involving renaming, extracting, splitting, or restructuring code
|
||||
|
||||
## Bind the repository first
|
||||
|
||||
Refactoring writes to disk. `rename` with `dry_run: false` edits files in
|
||||
whichever repository was resolved, so binding identity here is a safety gate,
|
||||
not bookkeeping.
|
||||
|
||||
Call `list_repos {}` before the first tool call. With one indexed repository,
|
||||
use the examples below as written. With more than one, pass `repo` on every
|
||||
call: an omitted `repo` normally errors, but under an MCP policy with a
|
||||
configured default it resolves to that default silently. If you cannot tell
|
||||
which repository is meant, stop and ask. Never run `rename` with
|
||||
`dry_run: false` until the preview in the same bound repository has been
|
||||
reviewed — its returned `file_path` values show which checkout is about to be
|
||||
written, so read them as a confirmation of identity.
|
||||
|
||||
`list_repos` is paginated, so page with `offset: pagination.nextOffset` until
|
||||
`hasMore` is false before concluding a repository is absent.
|
||||
|
||||
`detect_changes` takes `worktree` when you are editing a linked worktree the
|
||||
MCP server was not launched from; otherwise `git diff` runs in the wrong
|
||||
checkout and reports nothing changed, which reads as a verified refactor.
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
0. list_repos {} → Bind repo (and worktree)
|
||||
1. impact({target: "X", direction: "upstream"}) → Map all dependents
|
||||
2. query({search_query: "X"}) → Find execution flows involving X
|
||||
3. context({name: "X"}) → See all incoming/outgoing refs
|
||||
|
|
@ -29,7 +52,9 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
|||
### Rename Symbol
|
||||
|
||||
```
|
||||
- [ ] list_repos {} — bind repo; explicit repo when >1 indexed, ask if ambiguous
|
||||
- [ ] rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
|
||||
- [ ] Confirm the previewed file paths are in the bound repository/worktree
|
||||
- [ ] Review graph edits (high confidence) and text_search edits (review carefully)
|
||||
- [ ] If satisfied: rename({..., dry_run: false}) — apply edits
|
||||
- [ ] detect_changes() — verify only expected files changed
|
||||
|
|
@ -39,6 +64,7 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
|||
### Extract Module
|
||||
|
||||
```
|
||||
- [ ] list_repos {} — bind repo; explicit repo when >1 indexed, ask if ambiguous
|
||||
- [ ] context({name: target}) — see all incoming/outgoing refs
|
||||
- [ ] impact({target, direction: "upstream"}) — find all external callers
|
||||
- [ ] Define new module interface
|
||||
|
|
@ -50,6 +76,7 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
|||
### Split Function/Service
|
||||
|
||||
```
|
||||
- [ ] list_repos {} — bind repo; explicit repo when >1 indexed, ask if ambiguous
|
||||
- [ ] context({name: target}) — understand all callees
|
||||
- [ ] Group callees by responsibility
|
||||
- [ ] impact({target, direction: "upstream"}) — map callers to update
|
||||
|
|
@ -64,7 +91,7 @@ description: "Use when the user wants to rename, extract, split, move, or restru
|
|||
**rename** — automated multi-file rename:
|
||||
|
||||
```
|
||||
rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
|
||||
rename({symbol_name: "validateUser", new_name: "authenticateUser", repo: "my-app", dry_run: true})
|
||||
→ 12 edits across 8 files
|
||||
→ 10 graph edits (high confidence), 2 text_search edits (review)
|
||||
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
|
||||
|
|
@ -73,7 +100,7 @@ rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true
|
|||
**impact** — map all dependents first:
|
||||
|
||||
```
|
||||
impact({target: "validateUser", direction: "upstream"})
|
||||
impact({target: "validateUser", repo: "my-app", direction: "upstream"})
|
||||
→ d=1: loginHandler, apiMiddleware, testUtils
|
||||
→ Affected Processes: LoginFlow, TokenRefresh
|
||||
```
|
||||
|
|
@ -92,6 +119,9 @@ listing was capped) means the result is short of the truth: a short or empty
|
|||
list is not proof that only the expected files changed. Re-run it rather than
|
||||
treat the refactor as verified.
|
||||
|
||||
A wrong-worktree zero carries neither flag and is indistinguishable from a
|
||||
clean verification, so confirm the diffed checkout is the one you edited.
|
||||
|
||||
**cypher** — custom reference queries:
|
||||
|
||||
```cypher
|
||||
|
|
@ -107,20 +137,28 @@ RETURN caller.name, caller.filePath ORDER BY caller.filePath
|
|||
| Cross-area refs | Use detect_changes after to verify scope |
|
||||
| String/dynamic refs | query to find them |
|
||||
| External/public API | Version and deprecate properly |
|
||||
| Same name in another indexed repo | Bind `repo`; verify previewed paths before applying |
|
||||
|
||||
## Example: Rename `validateUser` to `authenticateUser`
|
||||
|
||||
```
|
||||
1. rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
|
||||
0. list_repos {}
|
||||
→ total: 2 (my-app, billing-api) — both define validateUser, so bind explicitly
|
||||
|
||||
1. rename({symbol_name: "validateUser", new_name: "authenticateUser", repo: "my-app", dry_run: true})
|
||||
→ 12 edits: 10 graph (safe), 2 text_search (review)
|
||||
→ Files: validator.ts, login.ts, middleware.ts, config.json...
|
||||
|
||||
2. Review text_search edits (config.json: dynamic reference!)
|
||||
|
||||
3. rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
|
||||
3. rename({symbol_name: "validateUser", new_name: "authenticateUser", repo: "my-app", dry_run: false})
|
||||
→ Applied 12 edits across 8 files
|
||||
|
||||
4. detect_changes({scope: "all"})
|
||||
4. detect_changes({scope: "all", repo: "my-app"})
|
||||
→ Affected: LoginFlow, TokenRefresh
|
||||
→ Risk: MEDIUM — run tests for these flows
|
||||
Repository: my-app (/abs/path/my-app) Worktree: same Index: current
|
||||
```
|
||||
|
||||
With a single indexed repository, step 0 returns `total: 1` and the `repo`
|
||||
argument drops out of every call above.
|
||||
|
|
|
|||
12
.gitattributes
vendored
12
.gitattributes
vendored
|
|
@ -15,3 +15,15 @@
|
|||
*.so binary
|
||||
*.dll binary
|
||||
*.dylib binary
|
||||
|
||||
# TypeScript sources are always text for diff purposes. Git's binary
|
||||
# heuristic fires when EITHER blob in a pair carries a NUL, so a source
|
||||
# file that carried one on a base commit still renders as "Binary files
|
||||
# differ" — with no hunks and no inline comments — long after the byte
|
||||
# itself is gone from the working tree. A head-side guard cannot see
|
||||
# that, by construction. This does not mark the files binary or change
|
||||
# how they are stored; it only stops the heuristic from hiding a diff.
|
||||
*.ts diff
|
||||
*.tsx diff
|
||||
*.mts diff
|
||||
*.cts diff
|
||||
|
|
|
|||
17
.github/actions/setup-gitnexus-web/action.yml
vendored
17
.github/actions/setup-gitnexus-web/action.yml
vendored
|
|
@ -11,12 +11,19 @@ runs:
|
|||
cache: npm
|
||||
cache-dependency-path: gitnexus-web/package-lock.json
|
||||
|
||||
- name: Build gitnexus-shared
|
||||
run: npm install && npm run build
|
||||
shell: bash
|
||||
working-directory: gitnexus-shared
|
||||
|
||||
- name: Install web dependencies
|
||||
run: npm ci
|
||||
shell: bash
|
||||
working-directory: gitnexus-web
|
||||
env:
|
||||
# Browsers are installed explicitly by e2e. Typecheck only needs types.
|
||||
PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD: '1'
|
||||
|
||||
# Compile shared with the web package's TypeScript 5. Do not npm-ci
|
||||
# gitnexus-shared (TypeScript 7 optional-platform install, ~7 minutes).
|
||||
- name: Build gitnexus-shared
|
||||
# node + lib/tsc.js — same on Windows/macOS/Linux. Do not use .bin/tsc
|
||||
# (tsc.cmd on Windows; execFileSync cannot launch .cmd without a shell).
|
||||
run: node ../gitnexus-web/node_modules/typescript/lib/tsc.js
|
||||
shell: bash
|
||||
working-directory: gitnexus-shared
|
||||
|
|
|
|||
31
.github/actions/setup-gitnexus/action.yml
vendored
31
.github/actions/setup-gitnexus/action.yml
vendored
|
|
@ -6,6 +6,13 @@ inputs:
|
|||
description: Whether to run npm run build after install
|
||||
required: false
|
||||
default: 'false'
|
||||
lifecycle-scripts:
|
||||
description: >
|
||||
Run npm lifecycle scripts (prepare/postinstall) during gitnexus npm ci.
|
||||
Typecheck-only and pack-only jobs should set this to false: they do not
|
||||
need dist/ or native grammar builds.
|
||||
required: false
|
||||
default: 'true'
|
||||
|
||||
runs:
|
||||
using: composite
|
||||
|
|
@ -16,16 +23,30 @@ runs:
|
|||
cache: npm
|
||||
cache-dependency-path: gitnexus/package-lock.json
|
||||
|
||||
- name: Build gitnexus-shared
|
||||
run: npm install && npm run build
|
||||
shell: bash
|
||||
working-directory: gitnexus-shared
|
||||
|
||||
# Do not npm-ci gitnexus-shared. Its TypeScript 7 install is a 7-minute
|
||||
# stall (optional platform packages) and is not in the CLI npm cache.
|
||||
# prepare/build.js compiles shared with gitnexus's tsc; typecheck does
|
||||
# the same below after an ignore-scripts install.
|
||||
- name: Install dependencies
|
||||
if: ${{ inputs.lifecycle-scripts != 'false' }}
|
||||
run: npm ci
|
||||
shell: bash
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Install dependencies
|
||||
if: ${{ inputs.lifecycle-scripts == 'false' }}
|
||||
run: npm ci --ignore-scripts
|
||||
shell: bash
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Build gitnexus-shared
|
||||
if: ${{ inputs.lifecycle-scripts == 'false' }}
|
||||
# node + lib/tsc.js — same on Windows/macOS/Linux. Do not use .bin/tsc
|
||||
# (tsc.cmd on Windows; execFileSync cannot launch .cmd without a shell).
|
||||
run: node ../gitnexus/node_modules/typescript/lib/tsc.js
|
||||
shell: bash
|
||||
working-directory: gitnexus-shared
|
||||
|
||||
- name: Build
|
||||
if: ${{ inputs.build == 'true' }}
|
||||
run: npm run build
|
||||
|
|
|
|||
|
|
@ -69,6 +69,7 @@ GRAMMARS: dict[str, tuple[str, str, str]] = {
|
|||
# Vendored parsers — kept here so the upstream coords for drift
|
||||
# detection are co-located with every other grammar's coords.
|
||||
"tree-sitter-proto": ("coder3101/tree-sitter-proto", "main", "src/parser.c"),
|
||||
"tree-sitter-zig": ("tree-sitter-grammars/tree-sitter-zig", "master", "src/parser.c"),
|
||||
}
|
||||
|
||||
# npm-installed grammars deliberately held below npm latest (surfaced so reviewers
|
||||
|
|
|
|||
|
|
@ -9,8 +9,8 @@ which is deliberately dependency-free so it runs on any vanilla runner. Run with
|
|||
(pytest also discovers ``unittest.TestCase`` classes, so a future pytest CI job
|
||||
picks these up unchanged.)
|
||||
|
||||
These tests lock in the #858 fix: the 5 vendored grammars
|
||||
(c/swift/kotlin/dart/proto) are classified from the shared manifest
|
||||
These tests lock in the #858 fix: the 6 vendored grammars
|
||||
(c/swift/kotlin/dart/proto/zig) are classified from the shared manifest
|
||||
(.github/vendored-grammars.json), their ABI is read from gitnexus/vendor/<name>,
|
||||
and the report never renders a bare ``?`` placeholder. All network is mocked.
|
||||
"""
|
||||
|
|
@ -192,7 +192,7 @@ class AssertCurrent(TestCase):
|
|||
def test_assert_current_is_network_free_and_passes(self):
|
||||
report, code = self._run_assert_current() # raises if any urlopen fires
|
||||
self.assertEqual(code, 0)
|
||||
# All 5 vendored grammars are introspected from the repo (ABI 14), not skipped.
|
||||
# All 6 vendored grammars are introspected from the repo (ABI 14), not skipped.
|
||||
for name in readiness.VENDORED_NAMES:
|
||||
self.assertIn(f"{name}: vendored ABI", report)
|
||||
|
||||
|
|
@ -324,6 +324,7 @@ class ReportRendering(TestCase):
|
|||
# which is what removes the old "? (fetch failed)" for tree-sitter-proto.
|
||||
self.assertNotIn("tree-sitter-proto", _render_report.last_npm_calls)
|
||||
self.assertNotIn("tree-sitter-dart", _render_report.last_npm_calls)
|
||||
self.assertNotIn("@tree-sitter-grammars/tree-sitter-zig", _render_report.last_npm_calls)
|
||||
self.assertNotIn("Could not check", self.report)
|
||||
self.assertNotIn("fetch failed", self.report)
|
||||
|
||||
|
|
@ -343,11 +344,11 @@ class ReportRendering(TestCase):
|
|||
cells = [c.strip() for c in self._matrix_row("tree-sitter-swift").strip().strip("|").split("|")]
|
||||
self.assertEqual(cells[6], "n/a") # Upstream ABI column
|
||||
|
||||
def test_row_diff_regex_captures_all_fifteen_grammar_statuses(self):
|
||||
def test_row_diff_regex_captures_all_grammar_statuses(self):
|
||||
# The change-detection bot keys on this regex: group 1 = grammar name,
|
||||
# group 2 = the Status cell ONLY (not the whole tail). It must match every
|
||||
# row after the format change so status transitions keep being detected.
|
||||
self.assertEqual(len(self.rows), 15)
|
||||
self.assertEqual(len(self.rows), len(readiness.GRAMMARS))
|
||||
for name in readiness.VENDORED_NAMES:
|
||||
self.assertIn(name, self.rows)
|
||||
# group 2 is the Status cell — held c renders exactly "Vendored — held",
|
||||
|
|
|
|||
4
.github/vendored-grammars.json
vendored
4
.github/vendored-grammars.json
vendored
|
|
@ -22,6 +22,10 @@
|
|||
"proto": {
|
||||
"name": "tree-sitter-proto",
|
||||
"upstream": { "github": "coder3101/tree-sitter-proto" }
|
||||
},
|
||||
"zig": {
|
||||
"name": "tree-sitter-zig",
|
||||
"upstream": { "npm": "@tree-sitter-grammars/tree-sitter-zig" }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -20,6 +20,11 @@ name: Build tree-sitter prebuilds
|
|||
# - tree-sitter-swift (vendored source; built from gitnexus/vendor/ — its
|
||||
# prebuilds were originally upstream-shipped, now
|
||||
# GitNexus-cross-built like the rest for uniformity)
|
||||
# - tree-sitter-zig (vendored source; built from gitnexus/vendor/ — moved
|
||||
# off npm optionalDependency so `npm i -g gitnexus`
|
||||
# no longer warns on peerOptional tree-sitter@^0.22.1.
|
||||
# Upstream linux-arm64 prebuild is a mispackaged
|
||||
# x86-64 binary; this workflow rebuilds all six.)
|
||||
#
|
||||
# Output: gitnexus/vendor/<grammar>/prebuilds/<platform-arch>/<grammar>.node for
|
||||
# all 6 targets ({linux,darwin,win32}-{x64,arm64}). tree-sitter grammars are
|
||||
|
|
@ -33,8 +38,9 @@ name: Build tree-sitter prebuilds
|
|||
# OR an edit to the grammar's build-affecting source (parser.c / grammar.js /
|
||||
# binding.gyp / scanner / bindings). The `guard` job is the real gate (it
|
||||
# diffs BOTH the recorded version AND the source files vs the PR base); the
|
||||
# `paths:` filter below keeps ordinary code PRs at ZERO matrix time and
|
||||
# excludes the prebuilds the job commits back, so it never retriggers itself.
|
||||
# `paths:` filter below keeps ordinary code PRs at ZERO matrix time. PR
|
||||
# filters see the cumulative diff, so the guard separately skips updates
|
||||
# containing only the prebuilds the job commits back.
|
||||
# Net effect: an ordinary code PR triggers nothing; touching one grammar's source
|
||||
# costs exactly one matrix run for that grammar. Delivery of the rebuilt binaries:
|
||||
# - same-repo PR -> committed straight onto the PR's own branch (in the SAME PR);
|
||||
|
|
@ -55,7 +61,7 @@ on:
|
|||
workflow_dispatch:
|
||||
inputs:
|
||||
grammars:
|
||||
description: 'Comma-separated grammar shortnames to build (c,dart,proto,kotlin,swift), or "all".'
|
||||
description: 'Comma-separated grammar shortnames to build (c,dart,proto,kotlin,swift,zig), or "all".'
|
||||
required: false
|
||||
type: string
|
||||
default: 'all'
|
||||
|
|
@ -80,13 +86,14 @@ on:
|
|||
# Any build-affecting change under a vendored grammar triggers a rebuild —
|
||||
# not just a version bump — so editing the vendored source (parser.c,
|
||||
# grammar.js, binding.gyp, scanner, bindings) re-cuts the prebuilds too.
|
||||
# The prebuilds we commit back are EXCLUDED (negated last) so the bot's own
|
||||
# in-PR commit can never retrigger this workflow (no build->commit->build loop).
|
||||
# Excludes PRs containing only prebuilds. A source PR still matches after a
|
||||
# bot commit because PR filters use the cumulative diff; `guard` stops the
|
||||
# build->commit->build loop using the synchronize event's before/head diff.
|
||||
- 'gitnexus/vendor/tree-sitter-*/**'
|
||||
- '!gitnexus/vendor/tree-sitter-*/prebuilds/**'
|
||||
# Self-test: re-run the guard if a future grammar pin is reintroduced in
|
||||
# the main package.json (optionalDependencies fallback). No-op otherwise —
|
||||
# all five grammars are now fully vendored (kotlin included).
|
||||
# all six grammars are now fully vendored (kotlin and zig included).
|
||||
- 'gitnexus/package.json'
|
||||
# Self-test: re-run the guard (normally a no-op) when the recipe changes.
|
||||
- '.github/workflows/build-tree-sitter-prebuilds.yml'
|
||||
|
|
@ -123,6 +130,9 @@ jobs:
|
|||
id: decide
|
||||
env:
|
||||
EVENT: ${{ github.event_name }}
|
||||
ACTION: ${{ github.event.action }}
|
||||
BEFORE_SHA: ${{ github.event.before }}
|
||||
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
|
||||
# Untrusted dispatch inputs — read via env only, validated in JS.
|
||||
INPUT_GRAMMARS: ${{ inputs.grammars }}
|
||||
INPUT_REF: ${{ inputs.ref }}
|
||||
|
|
@ -131,7 +141,7 @@ jobs:
|
|||
run: |
|
||||
set -euo pipefail
|
||||
node --input-type=module - <<'NODE'
|
||||
import { execSync } from 'node:child_process';
|
||||
import { execFileSync, execSync } from 'node:child_process';
|
||||
import fs from 'node:fs';
|
||||
import { appendFileSync } from 'node:fs';
|
||||
|
||||
|
|
@ -157,6 +167,11 @@ jobs:
|
|||
// so it builds from gitnexus/vendor/ like dart/proto. Its prebuilds
|
||||
// were originally upstream-shipped; rebuilding them here unifies it.
|
||||
swift: { name: 'tree-sitter-swift', kind: 'vendored' },
|
||||
// zig is vendored WITH its source (parser.c/binding.gyp). Moved off
|
||||
// the npm optionalDependency so published installs no longer warn
|
||||
// on peerOptional tree-sitter@^0.22.1. Upstream linux-arm64
|
||||
// prebuild is a mispackaged x86-64 binary; rebuild here.
|
||||
zig: { name: 'tree-sitter-zig', kind: 'vendored' },
|
||||
};
|
||||
const PLATFORMS = [
|
||||
{ platform_arch: 'linux-x64', os: 'ubuntu-24.04' },
|
||||
|
|
@ -187,6 +202,30 @@ jobs:
|
|||
const event = process.env.EVENT;
|
||||
const force = process.env.FORCE === 'true';
|
||||
|
||||
// PR path filters and the source/version checks below see the entire
|
||||
// PR, so excluding prebuilds there does NOT prevent a rebuild loop.
|
||||
// Check the whole push (not HEAD^ or the author's identity): a push
|
||||
// containing source edits followed by a binary commit must still build.
|
||||
if (event === 'pull_request' && process.env.ACTION === 'synchronize') {
|
||||
const before = process.env.BEFORE_SHA;
|
||||
const head = process.env.HEAD_SHA;
|
||||
for (const sha of [before, head]) {
|
||||
if (!sha || !/^[0-9a-fA-F]{40}$/.test(sha)) {
|
||||
throw new Error('synchronize requires valid before/head SHAs; refusing an unbounded rebuild');
|
||||
}
|
||||
}
|
||||
// Fail closed if either commit is unavailable. Never fall back to
|
||||
// the cumulative PR diff, which would re-enable the loop.
|
||||
const changed = execFileSync('git', [
|
||||
'diff', '--name-only', '--no-renames', '-z', before, head, '--',
|
||||
], { encoding: 'utf8' }).split('\0').filter(Boolean);
|
||||
if (changed.every((p) => /^gitnexus\/vendor\/tree-sitter-[^/]+\/prebuilds\//.test(p))) {
|
||||
appendFileSync(process.env.GITHUB_OUTPUT, 'any=false\nmatrix={"include":[]}\n');
|
||||
console.log('::notice::Push changes only prebuild outputs (or no files) — skipping native matrix.');
|
||||
process.exit(0);
|
||||
}
|
||||
}
|
||||
|
||||
// Select which grammar shortnames are in play.
|
||||
let selected;
|
||||
if (event === 'workflow_dispatch') {
|
||||
|
|
@ -243,9 +282,9 @@ jobs:
|
|||
} else {
|
||||
// pull_request: build when the recorded version changed OR any
|
||||
// build-affecting source file under the vendored grammar changed vs
|
||||
// the PR base. The prebuilds/ subtree is excluded from the diff so
|
||||
// the bot's own in-PR commit (which adds ONLY prebuilds) never reads
|
||||
// as a source change — this is the other half of the no-loop guard.
|
||||
// the PR base. Exclude generated outputs from build inputs; the
|
||||
// synchronize check above prevents rebuilding the original source
|
||||
// change after every generated-prebuild commit.
|
||||
const base = recordedVersion(baseRoot, name);
|
||||
const versionChanged = !!head && head !== base;
|
||||
let sourceChanged = false;
|
||||
|
|
@ -467,11 +506,14 @@ jobs:
|
|||
proto: "syntax = \"proto3\";\nmessage M { int32 id = 1; }",
|
||||
kotlin: "fun main() { println(\"hi\") }",
|
||||
swift: "func greet() { print(\"hi\") }",
|
||||
zig: "pub fn main() void {}",
|
||||
};
|
||||
const src = snippets[process.env.GRAMMAR];
|
||||
if (!src) throw new Error("no validate snippet for grammar: " + process.env.GRAMMAR);
|
||||
const lang = require("node-gyp-build")(process.cwd());
|
||||
const Parser = require("tree-sitter");
|
||||
const p = new Parser(); p.setLanguage(lang);
|
||||
const tree = p.parse(snippets[process.env.GRAMMAR]);
|
||||
const tree = p.parse(src);
|
||||
if (!tree || !tree.rootNode || tree.rootNode.hasError) {
|
||||
throw new Error("parse failed/error: " + (tree && tree.rootNode && tree.rootNode.type));
|
||||
}
|
||||
|
|
@ -567,7 +609,7 @@ jobs:
|
|||
NODE
|
||||
|
||||
- name: Attest build provenance (SLSA)
|
||||
uses: actions/attest-build-provenance@0f67c3f4856b2e3261c31976d6725780e5e4c373 # v4.1.1
|
||||
uses: actions/attest-build-provenance@4d101475d8b20a2381f78447822ac1eab6504dd8 # v4.2.2
|
||||
with:
|
||||
subject-path: 'gitnexus/vendor/tree-sitter-*/prebuilds/**/*.node'
|
||||
|
||||
|
|
|
|||
17
.github/workflows/ci-quality.yml
vendored
17
.github/workflows/ci-quality.yml
vendored
|
|
@ -9,7 +9,9 @@ permissions:
|
|||
jobs:
|
||||
format:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 5
|
||||
# Same root npm ci as lint. A cold install already took 4m19s here and
|
||||
# canceled prettier at the 5-minute job cap; lint needed 7m41s the same run.
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
||||
with:
|
||||
|
|
@ -19,7 +21,7 @@ jobs:
|
|||
node-version: 22
|
||||
cache: npm
|
||||
cache-dependency-path: package-lock.json
|
||||
- run: npm ci
|
||||
- run: npm ci --ignore-scripts
|
||||
- run: npx prettier --check .
|
||||
|
||||
lint:
|
||||
|
|
@ -34,7 +36,7 @@ jobs:
|
|||
node-version: 22
|
||||
cache: npm
|
||||
cache-dependency-path: package-lock.json
|
||||
- run: npm ci
|
||||
- run: npm ci --ignore-scripts
|
||||
- run: npx eslint .
|
||||
|
||||
typecheck:
|
||||
|
|
@ -44,13 +46,20 @@ jobs:
|
|||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
||||
with:
|
||||
persist-credentials: false
|
||||
# tsc --noEmit reads source + gitnexus-shared/dist. Skip prepare/postinstall
|
||||
# so a cold shared install cannot eat the 10-minute budget on a second tsc.
|
||||
- uses: ./.github/actions/setup-gitnexus
|
||||
with:
|
||||
lifecycle-scripts: 'false'
|
||||
- run: npx tsc --noEmit
|
||||
working-directory: gitnexus
|
||||
|
||||
typecheck-web:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
# Cold gitnexus-web npm ci is several minutes (mermaid/langchain/playwright).
|
||||
# A 10-minute cancel prevents setup-node from saving the cache, so the next
|
||||
# run is cold again.
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
||||
with:
|
||||
|
|
|
|||
123
.github/workflows/ci-tests.yml
vendored
123
.github/workflows/ci-tests.yml
vendored
|
|
@ -24,8 +24,14 @@ jobs:
|
|||
shard: ${{ fromJSON(needs.shard-plan.outputs.cov_shards) }}
|
||||
# Fail loudly (don't silently skip) if the FTS extension is unavailable, so
|
||||
# FTS-dependent lbug integration suites are guaranteed to run in CI.
|
||||
# Same contract for Zig's vendored grammar: this runner is linux-x64,
|
||||
# which vendor/tree-sitter-zig ships a prebuild for, so an absent grammar
|
||||
# here is a packaging regression and not an unsupported platform. Without
|
||||
# it every Zig suite skips and the job is green having never executed the
|
||||
# native Zig parser once.
|
||||
env:
|
||||
GITNEXUS_REQUIRE_FTS: '1'
|
||||
GITNEXUS_REQUIRE_ZIG: '1'
|
||||
steps:
|
||||
# persist-credentials: false — runs tests + uploads a blob artifact; the
|
||||
# default-persisted token must not be capturable through it (zizmor
|
||||
|
|
@ -278,8 +284,16 @@ jobs:
|
|||
shell: bash
|
||||
run: python3 .github/scripts/check-tree-sitter-upgrade-readiness.py --assert-current
|
||||
|
||||
# GITNEXUS_REQUIRE_ZIG=1: every OS in this matrix has a committed
|
||||
# vendored tree-sitter-zig prebuild (linux-arm64 is rebuilt by the
|
||||
# prebuild workflow; ubuntu/windows/macos latest are x64/arm64 with
|
||||
# shipped binaries), so the smoke's "optional grammar may be absent"
|
||||
# exemption is revoked here and an ABI-broken Zig binding fails the
|
||||
# job instead of being accepted as a clean absence.
|
||||
- name: Run parser-loader ABI load-smoke (dynamic)
|
||||
run: npx vitest run test/unit/parser-loader-abi.test.ts
|
||||
env:
|
||||
GITNEXUS_REQUIRE_ZIG: '1'
|
||||
working-directory: gitnexus
|
||||
|
||||
# End-to-end smoke test for the #1728 packaging fix: pack the published
|
||||
|
|
@ -295,7 +309,9 @@ jobs:
|
|||
matrix:
|
||||
os: [windows-latest, ubuntu-latest]
|
||||
runs-on: ${{ matrix.os }}
|
||||
timeout-minutes: 15
|
||||
# Windows pack + web install regularly exceeds 15 minutes when setup also
|
||||
# runs prepare/postinstall/build before prepack compiles the same tree again.
|
||||
timeout-minutes: 20
|
||||
steps:
|
||||
# persist-credentials: false — this job runs npm pack + npm install -g
|
||||
# from a tarball and never pushes back; the token in .git/config would
|
||||
|
|
@ -304,9 +320,20 @@ jobs:
|
|||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
||||
with:
|
||||
persist-credentials: false
|
||||
# Skip prepare/postinstall/build here. `npm pack` runs prepack, which
|
||||
# compiles CLI + web into the tarball this job actually installs.
|
||||
- uses: ./.github/actions/setup-gitnexus
|
||||
with:
|
||||
build: 'true'
|
||||
lifecycle-scripts: 'false'
|
||||
|
||||
# `npm pack` runs prepack, which builds the web UI into gitnexus/web/
|
||||
# so the tarball matches what `npm publish` ships. Install those deps
|
||||
# here, in their own visible step, rather than letting build.js do it
|
||||
# from inside an execSync.
|
||||
- name: Install gitnexus-web dependencies
|
||||
shell: bash
|
||||
run: npm ci
|
||||
working-directory: gitnexus-web
|
||||
|
||||
- name: Pack gitnexus tarball
|
||||
shell: bash
|
||||
|
|
@ -348,6 +375,10 @@ jobs:
|
|||
fi
|
||||
echo "Installed package at: $INSTALLED"
|
||||
|
||||
# The npm package contract includes the built web UI. Validate the
|
||||
# installed artifact, not just the source workflow that produced it.
|
||||
node "$INSTALLED/scripts/assert-web-assets.mjs" "$INSTALLED/web"
|
||||
|
||||
# #836 invariant: no node_modules/ or build/ under any vendor/*.
|
||||
BAD=$(find "$INSTALLED/vendor" \( -name node_modules -o -name build \) -print 2>/dev/null || true)
|
||||
if [ -n "$BAD" ]; then
|
||||
|
|
@ -407,9 +438,6 @@ jobs:
|
|||
node-version: '22'
|
||||
cache: npm
|
||||
cache-dependency-path: gitnexus/package-lock.json
|
||||
- name: Build gitnexus-shared
|
||||
run: npm ci && npm run build
|
||||
working-directory: gitnexus-shared
|
||||
- name: Install and build gitnexus
|
||||
shell: bash
|
||||
run: |
|
||||
|
|
@ -481,6 +509,20 @@ jobs:
|
|||
node --import tsx bench/python-scope/import-target-fingerprint.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Java wildcard-static route constant guards (#3110)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: named-import control vs wildcard materialization;
|
||||
# fingerprints bindings and guards scaling + absolute wall time.
|
||||
run: node --import tsx bench/java-wildcard-route-constants/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Kotlin package-star route constant guards (#3110)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: explicit-import control vs package-star folding;
|
||||
# fingerprints route facts and guards scaling + widening overhead.
|
||||
run: node --import tsx bench/kotlin-star-route-constants/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Cross-language scope-capture fingerprint + scaling guards
|
||||
# Runs even after an earlier guard fails (#2895). Every step here was
|
||||
# fail-fast, so the FIRST failing --check aborted the job and every guard
|
||||
|
|
@ -509,6 +551,31 @@ jobs:
|
|||
run: node --import tsx bench/callable-value-flow/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Java Lombok accessor synthesis guards (#2885)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: no-Lombok vs Lombok-heavy corpora; fingerprint over
|
||||
# synthetic Method ids; scaling + widening overhead budgets.
|
||||
run: node --import tsx bench/java-lombok-synthesis/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Kotlin JVM accessor synthesis guards (#2885)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: no-property vs data-class corpora; fingerprint over
|
||||
# synthetic Method ids; scaling + widening overhead budgets.
|
||||
run: node --import tsx bench/kotlin-jvm-accessors/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Kotlin Spring config-consumer capture guards (#2412)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: explicit-import control vs wildcard-import feature path;
|
||||
# fingerprints @Value / @ConfigurationProperties facts and guards scaling
|
||||
# + widening overhead. The parity check is the regression gate: each file
|
||||
# declares a sibling nested type named `Value`, which must not suppress
|
||||
# the imported Spring annotation (file-wide shadowing dropped 2 of every
|
||||
# 3 facts on this corpus).
|
||||
run: node --import tsx bench/spring-config-bindings/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Re-export closure scaling guards (#2864)
|
||||
# Build-free: asserts buildReexportClosures stays linear in chain depth
|
||||
# and within an absolute ceiling on a wide package corpus. #2864 changed
|
||||
|
|
@ -616,8 +683,8 @@ jobs:
|
|||
# where the measured overshoot is at most 6.3%. The estimator was fixed
|
||||
# rather than the budget widened; distributions in _arms_note.
|
||||
# The Kotlin arm here is a second corpus, not a replacement for the
|
||||
# kotlin-import-target bench below, which carries tie-break probes (both
|
||||
# file-set iteration orders, the four-tier cascade) this one does not.
|
||||
# kotlin-import-target bench below, which carries declared-package
|
||||
# correctness probes this shared corpus does not.
|
||||
# It sits with the other resolver-index guards rather than at the end of
|
||||
# the job: parking a new gate last is not safety, it is the slot least
|
||||
# likely to execute (#2895 measured the last two guards running zero
|
||||
|
|
@ -629,15 +696,13 @@ jobs:
|
|||
run: node --expose-gc --import tsx bench/import-target/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Kotlin import-resolution identity + scaling guards
|
||||
- name: Kotlin declared-package import correctness + scaling guards
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts resolveKotlinImportTarget resolves an unchanged
|
||||
# file set (fingerprint, in both file-set iteration orders — every
|
||||
# tie-break in that resolver is expressed only through iteration order)
|
||||
# and that per-import cost stays independent of workspace size. The
|
||||
# pre-index implementation scores 3.737 on this corpus against 0.99 for
|
||||
# the index, so the gate separates them by a wide margin. Rationale and
|
||||
# history: see the header of bench/kotlin-import-target/measure.mjs.
|
||||
# Build-free: fingerprints declared-package evidence, external decoy
|
||||
# rejection, top-level/member/wildcard imports, overload sets and root
|
||||
# packages, then guards one package-index build per workspace against
|
||||
# file-count and path-depth scaling. Rationale and history: see the
|
||||
# header of bench/kotlin-import-target/measure.mjs.
|
||||
run: node --import tsx bench/kotlin-import-target/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
|
|
@ -680,6 +745,13 @@ jobs:
|
|||
run: node --import tsx bench/scope-emission/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: Zig cross-file static-gating guards (#3162)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: fingerprints cross-file dead-call classification and
|
||||
# guards the workspace enrichment pass across file-count scaling.
|
||||
run: node --import tsx bench/zig-cross-file-resolution/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
- name: CFG construction time / disk / memory guards (#2081 M1)
|
||||
if: ${{ !cancelled() }}
|
||||
# Build-free: asserts collectFunctionCfgs output is unchanged
|
||||
|
|
@ -713,19 +785,22 @@ jobs:
|
|||
|
||||
- name: Cross-language pipeline benchmarks (GITNEXUS_BENCH, serial)
|
||||
if: ${{ !cancelled() }}
|
||||
# cpp-adl-benchmark.test.ts is not a `*-pipeline-benchmark.test.ts` but
|
||||
# belongs here for the same reason: it is skipIf-gated on GITNEXUS_BENCH,
|
||||
# so it had never run in CI and the PR #1990 ADL emit-scaling guard it
|
||||
# holds was dead. ~45s of test time.
|
||||
# cpp-adl-benchmark.test.ts and csharp-razor-view-components-benchmark.test.ts
|
||||
# are not `*-pipeline-benchmark.test.ts` files but belong here for the
|
||||
# same reason: they are skipIf-gated on GITNEXUS_BENCH, so the scaling
|
||||
# guards they hold never run in the main coverage job.
|
||||
env:
|
||||
GITNEXUS_BENCH: '1'
|
||||
run: >-
|
||||
npx vitest run --no-file-parallelism
|
||||
test/integration/cobol-pipeline-benchmark.test.ts
|
||||
test/integration/csharp-pipeline-benchmark.test.ts
|
||||
test/integration/csharp-razor-view-components-benchmark.test.ts
|
||||
test/integration/cpp-adl-benchmark.test.ts
|
||||
test/integration/data-route-table-benchmark.test.ts
|
||||
test/integration/instance-ownership-pipeline-benchmark.test.ts
|
||||
test/integration/spring-bean-resource-benchmark.test.ts
|
||||
test/integration/spring-dynamic-lookup-benchmark.test.ts
|
||||
test/integration/rust-pipeline-benchmark.test.ts
|
||||
test/integration/php-pipeline-benchmark.test.ts
|
||||
test/integration/ruby-pipeline-benchmark.test.ts
|
||||
|
|
@ -781,7 +856,7 @@ jobs:
|
|||
run: |
|
||||
set -euo pipefail
|
||||
sudo apt-get update
|
||||
sudo apt-get install --yes --no-install-recommends bubblewrap socat
|
||||
sudo apt-get install --yes --no-install-recommends bubblewrap ripgrep socat
|
||||
apparmor_userns=/proc/sys/kernel/apparmor_restrict_unprivileged_userns
|
||||
if [[ -r "${apparmor_userns}" ]] && [[ "$(<"${apparmor_userns}")" == '1' ]]; then
|
||||
sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
|
||||
|
|
@ -804,11 +879,6 @@ jobs:
|
|||
"${canary_runtime}/node_modules/@anthropic-ai/claude-code/package.json"
|
||||
test "$("${canary_runtime}/node_modules/@anthropic-ai/claude-code-linux-x64/claude" --version)" = \
|
||||
'2.1.214 (Claude Code)'
|
||||
- name: Build pinned shared runtime
|
||||
run: |
|
||||
npm ci
|
||||
npm run build
|
||||
working-directory: gitnexus-shared
|
||||
- name: Install and build pinned GitNexus runtime
|
||||
run: |
|
||||
npm ci
|
||||
|
|
@ -844,5 +914,6 @@ jobs:
|
|||
- name: Prove Windows process-tree ownership
|
||||
run: >-
|
||||
uv run --locked --extra dev python -m pytest
|
||||
tests/test_process_control.py -q
|
||||
tests/test_process_control.py tests/test_model_gateway.py
|
||||
-k "not locked_litellm" -q
|
||||
working-directory: eval
|
||||
|
|
|
|||
10
.github/workflows/codeql.yml
vendored
10
.github/workflows/codeql.yml
vendored
|
|
@ -48,7 +48,7 @@ jobs:
|
|||
persist-credentials: false
|
||||
|
||||
- name: Initialize CodeQL
|
||||
uses: github/codeql-action/init@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
uses: github/codeql-action/init@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
|
||||
with:
|
||||
languages: ${{ matrix.language }}
|
||||
queries: security-and-quality
|
||||
|
|
@ -71,8 +71,14 @@ jobs:
|
|||
# deliberately contain use-before-init / unused-variable shapes).
|
||||
- '**/test/fixtures/**'
|
||||
- '**/test/**/fixtures/**'
|
||||
# GET /api/grep intentionally builds RegExp from the query string
|
||||
# (literal=1 escapes). ReDoS is handled by worker terminate() —
|
||||
# see SECURITY.md. Inline codeql[] comments do not clear the
|
||||
# GitHub PR CodeQL gate, so this file is excluded to avoid
|
||||
# re-filing js/regex-injection on every push of the same line.
|
||||
- 'gitnexus/src/server/grep-params.ts'
|
||||
|
||||
- name: Perform CodeQL Analysis
|
||||
uses: github/codeql-action/analyze@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
uses: github/codeql-action/analyze@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
|
||||
with:
|
||||
category: '/language:${{ matrix.language }}'
|
||||
|
|
|
|||
6
.github/workflows/docker.yml
vendored
6
.github/workflows/docker.yml
vendored
|
|
@ -141,7 +141,7 @@ jobs:
|
|||
uses: docker/setup-qemu-action@96fe6ef7f33517b61c61be40b68a1882f3264fb8 # v4.2.0
|
||||
|
||||
- name: Set up Docker Buildx
|
||||
uses: docker/setup-buildx-action@bb05f3f5519dd87d3ba754cc423b652a5edd6d2c # v4.2.0
|
||||
uses: docker/setup-buildx-action@37fe631027851001ddb9b187196cc803df7f5f0e # v4.3.0
|
||||
|
||||
- name: Install Cosign
|
||||
uses: sigstore/cosign-installer@6f9f17788090df1f26f669e9d70d6ae9567deba6 # v4.1.2
|
||||
|
|
@ -256,7 +256,7 @@ jobs:
|
|||
# pulling from either GHCR or Docker Hub see the same provenance.
|
||||
- name: Generate build provenance attestation (GHCR)
|
||||
if: ${{ github.event_name != 'pull_request' && !inputs.dry_run }}
|
||||
uses: actions/attest-build-provenance@0f67c3f4856b2e3261c31976d6725780e5e4c373 # v4.1.1
|
||||
uses: actions/attest-build-provenance@4d101475d8b20a2381f78447822ac1eab6504dd8 # v4.2.2
|
||||
with:
|
||||
subject-name: ghcr.io/${{ github.repository_owner }}/${{ matrix.image.slug }}
|
||||
subject-digest: ${{ steps.build.outputs.digest }}
|
||||
|
|
@ -264,7 +264,7 @@ jobs:
|
|||
|
||||
- name: Generate build provenance attestation (Docker Hub)
|
||||
if: ${{ github.event_name != 'pull_request' && !inputs.dry_run }}
|
||||
uses: actions/attest-build-provenance@0f67c3f4856b2e3261c31976d6725780e5e4c373 # v4.1.1
|
||||
uses: actions/attest-build-provenance@4d101475d8b20a2381f78447822ac1eab6504dd8 # v4.2.2
|
||||
with:
|
||||
subject-name: docker.io/akonlabs/${{ matrix.image.slug }}
|
||||
subject-digest: ${{ steps.build.outputs.digest }}
|
||||
|
|
|
|||
360
.github/workflows/gitnexus-skill-evolution.yml
vendored
360
.github/workflows/gitnexus-skill-evolution.yml
vendored
|
|
@ -4,17 +4,22 @@
|
|||
# overlay. The gate is evidence FOR a PR, never a bypass of one — nothing
|
||||
# merges without review.
|
||||
#
|
||||
# Activation checklist (the scheduled lane is OFF by default).
|
||||
# [ ] Configure the repository secret GITNEXUS_BENCH_AUTH_TOKEN (an Anthropic
|
||||
# API key — benchmark sessions bill real usage; the Claude Code OAuth
|
||||
# subscription token does not work here).
|
||||
# [ ] Configure the RELEASE_APP_ID and RELEASE_APP_PRIVATE_KEY secrets (the
|
||||
# Activation and operations checklist.
|
||||
# [x] Configure at least one model secret on the `gitnexus-evolution`
|
||||
# Environment: GITNEXUS_BENCH_ANTHROPIC_API_KEY (Anthropic API key — not
|
||||
# the Claude Code OAuth token; legacy GITNEXUS_BENCH_AUTH_TOKEN is still
|
||||
# accepted) and/or GITNEXUS_BENCH_OPENAI_API_KEY. Sessions bill real usage.
|
||||
# OpenAI keys are not native to Claude Code; the loop starts a loopback
|
||||
# LiteLLM proxy and keeps the OpenAI key off the sandboxed agent. With only
|
||||
# the OpenAI secret, or with provider=openai, dispatch-time Claude model
|
||||
# defaults are gpt-5.6-sol with xhigh reasoning effort.
|
||||
# [x] Configure the RELEASE_APP_ID and RELEASE_APP_PRIVATE_KEY secrets (the
|
||||
# App that opens the promotion PR). The Mint-App-Token step hard-fails
|
||||
# without them once a promotion is detected. Verify the App installation
|
||||
# is scoped to this repo with only Contents: RW + Pull requests: RW.
|
||||
# [x] Create the protected Environment `gitnexus-evolution` with a
|
||||
# deployment-branch rule restricting it to `main`, and ideally scope the
|
||||
# three secrets above to that Environment. workflow_dispatch runs this
|
||||
# four secrets above to that Environment. workflow_dispatch runs this
|
||||
# workflow (and eval/workflow_bench/evolve.py) from the *dispatched ref*,
|
||||
# so this server-side rule — not a code-side guard the branch could edit
|
||||
# away — is what stops a non-main branch from running with the secrets.
|
||||
|
|
@ -40,13 +45,31 @@
|
|||
# most weekly. Revisit if run frequency increases or the threat model
|
||||
# changes; stopping already bounds the exposure window to the job's own
|
||||
# runtime on 1 day out of 7.
|
||||
# [ ] Run workflow_dispatch once and confirm: containment preflight passes,
|
||||
# [ ] Install and verify the runner survival policy below before enabling
|
||||
# scheduled runs. A run
|
||||
# spans ~15h and apt-daily-upgrade.timer fires daily (~06:34), so every
|
||||
# scheduled run crosses it. On 2026-08-02 unattended-upgrades upgraded
|
||||
# openssl at 07:54:02 and needrestart restarted the Actions runner five
|
||||
# seconds later: the job went to Canceled, and a cancelled job skips even
|
||||
# `if: always()`, so the evidence artifact died with it. Keep installing
|
||||
# updates, but never let them restart services here:
|
||||
# /etc/needrestart/conf.d/90-gitnexus-evolution.conf
|
||||
# $nrconf{restart} = 'l';
|
||||
# A drop-in, so a needrestart package upgrade cannot clobber it. Nothing
|
||||
# is left unpatched in practice — the box is stopped between runs, so the
|
||||
# new binaries take effect at the next boot.
|
||||
# [x] Run workflow_dispatch once and confirm: containment preflight passes,
|
||||
# the benchmark completes inside the job timeout, the results artifact
|
||||
# uploads, and a promotion (if any) opens a well-formed PR.
|
||||
# [ ] Set the repository variable GITNEXUS_EVOLUTION_ENABLED=true.
|
||||
# Roll back by setting that variable to false. Note: workflow_dispatch always
|
||||
# runs the full benchmark loop regardless of GITNEXUS_EVOLUTION_ENABLED and
|
||||
# bills real API usage on GITNEXUS_BENCH_AUTH_TOKEN.
|
||||
# uploads, and a promotion (if any) opens a well-formed PR. Run
|
||||
# 29907431284 (2026-07-22) went green end to end in 14h45m and reached a
|
||||
# gate decision (`insufficient_evidence`, no promotion).
|
||||
# [ ] After resizing the runner, prove a manual workers=3 run has zero excluded
|
||||
# runs and does not stretch the 48-minute serial mean toward the session
|
||||
# ceiling; then set GITNEXUS_EVOLUTION_WORKERS=3 and
|
||||
# GITNEXUS_EVOLUTION_ENABLED=true for scheduled runs. Scheduled runs
|
||||
# require both values, so leaving workers unset/1 is an immediate rollback;
|
||||
# workflow_dispatch remains available for the proof and bills real API
|
||||
# usage on GITNEXUS_BENCH_ANTHROPIC_API_KEY or GITNEXUS_BENCH_OPENAI_API_KEY.
|
||||
name: GitNexus skill evolution
|
||||
|
||||
on:
|
||||
|
|
@ -68,21 +91,51 @@ on:
|
|||
required: false
|
||||
default: '3'
|
||||
type: string
|
||||
workers:
|
||||
description: 'Benchmark cells of one task to run at once — raise only to match the runner’s vCPUs'
|
||||
required: false
|
||||
default: '1'
|
||||
type: string
|
||||
model:
|
||||
description: 'Model for the benchmark arms (match the model your skill users run)'
|
||||
required: false
|
||||
default: 'claude-sonnet-5'
|
||||
default: 'gpt-5.6-sol'
|
||||
type: string
|
||||
proposer_model:
|
||||
description: 'Model for the proposer/diagnosis session — a stronger model is fine (one session per generation)'
|
||||
required: false
|
||||
default: 'claude-opus-4-8'
|
||||
default: 'gpt-5.6-sol'
|
||||
type: string
|
||||
effort:
|
||||
description: 'Reasoning effort for every proposer and benchmark session'
|
||||
required: false
|
||||
default: xhigh
|
||||
type: choice
|
||||
options:
|
||||
- low
|
||||
- medium
|
||||
- high
|
||||
- xhigh
|
||||
- max
|
||||
provider:
|
||||
description: 'Model backend. auto uses Anthropic when that secret exists; openai forces the loopback OpenAI gateway even if an Anthropic key is also configured.'
|
||||
required: false
|
||||
default: openai
|
||||
type: choice
|
||||
options:
|
||||
- auto
|
||||
- openai
|
||||
- anthropic
|
||||
include_expensive:
|
||||
description: 'Include tasks marked expensive: true'
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
seed_from_previous:
|
||||
description: "Seed the proposer with the previous run's evidence and rejected proposal. Turn off to start from a blank slate — required when the earlier evidence is not trustworthy (e.g. produced before a harness-integrity fix), since a tainted proposal would otherwise propagate into every later generation."
|
||||
required: false
|
||||
default: true
|
||||
type: boolean
|
||||
|
||||
concurrency:
|
||||
group: ${{ github.workflow }}
|
||||
|
|
@ -97,31 +150,94 @@ jobs:
|
|||
github.repository == 'abhigyanpatwari/GitNexus' &&
|
||||
(
|
||||
github.event_name == 'workflow_dispatch' ||
|
||||
vars.GITNEXUS_EVOLUTION_ENABLED == 'true'
|
||||
(
|
||||
vars.GITNEXUS_EVOLUTION_ENABLED == 'true' &&
|
||||
vars.GITNEXUS_EVOLUTION_WORKERS == '3'
|
||||
)
|
||||
)
|
||||
runs-on: [self-hosted, linux, x64, gitnexus-evolution]
|
||||
# Gate promotion runs on a protected Environment. An admin must attach a
|
||||
# deployment-branch rule (main only) and ideally scope the three secrets to
|
||||
# it — server-side enforcement a dispatched non-main ref cannot bypass by
|
||||
# deployment-branch rule (main only) and ideally scope the model and App
|
||||
# secrets to it — server-side enforcement a dispatched non-main ref cannot bypass by
|
||||
# editing its own workflow copy. See the activation checklist above.
|
||||
environment: gitnexus-evolution
|
||||
timeout-minutes: 1440 # self-hosted ceiling is 5 days (7200min); 24h is a generous margin over a single-generation serial run
|
||||
# Three budgets have to nest, longest first, or the evidence is lost:
|
||||
# EventBridge instance uptime (24h from ~02:45)
|
||||
# > this job timeout (21h)
|
||||
# > the benchmark step timeout (19h, set on the step below)
|
||||
# A job-level timeout CANCELS the job, so the upload step never runs and a
|
||||
# multi-hour generation's evidence dies with it; a step-level timeout only
|
||||
# fails that step, and `if: always()` still uploads what the sweep wrote.
|
||||
# The instance must outlive the job for the same reason — when the box
|
||||
# stops the runner just disappears mid-step. Scheduled runs can start well
|
||||
# after the cron (the 2026-08-01 run was queued 65min late), so the job
|
||||
# budget has to absorb that delay and still land inside the uptime window.
|
||||
timeout-minutes: 1260
|
||||
permissions:
|
||||
contents: read # The promotion PR uses a short-lived App token minted below.
|
||||
actions: read # Read the previous run's evidence artifact to seed the proposer.
|
||||
env:
|
||||
GENERATIONS: ${{ inputs.generations || '1' }}
|
||||
RUNS: ${{ inputs.runs || '3' }}
|
||||
MODEL: ${{ inputs.model || 'claude-sonnet-5' }}
|
||||
PROPOSER_MODEL: ${{ inputs.proposer_model || 'claude-opus-4-8' }}
|
||||
# A manual input wins; scheduled runs use the repository rollout knob.
|
||||
# Both fall back to serial — see workflow_bench.runner --workers for why.
|
||||
WORKERS: ${{ inputs.workers || vars.GITNEXUS_EVOLUTION_WORKERS || '1' }}
|
||||
MODEL: ${{ inputs.model || 'gpt-5.6-sol' }}
|
||||
PROPOSER_MODEL: ${{ inputs.proposer_model || 'gpt-5.6-sol' }}
|
||||
EFFORT: ${{ inputs.effort || 'xhigh' }}
|
||||
PROVIDER: ${{ inputs.provider || 'openai' }}
|
||||
INCLUDE_EXPENSIVE: ${{ inputs.include_expensive && '1' || '' }}
|
||||
steps:
|
||||
- name: Require the benchmark auth secret
|
||||
env:
|
||||
HAS_TOKEN: ${{ secrets.GITNEXUS_BENCH_AUTH_TOKEN != '' }}
|
||||
HAS_ANTHROPIC: ${{ secrets.GITNEXUS_BENCH_ANTHROPIC_API_KEY != '' || secrets.GITNEXUS_BENCH_AUTH_TOKEN != '' }}
|
||||
HAS_OPENAI: ${{ secrets.GITNEXUS_BENCH_OPENAI_API_KEY != '' }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
if [[ "${HAS_TOKEN}" != 'true' ]]; then
|
||||
echo '::error::GITNEXUS_BENCH_AUTH_TOKEN is not configured. The evolution loop runs real benchmark sessions and needs an Anthropic API key (not the Claude Code OAuth token).'
|
||||
if [[ "${HAS_ANTHROPIC}" != 'true' && "${HAS_OPENAI}" != 'true' ]]; then
|
||||
echo '::error::Configure GITNEXUS_BENCH_ANTHROPIC_API_KEY (Anthropic API key, not the Claude Code OAuth token) and/or GITNEXUS_BENCH_OPENAI_API_KEY. The evolution loop runs real benchmark sessions.'
|
||||
exit 1
|
||||
fi
|
||||
case "${PROVIDER}" in
|
||||
openai)
|
||||
if [[ "${HAS_OPENAI}" != 'true' ]]; then
|
||||
echo '::error::provider=openai requires GITNEXUS_BENCH_OPENAI_API_KEY on the gitnexus-evolution environment.'
|
||||
exit 1
|
||||
fi
|
||||
;;
|
||||
anthropic)
|
||||
if [[ "${HAS_ANTHROPIC}" != 'true' ]]; then
|
||||
echo '::error::provider=anthropic requires GITNEXUS_BENCH_ANTHROPIC_API_KEY on the gitnexus-evolution environment.'
|
||||
exit 1
|
||||
fi
|
||||
;;
|
||||
auto)
|
||||
;;
|
||||
*)
|
||||
echo "::error::Unknown provider '${PROVIDER}' (expected auto, openai, or anthropic)."
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
|
||||
- name: Verify runner survival policy
|
||||
run: |
|
||||
set -euo pipefail
|
||||
needrestart_policy=/etc/needrestart/conf.d/90-gitnexus-evolution.conf
|
||||
needrestart_line="\$nrconf{restart} = 'l';"
|
||||
if [[ ! -r "${needrestart_policy}" ]] || ! grep -Fqx "${needrestart_line}" "${needrestart_policy}"; then
|
||||
echo "::error::${needrestart_policy} must contain: ${needrestart_line}"
|
||||
exit 1
|
||||
fi
|
||||
# The runner sets job processes to 500; the host oom-guard rewrites
|
||||
# them to -900. Read once and the check loses that race.
|
||||
oom_score_adjustment="$(</proc/self/oom_score_adj)"
|
||||
deadline=$((SECONDS + 5))
|
||||
while (( oom_score_adjustment > -900 && SECONDS < deadline )); do
|
||||
sleep 0.05
|
||||
oom_score_adjustment="$(</proc/self/oom_score_adj)"
|
||||
done
|
||||
if (( oom_score_adjustment > -900 )); then
|
||||
echo "::error::Runner.Worker descendants require OOMScoreAdjust=-900 or stronger; effective value is ${oom_score_adjustment}."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
|
|
@ -145,11 +261,26 @@ jobs:
|
|||
enable-cache: true
|
||||
cache-dependency-glob: eval/uv.lock
|
||||
|
||||
- name: Fetch pinned Compound Engineering review comparator
|
||||
env:
|
||||
CE_COMMIT: 3ad9b51bceecf0158e590c882034d0398dbb9c5c
|
||||
run: |
|
||||
set -euo pipefail
|
||||
destination="${RUNNER_TEMP}/compound-engineering-plugin"
|
||||
rm -rf "${destination}"
|
||||
git clone --filter=blob:none --no-checkout \
|
||||
https://github.com/EveryInc/compound-engineering-plugin.git "${destination}"
|
||||
git -C "${destination}" checkout --detach "${CE_COMMIT}"
|
||||
test "$(git -C "${destination}" rev-parse HEAD)" = "${CE_COMMIT}"
|
||||
|
||||
- name: Install sandbox runtime and pinned Claude CLI
|
||||
run: |
|
||||
set -euo pipefail
|
||||
sudo apt-get update
|
||||
sudo apt-get install --yes --no-install-recommends bubblewrap socat
|
||||
# This box is stopped six days a week, so persistent apt timers can
|
||||
# begin their catch-up run shortly after boot. Wait for dpkg instead
|
||||
# of racing the same package lock and failing the weekly lane.
|
||||
sudo apt-get -o DPkg::Lock::Timeout=600 update
|
||||
sudo apt-get -o DPkg::Lock::Timeout=600 install --yes --no-install-recommends bubblewrap ripgrep socat
|
||||
apparmor_userns=/proc/sys/kernel/apparmor_restrict_unprivileged_userns
|
||||
if [[ -r "${apparmor_userns}" ]] && [[ "$(<"${apparmor_userns}")" == '1' ]]; then
|
||||
sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
|
||||
|
|
@ -173,6 +304,16 @@ jobs:
|
|||
test "$("${canary_runtime}/node_modules/@anthropic-ai/claude-code-linux-x64/claude" --version)" = \
|
||||
'2.1.214 (Claude Code)'
|
||||
|
||||
- name: Verify contained review execution before paid sessions
|
||||
working-directory: eval
|
||||
env:
|
||||
GITNEXUS_REQUIRE_BWRAP_CANARY: '1'
|
||||
GITNEXUS_REQUIRE_CLAUDE_CANARY: '1'
|
||||
CLAUDE_CANARY_BIN: ${{ runner.temp }}/claude-canary/node_modules/@anthropic-ai/claude-code-linux-x64/claude
|
||||
run: |
|
||||
set -euo pipefail
|
||||
uv run --locked --extra dev python -m pytest tests/test_proposer_sandbox.py -q
|
||||
|
||||
- name: Install monorepo root dependencies
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
|
@ -201,50 +342,169 @@ jobs:
|
|||
- name: Point the benchmark task repo at the checkout
|
||||
run: |
|
||||
set -euo pipefail
|
||||
# tasks.scenarios.yaml addresses the target repo as ~/GitNexus (the
|
||||
# tasks.review.scenarios.yaml addresses the target repo as ~/GitNexus (the
|
||||
# developer-local convention). On the runner the repo is the checkout
|
||||
# at ${GITHUB_WORKSPACE}; link it so runner_tasks.py can resolve the
|
||||
# task `repo` path. The benchmark only clones the repo (copy-on-write)
|
||||
# and mounts dependencies read-only, so the checkout is never mutated.
|
||||
if [[ -e "${HOME}/GitNexus" && ! -L "${HOME}/GitNexus" ]]; then
|
||||
echo '::error::~/GitNexus exists and is not a symlink; refusing to place the checkout inside it.'
|
||||
exit 1
|
||||
fi
|
||||
ln -sfn "${GITHUB_WORKSPACE}" "${HOME}/GitNexus"
|
||||
# The review corpus pins historical object ids. Fetch main so those
|
||||
# objects are present even when actions/checkout selected another ref.
|
||||
git -C "${GITHUB_WORKSPACE}" fetch --no-tags --quiet \
|
||||
"${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git" \
|
||||
'+refs/heads/main:refs/remotes/origin/main'
|
||||
baseline_sha="$(git -C "${GITHUB_WORKSPACE}" rev-parse --verify 'refs/remotes/origin/main^{commit}')"
|
||||
echo "Fetched review corpus history at ${baseline_sha}"
|
||||
|
||||
- name: Seed the proposer with the previous run's evidence
|
||||
id: seed
|
||||
# Scheduled runs always seed; a dispatch can opt out to start clean.
|
||||
if: github.event_name != 'workflow_dispatch' || inputs.seed_from_previous
|
||||
# Best-effort seeding must not consume the benchmark's budget. This
|
||||
# step walks up to 10 prior runs and every iteration blocks on network
|
||||
# it does not control (`gh run download` of a multi-hundred-megabyte
|
||||
# artifact). Unbounded, a wedged download sits here until the 21h job
|
||||
# timeout CANCELS the job — and a cancelled job skips even
|
||||
# `if: always()`, so the sweep never starts and nothing is uploaded.
|
||||
# Bounding the step instead fails it in minutes, which is a loud,
|
||||
# cheap, re-runnable failure rather than a silent 21h loss. 15 minutes
|
||||
# is an order of magnitude above the observed walk (well under a
|
||||
# minute) and a rounding error against the 19h sweep it protects.
|
||||
timeout-minutes: 15
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
# Without this the weekly lane is memoryless: `--seed-results` is the
|
||||
# only way a run sees what already lost (evolve stages the prior
|
||||
# proposal when present and summarizes promotion.json when present),
|
||||
# and with the default --generations 1 there is no earlier generation
|
||||
# in-process to supply it. Every Saturday would otherwise propose
|
||||
# from a blank slate and could re-propose the same rejected candidate
|
||||
# forever. Best-effort by design: a first run, an expired artifact,
|
||||
# or a download failure must not cost a whole generation.
|
||||
if ! command -v gh >/dev/null; then
|
||||
echo '::warning::gh is not installed on this runner — proposing without prior evidence. The promotion-PR step needs gh too.'
|
||||
exit 0
|
||||
fi
|
||||
if ! previous_runs="$(gh run list \
|
||||
--repo "${GITHUB_REPOSITORY}" \
|
||||
--workflow gitnexus-skill-evolution.yml \
|
||||
--branch main \
|
||||
--status completed \
|
||||
--limit 10 \
|
||||
--json databaseId \
|
||||
--jq "map(.databaseId) | map(select(. != ${GITHUB_RUN_ID})) | .[]")"; then
|
||||
echo '::warning::Prior workflow runs could not be listed; proposing without prior evidence.'
|
||||
exit 0
|
||||
fi
|
||||
if [[ -z "${previous_runs}" ]]; then
|
||||
echo 'No prior completed run to seed from; the proposer starts from the learnings queue only.'
|
||||
exit 0
|
||||
fi
|
||||
seed_root="${RUNNER_TEMP}/wfseed"
|
||||
rm -rf "${seed_root}"
|
||||
install -d -m 0700 "${seed_root}"
|
||||
seed=''
|
||||
# Failed sweeps deliberately upload partial evidence, so "completed"
|
||||
# is the right population. Walk newest-first until one still-retained
|
||||
# artifact actually contains benchmark rows; an empty latest run must
|
||||
# not hide an older useful one.
|
||||
for previous in ${previous_runs}; do
|
||||
if [[ ! "${previous}" =~ ^[0-9]+$ ]]; then
|
||||
echo "::warning::Ignoring malformed prior run id: ${previous}"
|
||||
continue
|
||||
fi
|
||||
run_root="${seed_root}/${previous}"
|
||||
install -d -m 0700 "${run_root}"
|
||||
if ! gh run download "${previous}" --repo "${GITHUB_REPOSITORY}" --dir "${run_root}"; then
|
||||
echo "::warning::Evidence from run ${previous} could not be downloaded (expired or absent); trying an older run."
|
||||
continue
|
||||
fi
|
||||
unsafe="$(find "${run_root}" ! -type d ! -type f -print -quit)"
|
||||
if [[ -n "${unsafe}" ]]; then
|
||||
echo "::warning::Run ${previous} contains a non-regular artifact entry; trying an older run."
|
||||
continue
|
||||
fi
|
||||
# upload-artifact normalizes directories/files to 0755/0644, while
|
||||
# the evidence reader deliberately requires transcript paths to be
|
||||
# owner-only. Restore that trust-boundary invariant after download.
|
||||
if ! chmod -R go-rwx "${run_root}"; then
|
||||
echo "::warning::Evidence permissions from run ${previous} could not be restricted; trying an older run."
|
||||
continue
|
||||
fi
|
||||
# The artifact holds gen-N/bench/{results.jsonl,promotion.json,...};
|
||||
# the highest generation is the one that actually reached the gate.
|
||||
latest="$(find "${run_root}" -type f -path '*/gen-*/bench/results.jsonl' | sort -V | tail -1)"
|
||||
if [[ -z "${latest}" || -L "${latest}" || ! -f "${latest}" ]]; then
|
||||
echo "::warning::Run ${previous} uploaded no usable benchmark results; trying an older run."
|
||||
continue
|
||||
fi
|
||||
# Existence is insufficient: an interrupted run may leave an empty,
|
||||
# malformed, or session/infra-only JSONL. Reuse the same bounded
|
||||
# selection and transcript/digest preflight the proposer will use,
|
||||
# so an unusable newer run cannot hide an older useful one.
|
||||
if uv run --project eval --locked --extra dev python -c \
|
||||
'from pathlib import Path; import json, sys, tempfile; from workflow_bench.evolve import load_jsonl, select_evidence, stage_proposer_evidence_bundle, summarize_gate; result = Path(sys.argv[1]); root = result.parent; rows = select_evidence(load_jsonl(result)); rows or sys.exit(10); promotion = root / "promotion.json"; gate = summarize_gate(json.loads(promotion.read_text())) if promotion.is_file() else []; prior = root.parent / "proposal.md"; prior = prior if prior.is_file() and not prior.is_symlink() else None; dest = Path(tempfile.mkdtemp(prefix="wfseed-preflight-")) / "bundle"; stage_proposer_evidence_bundle(dest, results_dir=root, evidence=rows, learnings=[], gate_summary=gate, prior_proposal=prior)' \
|
||||
"${latest}"; then
|
||||
:
|
||||
else
|
||||
usability_status=$?
|
||||
echo "::warning::Run ${previous} failed evidence preflight (exit ${usability_status}); trying an older run."
|
||||
continue
|
||||
fi
|
||||
seed="$(dirname "${latest}")"
|
||||
echo "Seeding the proposer from run ${previous}: ${seed}"
|
||||
break
|
||||
done
|
||||
if [[ -z "${seed}" ]]; then
|
||||
echo '::warning::No usable prior benchmark artifact found; proposing without prior evidence.'
|
||||
exit 0
|
||||
fi
|
||||
echo "seed=${seed}" >> "${GITHUB_OUTPUT}"
|
||||
|
||||
- name: Run the propose → benchmark → gate loop
|
||||
id: loop
|
||||
# Kill the sweep with time left in the job to upload what it produced.
|
||||
# See the budget nesting on the job above.
|
||||
timeout-minutes: 1140
|
||||
env:
|
||||
GITNEXUS_BENCH_AUTH_TOKEN: ${{ secrets.GITNEXUS_BENCH_AUTH_TOKEN }}
|
||||
GITNEXUS_BENCH_ANTHROPIC_API_KEY: ${{ secrets.GITNEXUS_BENCH_ANTHROPIC_API_KEY || secrets.GITNEXUS_BENCH_AUTH_TOKEN }}
|
||||
GITNEXUS_BENCH_OPENAI_API_KEY: ${{ secrets.GITNEXUS_BENCH_OPENAI_API_KEY }}
|
||||
# The step's stdout is a pipe, so CPython block-buffers it and a
|
||||
# multi-hour generation would report nothing until it exits (run
|
||||
# 29907431284 emitted every line at the same timestamp, 14h45m in).
|
||||
PYTHONUNBUFFERED: '1'
|
||||
SEED_RESULTS: ${{ steps.seed.outputs.seed }}
|
||||
EVOLUTION_PROFILE: review
|
||||
CE_PLUGIN_DIR: ${{ runner.temp }}/compound-engineering-plugin
|
||||
CE_PLUGIN_VERSION: 3.24.0
|
||||
run: |
|
||||
set -euo pipefail
|
||||
out_root="${RUNNER_TEMP}/wfevolve"
|
||||
echo "out_root=${out_root}" >> "${GITHUB_OUTPUT}"
|
||||
extra=()
|
||||
if [[ -n "${INCLUDE_EXPENSIVE}" ]]; then
|
||||
extra+=(--include-expensive)
|
||||
fi
|
||||
uv run --locked --extra dev python -m workflow_bench.evolve \
|
||||
--tasks workflow_bench/tasks.scenarios.yaml \
|
||||
--model "${MODEL}" \
|
||||
--proposer-model "${PROPOSER_MODEL}" \
|
||||
--generations "${GENERATIONS}" \
|
||||
--runs "${RUNS}" \
|
||||
--claude-bin "${RUNNER_TEMP}/claude-canary/node_modules/@anthropic-ai/claude-code-linux-x64/claude" \
|
||||
--out-root "${out_root}" \
|
||||
--apply \
|
||||
"${extra[@]}"
|
||||
./workflow_bench/run-evolution.sh --apply
|
||||
working-directory: eval
|
||||
|
||||
- name: Upload benchmark evidence
|
||||
if: always() && steps.loop.outputs.out_root != ''
|
||||
# Unconditional: the sweep writes results.jsonl and transcripts as it
|
||||
# goes, so a killed or failed generation still has evidence worth
|
||||
# keeping — and that is exactly the run whose evidence is needed.
|
||||
if: always()
|
||||
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||
with:
|
||||
name: gitnexus-evolution-${{ github.run_id }}-${{ github.run_attempt }}
|
||||
path: ${{ steps.loop.outputs.out_root }}
|
||||
# Addressed directly rather than carried from the sweep step: that is
|
||||
# the step whose death is the reason this upload matters, and a value
|
||||
# threaded from it would not be there when it counts.
|
||||
path: ${{ runner.temp }}/wfevolve
|
||||
retention-days: 14
|
||||
if-no-files-found: warn
|
||||
|
||||
- name: Detect and bound the applied promotion
|
||||
id: promotion
|
||||
env:
|
||||
OUT_ROOT: ${{ steps.loop.outputs.out_root }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
changed="$(git status --porcelain)"
|
||||
|
|
@ -259,7 +519,7 @@ jobs:
|
|||
while IFS= read -r line; do
|
||||
path="${line:3}"
|
||||
case "${path}" in
|
||||
.claude/skills/*|gitnexus/skills/*|gitnexus-claude-plugin/skills/*) ;;
|
||||
.claude/skills/gitnexus-review/*|gitnexus/skills/gitnexus-review/*|gitnexus-claude-plugin/skills/gitnexus-review/*|gitnexus-cursor-integration/skills/gitnexus-review/*) ;;
|
||||
*)
|
||||
echo "::error::Promotion touched a path outside the skill trees: ${path}"
|
||||
exit 1
|
||||
|
|
@ -273,7 +533,7 @@ jobs:
|
|||
# generation's decisions could surface in the PR body. The heredoc
|
||||
# uses a per-run random delimiter so a summary value that ever
|
||||
# contains the marker cannot close the block early and inject keys.
|
||||
promotion_file="$(find "${OUT_ROOT}" -name promotion.json | sort -V | tail -1)"
|
||||
promotion_file="$(find "${RUNNER_TEMP}/wfevolve" -name promotion.json | sort -V | tail -1)"
|
||||
delim="PROMOTION_EOF_$(openssl rand -hex 16)"
|
||||
{
|
||||
echo "summary<<${delim}"
|
||||
|
|
@ -316,7 +576,7 @@ jobs:
|
|||
git config user.name 'gitnexus-evolution[bot]'
|
||||
git config user.email 'gitnexus-evolution[bot]@users.noreply.github.com'
|
||||
git checkout -b "${branch}"
|
||||
git add .claude/skills gitnexus/skills gitnexus-claude-plugin/skills
|
||||
git add .claude/skills gitnexus/skills gitnexus-claude-plugin/skills gitnexus-cursor-integration/skills/gitnexus-review
|
||||
git commit -m 'feat(skills): promoted evolution overlay (gate-passed)'
|
||||
|
||||
# The App token reaches git through GIT_ASKPASS reading step env at
|
||||
|
|
|
|||
11
.github/workflows/publish.yml
vendored
11
.github/workflows/publish.yml
vendored
|
|
@ -396,14 +396,17 @@ jobs:
|
|||
# cache-poisoning audit). ~30s slower per release; runs rarely.
|
||||
package-manager-cache: false
|
||||
|
||||
- name: Build gitnexus-shared
|
||||
run: npm ci && npm run build
|
||||
working-directory: gitnexus-shared
|
||||
|
||||
- name: Install gitnexus dependencies
|
||||
run: npm ci
|
||||
working-directory: gitnexus
|
||||
|
||||
# The published tarball ships the web UI (`files: [... "web"]`), built
|
||||
# by prepack during `npm publish`. Install its deps in their own step
|
||||
# so a slow install is visible here instead of dying inside build.js.
|
||||
- name: Install gitnexus-web dependencies
|
||||
run: npm ci
|
||||
working-directory: gitnexus-web
|
||||
|
||||
# ── Stable-only: verify the tag and package.json agree ───────────────
|
||||
- name: Verify version consistency (stable)
|
||||
if: needs.route.outputs.mode == 'stable'
|
||||
|
|
|
|||
2
.github/workflows/scorecard.yml
vendored
2
.github/workflows/scorecard.yml
vendored
|
|
@ -53,6 +53,6 @@ jobs:
|
|||
retention-days: 5
|
||||
|
||||
- name: Upload to Security tab
|
||||
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
|
||||
with:
|
||||
sarif_file: results.sarif
|
||||
|
|
|
|||
3
.github/workflows/skill-sync.yml
vendored
3
.github/workflows/skill-sync.yml
vendored
|
|
@ -55,9 +55,6 @@ jobs:
|
|||
node-version: '22'
|
||||
cache: npm
|
||||
cache-dependency-path: gitnexus/package-lock.json
|
||||
- name: Build gitnexus-shared
|
||||
run: npm ci && npm run build
|
||||
working-directory: gitnexus-shared
|
||||
- name: Install gitnexus
|
||||
run: npm ci
|
||||
working-directory: gitnexus
|
||||
|
|
|
|||
4
.github/workflows/trivy.yml
vendored
4
.github/workflows/trivy.yml
vendored
|
|
@ -50,7 +50,7 @@ jobs:
|
|||
persist-credentials: false
|
||||
|
||||
- name: Setup Buildx
|
||||
uses: docker/setup-buildx-action@bb05f3f5519dd87d3ba754cc423b652a5edd6d2c # v4.2.0
|
||||
uses: docker/setup-buildx-action@37fe631027851001ddb9b187196cc803df7f5f0e # v4.3.0
|
||||
|
||||
- name: Build image (load locally for scan)
|
||||
uses: docker/build-push-action@53b7df96c91f9c12dcc8a07bcb9ccacbed38856a # v7.3.0
|
||||
|
|
@ -76,7 +76,7 @@ jobs:
|
|||
exit-code: '0'
|
||||
|
||||
- name: Upload to Security tab
|
||||
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
|
||||
with:
|
||||
sarif_file: trivy-${{ matrix.image.name }}.sarif
|
||||
category: trivy-${{ matrix.image.name }}
|
||||
|
|
|
|||
2
.github/workflows/workflow-lint.yml
vendored
2
.github/workflows/workflow-lint.yml
vendored
|
|
@ -76,7 +76,7 @@ jobs:
|
|||
continue-on-error: true
|
||||
|
||||
- name: Upload SARIF
|
||||
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
|
||||
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
|
||||
with:
|
||||
sarif_file: zizmor.sarif
|
||||
category: zizmor
|
||||
|
|
|
|||
2
.gitignore
vendored
2
.gitignore
vendored
|
|
@ -31,6 +31,8 @@ npm-debug.log*
|
|||
|
||||
# Testing
|
||||
coverage/
|
||||
.tmp-test/
|
||||
gitnexus/.tmp-test/
|
||||
|
||||
# Misc
|
||||
*.local
|
||||
|
|
|
|||
|
|
@ -1,2 +1,4 @@
|
|||
# Deleted README placeholder from PR #2458; no credential was present.
|
||||
c9fdab17f25ebaf332fba6e6ba55ee328f20fe66:README.md:curl-auth-header:348
|
||||
# Synthetic Kotlin Actuator fixture value from PR #3107; no credential was present.
|
||||
3951079300a18b14e79f5b5f5dd778ae19ced6e3:gitnexus/test/integration/spring-actuator-kotlin-runtime-pipeline.test.ts:generic-api-key:8
|
||||
|
|
|
|||
|
|
@ -119,10 +119,9 @@ This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 rela
|
|||
|
||||
- **MUST run impact analysis before editing.** Use `impact({target: "symbolName", direction: "upstream"})` (MCP) or `node .gitnexus/run.cjs impact "symbolName" --direction upstream --repo .` (CLI fallback); report callers, processes, and risk. Never substitute grep for graph analysis. For unified PDG impact, add `mode: "pdg"` with optional `line: <N>` — it returns statement-level `affectedStatements` over CDG + REACHING_DEF and inter-procedural symbols in `interproceduralByDepth`/`byDepth`; no-layer/degraded PDG results are UNKNOWN-risk notes (`--pdg` layer). CLI equivalent: `node .gitnexus/run.cjs impact "symbolName" --direction upstream --mode pdg --line <N> --repo .`.
|
||||
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). `partial: true` or `truncated: true` is not a clean check — a zero means unseen, not unaffected; re-run it. For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- MUST warn on HIGH/CRITICAL `risk` pre-edit; never use `riskSharedAxes` to waive a HIGH/CRITICAL `risk` warning. Compare File/symbol: MCP File omits axes; Graph-RAG expands File.
|
||||
- **MUST treat `risk: UNKNOWN` as unresolved, not as low.** An empty caller set is not evidence the symbol is unused — it can also mean the callers are not resolvable by the index (plain-object property access, dynamic dispatch, cross-language calls). `impact` pairs `UNKNOWN` with a `riskNote` saying so. Confirm with a text search before treating the symbol as safe to change or delete; do not proceed on the strength of a zero.
|
||||
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
|
||||
- **MUST use `query({search_query: "concept"})` for concepts/flows, `context({name: "symbolName"})` for a named symbol, or `impact` for blast radius, on read-only callers, dependencies, imports, or execution flow.** Graph first; text search only for empty/`UNKNOWN`/literals.
|
||||
- For security review, `explain({target: "fileOrSymbol"})` lists taint findings (source→sink flows; needs `analyze --pdg`).
|
||||
- For control/data dependence, `pdg_query({mode: "controls", target: "fileOrSymbol"})` answers "under what condition does X run?" (CDG, incl. guard clauses) and `pdg_query({mode: "flows", target, variable})` traces "where does variable Y flow?" (REACHING_DEF). `--pdg` layer.
|
||||
|
||||
|
|
@ -193,6 +192,6 @@ npx gitnexus serve # HTTP API on port 4747 (from any ind
|
|||
|
||||
### Gotchas
|
||||
|
||||
- `npm install` in `gitnexus/` triggers `prepare` (builds via `tsc`) and `postinstall` (materializes the vendored grammars into `node_modules/`, then prefers a committed prebuild per platform-arch and only source-builds when none matches). A C/C++ toolchain (`python3`, `make`, `g++`) is needed only for that source-build fallback.
|
||||
- The vendored grammars `tree-sitter-{c,dart,proto,swift,kotlin}` are handled uniformly: c is required; dart/proto/swift/kotlin are optional and skippable via `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1`. Install warnings appear only when no prebuild matches the platform-arch and no toolchain is present, and are non-fatal — only that language's parsing is unavailable.
|
||||
- `npm install` in `gitnexus/` triggers `prepare` (builds via `tsc`) and `postinstall` (`build-tree-sitter-grammars.cjs` activates committed prebuilds in place under `vendor/`, and only source-builds when none matches). A C/C++ toolchain (`python3`, `make`, `g++`) is needed only for that source-build fallback.
|
||||
- The vendored grammars `tree-sitter-{c,dart,proto,swift,kotlin,zig}` are handled uniformly: c is required; dart/proto/swift/kotlin/zig are optional and skippable via `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1`. Install warnings appear only when no prebuild matches the platform-arch and no toolchain is present, and are non-fatal — only that language's parsing is unavailable.
|
||||
- ESLint configured via `eslint.config.mjs` (TS, React Hooks, unused-imports). No `npm run lint` script; use `npx eslint .`. Prettier runs via lint-staged. CI checks both in `ci-quality.yml`.
|
||||
|
|
|
|||
|
|
@ -98,7 +98,7 @@ scan → structure → [springConfig, markdown, cobol] → parse → [routes, to
|
|||
| `markdown` | `markdown.ts` | `structure` | Section nodes, cross-link edges from .md/.mdx |
|
||||
| `cobol` | `cobol.ts` | `structure` | COBOL program/paragraph/section nodes (regex, no tree-sitter) |
|
||||
| `parse` | `parse.ts` + `parse-impl.ts` | `structure`, `markdown`, `cobol` | Symbol nodes, IMPORTS/CALLS/EXTENDS edges, extracted routes/tools/ORM queries |
|
||||
| `routes` | `routes.ts` | `parse` | Route nodes + HANDLES_ROUTE edges (Next.js, Expo, PHP, decorators, and JS/TS dispatch guards — see below) |
|
||||
| `routes` | `routes.ts` | `parse` | Route nodes + HANDLES_ROUTE edges (Next.js, Expo, PHP, decorators, and JS/TS static route sources — see below) |
|
||||
| `tools` | `tools.ts` | `parse` | Tool nodes + HANDLES_TOOL edges |
|
||||
| `orm` | `orm.ts` | `parse` | QUERIES edges (Prisma, Supabase) |
|
||||
| `crossFile` | `cross-file.ts` + `cross-file-impl.ts` | `parse`, `routes`, `tools`, `orm` | Cross-file type propagation in topological import order |
|
||||
|
|
@ -108,7 +108,7 @@ scan → structure → [springConfig, markdown, cobol] → parse → [routes, to
|
|||
| `pruneLocalSymbols` | `prune-local-symbols.ts` | `scopeResolution` | Drops inert block-local `Const`/`Variable`/`Static` nodes (only a `File→DEFINES` edge) post-resolution |
|
||||
| `mro` | `mro.ts` | `crossFile`, `scopeResolution`, `pruneLocalSymbols`, `structure` | METHOD_OVERRIDES + METHOD_IMPLEMENTS edges |
|
||||
| `springAopInheritance` | `spring-aop.ts` | `springAop`, `mro` | Propagates declarative behavior through class/interface inheritance decisions |
|
||||
| `di` | `di.ts` | `mro` | INJECTS edges from consumer Classes or factory Methods to provider Classes/declaration CodeElements (framework-neutral DI resolution; per-language matchers registered in `di-extractors/`) |
|
||||
| `di` | `di.ts` | `mro` | INJECTS edges from consumer Classes, factory Methods, or AST-captured programmatic lookup callables to provider Classes/declaration CodeElements (framework-neutral DI resolution; per-language matchers registered in `di-extractors/`) |
|
||||
| `communities` | `communities.ts` | `mro`, `pruneLocalSymbols`, `structure` | Community nodes + MEMBER_OF edges (Leiden algorithm) |
|
||||
| `processes` | `processes.ts` | `communities`, `routes`, `tools`, `pruneLocalSymbols`, `structure` | Process nodes + STEP_IN_PROCESS edges |
|
||||
|
||||
|
|
@ -174,7 +174,7 @@ converging on the routes phase's `(method, url)` registry:
|
|||
| Filesystem convention | path → URL, no parsing | Next.js `app/`, Expo, PHP |
|
||||
| Single-file framework route | `isRouteFile` + worker extraction | Laravel `routes/*.php` |
|
||||
| Cross-file framework route | `discoverRootRouteFiles` + `extractRoutes` | Django `urlpatterns` |
|
||||
| AST-level route in a normal file | `extractDecoratorRoutes` | Spring, FastAPI, NestJS, **JS/TS dispatch guards** |
|
||||
| AST-level route in a normal file | `extractDecoratorRoutes` | Spring, FastAPI, NestJS (`@Controller` + `@Get`/`@Post`/…; URLs are controller-relative — `setGlobalPrefix` and URI versioning live in the bootstrap file and are not applied), **JS/TS dispatch guards and static data route tables** |
|
||||
|
||||
The last row is the one whose name undersells it. A route is DECLARED by a
|
||||
decorator, but it can also be **inferred** from a raw `node:http` server's own
|
||||
|
|
@ -185,6 +185,15 @@ handler resolution are shared with decorator routes, and
|
|||
`ExtractedDecoratorRoute.source` carries the provenance difference through to
|
||||
the `HANDLES_ROUTE` edge.
|
||||
|
||||
JS/TS data route tables share that transport when a route-named array contains
|
||||
direct object literals with static `path`, `method`, and `handler` fields and a
|
||||
same-scope `for...of` dispatcher positively compares the path and method before
|
||||
directly invoking the handler. Dynamic values, computed keys, spreads,
|
||||
inline/called handlers, unknown verbs, and ambiguous handler bindings are
|
||||
suppressed. Bare import aliases and single-level member handlers are attributed
|
||||
only through declared import and owner provenance; an unproven receiver never
|
||||
falls back to a global name guess.
|
||||
|
||||
That extractor is deliberately **precision-weighted**: `route_map` presents its
|
||||
output as fact, so a `startsWith` namespace test, a bare `pathname === '/'`
|
||||
without a verb, and any regex it cannot translate exactly are all dropped rather
|
||||
|
|
@ -394,6 +403,7 @@ Each language implements `LanguageProvider` (`language-provider.ts`). Key fields
|
|||
| `typeConfig` | Type annotation extraction rules |
|
||||
| `mroStrategy` | `first-wins` / `c3` / `none` |
|
||||
| `descriptionExtractor` | Optional hook returning a symbol's doc-comment text as its `description`; feeds the embedding metadata header so doc-only terms are semantically searchable (issue #2270). Most languages register `createLeadingDocDescriptionExtractor` (shared, language-neutral; per-language comment/wrapper config passed at the call site) |
|
||||
| `definitionPropertiesExtractor` | Optional language-owned hook for structured, clone-safe definition metadata. Shared ingestion persists these properties opaquely; the owning provider supplies the extraction semantics. |
|
||||
|
||||
16 providers in `languages/index.ts` via `satisfies Record<SupportedLanguages, LanguageProvider>` — missing a language is a compile error.
|
||||
|
||||
|
|
|
|||
|
|
@ -70,10 +70,9 @@ This project is indexed by GitNexus as **GitNexus** (248612 symbols, 565510 rela
|
|||
|
||||
- **MUST run impact analysis before editing.** Use `impact({target: "symbolName", direction: "upstream"})` (MCP) or `node .gitnexus/run.cjs impact "symbolName" --direction upstream --repo .` (CLI fallback); report callers, processes, and risk. Never substitute grep for graph analysis. For unified PDG impact, add `mode: "pdg"` with optional `line: <N>` — it returns statement-level `affectedStatements` over CDG + REACHING_DEF and inter-procedural symbols in `interproceduralByDepth`/`byDepth`; no-layer/degraded PDG results are UNKNOWN-risk notes (`--pdg` layer). CLI equivalent: `node .gitnexus/run.cjs impact "symbolName" --direction upstream --mode pdg --line <N> --repo .`.
|
||||
- **MUST analyze graph changes before committing.** Use `detect_changes({scope: "all"})` (MCP) or `node .gitnexus/run.cjs detect-changes --scope all --repo .` (CLI fallback). `partial: true` or `truncated: true` is not a clean check — a zero means unseen, not unaffected; re-run it. For regression review: `detect_changes({scope: "compare", base_ref: "main"})` or `node .gitnexus/run.cjs detect-changes --scope compare --base-ref "main" --repo .`.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- MUST warn on HIGH/CRITICAL `risk` pre-edit; never use `riskSharedAxes` to waive a HIGH/CRITICAL `risk` warning. Compare File/symbol: MCP File omits axes; Graph-RAG expands File.
|
||||
- **MUST treat `risk: UNKNOWN` as unresolved, not as low.** An empty caller set is not evidence the symbol is unused — it can also mean the callers are not resolvable by the index (plain-object property access, dynamic dispatch, cross-language calls). `impact` pairs `UNKNOWN` with a `riskNote` saying so. Confirm with a text search before treating the symbol as safe to change or delete; do not proceed on the strength of a zero.
|
||||
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
|
||||
- **MUST use `query({search_query: "concept"})` for concepts/flows, `context({name: "symbolName"})` for a named symbol, or `impact` for blast radius, on read-only callers, dependencies, imports, or execution flow.** Graph first; text search only for empty/`UNKNOWN`/literals.
|
||||
- For security review, `explain({target: "fileOrSymbol"})` lists taint findings (source→sink flows; needs `analyze --pdg`).
|
||||
- For control/data dependence, `pdg_query({mode: "controls", target: "fileOrSymbol"})` answers "under what condition does X run?" (CDG, incl. guard clauses) and `pdg_query({mode: "flows", target, variable})` traces "where does variable Y flow?" (REACHING_DEF). `--pdg` layer.
|
||||
|
||||
|
|
|
|||
|
|
@ -51,8 +51,9 @@ RUN npm run postinstall --prefix gitnexus
|
|||
# node:22-bookworm-slim
|
||||
FROM node:22-bookworm-slim@sha256:9f6d5975c7dca860947d3915877f85607946403fc55349f39b4bc3688448bb6e AS runtime
|
||||
|
||||
# curl for the healthcheck; git for cloning; ca-certificates for TLS verification.
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends curl git ca-certificates && rm -rf /var/lib/apt/lists/* \
|
||||
# curl for the healthcheck; git for cloning; procps for watch process identity;
|
||||
# ca-certificates for TLS verification.
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends curl git procps ca-certificates && rm -rf /var/lib/apt/lists/* \
|
||||
&& rm -rf /usr/local/lib/node_modules/npm \
|
||||
&& rm -rf /usr/local/lib/node_modules/corepack \
|
||||
&& rm -f /usr/local/bin/npm /usr/local/bin/npx /usr/local/bin/corepack
|
||||
|
|
@ -120,6 +121,7 @@ USER node
|
|||
|
||||
# The web UI defaults to http://localhost:4747 - keep that contract.
|
||||
ENV GITNEXUS_HOME=/data/gitnexus \
|
||||
GITNEXUS_NO_UPDATE_NOTIFIER=1 \
|
||||
NODE_ENV=production \
|
||||
PORT=4747
|
||||
|
||||
|
|
|
|||
|
|
@ -37,7 +37,7 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
|
|||
### Index seems corrupt or "incremental" is misbehaving
|
||||
|
||||
- **Trigger:** `analyze` produces unexpected results, or `incrementalInProgress` is set in the index metadata (`.gitnexus/gitnexus.json` / legacy `meta.json`), or the index is in a half-state after a crash.
|
||||
- **Do:** `npx gitnexus analyze --force` to rebuild from scratch. The dirty-flag check forces this automatically when a previous incremental run didn't complete cleanly, but `--force` is the manual escape hatch. A dirty-flag recovery rebuild parks the interrupted run's sidecars beside the DB as `lbug.wal.dirty-recovery` / `lbug.shadow.dirty-recovery` for post-mortem debugging — harmless, and removable with `npx gitnexus clean --lbug-sidecars`. Safe to delete the `.gitnexus/parse-cache/` directory (and any legacy `.gitnexus/parse-cache.json`) at any time — content-addressed, will be regenerated.
|
||||
- **Do:** `npx gitnexus analyze --force` to rebuild the graph and FTS indexes. This may reuse unchanged parser output; when debugging parser/capture changes, use `npx gitnexus analyze --no-parse-cache` to rebuild that output too. The dirty-flag check forces the graph rebuild automatically when a previous incremental run didn't complete cleanly. A dirty-flag recovery rebuild parks the interrupted run's sidecars beside the DB as `lbug.wal.dirty-recovery` / `lbug.shadow.dirty-recovery` for post-mortem debugging — harmless, and removable with `npx gitnexus clean --lbug-sidecars`. Safe to delete the `.gitnexus/parse-cache/` directory (and any legacy `.gitnexus/parse-cache.json`) at any time — content-addressed, will be regenerated.
|
||||
- **Why:** Incremental writeback is selective DB row replacement; if the on-disk state is inconsistent for any reason, a full rebuild is the cheapest path back to a known-good index.
|
||||
|
||||
### Embeddings vanished after analyze
|
||||
|
|
@ -52,6 +52,12 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
|
|||
- **Do:** Re-run plain `npx gitnexus analyze` — no `--embeddings` flag needed. A retained `embeddingCheckpoint` in the index metadata forces embedding generation for exactly the pending nodes regardless of flags, and clears once they succeed. `--drop-embeddings` abandons the pending nodes instead of retrying them; `--force` also discards the checkpoint (with a warning) and rebuilds without resuming it.
|
||||
- **Why:** A long analyze run against a flaky HTTP embedding endpoint tolerates bounded sub-batch failures instead of aborting the whole run: it deletes the affected nodes' embedding rows (so they hold zero rows, never a partial set) and records those nodes as pending in `embeddingCheckpoint`. `stats.embeddings` stays an honest, non-zero count of everything that did succeed, so this state never trips the "Embeddings vanished" Sign above — `embedding-checkpoint-pending` is the only reliable signal.
|
||||
|
||||
### Scope extraction is incomplete
|
||||
|
||||
- **Trigger:** `npx gitnexus status` reports `incompleteReasons: ["scope-extraction-failed"]` when files were omitted, or `incompleteReasons: ["scope-extraction-unverified"]` when the index predates the completeness receipt or its metadata is unreadable. `impact`/`context` reports the same uncertainty as `epistemic: "lower-bound"`; confirmed omissions set `causes.scopeExtractionFiles > 0`.
|
||||
- **Do:** Re-run `npx gitnexus analyze` (`--force` for a full graph rebuild). If the reason persists, inspect the scope-extraction warnings and treat impact counts as floors until the affected source is supported or corrected.
|
||||
- **Why:** Parsing continued, but scope captures for the reported file count could not be produced even after the main-thread fallback. Calls, inheritance, imports, or accesses originating there may therefore be absent from the graph.
|
||||
|
||||
### Analyze reports INCOMPLETE with a collapsed graph write
|
||||
|
||||
- **Trigger:** `npx gitnexus status` reports `incompleteReasons: ["graph-write-collapsed"]`; the analyze summary printed `Repository indexed INCOMPLETELY` naming an expected and a persisted relationship count, and the CLI exited non-zero.
|
||||
|
|
@ -67,8 +73,8 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
|
|||
### Wrong repo in multi-repo setups
|
||||
|
||||
- **Trigger:** Query/impact results belong to another project.
|
||||
- **Do:** Call `list_repos`, then pass `repo` on subsequent tools.
|
||||
- **Why:** Default target is ambiguous when multiple repos are registered.
|
||||
- **Do:** Confirm an MCP default is configured or the GitNexus process was launched inside the intended registered path without crossing into an unindexed nested Git checkout. Otherwise call `list_repos`, then pass `repo` on subsequent tools; pass it for mutating tools when multiple repos are registered and no MCP default exists.
|
||||
- **Why:** Read-only tools derive their default from MCP configuration or a process cwd that stays within one registered Git boundary. Outside those paths the target remains ambiguous, and mutating tools stay explicit unless configuration supplies the target.
|
||||
|
||||
### LadybugDB lock / "database busy"
|
||||
|
||||
|
|
|
|||
150
README.md
150
README.md
|
|
@ -1,4 +1,4 @@
|
|||
# GitNexus
|
||||
# GitNexus (Akon Labs)
|
||||
|
||||
**⚠️ Important Notice:** GitNexus has NO official cryptocurrency, token, or coin. Any token/coin using the GitNexus name on Pump.fun or any other platform is **not affiliated with, endorsed by, or created by** this project or its maintainers. Do not purchase any cryptocurrency claiming association with GitNexus.
|
||||
|
||||
|
|
@ -179,7 +179,7 @@ flowchart TB
|
|||
| `group_list` | List configured repository groups |
|
||||
| `group_sync` | Rebuild a group's Contract Registry and cross-repo links |
|
||||
|
||||
> Per-repo tools take an optional `repo` parameter (omit it when only one repo is indexed) and an optional `branch` for indexes pinned with `gitnexus analyze --branch`. Omitting `branch` queries the workspace index, which follows your checked-out working tree — switching branches and re-running `gitnexus analyze` updates it incrementally. `explain` and `pdg_query` need an index built with `gitnexus analyze --pdg`.
|
||||
> Per-repo read-only tools take an optional `repo` parameter. Omit it when only one repo is indexed, an MCP default is configured, or the GitNexus process cwd is inside a registered path without crossing into an unindexed nested Git checkout; otherwise pass it explicitly. Mutating tools require `repo` when multiple repos are indexed and no MCP default exists. Per-repo tools also take an optional `branch` for indexes pinned with `gitnexus analyze --branch`. Omitting `branch` queries the workspace index, which follows your checked-out working tree — switching branches and re-running `gitnexus analyze` updates it incrementally. `explain` and `pdg_query` need an index built with `gitnexus analyze --pdg`.
|
||||
|
||||
### Resources for instant context
|
||||
|
||||
|
|
@ -384,6 +384,7 @@ Everyday commands:
|
|||
```bash
|
||||
gitnexus setup # Configure MCP for detected editors (one-time; -c to select)
|
||||
gitnexus analyze [path] # Index a repository (or update a stale index)
|
||||
gitnexus analyze [path] --watch # Watch local files and serialize incremental refreshes
|
||||
gitnexus mcp # Start MCP server (stdio) — serves all indexed repos
|
||||
gitnexus serve # Start local HTTP server (multi-repo) for web UI connection
|
||||
gitnexus eval-server # Start lightweight evaluation HTTP tools (loopback by default)
|
||||
|
|
@ -396,6 +397,28 @@ gitnexus uninstall # Preview removal of GitNexus MCP/skills/hooks
|
|||
|
||||
You can also query the graph directly from the terminal — `gitnexus query`, `context`, `impact`, `trace`, `cypher`, `detect-changes`, and `check` mirror the MCP tools of the same names, and `gitnexus doctor` prints runtime platform capabilities.
|
||||
|
||||
`gitnexus analyze --watch` requires a Git repository. It runs one initial
|
||||
analysis, then debounces scanner-admitted working-tree changes for 300 ms by
|
||||
default and applies serialized incremental refreshes. Events arriving during a
|
||||
refresh remain queued, and retryable failures retain the same batch with bounded
|
||||
backoff. Invalid `.gitnexusrc` or ignore-file reloads pause ordinary refreshes
|
||||
until the control file is fixed. Stop the watcher with Ctrl+C.
|
||||
|
||||
Watch mode accepts `--debounce`, `--workers`, `--worker-timeout`,
|
||||
`--max-file-size`, `--branch`, `--pdg`, `--name`, `--allow-duplicate-name`, and
|
||||
`--verbose`. Explicit one-shot options such as `--force`, `--repair-fts`,
|
||||
embedding flags, `--skills`, `--self-commit`, `--index-only`, and `--skip-git`
|
||||
are rejected. Unsupported defaults from `.gitnexusrc` are ignored with a
|
||||
warning rather than making an otherwise valid repository unwatchable.
|
||||
|
||||
POSIX requests clone-first copy-and-swap publication when the live index has no
|
||||
orphan sidecars. Windows and sidecar fallback runs update in place: failures
|
||||
known to occur before writes are retried, while a failure that may have mutated
|
||||
the live index stops the watcher. Watch mode does not pull remotes. Running MCP
|
||||
and `serve` processes reopen a newly published index automatically; MCP observes
|
||||
the replacement on its next tool call, typically within five seconds, so no
|
||||
restart is required.
|
||||
|
||||
<details>
|
||||
<summary><strong>Authenticated <code>eval-server</code> binding</strong></summary>
|
||||
|
||||
|
|
@ -413,7 +436,8 @@ The token may be set in the shell, `.env.local`, or `.env` in the working direct
|
|||
<summary><strong>All <code>analyze</code> flags</strong></summary>
|
||||
|
||||
```bash
|
||||
gitnexus analyze --force # Full rebuild: re-parse + graph rebuild + FTS rebuild
|
||||
gitnexus analyze --force # Full graph + FTS rebuild (reuses unchanged parser output)
|
||||
gitnexus analyze --no-parse-cache # Full rebuild that re-parses every source file
|
||||
gitnexus analyze --repair-fts # Fast path: rebuild/verify only FTS indexes on existing index data
|
||||
gitnexus analyze --skills # Generate repo-specific skill files from detected communities
|
||||
gitnexus analyze --skip-embeddings # Skip embedding generation (faster)
|
||||
|
|
@ -426,10 +450,20 @@ gitnexus analyze --verbose # Log skipped files when parsers are unavailabl
|
|||
gitnexus analyze --worker-timeout 60 # Increase worker idle timeout for slow parses
|
||||
gitnexus analyze --workers <n> # Parse worker pool size (>=1; default: cores-1, capped at 16,
|
||||
# auto-sized to the repo). 0 is rejected — there is no sequential mode.
|
||||
gitnexus analyze --spring-actuator ./actuator # Enrich with local Spring Boot Actuator JSON snapshots
|
||||
gitnexus analyze --asyncapi-spec ./docs/asyncapi # Resolve broker addresses from AsyncAPI 3.x documents
|
||||
gitnexus analyze --wal-checkpoint-threshold 67108864 # LadybugDB WAL auto-checkpoint threshold in bytes
|
||||
# (default 67108864 = 64 MiB; -1 keeps Ladybug stock ~16 MiB)
|
||||
```
|
||||
|
||||
`--spring-actuator` is explicitly opt-in and accepts either a JSON bundle keyed by `mappings`, `beans`, `conditions`, `configprops`, and/or `env`, or a directory containing endpoint-named JSON files. It confirms matching static nodes and adds conservative runtime-only routes, beans, and property keys. The configured input is excluded from source scanning; only normalized repository-relative exclusions are retained for future scans, never absolute paths. Env/configprops values, origins, condition messages, and source names are never persisted or printed. Because snapshots are external runtime state, an enabled run always rebuilds; the first later run without the option rebuilds once to remove runtime evidence. The same path can be set as `springActuator` in `.gitnexusrc`.
|
||||
|
||||
`--asyncapi-spec` is explicitly opt-in and accepts a directory of AsyncAPI documents or a single document; the path is resolved against the repository root, so a committed `docs/asyncapi` and an absolute cache written by something else both work. Each `operations[]` entry of an **AsyncAPI 3.x** document can contribute a `Destination` node keyed by broker and address, with `action: send` emitting `PUBLISHES_TO` and `action: receive` emitting `CONSUMES_FROM`, so a document and source code that name one address on one broker land on the same node. Edges start at the document, not at a callable — a document states that the service talks to an address, not which method does — and no address a document names is ever attached to an unresolved source site.
|
||||
|
||||
An operation must name a protocol, either through its own `bindings` or through the `servers[].protocol` of the servers its channel resolves to (a channel that lists no `servers` resolves to all of them); operations that name none are refused, as are operations whose two readings name different brokers, and channels that inherit a multi-protocol server set without choosing. HTTP and WebSocket documents are refused for destination minting: there the host rather than the address names the place, and an HTTP endpoint is already modelled as a `Route`. A parameterized address — a channel declaring `parameters`, or an address containing `{` — is refused rather than keyed: two services publishing `{env}.orders` share a pattern, not a queue. AsyncAPI **2.x is refused** under its own counted reason and never mapped, because its `publish`/`subscribe` are inverted relative to 3.x `send`/`receive` and a naive mapping would reverse the async graph while leaving it connected. Every refusal is counted, and a configured path that yields nothing is reported rather than passed over in silence.
|
||||
|
||||
Like Actuator snapshots, documents are external to git freshness — replacing one moves no commit and dirties no file — so an enabled run always rebuilds, and the first later run without the option rebuilds once to remove document-derived evidence. There is no glob-based auto-discovery, and the option is unsupported with `--watch`.
|
||||
|
||||
If `analyze` reports a worker parse timeout on a large or unusual repository, it keeps running and falls back safely. To give slow worker jobs more time, use `--worker-timeout 60` or set `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS=60000`. For very large files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` controls the worker job byte budget.
|
||||
|
||||
**Embeddings node limit** — `gitnexus analyze --embeddings` generates semantic search vectors with a default 50,000-node safety cap to protect memory on large repositories:
|
||||
|
|
@ -444,6 +478,46 @@ If embeddings are skipped on a large repository, the indexed graph likely exceed
|
|||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>Keep remote repositories indexed with <code>gitnexus auto-sync</code></strong></summary>
|
||||
|
||||
`gitnexus auto-sync` clones or pulls configured repositories, analyzes new commits, and optionally syncs their group. It runs once immediately, then repeats on the configured interval. It runs in the foreground; use your process manager if it must survive a shell session. `gitnexus watch` is reserved and prints this split; it does not start auto-sync or local file watching.
|
||||
|
||||
```bash
|
||||
# 1. Create the config once. It never overwrites an existing file.
|
||||
gitnexus auto-sync init
|
||||
|
||||
# 2. Edit $GITNEXUS_HOME/watch_config.yml, then start it.
|
||||
gitnexus auto-sync start # `gitnexus auto-sync` is equivalent
|
||||
gitnexus auto-sync status
|
||||
gitnexus auto-sync restart # Required after config changes
|
||||
gitnexus auto-sync stop
|
||||
gitnexus auto-sync reset # Clear failure state; leaves clones and indexes intact
|
||||
```
|
||||
|
||||
`GITNEXUS_HOME` defaults to `~/.gitnexus`. A minimal configuration:
|
||||
|
||||
```yaml
|
||||
sync_interval_minutes: 10
|
||||
analyze_timeout: 5m
|
||||
projects:
|
||||
- local_path: /absolute/path/to/clones
|
||||
branches: [main, master]
|
||||
overwrite_local_changes: false
|
||||
remote_urls:
|
||||
- git@github.com:owner/repo.git
|
||||
```
|
||||
|
||||
- `sync_interval_minutes` must be at least `5`; `local_path` must be an absolute path. Clones are stored below it as `host/namespace/repo`.
|
||||
- Remote URLs must use SSH SCP form and are limited to GitHub, GitLab, or Gitee.
|
||||
- `branches` are tried in order. The legacy `branch` field is supported, but do not set both.
|
||||
- Analysis runs in an isolated worker; `analyze_timeout` defaults to, and cannot exceed, half of `sync_interval_minutes`. Timeout and `auto-sync stop` request safe cancellation; a worker in native work exits after reaching a JS-visible safe point. Until then, auto-sync reports `cancelling` or `stopping` and retains ownership so another auto-sync cannot take over, for up to 5 seconds — after that the parent stops waiting and leaves the worker to exit on its own rather than killing it mid-write. This behavior is the same on macOS and Windows. `overwrite_local_changes` defaults to `false`, so a dirty local clone is skipped rather than overwritten; setting it to `true` also deletes untracked files in the clone, while keeping ignored paths.
|
||||
- Add `group_name` only after creating that group with `gitnexus group create <name>`. Partial clone output is isolated and removed after 14 days.
|
||||
|
||||
See the [full auto-sync configuration and runtime reference](gitnexus/README.md#gitnexus-auto-sync) for concurrency, timeouts, failure thresholds, and runtime files.
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>Repository groups</strong> (multi-repo / monorepo service tracking)</summary>
|
||||
|
||||
|
|
@ -477,6 +551,7 @@ Commit a `.gitnexusrc` JSON file at the repo root to preconfigure recurring `ana
|
|||
"skipContextFiles": true, // alias of skipAgentsMd: keep your own AGENTS.md/CLAUDE.md
|
||||
"skipSkills": true, // don't install standard skill files under .claude/skills/ and .agents/skills/
|
||||
"embeddings": true, // generate embeddings by default
|
||||
"springActuator": "./actuator", // optional local runtime snapshot directory or bundle
|
||||
"workerTimeout": 60,
|
||||
}
|
||||
```
|
||||
|
|
@ -491,7 +566,7 @@ Notes:
|
|||
|
||||
- The default branch is resolved as: `--default-branch` > `.gitnexusrc` `defaultBranch`/`branch` > auto-detected `origin/HEAD` > `main`.
|
||||
- `skipContextFiles` / `skipAiContext` are aliases for `skipAgentsMd` — they skip the `AGENTS.md` / `CLAUDE.md` block only. They do **not** imply `skipSkills`. `indexOnly` is the stronger option that skips all file injection.
|
||||
- Supported keys: `defaultBranch` (`branch`), `skipAgentsMd` (`skipContextFiles`, `skipAiContext`), `skipSkills`, `indexOnly`, `stats`/`noStats`, `embeddings`, `dropEmbeddings`, `name`, `allowDuplicateName`, `maxFileSize`, `workerTimeout`, `walCheckpointThreshold`, `workers`, `embeddingThreads`, `embeddingBatchSize`, `embeddingSubBatchSize`, `embeddingDevice`.
|
||||
- Supported keys: `defaultBranch` (`branch`), `skipAgentsMd` (`skipContextFiles`, `skipAiContext`), `skipSkills`, `indexOnly`, `stats`/`noStats`, `embeddings`, `dropEmbeddings`, `name`, `allowDuplicateName`, `maxFileSize`, `workerTimeout`, `walCheckpointThreshold`, `workers`, `springActuator`, `embeddingThreads`, `embeddingBatchSize`, `embeddingSubBatchSize`, `embeddingDevice`.
|
||||
- The file is JSON only. Unknown keys and invalid values fail fast with an actionable error before analysis starts.
|
||||
|
||||
</details>
|
||||
|
|
@ -501,36 +576,39 @@ Notes:
|
|||
|
||||
Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max-file-size`, `--verbose`). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over env vars; env vars take precedence over built-in defaults.
|
||||
|
||||
| Variable | Default | Effect | Tune when… |
|
||||
| ----------------------------------------------- | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size (must be ≥ 1). Equivalent to `--workers <n>`. The worker pool is the sole parse path — there is no sequential parser, so `0` is rejected with an actionable error (the pool self-heals via quarantine + respawn). | Constrained containers (cgroup CPU limits) or CI runners with explicit quotas. To narrow down a worker crash set `1` for a single-worker pool — not `0`. |
|
||||
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
|
||||
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
|
||||
| `GITNEXUS_AUTH_TOKEN` | unset | Bearer token required when `eval-server` binds beyond loopback. May also be read from `.env.local` or `.env`; shell values take precedence. | Exposing the evaluation HTTP tools to a container, VM, or LAN. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED` | unset | When `1`, emits `[deferred-profile]` timing/progress logs for the post-chunk deferred resolution band (imports → heritage → buildHeritageMap → legacy call resolution). Implied by `GITNEXUS_VERBOSE`. | Diagnosing analyze stalls in "Resolving calls (all chunks)" on large Java/Kotlin repos (issue #1741) without the full verbose ingestion noise. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED_SLOW_MS` | `3000` (verbose) / `5000` | Per-file threshold in ms above which `processCallsFromExtracted` emits a `slow file …` log line. Parsed via `Number()`: accepts integers (`5000`), scientific notation (`2.5e3`), decimals (`.5`), and hex (`0x10`). Non-finite or non-positive values fall back to the default. | Hunting a few outlier files dominating the deferred call-resolution stage; lower to surface more, raise to focus only on the worst. |
|
||||
| `PROF_LBUG_LOAD` | unset | When `1`, emits one `[lbug-load prof]` summary line per `loadGraphToLbug` call breaking the graph-DB persistence wall into stages (`csv-emit` / `copy-nodes` / `copy-rels` / `fallback` / `total`) plus node & edge counts. Zero-cost when unset. | Attributing large-repo analyze wall time across CSV generation vs. LadybugDB `COPY` (issue #2203) — the analyze "emit" timing is the scope-resolution bucket, not this DB-write path. |
|
||||
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
|
||||
| `GITNEXUS_WORKER_READY_TIMEOUT_MS` | `5000` | Startup budget in milliseconds for a parse worker to load its grammar bindings and report `{type:'ready'}`. Slots that miss it are treated as startup crashes. | Slow or heavily loaded hosts where a full pool cold-starting concurrently needs more than 5s, and analyze aborts with "did not report ready within 5000ms". |
|
||||
| `GITNEXUS_FTS_STEMMER` | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` for matching repository comments. Re-run `gitnexus analyze --repair-fts` after changing it. | Keyword search quality is poor for non-English comments or identifiers under English stemming. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold in bytes. Equivalent to `--wal-checkpoint-threshold <bytes>`. `-1` keeps LadybugDB's stock threshold (~16 MiB). Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. | You need a larger or smaller WAL auto-checkpoint threshold for your analyze workload. |
|
||||
| `GITNEXUS_LBUG_BUFFER_POOL_SIZE` | min(2 GiB, 80% RAM) | LadybugDB buffer-pool ceiling in bytes for every GitNexus database (analyze, MCP server, serve, group bridges). `0` restores LadybugDB's native unbounded default of 80% of system RAM; invalid values warn and fall back to the default (#2557). During `analyze` the pool is right-sized to the graph, scaled on non-4 KiB-page hosts by the page-size granule ratio up to min(2 GiB × pageSize/4 KiB, 80% RAM) (#2631); this env var overrides all of that as an absolute value. | A long-lived `gitnexus mcp` or a big incremental `analyze` uses too much memory, or a huge repo's working set genuinely needs a pool larger than 2 GiB. |
|
||||
| `GITNEXUS_LBUG_MAX_DB_SIZE` | `17179869184` (16 GiB) | Maximum size in bytes of a single LadybugDB database file — an mmap/disk-address-space ceiling, not a memory limit (it does not constrain the buffer pool). Invalid values silently fall back to the default. | Indexing a genuinely huge monorepo whose on-disk graph index approaches 16 GiB. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
|
||||
| `GITNEXUS_WORKER_SHUTDOWN_DRAIN_MS` | `30000` | Max wait at pool shutdown for a retired worker still inside native code. The worker is terminated at its next JS-safe point instead of mid-native-call (which aborts the whole process with `Napi::Error`, #2432); on expiry it is left running, unref'd, and terminated when it surfaces. | Shutdown latency matters more than draining a wedged worker (lower), or a legitimately-slow native grammar needs longer to surface (raise). |
|
||||
| `GITNEXUS_CPP_CAPTURE_BUDGET_MS` | `20000` | Per-file wall-clock budget for C++ capture extraction. On breach the file keeps the captures accumulated so far and logs a warning — the worker returns to JS instead of stalling in native-heavy loops (#2432). `0` expires immediately. | Pathological generated C++ that still exceeds the budget after the indexed lookups; raise for completeness, lower to fail-fast. |
|
||||
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Chunk boundary used for cache-key composition and dispatch. Smaller = finer-grained cache hits but more dispatch overhead. | Tuning incremental-analyze cache behavior on monorepos. |
|
||||
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
|
||||
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips the vendored grammar materialize for `tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`, and `tree-sitter-kotlin` at install time (and the Dart/Proto source builds). Those four won't be parsed; the install still succeeds. | Installing on a host without a C++ toolchain or where the vendored prebuilds don't match; willing to skip Dart/Proto/Swift/Kotlin parsing. |
|
||||
| `GITNEXUS_MCP_READ_ONLY` | unset | Set to `1` to expose only proven single-repository read tools and resources; `0` disables the policy and any other value fails startup. | The MCP server runs in an environment where graph mutation, raw Cypher, and cross-repository group routing must be unavailable. |
|
||||
| `GITNEXUS_MCP_ALLOWED_REPOS` | unset | Comma-separated allowlist of canonical indexed repository names or absolute paths. Invalid, ambiguous, or blank entries fail startup. | One MCP process must expose only a bounded subset of the repositories in the global registry. |
|
||||
| `GITNEXUS_MCP_DEFAULT_REPO` | unset | Canonical indexed repository name or absolute path used when a tool or resource omits its repository. Must belong to the allowlist when one is set. | Several repositories are available but unqualified MCP calls should resolve deterministically. |
|
||||
| `GITNEXUS_MCP_DEFAULT_MAX_TOKENS` | unset | Default positive-integer response budget for MCP `query`, `context`, and `impact`, estimated at four UTF-8 bytes per token. Explicit `maxTokens` wins. | Long MCP responses consume too much model context and callers cannot reliably add a per-request budget. |
|
||||
| `GITNEXUS_PUBLIC_ORIGIN` | unset | The single browser origin `serve` is reached through, added to the CORS allowlist and to the write-route origin guard. A wildcard bind (`0.0.0.0`) has no host identity, so without this the server's own UI is refused. **Setting it currently refuses to start:** `serve` has no authentication, requests carrying no `Origin` header already reach `POST /api/analyze` and `DELETE /api/repo`, and this is the setting that would admit browser writes on top of that. Matching rules for when the gate lifts: the hostname must match exactly, and so must the scheme. A value with no scheme (`app.example.com`) means `https`, since a bare host comes from platform service discovery and those terminate TLS; spell out `http://app.example.com` for plain HTTP. An explicit port must match; with no port, any port on that hostname is accepted. Anything that is not one reachable host (a list, `*`, a bare port number, a `:0` port, a trailing dot) warns at startup and allows nothing. | `gitnexus serve` runs behind a reverse proxy or on a wildcard bind, and the UI's index/delete requests return `origin_not_allowed`. |
|
||||
| Variable | Default | Effect | Tune when… |
|
||||
| ----------------------------------------------- | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size (must be ≥ 1). Equivalent to `--workers <n>`. The worker pool is the sole parse path — there is no sequential parser, so `0` is rejected with an actionable error (the pool self-heals via quarantine + respawn). | Constrained containers (cgroup CPU limits) or CI runners with explicit quotas. To narrow down a worker crash set `1` for a single-worker pool — not `0`. |
|
||||
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
|
||||
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
|
||||
| `GITNEXUS_ANALYZER_IDENTITY_IN_PROCESS_GUARDS` | unset | When truthy (`1`/`true`/`yes`), forces in-process cache-guard validation once a batch has ≥128 requests. In-process mode also auto-selects when `packageRoot`/`buildRoot` fail `W_OK` with `EACCES`/`EROFS`. Otherwise those large batches use a Node subprocess probe. Batches under 128 always stay in-process. | Trusted or read-only installs where two identity subprocess spawns per analyze dominate wall time; leave unset to keep the default isolation path on writable trees. |
|
||||
| `GITNEXUS_RESOLVE_DEF_GRAPH_ID_MEMO` | on (unset) | Memoizes `resolveDefGraphId` per `nodeLookup` instance (WeakMap). Enabled by default. Set to `0`/`false`/`off`/`no` to disable and recompute on every call (debug / bisect memo bugs). | Suspecting stale graph-id resolution after a lookup rebuild, or comparing memo vs uncached cost on a large index. |
|
||||
| `GITNEXUS_AUTH_TOKEN` | unset | Bearer token required when `eval-server` binds beyond loopback. May also be read from `.env.local` or `.env`; shell values take precedence. | Exposing the evaluation HTTP tools to a container, VM, or LAN. |
|
||||
| `GITNEXUS_MCP_AUTH_TOKEN` | unset | Bearer token for the dedicated `gitnexus mcp --http` server, for a **directly reachable** `gitnexus serve` `/api/mcp` route, and for the `docker-server` / web proxy in front of one. A non-loopback dedicated MCP bind requires it; `serve` enables protocol-layer MCP auth when it is set. Behind a proxy, set the **same** value on both services: the proxy spends the edge `GITNEXUS_SERVE_AUTH_TOKEN`, then replaces `Authorization` with this token on `/api/mcp` only. | Dedicated MCP, a `serve` the client can reach directly, or a proxied deploy (Render Blueprint) where the backend runs protocol-layer MCP auth — configure it on the proxy too. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED` | unset | When `1`, emits `[deferred-profile]` timing/progress logs for the post-chunk deferred resolution band (imports → heritage → buildHeritageMap → legacy call resolution). Implied by `GITNEXUS_VERBOSE`. | Diagnosing analyze stalls in "Resolving calls (all chunks)" on large Java/Kotlin repos (issue #1741) without the full verbose ingestion noise. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED_SLOW_MS` | `3000` (verbose) / `5000` | Per-file threshold in ms above which `processCallsFromExtracted` emits a `slow file …` log line. Parsed via `Number()`: accepts integers (`5000`), scientific notation (`2.5e3`), decimals (`.5`), and hex (`0x10`). Non-finite or non-positive values fall back to the default. | Hunting a few outlier files dominating the deferred call-resolution stage; lower to surface more, raise to focus only on the worst. |
|
||||
| `PROF_LBUG_LOAD` | unset | When `1`, emits one `[lbug-load prof]` summary line per `loadGraphToLbug` call breaking the graph-DB persistence wall into stages (`csv-emit` / `copy-nodes` / `copy-rels` / `fallback` / `total`) plus node & edge counts. Zero-cost when unset. | Attributing large-repo analyze wall time across CSV generation vs. LadybugDB `COPY` (issue #2203) — the analyze "emit" timing is the scope-resolution bucket, not this DB-write path. |
|
||||
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
|
||||
| `GITNEXUS_WORKER_READY_TIMEOUT_MS` | `5000` | Startup budget in milliseconds for a parse worker to load its grammar bindings and report `{type:'ready'}`. Slots that miss it are treated as startup crashes. | Slow or heavily loaded hosts where a full pool cold-starting concurrently needs more than 5s, and analyze aborts with "did not report ready within 5000ms". |
|
||||
| `GITNEXUS_FTS_STEMMER` | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` for matching repository comments. Re-run `gitnexus analyze --repair-fts` after changing it. | Keyword search quality is poor for non-English comments or identifiers under English stemming. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold in bytes. Equivalent to `--wal-checkpoint-threshold <bytes>`. `-1` keeps LadybugDB's stock threshold (~16 MiB). Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. | You need a larger or smaller WAL auto-checkpoint threshold for your analyze workload. |
|
||||
| `GITNEXUS_LBUG_BUFFER_POOL_SIZE` | min(2 GiB, 80% RAM) | LadybugDB buffer-pool ceiling in bytes for every GitNexus database (analyze, MCP server, serve, group bridges). `0` restores LadybugDB's native unbounded default of 80% of system RAM; invalid values warn and fall back to the default (#2557). During `analyze` the pool is right-sized to the graph, scaled on non-4 KiB-page hosts by the page-size granule ratio up to min(2 GiB × pageSize/4 KiB, 80% RAM) (#2631); this env var overrides all of that as an absolute value. | A long-lived `gitnexus mcp` or a big incremental `analyze` uses too much memory, or a huge repo's working set genuinely needs a pool larger than 2 GiB. |
|
||||
| `GITNEXUS_LBUG_MAX_DB_SIZE` | `17179869184` (16 GiB) | Maximum size in bytes of a single LadybugDB database file — an mmap/disk-address-space ceiling, not a memory limit (it does not constrain the buffer pool). Invalid values silently fall back to the default. | Indexing a genuinely huge monorepo whose on-disk graph index approaches 16 GiB. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
|
||||
| `GITNEXUS_WORKER_SHUTDOWN_DRAIN_MS` | `30000` | Max wait at pool shutdown for a retired worker still inside native code. The worker is terminated at its next JS-safe point instead of mid-native-call (which aborts the whole process with `Napi::Error`, #2432); on expiry it is left running, unref'd, and terminated when it surfaces. | Shutdown latency matters more than draining a wedged worker (lower), or a legitimately-slow native grammar needs longer to surface (raise). |
|
||||
| `GITNEXUS_CPP_CAPTURE_BUDGET_MS` | `20000` | Per-file wall-clock budget for C++ capture extraction. On breach the file keeps the captures accumulated so far and logs a warning — the worker returns to JS instead of stalling in native-heavy loops (#2432). `0` expires immediately. | Pathological generated C++ that still exceeds the budget after the indexed lookups; raise for completeness, lower to fail-fast. |
|
||||
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Per-bucket byte budget for parse-cache packing. Files are grouped by `(language, hash(path) mod 128)`; packs inside a bucket are cut at this limit. Smaller = finer-grained invalidation and more dispatch. Default is always 2 MiB and no longer scales with worker count. | Tuning incremental-analyze cache invalidation on monorepos without changing `--workers`. |
|
||||
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
|
||||
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips the vendored grammar materialize for `tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`, and `tree-sitter-kotlin` at install time (and the Dart/Proto source builds). Those four won't be parsed; the install still succeeds. | Installing on a host without a C++ toolchain or where the vendored prebuilds don't match; willing to skip Dart/Proto/Swift/Kotlin parsing. |
|
||||
| `GITNEXUS_MCP_READ_ONLY` | unset | Set to `1` to expose only proven single-repository read tools and resources; `0` disables the policy and any other value fails startup. | The MCP server runs in an environment where graph mutation, raw Cypher, and cross-repository group routing must be unavailable. |
|
||||
| `GITNEXUS_MCP_ALLOWED_REPOS` | unset | Comma-separated allowlist of canonical indexed repository names or absolute paths. Invalid, ambiguous, or blank entries fail startup. | One MCP process must expose only a bounded subset of the repositories in the global registry. |
|
||||
| `GITNEXUS_MCP_DEFAULT_REPO` | unset | Canonical indexed repository name or absolute path used when a tool or resource omits its repository. Must belong to the allowlist when one is set. | Several repositories are available but unqualified MCP calls should resolve deterministically. |
|
||||
| `GITNEXUS_MCP_DEFAULT_MAX_TOKENS` | unset | Default positive-integer response budget for MCP `query`, `context`, and `impact`, estimated at four UTF-8 bytes per token. Explicit `maxTokens` wins. | Long MCP responses consume too much model context and callers cannot reliably add a per-request budget. |
|
||||
| `GITNEXUS_PUBLIC_ORIGIN` | unset | The single browser origin `serve` is reached through, added to the CORS allowlist and to the write-route origin guard. A wildcard bind (`0.0.0.0`) has no host identity, so without this the server's own UI is refused. **Setting it currently refuses to start:** `serve` has no authentication, requests carrying no `Origin` header already reach `POST /api/analyze` and `DELETE /api/repo`, and this is the setting that would admit browser writes on top of that. Matching rules for when the gate lifts: the hostname must match exactly, and so must the scheme. A value with no scheme (`app.example.com`) means `https`, since a bare host comes from platform service discovery and those terminate TLS; spell out `http://app.example.com` for plain HTTP. An explicit port must match; with no port, any port on that hostname is accepted. Anything that is not one reachable host (a list, `*`, a bare port number, a `:0` port, a trailing dot) warns at startup and allows nothing. | `gitnexus serve` runs behind a reverse proxy or on a wildcard bind, and the UI's index/delete requests return `origin_not_allowed`. |
|
||||
| `GITNEXUS_TRUST_PROXY` | `loopback, linklocal, uniquelocal` | Express `trust proxy` value — which upstream hops may set `X-Forwarded-*`, and so what the per-IP rate limiter reads as the client IP. Set it to the exact number of proxies you control. Every hop past that is one more entry of the chain the caller gets to write. `false`/`no`/`off` (and a `0` hop count) trust no hop; a proxy list Express can compile (`loopback`, `10.0.0.0/8, 127.0.0.1`) names them instead. `true`/`yes`/`on` is **rejected**: it reads the client-controlled leftmost `X-Forwarded-For` entry, so a spoofed chain earns a fresh rate-limit key per request, and express-rate-limit rejects it too (`ERR_ERL_PERMISSIVE_TRUST_PROXY`). Counts above `16` are rejected as well, as a sanity ceiling rather than a safety boundary. Any invalid value warns and falls back to the default. Bind non-loopback with this unset and `serve` warns: a load balancer outside the private ranges is untrusted, so every request keys to the balancer and the per-IP limit becomes one shared limit. | `serve` sits behind a load balancer outside the private ranges (AWS ALB, Cloudflare, CGNAT), where every request otherwise collapses to the proxy hop and rate limiting goes global. |
|
||||
|
||||
</details>
|
||||
|
|
@ -580,6 +658,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
|
|||
| C | — | — | ✓ | — | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| C++ | — | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Dart | ✓ | — | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
||||
| Zig | ✓ | — | ✓ | — | ✓ | ✓ | ✓ | — | ✓ |
|
||||
|
||||
**Imports** — cross-file import resolution · **Named Bindings** — `import { X as Y }` / re-export tracking · **Exports** — public/exported symbol detection · **Heritage** — class inheritance, interfaces, mixins · **Type Annotations** — explicit type extraction for receiver resolution · **Constructor Inference** — infer receiver type from constructor calls (`self`/`this` resolution included for all languages) · **Config** — language toolchain config parsing (tsconfig, go.mod, etc.) · **Frameworks** — AST-based framework pattern detection · **Entry Points** — entry point scoring heuristics
|
||||
|
||||
|
|
@ -589,7 +668,7 @@ GitNexus builds a complete knowledge graph of your codebase through a multi-phas
|
|||
|
||||
GitNexus uses a **global registry** so one MCP server can serve multiple indexed repos. No per-project MCP config needed — set it up once and it works everywhere.
|
||||
|
||||
Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. LadybugDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). If only one repo is indexed, the `repo` parameter is optional on all tools — agents don't need to change anything.
|
||||
Each `gitnexus analyze` stores the index in `.gitnexus/` inside the repo (portable, gitignored) and registers a pointer in `~/.gitnexus/registry.json`. When an AI agent starts, the MCP server reads the registry and can serve any indexed repo. LadybugDB connections are opened lazily on first query and evicted after 5 minutes of inactivity (max 5 concurrent). Read-only tools can omit `repo` when only one repo is indexed, an MCP default is configured, or the GitNexus process cwd is inside a registered path without crossing into an unindexed nested Git checkout. Outside those paths—and for mutating tools with multiple indexed repos and no MCP default—pass `repo` explicitly.
|
||||
|
||||
<details>
|
||||
<summary><strong>Architecture diagram</strong></summary>
|
||||
|
|
@ -767,6 +846,7 @@ gitnexus wiki
|
|||
# Use a custom model or provider (default model: minimax/minimax-m2.5)
|
||||
gitnexus wiki --model gpt-4o
|
||||
gitnexus wiki --base-url https://api.anthropic.com/v1
|
||||
gitnexus wiki --provider grok # local Grok Build CLI (uses `grok login`, no API key)
|
||||
|
||||
# Use the Atlas Cloud preset
|
||||
export ATLASCLOUD_API_KEY=<key>
|
||||
|
|
|
|||
11
RUNBOOK.md
11
RUNBOOK.md
|
|
@ -46,6 +46,17 @@ npx gitnexus status
|
|||
npx gitnexus list
|
||||
```
|
||||
|
||||
**Scope extraction incomplete:** `npx gitnexus status` reports
|
||||
`incompleteReasons: ["scope-extraction-failed"]` when one or more files still
|
||||
lack scope captures after the worker and fallback passes. `impact` and `context`
|
||||
then report a lower bound with `causes.scopeExtractionFiles` set to the affected
|
||||
file count. Re-run `npx gitnexus analyze --force`; if the reason remains, inspect
|
||||
the scope-extraction warnings for the unsupported or malformed source file.
|
||||
Every pre-existing index remains unverified until it is analyzed once by a
|
||||
version that writes the completeness receipt. An older index or unreadable completeness record reports
|
||||
`incompleteReasons: ["scope-extraction-unverified"]`; re-analyze it before treating
|
||||
empty impact results as exact.
|
||||
|
||||
---
|
||||
|
||||
## Embeddings
|
||||
|
|
|
|||
|
|
@ -59,11 +59,16 @@ The `render.yaml` Blueprint (see the README's **Deploy to Render**) puts `gitnex
|
|||
- **The generated `GITNEXUS_SERVE_AUTH_TOKEN` is the only access control.** The proxy rejects any `/api/*` request without it with a `401` before forwarding. Rotate it by editing the environment variable on the `gitnexus-web` service and redeploying.
|
||||
- **The CSRF guard is inert on this path.** The proxy strips `Origin` before forwarding, so the server's write-origin guard does nothing for proxied traffic — it passes `Origin`-less requests through by design. The token is not a second layer behind the guard.
|
||||
- **Anyone holding the token can read every indexed repo's source.** These routes carry no origin guard, and the first three carry no rate limiter either: `GET /api/repos`, `GET /api/graph`, `POST /api/query`, `GET /api/file`, `GET /api/grep`. Whoever has the token can also index and delete repositories.
|
||||
- **`POST /api/mcp` rides the same path.** `serve` mounts the MCP handler via `mountMCPEndpoints`, and `createStreamableHttpHandler` is called with no `authToken` — a **pre-existing** gap in `serve` itself, not something this deploy introduces. On Render it is closed only by the edge token and the private network. A `serve` bound directly to a public interface has no such cover.
|
||||
- **`POST /api/mcp` rides the same path.** When `GITNEXUS_MCP_AUTH_TOKEN` is set on the backend, `serve` protects `/api/mcp` with the same constant-time Bearer check as the dedicated HTTP MCP server, before parsing the request body. The Render Blueprint does not set a backend MCP token by default. To enable it behind the proxy, set the **same** `GITNEXUS_MCP_AUTH_TOKEN` on both the `gitnexus-web` proxy and the `gitnexus-server` backend: the proxy consumes the edge `GITNEXUS_SERVE_AUTH_TOKEN`, then replaces `Authorization` with the MCP token on `/api/mcp` (and its subpaths) only — the edge credential is never forwarded, and other `/api/*` routes stay stripped. Configuring it on the backend alone makes every proxied MCP request `401`.
|
||||
- **A directly reachable `serve` still needs an explicit control.** If neither `GITNEXUS_MCP_AUTH_TOKEN` nor an authenticated edge/private-network boundary is present, `/api/mcp` is unauthenticated. Do not bind that topology to a LAN or public interface: MCP readers can access indexed source and graph context.
|
||||
- **Rate limits bound cost, not access.** They cap what a token holder can spend; they do not decide who gets in.
|
||||
|
||||
Do not hand the URL out as a public demo. A token holder has read access to everything the deploy has indexed.
|
||||
|
||||
### `/api/grep` regex semantics and residual ReDoS exposure
|
||||
|
||||
`GET /api/grep` executes caller-supplied patterns as real regular expressions (with an optional path-substring `fileFilter` and `caseSensitive` flag) to honor the web chat's grep tool contract; `literal=1` restores the older escaped-substring mode. Mitigations: a 200-character pattern cap, line-by-line matching, a max-200 result cap, and a 5-second wall-clock budget. Matching runs in a `worker_threads` worker so a catastrophic pattern (e.g. `(a+)+$`) can be killed with `terminate()` when the budget expires — the parent event loop (other routes + SSE) stays responsive. A timed-out scan returns partial results with `timedOut: true`; the web grep tool surfaces that flag so an agent does not treat a cut-off scan as exhaustive. CodeQL still flags constructing a `RegExp` from the query string; that is the advertised contract, not accidental injection. Hosted deploys continue to gate the route behind the edge token.
|
||||
|
||||
## Automated Scans Running in CI
|
||||
|
||||
This repository runs the following scans automatically. Findings appear under the repository's **Security → Code scanning** tab.
|
||||
|
|
|
|||
|
|
@ -112,6 +112,14 @@ const upstreamOrigin = upstreamBase ? new URL(upstreamBase).origin : null;
|
|||
// (gitnexus/src/mcp/http-transport.ts).
|
||||
const authToken = process.env.GITNEXUS_SERVE_AUTH_TOKEN?.trim() || null;
|
||||
|
||||
// The protocol-layer credential the upstream `serve` expects on /api/mcp when it
|
||||
// runs with MCP Bearer auth enabled. Set it to the SAME value on both services:
|
||||
// the edge token is spent here and replaced with this one for MCP requests only
|
||||
// (see proxyToUpstream). Unset — the default — means no injection, so a backend
|
||||
// without MCP auth is unaffected. Blank-is-absent follows resolveAuthToken
|
||||
// (gitnexus/src/mcp/http-transport.ts). Never logged.
|
||||
const mcpAuthToken = process.env.GITNEXUS_MCP_AUTH_TOKEN?.trim() || null;
|
||||
|
||||
// Mirrors the non-loopback refusal in http-transport.ts (startMcpHttpServer),
|
||||
// relocated because the trust boundary is here: an unguarded `serve` behind a
|
||||
// private service is legitimate, an unguarded public proxy is not.
|
||||
|
|
@ -341,11 +349,17 @@ async function proxyToUpstream(req, res) {
|
|||
// talks to this same-origin web service.
|
||||
delete headers.origin;
|
||||
delete headers.referer;
|
||||
// The edge token is spent here. `serve` reads no Authorization header
|
||||
// (gitnexus/src/server/mcp-http.ts mounts /api/mcp unguarded), so forwarding
|
||||
// it would only copy a live credential into another service's logs. Pinned by
|
||||
// test.
|
||||
// The edge token is spent here and must never be forwarded: copying
|
||||
// Authorization would put a live credential into another service's logs. So
|
||||
// drop it unconditionally first, then — for the MCP route alone, and only
|
||||
// when a backend token is configured — replace it with that separate
|
||||
// protocol credential. Unset GITNEXUS_MCP_AUTH_TOKEN (the default) leaves
|
||||
// every request stripped, as before. The scope is the normalized pathname,
|
||||
// so a query string can't widen it and /api/mcpfoo doesn't qualify.
|
||||
delete headers.authorization;
|
||||
const upstreamPath = upstream.pathname;
|
||||
const isMcpRoute = upstreamPath === '/api/mcp' || upstreamPath.startsWith('/api/mcp/');
|
||||
if (isMcpRoute && mcpAuthToken) headers.authorization = `Bearer ${mcpAuthToken}`;
|
||||
headers.host = upstream.host;
|
||||
// Replace, never forward, the inbound chain (see clientAddressFor).
|
||||
const clientAddress = clientAddressFor(req);
|
||||
|
|
|
|||
|
|
@ -271,6 +271,12 @@ it('does not inject config into static assets', async () => {
|
|||
const TEST_AUTH_TOKEN = 'proxy-test-token-0123456789abcdefghij';
|
||||
const TEST_BEARER = `Bearer ${TEST_AUTH_TOKEN}`;
|
||||
|
||||
// The protocol token the upstream expects on /api/mcp. Deliberately unlike the
|
||||
// edge token, so "injected the backend credential" and "forwarded the edge one"
|
||||
// can never both satisfy an assertion.
|
||||
const TEST_MCP_TOKEN = 'backend-mcp-token-0123456789abcdefghij';
|
||||
const TEST_MCP_BEARER = `Bearer ${TEST_MCP_TOKEN}`;
|
||||
|
||||
// rawRequest never sends credentials; apiRequest does. In a file whose subject
|
||||
// is who gets let through, no test should pass because a helper quietly
|
||||
// authenticated for it.
|
||||
|
|
@ -376,6 +382,11 @@ async function withProxy(
|
|||
const proc = spawnServerWithEnv(dir, port, {
|
||||
GITNEXUS_UPSTREAM_URL: schemeless ? target : `http://${target}`,
|
||||
GITNEXUS_SERVE_AUTH_TOKEN: TEST_AUTH_TOKEN,
|
||||
// An ambient GITNEXUS_MCP_AUTH_TOKEN in the developer's shell would make the
|
||||
// proxy inject one on /api/mcp, so drop it: spawn omits undefined entries,
|
||||
// which unsets the inherited value. A test that wants injection sets it via
|
||||
// `env` below.
|
||||
GITNEXUS_MCP_AUTH_TOKEN: undefined,
|
||||
...env,
|
||||
});
|
||||
proc.stderr.setEncoding('utf8');
|
||||
|
|
@ -969,8 +980,9 @@ it('forwards an /api/* request that carries the correct token', async () => {
|
|||
});
|
||||
|
||||
it('strips the Authorization header instead of forwarding the edge token', async () => {
|
||||
// The token is spent at this hop. `serve` reads no Authorization header, so
|
||||
// forwarding would only copy a live credential into another service's logs.
|
||||
// The edge credential is spent and stripped at this hop. Forwarding it
|
||||
// would copy a live credential into another service's logs. With no
|
||||
// GITNEXUS_MCP_AUTH_TOKEN configured — the default — nothing replaces it.
|
||||
await withProxy({}, async (port, ctx) => {
|
||||
const res = await apiRequest(port, '/api/mcp', { method: 'POST', body: '{}' });
|
||||
assert.equal(res.status, 200, 'the request itself must still be proxied');
|
||||
|
|
@ -978,6 +990,72 @@ it('strips the Authorization header instead of forwarding the edge token', async
|
|||
});
|
||||
});
|
||||
|
||||
// -- Upstream MCP token injection (GITNEXUS_MCP_AUTH_TOKEN) -----------------
|
||||
//
|
||||
// A backend running protocol-layer MCP auth expects its own Bearer on
|
||||
// /api/mcp, and the edge credential can't serve as one. Both services are
|
||||
// configured with the same GITNEXUS_MCP_AUTH_TOKEN; this hop spends the edge
|
||||
// token and substitutes the backend one, for that route only.
|
||||
|
||||
// Stands in for a `serve` with MCP Bearer auth enabled: only the exact backend
|
||||
// credential gets through, so a passing two-hop request proves what was sent.
|
||||
const mcpBackend = (req, res) => {
|
||||
if (req.headers.authorization !== TEST_MCP_BEARER) {
|
||||
res.writeHead(401, { 'Content-Type': 'application/json; charset=utf-8' });
|
||||
res.end('{"error":"unauthorized"}');
|
||||
return;
|
||||
}
|
||||
res.writeHead(200, { 'Content-Type': 'application/json; charset=utf-8' });
|
||||
res.end('{"ok":true}');
|
||||
};
|
||||
|
||||
it('treats a blank GITNEXUS_MCP_AUTH_TOKEN as unset and still strips', async () => {
|
||||
const env = { GITNEXUS_MCP_AUTH_TOKEN: ' ' };
|
||||
await withProxy({ env }, async (port, ctx) => {
|
||||
const res = await apiRequest(port, '/api/mcp', { method: 'POST', body: '{}' });
|
||||
assert.equal(res.status, 200);
|
||||
assert.equal(ctx.received.headers.authorization, undefined);
|
||||
});
|
||||
});
|
||||
|
||||
it('replaces the edge credential with the upstream MCP token on /api/mcp', async () => {
|
||||
const env = { GITNEXUS_MCP_AUTH_TOKEN: TEST_MCP_TOKEN };
|
||||
await withProxy({ upstream: mcpBackend, env }, async (port, ctx) => {
|
||||
const res = await apiRequest(port, '/api/mcp', { method: 'POST', body: '{}' });
|
||||
assert.equal(res.status, 200, 'a backend that demands the MCP token must accept this hop');
|
||||
assert.equal(ctx.received.headers.authorization, TEST_MCP_BEARER);
|
||||
assert.notEqual(
|
||||
ctx.received.headers.authorization,
|
||||
TEST_BEARER,
|
||||
'the edge credential must never be forwarded',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
it('injects the upstream MCP token on /api/mcp subpaths and ignores the query string', async () => {
|
||||
const env = { GITNEXUS_MCP_AUTH_TOKEN: TEST_MCP_TOKEN };
|
||||
await withProxy({ upstream: mcpBackend, env }, async (port, ctx) => {
|
||||
for (const path of ['/api/mcp/messages', '/api/mcp?session=abc']) {
|
||||
const res = await apiRequest(port, path, { method: 'POST', body: '{}' });
|
||||
assert.equal(res.status, 200, `${path} must reach the MCP backend authenticated`);
|
||||
assert.equal(ctx.received.headers.authorization, TEST_MCP_BEARER, path);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
it('leaves non-MCP routes stripped when an upstream MCP token is configured', async () => {
|
||||
// /api/mcpfoo shares a prefix with the MCP route but is not it, and a plain
|
||||
// API route never carries a protocol credential.
|
||||
const env = { GITNEXUS_MCP_AUTH_TOKEN: TEST_MCP_TOKEN };
|
||||
await withProxy({ env }, async (port, ctx) => {
|
||||
for (const path of ['/api/mcpfoo', '/api/health']) {
|
||||
const res = await apiRequest(port, path);
|
||||
assert.equal(res.status, 200);
|
||||
assert.equal(ctx.received.headers.authorization, undefined, path);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
it('never gates static assets behind the token', async () => {
|
||||
// The UI has to load before it can prompt for a token.
|
||||
await withProxy({}, async (port, ctx) => {
|
||||
|
|
|
|||
312
docs/plans/2026-08-28-gitnexus-plan-impact-file-risk.md
Normal file
312
docs/plans/2026-08-28-gitnexus-plan-impact-file-risk.md
Normal file
|
|
@ -0,0 +1,312 @@
|
|||
# GitNexus Engineering Plan
|
||||
|
||||
> Task: Fix #3075 — File `impact` risk is not comparable to Function/Method risk.
|
||||
> Evidence verified at commit `6bff33d14cbfe1e7b4f04bca51507e9f64ef579c` (`feat/kotlin-const-resolver`); GitNexus index 129 commits behind, refresh skipped: full-repo `--index-only --pdg` rebuild is impractical this session. Scorer and schema claims are `[verified]` from source; live inversion numbers are `[graph]` on the stale index.
|
||||
|
||||
## 1. Objective
|
||||
|
||||
Make File vs symbol `impact.risk` honest for consumers: either they can tell the scales differ, or they can compare on a shared two-axis score. Do **not** DEFINES-bridge processes/modules onto File targets (issue reporter sampled 8/10 one-importer files jumping to HIGH/CRITICAL). Do **not** retune Function HIGH/CRITICAL thresholds (agent warn-before-edit).
|
||||
|
||||
Acceptance:
|
||||
|
||||
- A File with a wider blast radius than a Function in the same file no longer looks “safer” when a consumer only reads `risk`, **or** the result states that `risk` is not comparable across kinds and offers `riskSharedAxes` for comparison.
|
||||
- File targets still cannot trip HIGH/CRITICAL via `processes_affected` / `modules_affected` unless those axes become real in the index (they are not today).
|
||||
- Existing Function/Method labels under the current four-axis ladder stay the same for the same inputs.
|
||||
- MCP `riskNote` remains UNKNOWN-only (`tools.ts` contract).
|
||||
|
||||
## 2. Current Behaviour
|
||||
|
||||
Callgraph `impact` ends in `LocalBackend._runImpactBFS` (`gitnexus/src/mcp/local/local-backend.ts`). After BFS it enriches impacted ids with `STEP_IN_PROCESS` and `MEMBER_OF`, then scores:
|
||||
|
||||
```7720:7738:gitnexus/src/mcp/local/local-backend.ts
|
||||
} else if (
|
||||
directCount >= 30 ||
|
||||
processCount >= 5 ||
|
||||
moduleCount >= 5 ||
|
||||
impacted.length >= 200
|
||||
) {
|
||||
risk = 'CRITICAL';
|
||||
} else if (
|
||||
directCount >= 15 ||
|
||||
processCount >= 3 ||
|
||||
moduleCount >= 3 ||
|
||||
impacted.length >= 100
|
||||
) {
|
||||
risk = 'HIGH';
|
||||
} else if (directCount >= 5 || impacted.length >= 30) {
|
||||
risk = 'MEDIUM';
|
||||
} else {
|
||||
risk = 'LOW';
|
||||
}
|
||||
```
|
||||
|
||||
Empty upstream → `UNKNOWN` + `riskNote`. Downstream empty stays LOW. `skipEnrichment` (ambiguous probes) already scores on direct+total only. PDG mode forces `risk: UNKNOWN` (`composeUnifiedPdgImpactResult`) — out of scope.
|
||||
|
||||
File BFS walk is mostly File←IMPORTS File. Enrichment queries those File ids. Processes are CALLS traces (`process-processor.ts`); communities admit only Function/Class/Method/Interface (`isCommunitySymbol` in `community-processor.ts:412-416`). File is not in that set. `enrichCandidateLabels` UNION also **omits File**, so File `target.type` is often `""`; detect File via `id` prefix `File:`.
|
||||
|
||||
Web Graph RAG (`gitnexus-web/src/core/llm/tools.ts` ~1331–1346) duplicates the same ladder.
|
||||
|
||||
## 3. Relevant Architecture
|
||||
|
||||
| Layer | Role |
|
||||
|---|---|
|
||||
| Index | File never sources `STEP_IN_PROCESS` / `MEMBER_OF` by construction |
|
||||
| MCP `_runImpactBFS` | Blast radius + four-axis `risk` |
|
||||
| Ambiguous probes | `skipEnrichment` → 2-axis `risk` already |
|
||||
| `mergeRisk` | Group overlay; monotone in crossings; does not know target kind |
|
||||
| CLI `formatImpactResult` | Prints counts; **does not print `risk`** on the resolved callgraph path; JSON `impactCommand` still ships `risk` |
|
||||
| `ai-context.ts` / `tools.ts` | Agent contract: warn on HIGH/CRITICAL; `riskNote` UNKNOWN-only |
|
||||
| Web LLM `impact` | Same formula, prose `RISK:` line |
|
||||
|
||||
Modules: Local (MCP), Cli (format/docs), Group (`mergeRisk`), gitnexus-web LLM tools. Shared package `gitnexus-shared` is already a dependency of both CLI and web.
|
||||
|
||||
## 4. GitNexus Findings
|
||||
|
||||
- Primary: `_runImpactBFS` — d=1 `[graph]` `impact(target:_runImpactBFS, maxDepth:1, includeTests:true)`: `_impactImpl`, `impactByUid`. Production chain `[verified]`: `impact` → `_impactImpl` → `_runImpactBFS`; `impactByUid` skips per-symbol process lists but **not** aggregation (`skipPerSymbolEnrichment` only).
|
||||
- `LocalBackend.impact` d=1 `[graph]` `context`: `callTool`.
|
||||
- Duplicate scorer `[verified]` grep: `gitnexus-web/src/core/llm/tools.ts`.
|
||||
- `mergeRisk` `[verified]` callers in `src/`: only `runGroupImpact` (`cross-impact.ts:907`). Graph d=1 listed a test File (`impact-pdg-shape.test.ts`) and missed `runGroupImpact` — trust source.
|
||||
- Schema `[verified]`: `isCommunitySymbol` excludes File; `schema.ts` documents MEMBER_OF as Function/Class/Method/Interface only.
|
||||
- Live inversion `[graph]` stale index, `impact summaryOnly` on GitNexus:
|
||||
|
||||
| target | kind | impacted | direct | processes | modules | risk |
|
||||
|---|---|---|---|---|---|---|
|
||||
| `lbug-config.ts` | File | 54 | 12 | 0 | 0 | MEDIUM |
|
||||
| `openLbugConnection` | Function | 16 | 9 | 3 | 2 | HIGH |
|
||||
| `local-backend.ts` | File | 12 | 10 | 0 | 0 | MEDIUM |
|
||||
| `refreshRepos` | Method | 50 | 5 | 4 | 7 | CRITICAL |
|
||||
|
||||
- Clusters/processes resources `[graph]`: Local/Cli/Group sit in the impact path; process traces are function-stepped, not File-stepped.
|
||||
- Related tests `[verified]`: `test/unit/impact-pagination.test.ts` (CRITICAL from `direct=400`); `test/integration/impact-zero-caller-risk.test.ts` (`withTestLbugDB` seed — pattern to extend); `test/unit/eval-formatters.test.ts` (`formatImpactResult`); group `mergeRisk` tests.
|
||||
|
||||
## 5. Statement-Level PDG Findings
|
||||
|
||||
PDG unavailable (`pdg_query` on `_runImpactBFS`: “no PDG layer”). Recommend `node .gitnexus/run.cjs analyze --index-only --pdg` before any future statement-slice work. Control flow of the scorer is a straight if/else after enrichment; no hidden guards. `skipEnrichment` is the only branch that structurally zeros process/module counts besides File ids.
|
||||
|
||||
## 6. Proposed Changes
|
||||
|
||||
### 6.1 Extract `scoreImpactRisk` — `gitnexus-shared/src/impact-risk.ts` (new)
|
||||
|
||||
- **Responsibility:** Pure function: `{ direction, directCount, processCount, moduleCount, impactedCount, unusedAxes }` → `{ risk, riskSharedAxes, riskScale }`.
|
||||
- **Behaviour:** Existing UNKNOWN/CRITICAL/HIGH/MEDIUM/LOW thresholds unchanged when `unusedAxes` is empty. `riskSharedAxes` always scores as if `processCount=0` and `moduleCount=0` (UNKNOWN rule still applies). `riskScale.comparableAcrossKinds` is false iff `unusedAxes` is non-empty. `riskScale.unusedAxes` lists `{ axis, reason }`.
|
||||
- **Constraints:** Zero deps. Export from `gitnexus-shared/src/index.ts`. Do not put MCP types here.
|
||||
- **File detection:** caller passes unused axes; helper does not parse UIDs.
|
||||
|
||||
### 6.2 Wire MCP — `_runImpactBFS` in `local-backend.ts`
|
||||
|
||||
- After computing `processCount`/`moduleCount`, set `unusedAxes`:
|
||||
- target `id` starts with `File:` **or** `symType === 'File'` → processes + modules, reason `file-nodes-have-no-process-or-community-membership`;
|
||||
- `skipEnrichment` → same axes, reason `enrichment-skipped` (ambiguous probes).
|
||||
- Replace inline ladder with `scoreImpactRisk`.
|
||||
- Spread `riskScale` and `riskSharedAxes` on the result next to `risk`. Do **not** set `riskNote` for File.
|
||||
- Ambiguous candidate summaries: forward the new fields (probes already skip enrichment).
|
||||
- `target.type` for File: if still `""`, prefer `'File'` when `id` starts with `File:` (display-only; helps CLI).
|
||||
|
||||
### 6.3 Web duplicate — `gitnexus-web/src/core/llm/tools.ts`
|
||||
|
||||
- Import `scoreImpactRisk` from `gitnexus-shared`. Print `RISK:` from `risk`; if `!comparableAcrossKinds`, one extra line: not comparable to Function risk; shared-axes label is `riskSharedAxes`.
|
||||
|
||||
### 6.4 Agent/MCP contract copy
|
||||
|
||||
- `gitnexus/src/mcp/tools.ts` impact description: document `riskScale` / `riskSharedAxes`; keep `riskNote` UNKNOWN-only; say File `risk` is not comparable to symbol `risk`.
|
||||
- `gitnexus/src/cli/ai-context.ts`: HIGH/CRITICAL warning still applies; add: do not rank a File `MEDIUM` below a contained Function `HIGH` without `riskSharedAxes`.
|
||||
- `formatImpactResult`: on resolved callgraph results with `risk`, print `Risk: {risk}` and, when incomparable, `Shared-axes risk: {riskSharedAxes} (File/process axes unused)`.
|
||||
|
||||
### 6.5 Explicitly not changing
|
||||
|
||||
- DEFINES-bridge, community/process indexers, `mergeRisk` formula, PDG `UNKNOWN`, `detectChanges` `risk_level`, Function thresholds.
|
||||
|
||||
## 7. Implementation Sequence
|
||||
|
||||
1. Add `gitnexus-shared` helper + unit table (issue-shaped inputs + UNKNOWN + skipEnrichment). Shared package tests if present; otherwise `gitnexus/test/unit/impact-risk.test.ts` importing the helper.
|
||||
2. Switch `_runImpactBFS` + candidate probe payload. Tree still coherent: old `risk` values identical for Function fixtures.
|
||||
3. Integration seed in `impact-zero-caller-risk.test.ts` **or** new `impact-file-risk-scale.test.ts`: File with ≥5 File IMPORTS (MEDIUM on direct) vs Function with 3 process-member callers (HIGH); assert File `riskScale.comparableAcrossKinds === false`, Function true, File `riskSharedAxes === risk`, Function `riskSharedAxes` is LOW/MEDIUM while `risk` is HIGH.
|
||||
4. CLI formatter + `eval-formatters.test.ts`.
|
||||
5. `tools.ts` + `ai-context.ts` wording.
|
||||
6. Web import + a unit assertion on the printed RISK block if a test already covers that tool.
|
||||
7. `npx tsc --noEmit` in `gitnexus/` and `gitnexus-web/`; `cd gitnexus && npm run test:unit -- test/unit/impact-risk.test.ts test/unit/eval-formatters.test.ts`; integration file from step 3.
|
||||
|
||||
## 8. Test Strategy
|
||||
|
||||
| File | Scenarios |
|
||||
|---|---|
|
||||
| `gitnexus/test/unit/impact-risk.test.ts` (new) | Issue table: File(25,13,0,0)→MEDIUM; Function(15,2,4,2)→HIGH; shared-axes File MEDIUM vs Function LOW; empty upstream UNKNOWN; downstream empty LOW; skipEnrichment unused axes; CRITICAL via direct≥30 still works with unused process axes |
|
||||
| `gitnexus/test/integration/impact-file-risk-scale.test.ts` (new) | `withTestLbugDB` seed: `File:src/crypto.ts` ← 13 File IMPORTS, no File STEP_IN_PROCESS; `getEncryptionKey` with 2 CALLS from functions that have STEP_IN_PROCESS to 4 distinct Process nodes — reproduce inversion; assert new fields |
|
||||
| `gitnexus/test/integration/impact-zero-caller-risk.test.ts` | Unchanged UNKNOWN/`riskNote`; candidates may grow `riskScale` — assert still present only when UNKNOWN for `riskNote` |
|
||||
| `gitnexus/test/unit/impact-pagination.test.ts` | Hub CRITICAL unchanged |
|
||||
| `gitnexus/test/unit/eval-formatters.test.ts` | Resolved result prints Risk + shared-axes line for File-shaped `riskScale` |
|
||||
| Web | Only if an existing Graph RAG impact test snapshots `RISK:` |
|
||||
|
||||
Commands (exist in `gitnexus/package.json`): `npm run test:unit`, `npm test` (full vitest), `npx tsc --noEmit`. Web: `npm test`, `npx tsc -b --noEmit`. Integration needs `pretest:integration` / `npm run test:integration` (runs `scripts/build.js`).
|
||||
|
||||
## 9. Risk and Impact Analysis
|
||||
|
||||
Direct dependents of `_runImpactBFS` `[graph]`: `_impactImpl`, `impactByUid`. `_impactImpl` is the only d=1 of `impact` besides the method’s own class. Any JSON consumer of `impact` (MCP, CLI `output(result)`, group local leg) sees additive fields — compatible if they ignore unknowns.
|
||||
|
||||
- **HIGH workflow:** Function HIGH/CRITICAL unchanged. File still cannot reach HIGH via processes; a File with `direct≥15` or `total≥100` still can. Agents that compare File MEDIUM vs Function HIGH must start using `riskSharedAxes` or `riskScale`.
|
||||
- **Ambiguous `maxRisk`:** probes skip enrichment, so File vs Function candidates are already 2-axis there — inversion is weaker on that path.
|
||||
- **Group `mergeRisk`:** still compares incomparable File local `risk` to crossing count. Do not retune this PR; if a group File target is common, follow-up.
|
||||
- **Web:** browser bundle picks up `gitnexus-shared` export — confirm `gitnexus-shared` build/exports include the new file.
|
||||
- **Performance:** none (pure arithmetic after existing enrichment).
|
||||
- **Ladybug empty labels:** File detection must not rely on `symType` alone.
|
||||
|
||||
## 10. Files Expected to Change
|
||||
|
||||
| File | Symbols | Reason |
|
||||
|---|---|---|
|
||||
| `gitnexus-shared/src/impact-risk.ts` | `scoreImpactRisk` | New shared scorer |
|
||||
| `gitnexus-shared/src/index.ts` | exports | Public helper |
|
||||
| `gitnexus/src/mcp/local/local-backend.ts` | `_runImpactBFS`, ambiguous candidate map | Wire scorer + File unused axes |
|
||||
| `gitnexus/src/mcp/tools.ts` | `impact` description | Contract |
|
||||
| `gitnexus/src/cli/ai-context.ts` | generated Always Do | Agent warning |
|
||||
| `gitnexus/src/cli/eval-server.ts` | `formatImpactResult` | Print scale |
|
||||
| `gitnexus-web/src/core/llm/tools.ts` | web `impact` | Same formula |
|
||||
| `gitnexus/test/unit/impact-risk.test.ts` | — | Table tests |
|
||||
| `gitnexus/test/integration/impact-file-risk-scale.test.ts` | — | Seeded inversion |
|
||||
| `gitnexus/test/unit/eval-formatters.test.ts` | `formatImpactResult` | Formatter |
|
||||
|
||||
## 11. Reusable Implementation Context
|
||||
|
||||
```yaml
|
||||
implementation_context:
|
||||
task_summary: "Fix #3075: File impact.risk is a 2-axis score silently labelled on a 4-axis scale. Extract scoreImpactRisk; mark File/skipEnrichment axes unused; add riskScale + riskSharedAxes; do not DEFINES-bridge or retune Function thresholds."
|
||||
acceptance_criteria:
|
||||
- "File vs Function comparison is either labelled incomparable (riskScale) or done via riskSharedAxes"
|
||||
- "Function/Method risk for identical four-axis inputs unchanged"
|
||||
- "riskNote still UNKNOWN-only"
|
||||
- "Integration seed reproduces crypto.ts-style inversion and asserts the new fields"
|
||||
primary_symbols:
|
||||
- symbol: "_runImpactBFS"
|
||||
file: "gitnexus/src/mcp/local/local-backend.ts"
|
||||
lines: "6991-7888"
|
||||
role: "BFS + enrichment + inline risk ladder (replace ladder only)"
|
||||
- symbol: "scoreImpactRisk"
|
||||
file: "gitnexus-shared/src/impact-risk.ts"
|
||||
lines: "new"
|
||||
role: "Pure scorer + shared-axes + riskScale"
|
||||
- symbol: "formatImpactResult"
|
||||
file: "gitnexus/src/cli/eval-server.ts"
|
||||
lines: "305-641"
|
||||
role: "Human/LLM text surface for impact JSON"
|
||||
related_symbols:
|
||||
- symbol: "_impactImpl"
|
||||
relationship: "CALLS"
|
||||
relevance: "Resolves target, PDG vs callgraph, ambiguous skipEnrichment probes"
|
||||
- symbol: "impactByUid"
|
||||
relationship: "CALLS"
|
||||
relevance: "Group fan-out; keep skipPerSymbolEnrichment; still run aggregation"
|
||||
- symbol: "mergeRisk"
|
||||
relationship: "consumes risk string"
|
||||
relevance: "Do not change this PR"
|
||||
- symbol: "isCommunitySymbol"
|
||||
relationship: "index gate"
|
||||
relevance: "Why File modules_affected is always 0"
|
||||
- symbol: "composeUnifiedPdgImpactResult"
|
||||
relationship: "separate path"
|
||||
relevance: "PDG risk stays UNKNOWN"
|
||||
execution_path:
|
||||
- "impact / callTool → _impactImpl (resolve symbol, File id prefix File:)"
|
||||
- "_runImpactBFS: IMPORTS-heavy walk for File; CALLS walk for Function"
|
||||
- "Enrich STEP_IN_PROCESS / MEMBER_OF on impacted ids (empty for File ids)"
|
||||
- "scoreImpactRisk with unusedAxes for File or skipEnrichment"
|
||||
- "JSON to MCP/CLI; formatImpactResult for eval text; web LLM tools parallel path"
|
||||
pdg_constraints:
|
||||
- description: "No PDG layer on the planning index; scorer is post-enrichment arithmetic"
|
||||
affected_statements: []
|
||||
implementation_consequence: "Do not wait on PDG; do not change pdg impact risk"
|
||||
architectural_patterns:
|
||||
- pattern: "Additive optional JSON fields on impact (riskNote, epistemic, partial)"
|
||||
example_location: "gitnexus/src/mcp/local/local-backend.ts _runImpactBFS base object ~7754"
|
||||
usage_guidance: "Add riskScale/riskSharedAxes the same way; never overload riskNote"
|
||||
- pattern: "withTestLbugDB CREATE seed for impact contract"
|
||||
example_location: "gitnexus/test/integration/impact-zero-caller-risk.test.ts"
|
||||
usage_guidance: "Seed File IMPORTS + Function CALLS + Process membership separately"
|
||||
files_to_modify:
|
||||
- file: "gitnexus-shared/src/impact-risk.ts"
|
||||
symbols: ["scoreImpactRisk"]
|
||||
intended_change: "new pure scorer"
|
||||
- file: "gitnexus-shared/src/index.ts"
|
||||
symbols: []
|
||||
intended_change: "re-export"
|
||||
- file: "gitnexus/src/mcp/local/local-backend.ts"
|
||||
symbols: ["_runImpactBFS"]
|
||||
intended_change: "unusedAxes + helper; File type display"
|
||||
- file: "gitnexus/src/mcp/tools.ts"
|
||||
symbols: []
|
||||
intended_change: "document fields"
|
||||
- file: "gitnexus/src/cli/ai-context.ts"
|
||||
symbols: []
|
||||
intended_change: "agent comparability note"
|
||||
- file: "gitnexus/src/cli/eval-server.ts"
|
||||
symbols: ["formatImpactResult"]
|
||||
intended_change: "print risk + shared-axes when incomparable"
|
||||
- file: "gitnexus-web/src/core/llm/tools.ts"
|
||||
symbols: []
|
||||
intended_change: "import helper; extra prose line"
|
||||
tests:
|
||||
- file: "gitnexus/test/unit/impact-risk.test.ts"
|
||||
scenarios:
|
||||
- "File(25,13,0,0)+unused process/module → risk MEDIUM, comparableAcrossKinds false, riskSharedAxes MEDIUM"
|
||||
- "Function(15,2,4,2) → HIGH, riskSharedAxes LOW (direct 2, total 15)"
|
||||
- "upstream impactedCount 0 → UNKNOWN both fields"
|
||||
- "direct 400 → CRITICAL even with unused process axes"
|
||||
- file: "gitnexus/test/integration/impact-file-risk-scale.test.ts"
|
||||
scenarios:
|
||||
- "Seed File crypto.ts with 13 File importers vs getEncryptionKey with process-rich callers → inversion on risk, File incomparable, Function comparable"
|
||||
- file: "gitnexus/test/unit/eval-formatters.test.ts"
|
||||
scenarios:
|
||||
- "formatImpactResult includes Shared-axes risk when riskScale.comparableAcrossKinds is false"
|
||||
verification_commands:
|
||||
- "cd gitnexus && npx tsc --noEmit"
|
||||
- "cd gitnexus && npm run test:unit -- test/unit/impact-risk.test.ts test/unit/eval-formatters.test.ts test/unit/impact-pagination.test.ts"
|
||||
- "cd gitnexus && npm run test:integration -- test/integration/impact-file-risk-scale.test.ts test/integration/impact-zero-caller-risk.test.ts"
|
||||
- "cd gitnexus-web && npx tsc -b --noEmit"
|
||||
risks:
|
||||
- "Consumers that only read risk still see the inversion unless they adopt riskScale/riskSharedAxes — that is the chosen (explicit-scale) fix"
|
||||
- "File type often empty; must key unusedAxes off File: id prefix"
|
||||
- "gitnexus-shared export must reach the web bundle"
|
||||
assumptions:
|
||||
- "WHAT: File nodes never gain STEP_IN_PROCESS/MEMBER_OF without an indexer change. HOW: keep isCommunitySymbol and process traces as-is; tests seed File with zero such edges"
|
||||
- "WHAT: Additive JSON fields are backward compatible. HOW: existing tests that exact-match the full impact object may need to allow extra keys — grep expect(res).toEqual on impact results before landing"
|
||||
- "WHAT: HEAD 6bff33d is the pin; scorer line numbers ~7720. HOW: re-read the ladder if that hunk moved"
|
||||
open_questions:
|
||||
- "Whether GroupImpactResult should copy riskScale from local File targets (deferred unless tests already snapshot the full group object)"
|
||||
avoid:
|
||||
- "Do not DEFINES-bridge File→symbol processes/modules"
|
||||
- "Do not lower Function process/module HIGH/CRITICAL thresholds"
|
||||
- "Do not reuse riskNote for File incomparability"
|
||||
- "Do not change PDG impact risk or detectChanges risk_level"
|
||||
- "Do not treat labels(n)[0] or empty target.type as proof the node is not a File"
|
||||
- "Do not repeat full repository discovery"
|
||||
```
|
||||
|
||||
## 12. Assumptions and Open Questions
|
||||
|
||||
**Assumptions**
|
||||
|
||||
- Indexer will not start attaching File→Process/Community in this change (`isCommunitySymbol` stays). `[verified]` source; `[assumed]` future indexers.
|
||||
- Ignoring unknown JSON keys is safe for MCP clients; any `toEqual` goldens in-repo must be updated. `[assumed]` — grep during implement.
|
||||
- Stale-index inversion (`lbug-config.ts` vs `openLbugConnection`) is illustrative; the integration seed is the regression lock. `[graph]` vs `[verified]` seed.
|
||||
|
||||
**Open questions**
|
||||
|
||||
- Group `mergeRisk` + File local risk: copy `riskScale` onto `GroupImpactResult`? Default **no** unless a test breaks.
|
||||
- Class/Interface STEP_IN_PROCESS sparsity: out of scope (#3075 is File).
|
||||
- Printing `risk` on CLI formatted output is new (JSON already has it). Keep the extra lines short.
|
||||
|
||||
**Deferred**
|
||||
|
||||
- Recalibrated File-only HIGH thresholds.
|
||||
- Indexing File community membership.
|
||||
- DEFINES-bridge after a threshold RFC.
|
||||
- Related #2975 (docs vs scorer wording) except as touched by `tools.ts`.
|
||||
|
||||
## 13. Definition of Done
|
||||
|
||||
- [ ] `scoreImpactRisk` is the only callgraph ladder in MCP and web.
|
||||
- [ ] File (and skipEnrichment) results include `riskScale.comparableAcrossKinds === false` and `riskSharedAxes`.
|
||||
- [ ] Function four-axis HIGH/CRITICAL cases in unit tests still pass with the same labels.
|
||||
- [ ] Integration seed proves wider File blast + lower `risk` than a contained Function, and `riskSharedAxes` orders them without pretending processes existed on the File.
|
||||
- [ ] `riskNote` still absent unless `risk === 'UNKNOWN'`.
|
||||
- [ ] `tools.ts` + `ai-context.ts` state that File `risk` is not comparable to symbol `risk`.
|
||||
- [ ] `cd gitnexus && npx tsc --noEmit` and the named unit/integration commands pass; web typecheck passes.
|
||||
|
|
@ -19,14 +19,14 @@
|
|||
*
|
||||
* False-positive suppression:
|
||||
* - Skips calls whose receiver is a known non-tree-sitter library (`JSON`,
|
||||
* `URL`, `marked`, `Number`).
|
||||
* `URL`, `marked`, `Number`, `path`).
|
||||
* - Skips calls whose first argument is a string-literal (grammar-load smoke
|
||||
* tests like `_testParser.parse('service X { rpc Y (R) returns (R); }')`).
|
||||
* - Skips test files (`.test.ts`/`.test.tsx`/`.spec.ts`).
|
||||
* - Skips the `safe-parse.ts` helper itself.
|
||||
*/
|
||||
|
||||
const SKIPPED_RECEIVERS = new Set(['JSON', 'URL', 'marked', 'Number', 'Math']);
|
||||
const SKIPPED_RECEIVERS = new Set(['JSON', 'URL', 'marked', 'Number', 'Math', 'path']);
|
||||
|
||||
export default {
|
||||
meta: {
|
||||
|
|
@ -74,7 +74,7 @@ export default {
|
|||
// Receiver-text-shape skip: anything matching well-known JS APIs that
|
||||
// happen to have a `.parse(<expr>)` shape but aren't tree-sitter.
|
||||
if (
|
||||
/^(JSON|URL|marked|Number|Math|Date|globalThis\.JSON)\b/.test(receiverText) ||
|
||||
/^(JSON|URL|marked|Number|Math|Date|path|globalThis\.JSON)\b/.test(receiverText) ||
|
||||
/\bjson\.parse\b/i.test(receiverText)
|
||||
) {
|
||||
return;
|
||||
|
|
|
|||
|
|
@ -17,7 +17,6 @@ Usage:
|
|||
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
|
|
|||
|
|
@ -14,7 +14,6 @@ import subprocess
|
|||
import sys
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from constants import (
|
||||
|
|
|
|||
|
|
@ -6,7 +6,10 @@ readme = "README.md"
|
|||
requires-python = ">=3.11"
|
||||
dependencies = [
|
||||
"mini-swe-agent>=2.0.0",
|
||||
"litellm!=1.82.7,!=1.82.8,>=1.83.7",
|
||||
"litellm[proxy]!=1.82.7,!=1.82.8,>=1.99.0",
|
||||
"cryptography>=50.0.0",
|
||||
"python-multipart>=0.0.30",
|
||||
"restrictedpython>=8.3",
|
||||
"datasets>=3.0.0",
|
||||
"typer>=0.12.0",
|
||||
"rich>=13.0.0",
|
||||
|
|
|
|||
|
|
@ -27,12 +27,10 @@ import threading
|
|||
import time
|
||||
from itertools import product
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import typer
|
||||
import yaml
|
||||
from rich.console import Console
|
||||
from rich.live import Live
|
||||
from rich.table import Table
|
||||
|
||||
from utils.errors import is_debug_enabled, log_safe_exception
|
||||
|
|
|
|||
|
|
@ -6,15 +6,18 @@ import os
|
|||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
from contextlib import contextmanager
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from workflow_bench import evolve
|
||||
from workflow_bench import evolve, evolution
|
||||
from workflow_bench.runner_sessions import PARENT_EVENT_STREAM_SOURCE
|
||||
from workflow_bench.evolve import (
|
||||
build_parser,
|
||||
build_proposer_prompt,
|
||||
executed_benchmark_arms,
|
||||
generation_timeout_seconds,
|
||||
load_jsonl,
|
||||
proposer_evidence_entries,
|
||||
|
|
@ -25,10 +28,39 @@ from workflow_bench.evolve import (
|
|||
summarize_gate,
|
||||
validate_promotion_for_apply,
|
||||
)
|
||||
from workflow_bench.process_control import run_managed
|
||||
from workflow_bench.process_control import ManagedProcessResult, run_managed
|
||||
from workflow_bench.proposer_sandbox import pid_namespace_command, preflight_bubblewrap
|
||||
|
||||
|
||||
def test_runner_environment_does_not_forward_the_openai_key() -> None:
|
||||
from workflow_bench.model_gateway import credential_secrets
|
||||
|
||||
args = build_parser().parse_args(
|
||||
[
|
||||
"--tasks",
|
||||
"t.yaml",
|
||||
"--model",
|
||||
"gpt-4.1",
|
||||
"--anthropic-api-key",
|
||||
"loopback-master",
|
||||
"--openai-api-key",
|
||||
"sk-openai-secret",
|
||||
]
|
||||
)
|
||||
env = evolve.runner_environment(args)
|
||||
assert env["GITNEXUS_BENCH_ANTHROPIC_API_KEY"] == "loopback-master"
|
||||
assert "GITNEXUS_BENCH_AUTH_TOKEN" not in env
|
||||
assert "OPENAI_API_KEY" not in env
|
||||
assert "GITNEXUS_BENCH_OPENAI_API_KEY" not in env
|
||||
assert "sk-openai-secret" not in env.values()
|
||||
assert credential_secrets(args) == ["loopback-master", "sk-openai-secret"]
|
||||
|
||||
|
||||
def test_parser_keeps_the_legacy_auth_token_alias() -> None:
|
||||
args = build_parser().parse_args(["--tasks", "t.yaml", "--model", "pinned", "--auth-token", "alias-secret"])
|
||||
assert args.auth_token == "alias-secret"
|
||||
|
||||
|
||||
def row(**overrides):
|
||||
base = {
|
||||
"task": "demo-task",
|
||||
|
|
@ -101,9 +133,9 @@ def test_read_learnings_keeps_the_most_recent_entries(tmp_path):
|
|||
]
|
||||
path.write_text("\n".join(json.dumps(row) for row in rows) + "\n")
|
||||
assert read_learnings(path, cap=3) == [
|
||||
{"skill": "gitnexus-work", "n": 7},
|
||||
{"skill": "gitnexus-work", "n": 8},
|
||||
{"skill": "gitnexus-work", "n": 9},
|
||||
{"skill": "gitnexus-review", "n": 10},
|
||||
]
|
||||
|
||||
|
||||
|
|
@ -137,12 +169,30 @@ def test_build_proposer_prompt_carries_evidence_constraints_and_paths(tmp_path):
|
|||
assert "node .gitnexus/run.cjs analyze" in prompt
|
||||
assert "1 row(s) in /evidence/learnings.json" in prompt
|
||||
assert "1 selected row(s) in /evidence/selected-rows.json" in prompt
|
||||
assert "exact staged" in prompt
|
||||
assert "no full results.jsonl" in prompt
|
||||
assert "1 decision(s) in /evidence/gate-summary.json" in prompt
|
||||
assert "budget blown on reruns" not in prompt
|
||||
assert "verify-failed" not in prompt
|
||||
assert "~/.claude/projects" not in prompt
|
||||
|
||||
|
||||
def test_build_proposer_prompt_points_at_the_rejected_prior_proposal(tmp_path):
|
||||
common = {
|
||||
"results_dir": tmp_path / "bench",
|
||||
"evidence": [],
|
||||
"learnings": [],
|
||||
"gate_summary": ["candidate_workflow: keep_incumbent — cost regressed"],
|
||||
"overlay_dir": tmp_path / "overlay",
|
||||
"proposal_path": tmp_path / "proposal.md",
|
||||
"incumbent_arms": ["workflow"],
|
||||
}
|
||||
# The gate summary alone says a candidate lost, never what it proposed —
|
||||
# so without this line the proposer can re-propose the same prose forever.
|
||||
assert "/evidence/prior-proposal.md" in build_proposer_prompt(**common, prior_proposal=True)
|
||||
assert "/evidence/prior-proposal.md" not in build_proposer_prompt(**common)
|
||||
|
||||
|
||||
def test_build_proposer_prompt_first_generation_has_no_results_dir(tmp_path):
|
||||
prompt = build_proposer_prompt(
|
||||
results_dir=None,
|
||||
|
|
@ -166,6 +216,8 @@ def test_proposer_reads_only_digest_bound_transcripts_below_results(tmp_path, mo
|
|||
artifact = transcripts / "task-workflow-run0-session.jsonl"
|
||||
artifact.write_bytes(payload)
|
||||
artifact.chmod(0o600)
|
||||
patch = results / "demo-task-workflow-run0.patch"
|
||||
patch.write_text("diff --git a/a b/a\n")
|
||||
metadata = {
|
||||
"path": "transcripts/task-workflow-run0-session.jsonl",
|
||||
"sha256": hashlib.sha256(payload).hexdigest(),
|
||||
|
|
@ -185,7 +237,11 @@ def test_proposer_reads_only_digest_bound_transcripts_below_results(tmp_path, mo
|
|||
gate_summary=[],
|
||||
)
|
||||
|
||||
assert entries["transcript-0-0.jsonl"] == payload.decode()
|
||||
assert [json.loads(line) for line in entries["transcript-0-0.jsonl"].splitlines()] == [json.loads(payload)]
|
||||
assert entries["patch-0.diff"] == patch.read_text()
|
||||
staged_rows = entries["selected-rows.json"]
|
||||
assert staged_rows[0]["patch_file"] == "patch-0.diff"
|
||||
assert staged_rows[0]["transcript_files"] == ["transcript-0-0.jsonl"]
|
||||
assert "foreign host transcript" not in json.dumps(entries)
|
||||
|
||||
bad_digest = {**metadata, "sha256": "0" * 64}
|
||||
|
|
@ -198,6 +254,61 @@ def test_proposer_reads_only_digest_bound_transcripts_below_results(tmp_path, mo
|
|||
)
|
||||
|
||||
|
||||
def test_proposer_compacts_transcripts_as_complete_json_events(tmp_path):
|
||||
results = tmp_path / "results"
|
||||
transcripts = results / "transcripts"
|
||||
transcripts.mkdir(parents=True, mode=0o700)
|
||||
events = [
|
||||
{
|
||||
"type": "assistant",
|
||||
"message": {
|
||||
"content": [
|
||||
{
|
||||
"type": "thinking",
|
||||
"thinking": "analysis-" + ("x" * 100_000),
|
||||
"signature": "opaque-base64-signature",
|
||||
}
|
||||
]
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "result",
|
||||
"session_id": "session-1",
|
||||
"usage": {"input_tokens": 1, "output_tokens": 2},
|
||||
},
|
||||
]
|
||||
payload = "".join(json.dumps(event) + "\n" for event in events).encode()
|
||||
artifact = transcripts / "session.jsonl"
|
||||
artifact.write_bytes(payload)
|
||||
artifact.chmod(0o600)
|
||||
|
||||
entries = proposer_evidence_entries(
|
||||
results_dir=results,
|
||||
evidence=[
|
||||
row(
|
||||
transcript_artifacts=[
|
||||
{
|
||||
"path": "transcripts/session.jsonl",
|
||||
"sha256": hashlib.sha256(payload).hexdigest(),
|
||||
"bytes": len(payload),
|
||||
"source": PARENT_EVENT_STREAM_SOURCE,
|
||||
}
|
||||
]
|
||||
)
|
||||
],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
artifact_limit=8192,
|
||||
)
|
||||
|
||||
staged = entries["transcript-0-0.jsonl"]
|
||||
parsed = [json.loads(line) for line in staged.splitlines()]
|
||||
assert len(staged.encode()) <= 8192
|
||||
assert parsed[-1]["type"] == "result"
|
||||
assert parsed[0]["message"]["content"][0]["signature"] == "[OMITTED]"
|
||||
assert "opaque-base64-signature" not in staged
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="transcript symlink containment is POSIX-only")
|
||||
def test_proposer_rejects_symlink_and_foreign_transcript_artifacts(tmp_path):
|
||||
results = tmp_path / "results"
|
||||
|
|
@ -299,6 +410,258 @@ def test_proposer_bounds_transcript_metadata_per_row_and_globally_before_materia
|
|||
)
|
||||
|
||||
|
||||
def test_proposer_refuses_a_selected_row_with_no_transcript_reference():
|
||||
# Every selectable row comes from sum_sessions(), which always emits the
|
||||
# key, and select_evidence() drops the kinds a failed transcript
|
||||
# persistence produces (session-error, infra-error, evidence-unverified,
|
||||
# cleanup-failure). A selected row without a transcript is therefore
|
||||
# evidence lost between producer and proposer, not a row that had none.
|
||||
with pytest.raises(evolve.SandboxError, match="missing transcript_artifacts"):
|
||||
proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[row()],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
)
|
||||
with pytest.raises(evolve.SandboxError, match="carries no transcript artifact"):
|
||||
proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[row(transcript_artifacts=[])],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
)
|
||||
|
||||
|
||||
def test_proposer_stages_the_bounded_prior_proposal(tmp_path):
|
||||
proposal = tmp_path / "proposal.md"
|
||||
proposal.write_text("# rejected candidate\n\nTightened the plan budget.\n")
|
||||
proposal.chmod(0o600)
|
||||
|
||||
entries = proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
prior_proposal=proposal,
|
||||
)
|
||||
|
||||
assert entries["prior-proposal.md"] == proposal.read_text()
|
||||
|
||||
oversized = tmp_path / "oversized.md"
|
||||
oversized.write_bytes(b"x" * (evolve.MAX_EVIDENCE_FILE_BYTES + 4096))
|
||||
oversized.chmod(0o600)
|
||||
bounded = proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
prior_proposal=oversized,
|
||||
)
|
||||
assert len(bounded["prior-proposal.md"]) == evolve.MAX_EVIDENCE_FILE_BYTES
|
||||
|
||||
|
||||
def test_stage_proposer_evidence_bundle_drops_prior_proposal_to_fit_budget(tmp_path, monkeypatch, capsys):
|
||||
# Per-file caps alone can still exceed the aggregate budget; the helper must
|
||||
# drop the prior proposal instead of aborting the generation.
|
||||
monkeypatch.setattr(evolve, "MAX_BUNDLE_BYTES", 2048)
|
||||
monkeypatch.setattr("workflow_bench.proposer_sandbox.MAX_BUNDLE_BYTES", 2048)
|
||||
|
||||
prior = tmp_path / "proposal.md"
|
||||
prior.write_text("x" * 2500)
|
||||
prior.chmod(0o600)
|
||||
|
||||
from workflow_bench.proposer_sandbox import SandboxError, stage_evidence_bundle
|
||||
|
||||
oversized = evolve.proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
prior_proposal=prior,
|
||||
)
|
||||
with pytest.raises(SandboxError, match="total byte limit"):
|
||||
stage_evidence_bundle(tmp_path / "raw", oversized)
|
||||
|
||||
bundle = evolve.stage_proposer_evidence_bundle(
|
||||
tmp_path / "bundle",
|
||||
results_dir=None,
|
||||
evidence=[],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
prior_proposal=prior,
|
||||
)
|
||||
names = {path.name for path in bundle.iterdir()}
|
||||
assert "selected-rows.json" in names
|
||||
assert "prior-proposal.md" not in names
|
||||
logged = capsys.readouterr().out
|
||||
assert "trimmed proposer evidence" in logged
|
||||
assert "omitted prior proposal" in logged
|
||||
|
||||
|
||||
def test_stage_proposer_evidence_bundle_compacts_artifacts_before_dropping_rows(tmp_path, monkeypatch, capsys):
|
||||
monkeypatch.setattr(evolve, "MAX_BUNDLE_BYTES", 150_000)
|
||||
monkeypatch.setattr("workflow_bench.proposer_sandbox.MAX_BUNDLE_BYTES", 150_000)
|
||||
results = tmp_path / "results"
|
||||
transcripts = results / "transcripts"
|
||||
transcripts.mkdir(parents=True, mode=0o700)
|
||||
rows = []
|
||||
for index in range(2):
|
||||
payload = (
|
||||
json.dumps(
|
||||
{
|
||||
"type": "assistant",
|
||||
"message": {"content": [{"type": "text", "text": "x" * 100_000}]},
|
||||
}
|
||||
)
|
||||
+ "\n"
|
||||
+ json.dumps({"type": "result", "session_id": f"session-{index}"})
|
||||
+ "\n"
|
||||
).encode()
|
||||
transcript = transcripts / f"session-{index}.jsonl"
|
||||
transcript.write_bytes(payload)
|
||||
transcript.chmod(0o600)
|
||||
(results / f"task-{index}-workflow-run0.patch").write_bytes(b"p" * 100_000)
|
||||
rows.append(
|
||||
row(
|
||||
task=f"task-{index}",
|
||||
transcript_artifacts=[
|
||||
{
|
||||
"path": f"transcripts/session-{index}.jsonl",
|
||||
"sha256": hashlib.sha256(payload).hexdigest(),
|
||||
"bytes": len(payload),
|
||||
"source": PARENT_EVENT_STREAM_SOURCE,
|
||||
}
|
||||
],
|
||||
)
|
||||
)
|
||||
|
||||
bundle = evolve.stage_proposer_evidence_bundle(
|
||||
tmp_path / "bundle",
|
||||
results_dir=results,
|
||||
evidence=rows,
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
)
|
||||
|
||||
staged_rows = json.loads((bundle / "selected-rows.json").read_text())
|
||||
assert len(staged_rows) == 2
|
||||
assert all((bundle / staged["patch_file"]).is_file() for staged in staged_rows)
|
||||
logged = capsys.readouterr().out
|
||||
assert "artifact cap" in logged
|
||||
assert "dropped 0 row(s)" in logged
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="proposal containment checks are POSIX-only")
|
||||
def test_proposer_refuses_a_prior_proposal_that_lost_its_trust_boundary(tmp_path):
|
||||
outside = tmp_path / "outside.md"
|
||||
outside.write_text("attacker-controlled prose")
|
||||
linked = tmp_path / "linked-proposal.md"
|
||||
linked.symlink_to(outside)
|
||||
|
||||
with pytest.raises(evolve.SandboxError, match="regular non-symlink"):
|
||||
proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
prior_proposal=linked,
|
||||
)
|
||||
|
||||
# run_proposer copies the proposal out 0600; anything looser means the
|
||||
# bytes are no longer only the ones this driver wrote.
|
||||
shared = tmp_path / "shared-proposal.md"
|
||||
shared.write_text("proposal")
|
||||
shared.chmod(0o644)
|
||||
with pytest.raises(evolve.SandboxError, match="owner-only"):
|
||||
proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
prior_proposal=shared,
|
||||
)
|
||||
|
||||
with pytest.raises(evolve.SandboxError, match="unavailable"):
|
||||
proposer_evidence_entries(
|
||||
results_dir=None,
|
||||
evidence=[],
|
||||
learnings=[],
|
||||
gate_summary=[],
|
||||
prior_proposal=tmp_path / "absent.md",
|
||||
)
|
||||
|
||||
|
||||
def test_run_proposer_hides_the_hidden_harness_and_keeps_the_full_tool_surface(monkeypatch, tmp_path):
|
||||
"""The proposer writes the artifact the arms are scored with.
|
||||
|
||||
So its clone must be sanitized before the session starts — a proposer that
|
||||
can read eval/workflow_bench reads the task prompts and hidden oracles it
|
||||
is about to be graded against, and can encode the answers into the skill.
|
||||
"""
|
||||
|
||||
evidence = tmp_path / "evidence"
|
||||
evidence.mkdir()
|
||||
transcript_projects = tmp_path / "transcript-projects"
|
||||
transcript_projects.mkdir()
|
||||
captured: dict[str, object] = {}
|
||||
sanitized: list[Path] = []
|
||||
events: list[str] = []
|
||||
|
||||
def fake_make_worktree(_repo, _ref, destination):
|
||||
clone = destination / "clone"
|
||||
clone.mkdir()
|
||||
return clone
|
||||
|
||||
def fake_sanitize(clone):
|
||||
events.append("sanitize")
|
||||
sanitized.append(clone)
|
||||
return "0" * 40
|
||||
|
||||
class FakeSandbox:
|
||||
claude_bin = "claude"
|
||||
command_prefix: list[str] = []
|
||||
settings_json = '{"permissions":{"allow":["Read"]}}'
|
||||
|
||||
@property
|
||||
def transcript_projects(self):
|
||||
return transcript_projects
|
||||
|
||||
@contextmanager
|
||||
def fake_prepare_sandbox(**_kwargs):
|
||||
events.append("prepare")
|
||||
yield FakeSandbox()
|
||||
|
||||
def fake_run_claude(*_args, **kwargs):
|
||||
captured.update(kwargs)
|
||||
return {"ok": False, "error_kind": "session-error"}
|
||||
|
||||
monkeypatch.setattr(evolve.runner, "make_worktree", fake_make_worktree)
|
||||
monkeypatch.setattr(evolve.runner, "remove_clone", lambda _clone: None)
|
||||
monkeypatch.setattr(evolve, "sanitize_clone_for_hidden_oracles", fake_sanitize)
|
||||
monkeypatch.setattr(evolve, "prepare_sandbox", fake_prepare_sandbox)
|
||||
monkeypatch.setattr(evolve.runner, "run_claude", fake_run_claude)
|
||||
args = build_parser().parse_args(["--tasks", "tasks.yaml", "--model", "model"])
|
||||
|
||||
record = evolve.run_proposer(
|
||||
"prompt",
|
||||
args,
|
||||
overlay_dir=tmp_path / "overlay",
|
||||
proposal_path=tmp_path / "proposal.md",
|
||||
evidence_bundle=evidence,
|
||||
bwrap_bin=tmp_path / "bwrap",
|
||||
)
|
||||
|
||||
assert record["ok"] is False
|
||||
# Sanitization has to happen on the clone the session actually runs in,
|
||||
# and before the sandbox is prepared around it.
|
||||
assert events[:2] == ["sanitize", "prepare"]
|
||||
assert [clone.name for clone in sanitized] == ["clone"]
|
||||
# Not --bare: bare ignores --tools and would cost the proposer Grep/Glob.
|
||||
assert captured.get("bare", False) is False
|
||||
assert captured["allowed_tools"] == evolve.PROPOSER_ALLOWED_TOOLS
|
||||
assert captured["settings_json"] == FakeSandbox.settings_json
|
||||
|
||||
|
||||
def test_parser_defaults_match_the_gate_minimums():
|
||||
args = build_parser().parse_args(["--tasks", "t.yaml", "--model", "pinned"])
|
||||
assert args.runs == 3
|
||||
|
|
@ -368,7 +731,7 @@ def test_evolve_proposer_failure_returns_nonzero(monkeypatch, tmp_path):
|
|||
assert evolve.main() == 1
|
||||
|
||||
|
||||
def test_proposer_session_record_is_redacted_before_upload(monkeypatch, tmp_path):
|
||||
def test_proposer_session_record_is_redacted_before_upload(monkeypatch, tmp_path, capsys):
|
||||
tasks = tmp_path / "tasks.yaml"
|
||||
tasks.write_text(
|
||||
"""tasks:
|
||||
|
|
@ -397,7 +760,7 @@ def test_proposer_session_record_is_redacted_before_upload(monkeypatch, tmp_path
|
|||
"pinned-model",
|
||||
"--out-root",
|
||||
str(tmp_path / "out"),
|
||||
"--auth-token",
|
||||
"--anthropic-api-key",
|
||||
literal_token,
|
||||
],
|
||||
)
|
||||
|
|
@ -422,6 +785,84 @@ def test_proposer_session_record_is_redacted_before_upload(monkeypatch, tmp_path
|
|||
assert pattern_token not in written
|
||||
assert "[REDACTED]" in written
|
||||
|
||||
# The same record is printed one line later, and the driver's stdout is a
|
||||
# live CI log now that the sweep echoes it — same bar as the artifact.
|
||||
printed = capsys.readouterr().out
|
||||
assert "proposer session failed" in printed
|
||||
assert literal_token not in printed
|
||||
assert pattern_token not in printed
|
||||
assert "[REDACTED]" in printed
|
||||
|
||||
|
||||
def test_benchmark_failure_print_is_redacted(monkeypatch, tmp_path, capsys):
|
||||
tasks = tmp_path / "tasks.yaml"
|
||||
tasks.write_text(
|
||||
"""tasks:
|
||||
- id: demo
|
||||
class: test
|
||||
repo: .
|
||||
prompt: implement
|
||||
verify: "true"
|
||||
oracle:
|
||||
command: "true"
|
||||
files:
|
||||
- source: hidden.test.ts
|
||||
target: hidden.test.ts
|
||||
"""
|
||||
)
|
||||
overlay = tmp_path / "overlay"
|
||||
skill = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
|
||||
skill.parent.mkdir(parents=True)
|
||||
skill.write_text("candidate")
|
||||
literal_token = "secret-LITERAL-XYZ"
|
||||
pattern_token = "sk-ant-FAKEEXAMPLE0000"
|
||||
monkeypatch.setattr(
|
||||
sys,
|
||||
"argv",
|
||||
[
|
||||
"workflow_bench.evolve",
|
||||
"--tasks",
|
||||
str(tasks),
|
||||
"--model",
|
||||
"pinned-model",
|
||||
"--out-root",
|
||||
str(tmp_path / "out"),
|
||||
"--initial-overlay",
|
||||
str(overlay),
|
||||
"--anthropic-api-key",
|
||||
literal_token,
|
||||
],
|
||||
)
|
||||
monkeypatch.setattr(evolve.runner, "selected_task_bindings", lambda _tasks: [{"id": "demo"}])
|
||||
monkeypatch.setattr(evolve, "preflight_bubblewrap", lambda: tmp_path / "bwrap")
|
||||
monkeypatch.setattr(evolve, "require_claude_sandbox_helpers", lambda: None)
|
||||
monkeypatch.setattr(evolve, "resolve_incumbent_arms", lambda *_args, **_kwargs: ["workflow"])
|
||||
monkeypatch.setattr(evolve, "freeze_overlay", lambda _source, _destination: "d" * 64)
|
||||
monkeypatch.setattr(evolve, "committed_destination_base_digests", lambda _overlay: {})
|
||||
monkeypatch.setattr(evolve, "destination_base_digests", lambda _overlay: {})
|
||||
# The sweep is launched with GITNEXUS_BENCH_ANTHROPIC_API_KEY in its environment,
|
||||
# so its detail/stderr tail is as token-bearing as any session record.
|
||||
monkeypatch.setattr(
|
||||
evolve,
|
||||
"run_managed",
|
||||
lambda *_args, **_kwargs: ManagedProcessResult(
|
||||
state="exited",
|
||||
returncode=2,
|
||||
stdout_tail="",
|
||||
stderr_tail=f"ANTHROPIC_API_KEY={pattern_token}",
|
||||
duration_s=1.0,
|
||||
detail=f"sweep died with {literal_token}",
|
||||
),
|
||||
)
|
||||
|
||||
assert evolve.main() == 1
|
||||
|
||||
printed = capsys.readouterr().out
|
||||
assert "benchmark run failed" in printed
|
||||
assert literal_token not in printed
|
||||
assert pattern_token not in printed
|
||||
assert "[REDACTED]" in printed
|
||||
|
||||
|
||||
def test_runner_argv_pairs_each_incumbent_with_its_candidate(tmp_path):
|
||||
args = build_parser().parse_args(
|
||||
|
|
@ -432,6 +873,8 @@ def test_runner_argv_pairs_each_incumbent_with_its_candidate(tmp_path):
|
|||
"pinned",
|
||||
"--arms",
|
||||
"workflow",
|
||||
"--workers",
|
||||
"3",
|
||||
"--include-expensive",
|
||||
]
|
||||
)
|
||||
|
|
@ -455,11 +898,32 @@ def test_runner_argv_pairs_each_incumbent_with_its_candidate(tmp_path):
|
|||
assert str(tmp_path / "bench") in argv
|
||||
assert "pinned" in argv
|
||||
assert argv[argv.index("--proposer-model") + 1] == "pinned"
|
||||
assert argv[argv.index("--effort") + 1] == "xhigh"
|
||||
assert argv[argv.index("--workers") + 1] == "3"
|
||||
assert "--include-expensive" in argv
|
||||
assert json.loads(argv[argv.index("--task-bindings-json") + 1]) == task_bindings
|
||||
assert json.loads(argv[argv.index("--promotion-target-bases-json") + 1]) == target_bases
|
||||
|
||||
|
||||
def test_runner_argv_inserts_ce_review_for_review_overlay(tmp_path):
|
||||
args = build_parser().parse_args(
|
||||
["--tasks", "t.yaml", "--model", "pinned", "--arms", "review"]
|
||||
)
|
||||
overlay = tmp_path / "overlay"
|
||||
skill = overlay / ".claude" / "skills" / "gitnexus-review" / "SKILL.md"
|
||||
skill.parent.mkdir(parents=True)
|
||||
skill.write_text("candidate")
|
||||
argv = runner_argv(
|
||||
args,
|
||||
tmp_path / "bench",
|
||||
overlay,
|
||||
task_bindings=[{"id": "task"}],
|
||||
target_base_digests={},
|
||||
)
|
||||
arms = argv[argv.index("--arms") + 1 : argv.index("--promotion-metric")]
|
||||
assert arms == ["ce_review", "review", "candidate_review"]
|
||||
|
||||
|
||||
def test_runner_argv_omits_proposer_for_manual_overlay(tmp_path):
|
||||
args = build_parser().parse_args(["--tasks", "t.yaml", "--model", "pinned"])
|
||||
overlay = tmp_path / "overlay"
|
||||
|
|
@ -479,6 +943,26 @@ def test_runner_argv_omits_proposer_for_manual_overlay(tmp_path):
|
|||
assert "--proposer-model" not in argv
|
||||
|
||||
|
||||
def test_runner_argv_forwards_explicit_unsafe_backend(tmp_path):
|
||||
args = build_parser().parse_args(
|
||||
["--tasks", "t.yaml", "--model", "pinned", "--unsafe-no-bwrap"]
|
||||
)
|
||||
overlay = tmp_path / "overlay"
|
||||
skill = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
|
||||
skill.parent.mkdir(parents=True)
|
||||
skill.write_text("candidate")
|
||||
|
||||
argv = runner_argv(
|
||||
args,
|
||||
tmp_path / "bench",
|
||||
overlay,
|
||||
task_bindings=[{"id": "task"}],
|
||||
target_base_digests={},
|
||||
)
|
||||
|
||||
assert "--unsafe-no-bwrap" in argv
|
||||
|
||||
|
||||
def test_runner_argv_keeps_task_commit_pinned_when_ref_moves(tmp_path):
|
||||
repo = tmp_path / "task-repo"
|
||||
repo.mkdir()
|
||||
|
|
@ -510,7 +994,7 @@ def test_runner_argv_keeps_task_commit_pinned_when_ref_moves(tmp_path):
|
|||
"command": "true",
|
||||
"files": [
|
||||
{
|
||||
"source": "trivial-version-alias.oracle.test.ts",
|
||||
"source": "trivial-status-json-alias.oracle.test.ts",
|
||||
"target": "oracle.test.ts",
|
||||
}
|
||||
],
|
||||
|
|
@ -581,6 +1065,59 @@ def test_generation_timeout_budgets_three_task_workflow_pair():
|
|||
assert timeout >= 3 * (1 + 3 * paired_arm_cells) * evolve.WORKTREE_PREPARATION_TIMEOUT_SECONDS
|
||||
|
||||
|
||||
def test_executed_benchmark_arms_inserts_review_comparator() -> None:
|
||||
assert executed_benchmark_arms(["workflow"]) == ["workflow", "candidate_workflow"]
|
||||
assert executed_benchmark_arms(["review"]) == ["ce_review", "review", "candidate_review"]
|
||||
|
||||
|
||||
def test_generation_timeout_budgets_review_pair_plus_ce_comparator() -> None:
|
||||
timeout = generation_timeout_seconds(
|
||||
task_count=6,
|
||||
runs=1,
|
||||
session_timeout=3600,
|
||||
incumbent_arms=["review"],
|
||||
)
|
||||
|
||||
per_task_preparation = (
|
||||
evolve.TASK_BINDING_GIT_PHASES * evolve.GIT_COMMAND_TIMEOUT_SECONDS
|
||||
+ 2 * evolve.TASK_SNAPSHOT_TIMEOUT_SECONDS
|
||||
+ evolve.WORKTREE_PREPARATION_TIMEOUT_SECONDS
|
||||
+ evolve.GRAPH_SOURCE_PREPARATION_TIMEOUT_SECONDS
|
||||
+ evolve.GRAPH_BUILD_TIMEOUT_SECONDS
|
||||
+ 2 * evolve.GRAPH_QUERY_TIMEOUT_SECONDS
|
||||
+ evolve.CLEANUP_TIMEOUT_SECONDS
|
||||
)
|
||||
paired_arm_cells = 3
|
||||
session_slots = 3
|
||||
workspace_snapshot_slots = 3
|
||||
per_task_run = session_slots * (3600 + evolve.SESSION_FINALIZATION_TIMEOUT_SECONDS) + paired_arm_cells * (
|
||||
evolve.WORKTREE_PREPARATION_TIMEOUT_SECONDS
|
||||
+ evolve.ARM_ASSET_MATERIALIZATION_PHASES * evolve.TASK_SNAPSHOT_TIMEOUT_SECONDS
|
||||
+ evolve.SETUP_TIMEOUT_SECONDS
|
||||
+ 2 * 3600
|
||||
+ evolve.ARM_EVIDENCE_GIT_PHASES * evolve.GIT_COMMAND_TIMEOUT_SECONDS
|
||||
+ evolve.CLEANUP_TIMEOUT_SECONDS
|
||||
)
|
||||
per_task_run += workspace_snapshot_slots * evolve.TASK_SNAPSHOT_TIMEOUT_SECONDS
|
||||
per_task_run += evolve.CANDIDATE_OVERLAY_GIT_PHASES * evolve.GIT_COMMAND_TIMEOUT_SECONDS
|
||||
|
||||
assert timeout == (
|
||||
evolve.PROMOTION_BASE_TIMEOUT_SECONDS
|
||||
+ 6 * (per_task_preparation + per_task_run)
|
||||
+ evolve.DRIVER_OVERHEAD_SECONDS
|
||||
)
|
||||
|
||||
|
||||
def test_generation_timeout_rejects_unknown_arm() -> None:
|
||||
with pytest.raises(ValueError, match="unsupported evolution arm: mystery"):
|
||||
generation_timeout_seconds(
|
||||
task_count=1,
|
||||
runs=1,
|
||||
session_timeout=60,
|
||||
incumbent_arms=["mystery"],
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.skipif(sys.platform != "linux", reason="Bubblewrap PID namespaces require Linux")
|
||||
def test_outer_runner_pid_namespace_kills_setsid_descendant(tmp_path):
|
||||
try:
|
||||
|
|
@ -625,9 +1162,9 @@ def test_resolve_incumbent_arms_rejects_incomplete_and_extra_explicit_sets(tmp_p
|
|||
resolve_incumbent_arms(work, ["workflow"])
|
||||
|
||||
|
||||
def bound_task_fixture():
|
||||
def bound_task_fixture(task_id="task-a"):
|
||||
return {
|
||||
"id": "task",
|
||||
"id": task_id,
|
||||
"prompt_digest": "prompt",
|
||||
"oracle_digest": "a" * 64,
|
||||
"oracle_command_digest": "b" * 64,
|
||||
|
|
@ -638,37 +1175,50 @@ def bound_task_fixture():
|
|||
}
|
||||
|
||||
|
||||
def bound_task_fixtures():
|
||||
return [bound_task_fixture("task-a"), bound_task_fixture("task-impossible")]
|
||||
|
||||
|
||||
def promote_decision(**overrides):
|
||||
"""A real producer decision, including its recomputable paired metrics."""
|
||||
base = {"runs": 3, "valid_runs": 3, "excluded_runs": 0, "error_kinds": {}}
|
||||
results = {
|
||||
"task-a": {
|
||||
"workflow": {**base, "resolved": 3, "cost_usd": 1.0},
|
||||
"candidate_workflow": {**base, "resolved": 3, "cost_usd": 0.8},
|
||||
},
|
||||
"task-impossible": {
|
||||
"workflow": {**base, "resolved": 0, "cost_usd": 1.0},
|
||||
"candidate_workflow": {**base, "resolved": 0, "cost_usd": 1.0},
|
||||
},
|
||||
}
|
||||
decision = evolution.promotion_evidence(
|
||||
results,
|
||||
policy=evolution.promotion_policy(["candidate_workflow"]),
|
||||
model="bench-model",
|
||||
complete=True,
|
||||
)["decisions"][0]
|
||||
decision.update(overrides)
|
||||
return decision
|
||||
|
||||
|
||||
def promotion_fixture(*, decisions=None, expires_delta=timedelta(days=1)):
|
||||
now = datetime.now(UTC)
|
||||
return {
|
||||
"schema_version": 3,
|
||||
"schema_version": 6,
|
||||
"run_status": "complete",
|
||||
"generated_at": now.isoformat(),
|
||||
"evidence_expires_at": (now + expires_delta).isoformat(),
|
||||
"benchmark_model": "bench-model",
|
||||
"proposer_model": "proposer-model",
|
||||
"effort": "xhigh",
|
||||
"candidate_origin": "model-proposer",
|
||||
"candidate_overlay_digest": "digest",
|
||||
"target_base_digests": {"path": "base"},
|
||||
"required_candidate_arms": ["candidate_workflow"],
|
||||
"selected_tasks": [bound_task_fixture()],
|
||||
"policy": {
|
||||
"metric": "cost_usd",
|
||||
"min_runs": 3,
|
||||
"min_improvement_pct": 5.0,
|
||||
"max_task_regression_pct": 20.0,
|
||||
},
|
||||
"decisions": (
|
||||
decisions
|
||||
if decisions is not None
|
||||
else [
|
||||
{
|
||||
"incumbent_arm": "workflow",
|
||||
"candidate_arm": "candidate_workflow",
|
||||
"decision": "promote",
|
||||
"metric": "cost_usd",
|
||||
}
|
||||
]
|
||||
),
|
||||
"selected_tasks": bound_task_fixtures(),
|
||||
"policy": evolution.promotion_policy(["candidate_workflow"]),
|
||||
"decisions": decisions if decisions is not None else [promote_decision()],
|
||||
}
|
||||
|
||||
|
||||
|
|
@ -678,15 +1228,11 @@ def validate_fixture(promotion):
|
|||
overlay_digest="digest",
|
||||
benchmark_model="bench-model",
|
||||
proposer_model="proposer-model",
|
||||
selected_tasks=[bound_task_fixture()],
|
||||
effort="xhigh",
|
||||
selected_tasks=bound_task_fixtures(),
|
||||
target_base_digests={"path": "base"},
|
||||
required_candidate_arms=["candidate_workflow"],
|
||||
policy={
|
||||
"metric": "cost_usd",
|
||||
"min_runs": 3,
|
||||
"min_improvement_pct": 5.0,
|
||||
"max_task_regression_pct": 20.0,
|
||||
},
|
||||
policy=evolution.promotion_policy(["candidate_workflow"]),
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -695,41 +1241,85 @@ def test_promotion_apply_requires_one_promote_for_every_bound_arm():
|
|||
|
||||
for decisions in (
|
||||
[],
|
||||
[
|
||||
{
|
||||
"incumbent_arm": "workflow",
|
||||
"candidate_arm": "candidate_workflow",
|
||||
"decision": "keep_incumbent",
|
||||
"metric": "cost_usd",
|
||||
}
|
||||
],
|
||||
[
|
||||
{
|
||||
"incumbent_arm": "workflow",
|
||||
"candidate_arm": "candidate_workflow",
|
||||
"decision": "promote",
|
||||
"metric": "cost_usd",
|
||||
},
|
||||
{
|
||||
"incumbent_arm": "workflow",
|
||||
"candidate_arm": "candidate_workflow",
|
||||
"decision": "promote",
|
||||
"metric": "cost_usd",
|
||||
},
|
||||
],
|
||||
[
|
||||
{
|
||||
"incumbent_arm": "workflow_direct",
|
||||
"candidate_arm": "candidate_workflow_direct",
|
||||
"decision": "promote",
|
||||
"metric": "cost_usd",
|
||||
}
|
||||
],
|
||||
[promote_decision(decision="keep_incumbent")],
|
||||
[promote_decision(), promote_decision()],
|
||||
[promote_decision(incumbent_arm="workflow_direct", candidate_arm="candidate_workflow_direct")],
|
||||
):
|
||||
with pytest.raises(ValueError):
|
||||
validate_fixture(promotion_fixture(decisions=decisions))
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("overrides", "match"),
|
||||
[
|
||||
# An older decision relabeled as schema 5: the verdict without the
|
||||
# gated evidence base schema 5 promotes on.
|
||||
({"tasks": None, "ungated_tasks": None}, "no per-task gate evidence"),
|
||||
({"tasks": [{"task": "task-a"}]}, "malformed per-task gate evidence"),
|
||||
({"tasks": [{"task": "task-a", "gated": "yes"}]}, "malformed per-task gate evidence"),
|
||||
(
|
||||
{"tasks": [{"task": "task-a", "gated": True}, {"task": "task-a", "gated": False}]},
|
||||
"repeats a task",
|
||||
),
|
||||
(
|
||||
{"ungated_tasks": [], "tasks": [{"task": "fabricated", "gated": True}]},
|
||||
"does not match selected tasks",
|
||||
),
|
||||
({"ungated_tasks": None}, "missing its ungated task list"),
|
||||
# The verdict claims a full gate; the per-task rows say a task sat
|
||||
# outside it.
|
||||
({"ungated_tasks": []}, "disagree with its per-task evidence"),
|
||||
(
|
||||
{
|
||||
"ungated_tasks": ["task-a", "task-impossible"],
|
||||
"tasks": [
|
||||
{"task": "task-a", "gated": False},
|
||||
{"task": "task-impossible", "gated": False},
|
||||
],
|
||||
},
|
||||
"no gated task",
|
||||
),
|
||||
],
|
||||
)
|
||||
def test_promotion_apply_binds_the_schema_5_gate_evidence(overrides, match):
|
||||
decision = promote_decision()
|
||||
for field, value in overrides.items():
|
||||
if value is None:
|
||||
decision.pop(field)
|
||||
else:
|
||||
decision[field] = value
|
||||
|
||||
with pytest.raises(ValueError, match=match):
|
||||
validate_fixture(promotion_fixture(decisions=[decision]))
|
||||
|
||||
|
||||
def test_promotion_apply_rejects_a_gate_with_only_one_of_three_selected_tasks():
|
||||
selected = [*bound_task_fixtures(), bound_task_fixture("task-impossible-2")]
|
||||
decision = promote_decision(
|
||||
ungated_tasks=["task-impossible", "task-impossible-2"],
|
||||
tasks=[
|
||||
{"task": "task-a", "gated": True},
|
||||
{"task": "task-impossible", "gated": False},
|
||||
{"task": "task-impossible-2", "gated": False},
|
||||
],
|
||||
)
|
||||
promotion = promotion_fixture(decisions=[decision])
|
||||
promotion["selected_tasks"] = selected
|
||||
|
||||
with pytest.raises(ValueError, match="too thin a gated evidence base"):
|
||||
validate_promotion_for_apply(
|
||||
promotion,
|
||||
overlay_digest="digest",
|
||||
benchmark_model="bench-model",
|
||||
proposer_model="proposer-model",
|
||||
effort="xhigh",
|
||||
selected_tasks=selected,
|
||||
target_base_digests={"path": "base"},
|
||||
required_candidate_arms=["candidate_workflow"],
|
||||
policy=promotion["policy"],
|
||||
)
|
||||
|
||||
|
||||
def test_manual_initial_overlay_has_no_fictitious_proposer_model():
|
||||
promotion = promotion_fixture()
|
||||
promotion["proposer_model"] = None
|
||||
|
|
@ -740,7 +1330,8 @@ def test_manual_initial_overlay_has_no_fictitious_proposer_model():
|
|||
overlay_digest="digest",
|
||||
benchmark_model="bench-model",
|
||||
proposer_model=None,
|
||||
selected_tasks=[bound_task_fixture()],
|
||||
effort="xhigh",
|
||||
selected_tasks=bound_task_fixtures(),
|
||||
target_base_digests={"path": "base"},
|
||||
required_candidate_arms=["candidate_workflow"],
|
||||
policy=promotion["policy"],
|
||||
|
|
@ -751,7 +1342,7 @@ def test_manual_initial_overlay_has_no_fictitious_proposer_model():
|
|||
|
||||
def test_promotion_apply_rejects_pre_oracle_schema_and_missing_oracle_bindings():
|
||||
legacy = promotion_fixture()
|
||||
legacy["schema_version"] = 2
|
||||
legacy["schema_version"] = 3
|
||||
with pytest.raises(ValueError, match="unsupported schema"):
|
||||
validate_fixture(legacy)
|
||||
|
||||
|
|
@ -764,6 +1355,7 @@ def test_promotion_apply_rejects_pre_oracle_schema_and_missing_oracle_bindings()
|
|||
overlay_digest="digest",
|
||||
benchmark_model="bench-model",
|
||||
proposer_model="proposer-model",
|
||||
effort="xhigh",
|
||||
selected_tasks=[weak_task],
|
||||
target_base_digests={"path": "base"},
|
||||
required_candidate_arms=["candidate_workflow"],
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
"""Tests for MCPBridge._find_gitnexus_command() and subprocess spawn."""
|
||||
import subprocess
|
||||
import unittest
|
||||
from unittest.mock import MagicMock, call, patch
|
||||
from unittest.mock import MagicMock, patch
|
||||
|
||||
|
||||
class TestFindGitnexusCommand(unittest.TestCase):
|
||||
|
|
|
|||
435
eval/tests/test_model_gateway.py
Normal file
435
eval/tests/test_model_gateway.py
Normal file
|
|
@ -0,0 +1,435 @@
|
|||
"""Credential routing for the OpenAI loopback gateway."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import subprocess
|
||||
import os
|
||||
import json
|
||||
import signal
|
||||
import socket
|
||||
import time
|
||||
import sys
|
||||
import threading
|
||||
import urllib.request
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
from pathlib import Path
|
||||
from unittest import mock
|
||||
|
||||
import pytest
|
||||
import yaml
|
||||
|
||||
|
||||
from workflow_bench.model_gateway import (
|
||||
DEFAULT_GATEWAY_READY_TIMEOUT_S,
|
||||
GATEWAY_READY_TIMEOUT_ENV,
|
||||
GATEWAY_REQUEST_TIMEOUT_S,
|
||||
OpenAIGateway,
|
||||
gateway_ready_timeout_s,
|
||||
anthropic_api_key_from_environ,
|
||||
claude_gateway_model_env,
|
||||
is_openai_model,
|
||||
litellm_proxy_argv,
|
||||
openai_backend_model,
|
||||
openai_litellm_config,
|
||||
resolve_model_access,
|
||||
write_openai_litellm_config,
|
||||
)
|
||||
|
||||
|
||||
def test_supervisor_reports_proxy_failure_without_aborting_on_its_stdin_reader():
|
||||
supervisor = Path(__file__).resolve().parents[1] / "workflow_bench" / "gateway_supervisor.py"
|
||||
process = subprocess.Popen(
|
||||
[sys.executable, str(supervisor), sys.executable, "-c", "raise RuntimeError('proxy failed')"],
|
||||
stdin=subprocess.PIPE,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
)
|
||||
try:
|
||||
# Keep the owner pipe open: the proxy exits independently of its owner.
|
||||
process.wait(timeout=10)
|
||||
assert process.returncode == 1
|
||||
assert b"proxy failed" in process.stderr.read()
|
||||
finally:
|
||||
process.stdin.close()
|
||||
if process.poll() is None:
|
||||
process.kill()
|
||||
process.wait(timeout=5)
|
||||
|
||||
|
||||
def test_locked_litellm_translates_messages_to_offline_responses(monkeypatch, tmp_path):
|
||||
from workflow_bench import model_gateway
|
||||
|
||||
observed = []
|
||||
|
||||
class Upstream(BaseHTTPRequestHandler):
|
||||
def log_message(self, *args):
|
||||
pass
|
||||
|
||||
def do_POST(self):
|
||||
body = json.loads(self.rfile.read(int(self.headers["Content-Length"])))
|
||||
observed.append((self.path, self.headers.get("Authorization"), body))
|
||||
response = {
|
||||
"id": "resp_offline",
|
||||
"object": "response",
|
||||
"created_at": int(time.time()),
|
||||
"status": "completed",
|
||||
"model": "gpt-4.1",
|
||||
"error": None,
|
||||
"output": [
|
||||
{
|
||||
"id": "msg_offline",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"status": "completed",
|
||||
"content": [{"type": "output_text", "text": "offline pong", "annotations": []}],
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"input_tokens": 1,
|
||||
"output_tokens": 2,
|
||||
"total_tokens": 3,
|
||||
"input_tokens_details": {"cached_tokens": 0},
|
||||
"output_tokens_details": {"reasoning_tokens": 0},
|
||||
},
|
||||
}
|
||||
payload = json.dumps(response).encode()
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(payload)))
|
||||
self.end_headers()
|
||||
self.wfile.write(payload)
|
||||
|
||||
upstream = ThreadingHTTPServer(("127.0.0.1", 0), Upstream)
|
||||
worker = threading.Thread(target=upstream.serve_forever, daemon=True)
|
||||
worker.start()
|
||||
original = model_gateway.write_openai_litellm_config
|
||||
|
||||
def config(path, names):
|
||||
original(path, names)
|
||||
document = yaml.safe_load(path.read_text())
|
||||
for entry in document["model_list"]:
|
||||
entry["litellm_params"]["api_base"] = f"http://127.0.0.1:{upstream.server_port}/v1"
|
||||
path.write_text(yaml.safe_dump(document))
|
||||
return path
|
||||
|
||||
monkeypatch.setattr(model_gateway, "write_openai_litellm_config", config)
|
||||
try:
|
||||
with OpenAIGateway(
|
||||
openai_api_key="offline-upstream-secret",
|
||||
model_names=["gpt-4.1"],
|
||||
work_dir=tmp_path / "gateway",
|
||||
ready_timeout_s=60,
|
||||
) as gateway:
|
||||
port = gateway.port
|
||||
request = urllib.request.Request(
|
||||
gateway.base_url + "/v1/messages",
|
||||
data=json.dumps(
|
||||
{
|
||||
"model": "gpt-4.1",
|
||||
"max_tokens": 32,
|
||||
"messages": [{"role": "user", "content": "ping"}],
|
||||
}
|
||||
).encode(),
|
||||
headers={
|
||||
"Content-Type": "application/json",
|
||||
"x-api-key": gateway.auth_token,
|
||||
"anthropic-version": "2023-06-01",
|
||||
},
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=30) as response:
|
||||
translated = json.load(response)
|
||||
assert "offline pong" in json.dumps(translated)
|
||||
assert len(observed) == 1 and observed[0][0] == "/v1/responses"
|
||||
assert observed[0][1] == "Bearer offline-upstream-secret"
|
||||
assert gateway.auth_token != "offline-upstream-secret"
|
||||
with socket.socket() as client:
|
||||
assert client.connect_ex(("127.0.0.1", port)) != 0
|
||||
gateway.close() # ownership close is idempotent
|
||||
finally:
|
||||
upstream.shutdown()
|
||||
upstream.server_close()
|
||||
worker.join(timeout=5)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("termination", ["terminate", "kill"])
|
||||
@pytest.mark.parametrize("phase", ["ready", "startup"])
|
||||
def test_gateway_lifetime_ends_with_its_parent(tmp_path, termination, phase):
|
||||
ready = tmp_path / "ready.json"
|
||||
proxy_pid = tmp_path / "proxy-pid"
|
||||
proxy = tmp_path / "proxy.py"
|
||||
proxy.write_text(f"""import os,sys
|
||||
from pathlib import Path
|
||||
from http.server import BaseHTTPRequestHandler, HTTPServer
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
def do_GET(self):
|
||||
self.send_response({200 if phase == "ready" else 503}); self.end_headers()
|
||||
Path(sys.argv[2]).write_text(str(os.getpid()))
|
||||
HTTPServer(('127.0.0.1', int(sys.argv[1])), Handler).serve_forever()
|
||||
""")
|
||||
parent_code = f"""
|
||||
import json,os,sys,time,subprocess
|
||||
from pathlib import Path
|
||||
from workflow_bench import model_gateway
|
||||
model_gateway.litellm_proxy_argv=lambda **kwargs: [sys.executable, {str(proxy)!r}, str(kwargs['port']), {str(proxy_pid)!r}]
|
||||
gateway=model_gateway.OpenAIGateway(openai_api_key='offline-secret', model_names=['gpt-4.1'], work_dir=Path({str(tmp_path / "gateway")!r}), ready_timeout_s=10)
|
||||
if {phase == "startup"!r}:
|
||||
Path({str(ready)!r}).write_text(json.dumps({{'port':gateway.port}}))
|
||||
gateway.__enter__()
|
||||
if gateway._process.stdin is not None:
|
||||
assert not os.get_inheritable(gateway._process.stdin.fileno())
|
||||
Path({str(ready)!r}).write_text(json.dumps({{'port':gateway.port}}))
|
||||
time.sleep(60)
|
||||
"""
|
||||
parent = subprocess.Popen(
|
||||
[sys.executable, "-c", parent_code],
|
||||
cwd=Path(__file__).resolve().parents[1],
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
text=True,
|
||||
)
|
||||
try:
|
||||
deadline = time.monotonic() + 12
|
||||
while (not ready.exists() or not proxy_pid.exists()) and parent.poll() is None and time.monotonic() < deadline:
|
||||
time.sleep(0.02)
|
||||
if not ready.exists():
|
||||
_, stderr = parent.communicate(timeout=1)
|
||||
pytest.fail(stderr)
|
||||
port = json.loads(ready.read_text())["port"]
|
||||
getattr(parent, termination)()
|
||||
parent.wait(timeout=5)
|
||||
deadline = time.monotonic() + 15
|
||||
while time.monotonic() < deadline:
|
||||
with socket.socket() as client:
|
||||
if client.connect_ex(("127.0.0.1", port)) != 0:
|
||||
break
|
||||
time.sleep(0.02)
|
||||
else:
|
||||
pytest.fail("gateway port survived abrupt parent death")
|
||||
finally:
|
||||
if parent.poll() is None:
|
||||
parent.kill()
|
||||
parent.wait(timeout=5)
|
||||
if proxy_pid.exists():
|
||||
try:
|
||||
if os.name == "nt":
|
||||
os.kill(int(proxy_pid.read_text()), signal.SIGTERM)
|
||||
else:
|
||||
os.killpg(int(proxy_pid.read_text()), signal.SIGKILL)
|
||||
except OSError:
|
||||
# Successful ownership cleanup has already removed this process.
|
||||
pass
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("model", "expected"),
|
||||
[
|
||||
("gpt-4.1", True),
|
||||
("gpt-4o-mini", True),
|
||||
("openai/gpt-4.1", True),
|
||||
("o3", True),
|
||||
("o4-mini", True),
|
||||
("claude-sonnet-5", False),
|
||||
("free-coder", False),
|
||||
("pinned-model", False),
|
||||
],
|
||||
)
|
||||
def test_is_openai_model(model: str, expected: bool) -> None:
|
||||
assert is_openai_model(model) is expected
|
||||
|
||||
|
||||
def test_openai_litellm_config_routes_each_id_to_openai_and_env_key(tmp_path: Path) -> None:
|
||||
config = openai_litellm_config(["gpt-4.1", "openai/gpt-4.1-mini", "gpt-4.1"])
|
||||
assert [row["model_name"] for row in config["model_list"]] == ["gpt-4.1", "openai/gpt-4.1-mini"]
|
||||
assert config["model_list"][0]["litellm_params"]["model"] == "openai/gpt-4.1"
|
||||
assert config["model_list"][1]["litellm_params"]["model"] == "openai/gpt-4.1-mini"
|
||||
assert all(row["litellm_params"]["api_key"] == "os.environ/OPENAI_API_KEY" for row in config["model_list"])
|
||||
assert all(row["model_info"] == {"mode": "responses"} for row in config["model_list"])
|
||||
assert all(row["litellm_params"]["timeout"] == GATEWAY_REQUEST_TIMEOUT_S for row in config["model_list"])
|
||||
assert config["litellm_settings"]["request_timeout"] == GATEWAY_REQUEST_TIMEOUT_S
|
||||
path = write_openai_litellm_config(tmp_path / "litellm.yaml", ["gpt-4.1"])
|
||||
assert yaml.safe_load(path.read_text())["model_list"][0]["model_name"] == "gpt-4.1"
|
||||
if os.name != "nt":
|
||||
# Windows chmod exposes a read-only flag, not POSIX access bits.
|
||||
assert path.stat().st_mode & 0o777 == 0o600
|
||||
|
||||
|
||||
def test_resolve_model_access_starts_proxy_only_for_openai_ids() -> None:
|
||||
openai = resolve_model_access(
|
||||
auth_token=None,
|
||||
openai_api_key="sk-openai",
|
||||
base_url=None,
|
||||
models=["gpt-4.1", "gpt-4.1"],
|
||||
)
|
||||
assert openai.start_proxy is True
|
||||
assert openai.openai_api_key == "sk-openai"
|
||||
|
||||
anthropic = resolve_model_access(
|
||||
auth_token="sk-ant",
|
||||
openai_api_key="sk-openai",
|
||||
base_url=None,
|
||||
models=["claude-sonnet-5"],
|
||||
)
|
||||
assert anthropic.start_proxy is False
|
||||
|
||||
existing = resolve_model_access(
|
||||
auth_token="proxy-master",
|
||||
openai_api_key=None,
|
||||
base_url="http://127.0.0.1:4000",
|
||||
models=["free-coder"],
|
||||
)
|
||||
assert existing.start_proxy is False
|
||||
|
||||
|
||||
def test_resolve_model_access_rejects_openai_ids_without_a_key_and_mixed_providers() -> None:
|
||||
with pytest.raises(ValueError, match="GITNEXUS_BENCH_OPENAI_API_KEY"):
|
||||
resolve_model_access(
|
||||
auth_token="sk-ant",
|
||||
openai_api_key=None,
|
||||
base_url=None,
|
||||
models=["gpt-4.1"],
|
||||
)
|
||||
with pytest.raises(ValueError, match="mix"):
|
||||
resolve_model_access(
|
||||
auth_token=None,
|
||||
openai_api_key="sk-openai",
|
||||
base_url=None,
|
||||
models=["gpt-4.1", "claude-sonnet-5"],
|
||||
)
|
||||
with pytest.raises(ValueError, match="--base-url"):
|
||||
resolve_model_access(
|
||||
auth_token=None,
|
||||
openai_api_key=None,
|
||||
base_url="http://127.0.0.1:4000",
|
||||
models=["free-coder"],
|
||||
)
|
||||
|
||||
|
||||
def test_claude_gateway_aliases_pin_every_internal_tier_to_the_session_model() -> None:
|
||||
env = claude_gateway_model_env("gpt-4.1")
|
||||
assert env["ANTHROPIC_MODEL"] == "gpt-4.1"
|
||||
assert env["ANTHROPIC_DEFAULT_HAIKU_MODEL"] == "gpt-4.1"
|
||||
assert env["CLAUDE_CODE_SUBAGENT_MODEL"] == "gpt-4.1"
|
||||
# High-effort reasoning outlives Claude Code's default client timeout.
|
||||
assert env["API_TIMEOUT_MS"] == str(GATEWAY_REQUEST_TIMEOUT_S * 1000)
|
||||
|
||||
|
||||
def test_openai_gateway_never_leaves_proxy_output_on_an_undrained_pipe(tmp_path: Path) -> None:
|
||||
# Nothing reads the proxy's output after startup, so a pipe would block the
|
||||
# proxy once its request logs filled the buffer and hang every session.
|
||||
gateway = OpenAIGateway(
|
||||
openai_api_key="sk-openai-secret",
|
||||
model_names=["gpt-4.1"],
|
||||
work_dir=tmp_path / "gw",
|
||||
ready_timeout_s=0.1,
|
||||
)
|
||||
captured: dict[str, object] = {}
|
||||
|
||||
def fake_popen(argv, **kwargs):
|
||||
captured.update(kwargs)
|
||||
raise OSError("no proxy in this test")
|
||||
|
||||
with mock.patch.object(subprocess, "Popen", fake_popen):
|
||||
with pytest.raises(RuntimeError, match="failed to start the OpenAI LiteLLM gateway"):
|
||||
gateway.__enter__()
|
||||
|
||||
assert captured["stderr"] is subprocess.STDOUT
|
||||
assert captured["stdout"] is not subprocess.PIPE
|
||||
assert getattr(captured["stdout"], "name", "") == str(gateway.log_path)
|
||||
if os.name != "nt":
|
||||
assert gateway.log_path.stat().st_mode & 0o777 == 0o600
|
||||
|
||||
|
||||
def test_gateway_startup_budget_outlives_a_cold_litellm_import(monkeypatch, tmp_path: Path) -> None:
|
||||
# Importing LiteLLM takes ~17s on a cold container filesystem and the proxy
|
||||
# binds its port only afterwards, so a sub-20s budget fails as "connection
|
||||
# refused" on a proxy that was merely still starting.
|
||||
monkeypatch.delenv(GATEWAY_READY_TIMEOUT_ENV, raising=False)
|
||||
assert DEFAULT_GATEWAY_READY_TIMEOUT_S >= 60
|
||||
assert gateway_ready_timeout_s() == DEFAULT_GATEWAY_READY_TIMEOUT_S
|
||||
assert (
|
||||
OpenAIGateway(
|
||||
openai_api_key="sk-openai-secret",
|
||||
model_names=["gpt-4.1"],
|
||||
work_dir=tmp_path / "gw",
|
||||
).ready_timeout_s
|
||||
== DEFAULT_GATEWAY_READY_TIMEOUT_S
|
||||
)
|
||||
|
||||
monkeypatch.setenv(GATEWAY_READY_TIMEOUT_ENV, "42.5")
|
||||
assert gateway_ready_timeout_s() == 42.5
|
||||
|
||||
for bad in ("0", "-1", "soon", "nan", "inf", "-inf", "1e999"):
|
||||
monkeypatch.setenv(GATEWAY_READY_TIMEOUT_ENV, bad)
|
||||
with pytest.raises(ValueError, match=GATEWAY_READY_TIMEOUT_ENV):
|
||||
gateway_ready_timeout_s()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("timeout", [0.0, -1.0, float("nan"), float("inf"), float("-inf")])
|
||||
def test_gateway_rejects_invalid_explicit_readiness_budgets(tmp_path: Path, timeout: float) -> None:
|
||||
with pytest.raises(ValueError, match="finite and positive"):
|
||||
OpenAIGateway(
|
||||
openai_api_key="sk-offline-test",
|
||||
model_names=["gpt-4.1"],
|
||||
work_dir=tmp_path / "gw",
|
||||
ready_timeout_s=timeout,
|
||||
)
|
||||
|
||||
|
||||
def test_gateway_readiness_timeout_reports_the_proxy_log_and_the_override(tmp_path: Path) -> None:
|
||||
gateway = OpenAIGateway(
|
||||
openai_api_key="sk-openai-secret",
|
||||
model_names=["gpt-4.1"],
|
||||
work_dir=tmp_path / "gw",
|
||||
ready_timeout_s=0.1,
|
||||
)
|
||||
gateway.work_dir.mkdir(parents=True)
|
||||
gateway.log_path.write_text("ImportError: litellm proxy extras missing")
|
||||
|
||||
with pytest.raises(RuntimeError) as excinfo:
|
||||
gateway._wait_until_ready()
|
||||
|
||||
message = str(excinfo.value)
|
||||
assert "ImportError: litellm proxy extras missing" in message
|
||||
assert GATEWAY_READY_TIMEOUT_ENV in message
|
||||
|
||||
|
||||
def test_openai_backend_model_preserves_openai_prefix() -> None:
|
||||
assert openai_backend_model("gpt-4.1") == "openai/gpt-4.1"
|
||||
assert openai_backend_model("openai/gpt-4.1") == "openai/gpt-4.1"
|
||||
|
||||
|
||||
def test_litellm_proxy_argv_uses_console_script_not_python_module(tmp_path: Path, monkeypatch) -> None:
|
||||
# litellm 1.87 ships a console script and no litellm.__main__, so
|
||||
# `python -m litellm` dies before the health check. Pin the supported argv.
|
||||
# Under `uv run`, sys.executable is the base CPython — the script lives in
|
||||
# VIRTUAL_ENV/bin instead.
|
||||
python = tmp_path / "base" / "python"
|
||||
venv_bin = tmp_path / "venv" / "bin"
|
||||
python.parent.mkdir(parents=True)
|
||||
venv_bin.mkdir(parents=True)
|
||||
litellm = venv_bin / "litellm"
|
||||
python.write_text("#!/bin/sh\n")
|
||||
litellm.write_text("#!/bin/sh\n")
|
||||
python.chmod(0o755)
|
||||
litellm.chmod(0o755)
|
||||
monkeypatch.setenv("VIRTUAL_ENV", str(tmp_path / "venv"))
|
||||
monkeypatch.delenv("PATH", raising=False)
|
||||
config = tmp_path / "litellm.yaml"
|
||||
config.write_text("model_list: []\n")
|
||||
argv = litellm_proxy_argv(
|
||||
config=config,
|
||||
host="127.0.0.1",
|
||||
port=4010,
|
||||
python_executable=str(python),
|
||||
)
|
||||
assert argv[0] == str(litellm.resolve())
|
||||
assert "-m" not in argv
|
||||
assert argv[1:] == ["--config", str(config), "--host", "127.0.0.1", "--port", "4010"]
|
||||
|
||||
|
||||
def test_anthropic_api_key_prefers_the_named_env_and_keeps_the_legacy_alias(monkeypatch) -> None:
|
||||
monkeypatch.delenv("GITNEXUS_BENCH_ANTHROPIC_API_KEY", raising=False)
|
||||
monkeypatch.setenv("GITNEXUS_BENCH_AUTH_TOKEN", "legacy-secret")
|
||||
assert anthropic_api_key_from_environ() == "legacy-secret"
|
||||
monkeypatch.setenv("GITNEXUS_BENCH_ANTHROPIC_API_KEY", "named-secret")
|
||||
assert anthropic_api_key_from_environ() == "named-secret"
|
||||
|
|
@ -5,7 +5,9 @@ from __future__ import annotations
|
|||
import argparse
|
||||
import hashlib
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
|
|
@ -13,7 +15,13 @@ import pytest
|
|||
|
||||
from workflow_bench import oracle_assets, runner
|
||||
from workflow_bench.evolution import evaluate_candidate
|
||||
from workflow_bench.oracle_assets import capture_task_oracle, staged_task_oracle
|
||||
from workflow_bench.oracle_assets import (
|
||||
capture_task_oracle,
|
||||
require_hidden_harness_absent,
|
||||
review_case_setup_command,
|
||||
staged_task_oracle,
|
||||
with_hidden_harness_apply_exclude,
|
||||
)
|
||||
|
||||
|
||||
def oracle_task(*, command: str = "true", source: str = "oracle.test.ts") -> dict[str, object]:
|
||||
|
|
@ -52,6 +60,7 @@ def bench_args() -> argparse.Namespace:
|
|||
claude_bin="claude",
|
||||
timeout=5,
|
||||
model="pinned-model",
|
||||
effort="xhigh",
|
||||
base_url=None,
|
||||
auth_token=None,
|
||||
)
|
||||
|
|
@ -148,6 +157,73 @@ def test_clone_sanitization_prunes_harness_checkout_and_recoverable_history(tmp_
|
|||
assert git(clone, "status", "--porcelain=v1", "--untracked-files=all").stdout == ""
|
||||
|
||||
|
||||
def test_require_hidden_harness_absent_fails_closed_on_leftover_tree(tmp_path: Path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
hidden = clone / "eval" / "workflow_bench"
|
||||
hidden.mkdir(parents=True)
|
||||
(hidden / "review_cases").mkdir()
|
||||
|
||||
with pytest.raises(ValueError, match="hidden harness visible"):
|
||||
require_hidden_harness_absent(clone)
|
||||
|
||||
shutil.rmtree(hidden)
|
||||
require_hidden_harness_absent(clone)
|
||||
|
||||
|
||||
def test_hidden_harness_apply_exclude_is_idempotent() -> None:
|
||||
raw = "git apply eval/workflow_bench/review_cases/pr.patch && rm -rf eval/workflow_bench"
|
||||
once = with_hidden_harness_apply_exclude(raw)
|
||||
assert once == review_case_setup_command("pr.patch")
|
||||
assert with_hidden_harness_apply_exclude(once) == once
|
||||
assert with_hidden_harness_apply_exclude("true") == "true"
|
||||
|
||||
|
||||
def test_review_setup_skips_sanitized_harness_hunks(tmp_path: Path) -> None:
|
||||
repo = tmp_path / "repo"
|
||||
repo.mkdir()
|
||||
|
||||
def git(*args: str, check: bool = True) -> subprocess.CompletedProcess[str]:
|
||||
result = subprocess.run(
|
||||
["git", "-C", str(repo), *args],
|
||||
check=False,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
)
|
||||
if check and result.returncode != 0:
|
||||
pytest.fail(f"git {' '.join(args)} failed: {result.stderr}")
|
||||
return result
|
||||
|
||||
git("init", "--quiet", "--initial-branch=main")
|
||||
git("config", "user.name", "Review Setup")
|
||||
git("config", "user.email", "review-setup.invalid")
|
||||
(repo / "visible.py").write_text("old\n")
|
||||
hidden = repo / "eval" / "workflow_bench"
|
||||
hidden.mkdir(parents=True)
|
||||
(hidden / "learnings.jsonl").write_text("{}\n")
|
||||
git("add", "--all")
|
||||
git("commit", "--quiet", "-m", "base with harness file")
|
||||
(repo / "visible.py").write_text("new\n")
|
||||
(hidden / "learnings.jsonl").write_text("{}\nextra\n")
|
||||
patch = git("diff").stdout
|
||||
git("checkout", "--", ".")
|
||||
shutil.rmtree(hidden)
|
||||
patch_path = hidden / "review_cases" / "case.patch"
|
||||
patch_path.parent.mkdir(parents=True)
|
||||
patch_path.write_text(patch)
|
||||
|
||||
rejected = git("apply", "--check", str(patch_path.relative_to(repo)), check=False)
|
||||
assert rejected.returncode != 0
|
||||
assert "learnings.jsonl" in rejected.stderr
|
||||
|
||||
setup = with_hidden_harness_apply_exclude(
|
||||
"git apply eval/workflow_bench/review_cases/case.patch && rm -rf eval/workflow_bench"
|
||||
)
|
||||
applied = subprocess.run(["/bin/sh", "-lc", setup], cwd=repo, check=False, capture_output=True, text=True)
|
||||
assert applied.returncode == 0, applied.stderr
|
||||
assert (repo / "visible.py").read_text() == "new\n"
|
||||
assert not hidden.exists()
|
||||
|
||||
|
||||
def test_clone_sanitization_prunes_remote_history_when_head_never_had_harness(tmp_path: Path) -> None:
|
||||
source = tmp_path / "source"
|
||||
source.mkdir()
|
||||
|
|
@ -299,6 +375,26 @@ def test_vacuous_authored_test_cannot_self_certify_resolution(
|
|||
assert record["error_kind"] == "oracle-failed"
|
||||
|
||||
|
||||
def test_host_unsafe_oracle_executes_beside_candidate_and_is_removed(tmp_path: Path) -> None:
|
||||
from workflow_bench.proposer_sandbox import prepare_sandbox
|
||||
|
||||
source = tmp_path / "oracles"
|
||||
write_oracle(
|
||||
source,
|
||||
b"from pathlib import Path\nassert (Path(__file__).resolve().parents[2] / 'candidate.txt').read_text() == 'candidate'\n",
|
||||
)
|
||||
snapshot = capture_task_oracle(
|
||||
oracle_task(command='python3 "$GITNEXUS_BENCH_ORACLE_ROOT/nested/oracle.test.ts"'), root=source
|
||||
)
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
(clone / "candidate.txt").write_text("candidate")
|
||||
with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as session:
|
||||
passed, detail = runner._run_hidden_oracle(snapshot, clone, bench_args(), session)
|
||||
assert passed, detail
|
||||
assert not list(clone.glob(".wfbench-oracle-*"))
|
||||
|
||||
|
||||
def test_oracle_path_and_bytes_appear_only_after_the_model_session(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
tmp_path: Path,
|
||||
|
|
|
|||
|
|
@ -7,7 +7,12 @@ import os
|
|||
import signal
|
||||
import sys
|
||||
import time
|
||||
import threading
|
||||
import subprocess
|
||||
import json
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
|
|
@ -23,6 +28,100 @@ from workflow_bench.process_control import (
|
|||
PYTHON = sys.executable
|
||||
|
||||
|
||||
def test_cancel_before_spawn_never_starts_the_child(tmp_path):
|
||||
event = threading.Event()
|
||||
event.set()
|
||||
sentinel = tmp_path / "spawned"
|
||||
result = run_managed([PYTHON, "-c", f"open({str(sentinel)!r}, 'w').close()"], timeout=60, cancel_event=event)
|
||||
assert result.state == "cancelled"
|
||||
assert not sentinel.exists()
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="POSIX signal escalation")
|
||||
def test_cancel_after_spawn_kills_a_term_ignoring_child():
|
||||
event = threading.Event()
|
||||
started = time.monotonic()
|
||||
result = run_managed(
|
||||
[
|
||||
PYTHON,
|
||||
"-c",
|
||||
"import signal,time; signal.signal(signal.SIGTERM, signal.SIG_IGN); print('ready',flush=True); time.sleep(60)",
|
||||
],
|
||||
timeout=60,
|
||||
terminate_grace=0.1,
|
||||
cancel_event=event,
|
||||
stdout_observer=lambda chunk: event.set() if b"ready" in chunk else None,
|
||||
)
|
||||
assert result.state == "cancelled" and result.forced_kill
|
||||
assert not result.timed_out
|
||||
assert time.monotonic() - started < 5
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="POSIX parent signals")
|
||||
@pytest.mark.parametrize("signum", [signal.SIGINT, signal.SIGTERM])
|
||||
def test_signal_cancels_active_wave_before_assets_are_released(tmp_path, signum):
|
||||
ready = tmp_path / "child-pid"
|
||||
assets = tmp_path / "assets"
|
||||
assets.write_text("shared")
|
||||
command = (
|
||||
f"import os,time; from pathlib import Path; Path({str(ready)!r}).write_text(str(os.getpid())); time.sleep(60)"
|
||||
)
|
||||
script = f"""
|
||||
import json,sys
|
||||
from pathlib import Path
|
||||
from workflow_bench.process_control import cancellation_scope, run_managed
|
||||
from workflow_bench.runner import sweep_task_cells
|
||||
rows=[]
|
||||
def run(index, arm):
|
||||
if index == 0:
|
||||
return {{'resolved': True, 'error_kind': None}}
|
||||
result=run_managed([sys.executable, '-c', {command!r}], timeout=60)
|
||||
assert Path({str(assets)!r}).exists(), 'assets removed while a worker was active'
|
||||
return {{'resolved': False, 'error_kind': result.state}}
|
||||
with cancellation_scope(handle_signals=True) as event:
|
||||
streak, stopped=sweep_task_cells([(i, 'review') for i in range(10)], workers=2, run=run,
|
||||
on_start=lambda *args: None, on_record=lambda i,a,r: rows.append([i,r]),
|
||||
outage_streak=0, outage_limit=5, cancel_event=event)
|
||||
Path({str(assets)!r}).unlink()
|
||||
print(json.dumps({{'rows': rows, 'stopped': stopped}}))
|
||||
"""
|
||||
process = subprocess.Popen(
|
||||
[PYTHON, "-c", script],
|
||||
cwd=Path(__file__).resolve().parents[1],
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
text=True,
|
||||
)
|
||||
try:
|
||||
deadline = time.monotonic() + 10
|
||||
while not ready.exists() and process.poll() is None and time.monotonic() < deadline:
|
||||
time.sleep(0.01)
|
||||
assert ready.exists()
|
||||
child_pid = int(ready.read_text())
|
||||
started = time.monotonic()
|
||||
process.send_signal(signum)
|
||||
stdout, stderr = process.communicate(timeout=15)
|
||||
assert process.returncode == 0, stderr
|
||||
assert time.monotonic() - started < 15
|
||||
report = json.loads(stdout)
|
||||
assert report["stopped"] and [row[0] for row in report["rows"]] == [0, 1]
|
||||
assert report["rows"][0][1]["resolved"] is True
|
||||
assert report["rows"][1][1]["error_kind"] == "cancelled"
|
||||
with pytest.raises(ProcessLookupError):
|
||||
os.kill(child_pid, 0)
|
||||
assert not assets.exists()
|
||||
finally:
|
||||
if process.poll() is None:
|
||||
process.kill()
|
||||
process.wait(timeout=5)
|
||||
if ready.exists():
|
||||
try:
|
||||
os.killpg(int(ready.read_text()), signal.SIGKILL)
|
||||
except ProcessLookupError:
|
||||
# Successful cancellation has already reaped this process group.
|
||||
pass
|
||||
|
||||
|
||||
def test_managed_process_captures_normal_exit() -> None:
|
||||
result = run_managed(
|
||||
[PYTHON, "-c", "import sys; print('out'); print('err', file=sys.stderr)"],
|
||||
|
|
@ -124,6 +223,64 @@ def test_parent_stdout_capture_reports_overflow_without_stopping_drain() -> None
|
|||
assert result.stdout_tail.endswith("END")
|
||||
|
||||
|
||||
def test_echo_stdout_streams_child_progress_and_stays_off_by_default(capfd) -> None:
|
||||
command = [PYTHON, "-c", "import os; os.write(1, b'[task][arm][run 0] starting\\n')"]
|
||||
|
||||
quiet = run_managed(command, timeout=5)
|
||||
assert quiet.ok
|
||||
assert "starting" not in capfd.readouterr().err
|
||||
|
||||
echoed = run_managed(command, timeout=5, echo_stdout=True)
|
||||
assert echoed.ok
|
||||
captured = capfd.readouterr()
|
||||
assert "[task][arm][run 0] starting" in captured.err
|
||||
# Echoing is a passthrough, not a redirect: the tail stays intact for the
|
||||
# caller that reports it after the process ends.
|
||||
assert "starting" in echoed.stdout_tail
|
||||
|
||||
|
||||
def test_echo_reaches_the_log_while_the_child_is_still_running(tmp_path: Path, monkeypatch) -> None:
|
||||
"""Streaming has to be prompt, not merely eventual.
|
||||
|
||||
A sweep emits one line every ~45 minutes. Draining with `read(8192)` still
|
||||
delivers every byte, so the tail and the capture look correct — but nothing
|
||||
surfaces until the pipe closes, which turns a 15-hour job into a silent one
|
||||
and is the whole reason this passthrough exists.
|
||||
|
||||
The child here refuses to exit until the echoed line has been observed, so
|
||||
an implementation that only flushes at EOF deadlocks and fails on the
|
||||
timeout rather than passing on a technicality.
|
||||
"""
|
||||
released = tmp_path / "echo-observed"
|
||||
|
||||
class Sink:
|
||||
def write(self, data: bytes) -> int:
|
||||
if b"first-line" in data:
|
||||
released.write_text("go")
|
||||
return len(data)
|
||||
|
||||
def flush(self) -> None:
|
||||
pass
|
||||
|
||||
monkeypatch.setattr(process_control.sys, "stderr", SimpleNamespace(buffer=Sink()))
|
||||
script = """
|
||||
import pathlib, sys, time
|
||||
sys.stdout.write('first-line\\n')
|
||||
sys.stdout.flush()
|
||||
target = pathlib.Path(%r)
|
||||
for _ in range(400):
|
||||
if target.exists():
|
||||
break
|
||||
time.sleep(0.05)
|
||||
""" % str(released)
|
||||
|
||||
result = run_managed([PYTHON, "-c", script], timeout=15, echo_stdout=True)
|
||||
|
||||
assert released.exists(), "the line never reached the echo sink while the child ran"
|
||||
assert result.ok
|
||||
assert "first-line" in result.stdout_tail
|
||||
|
||||
|
||||
def test_incomplete_stdin_delivery_cannot_report_success() -> None:
|
||||
result = run_managed(
|
||||
[PYTHON, "-c", "import os,time; os.close(0); time.sleep(0.05)"],
|
||||
|
|
@ -453,3 +610,57 @@ def test_windows_normal_parent_with_grandchild_is_not_successful_evidence(tmp_pa
|
|||
assert result.forced_kill
|
||||
assert not result.ok
|
||||
assert not sentinel.exists()
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="POSIX process-group ownership canary")
|
||||
def test_concurrent_cells_reap_only_their_own_process_tree(tmp_path: Path) -> None:
|
||||
"""One cell timing out must not touch a sibling cell running beside it.
|
||||
|
||||
`run_managed` reaps by process group. Cells only ever ran one at a time
|
||||
before, so nothing exercised what happens when a `killpg` fires while other
|
||||
owned trees are alive — a leaked or shared pgid would take the siblings
|
||||
down with it, and the sweep would read that as two more excluded runs.
|
||||
"""
|
||||
survivor_sentinel = tmp_path / "survivor-finished"
|
||||
victim_sentinel = tmp_path / "victim-escaped"
|
||||
|
||||
# Each cell spawns a descendant, like a sandboxed session does.
|
||||
survivor = """
|
||||
import pathlib, subprocess, sys, time
|
||||
child = subprocess.Popen([sys.executable, '-c', "import time; time.sleep(2)"])
|
||||
time.sleep(1.0)
|
||||
pathlib.Path(%r).write_text('finished')
|
||||
child.wait()
|
||||
print('survivor-done', flush=True)
|
||||
""" % str(survivor_sentinel)
|
||||
victim = """
|
||||
import pathlib, signal, subprocess, sys, time
|
||||
signal.signal(signal.SIGTERM, signal.SIG_IGN)
|
||||
subprocess.Popen([
|
||||
sys.executable, '-c',
|
||||
"import signal,time,pathlib; signal.signal(signal.SIGTERM, signal.SIG_IGN); time.sleep(1.5); pathlib.Path(%r).write_text('escaped')"
|
||||
])
|
||||
while True:
|
||||
time.sleep(0.01)
|
||||
""" % str(victim_sentinel)
|
||||
|
||||
def cell(source: str, timeout: float):
|
||||
return run_managed([PYTHON, "-c", source], timeout=timeout, terminate_grace=0.1)
|
||||
|
||||
with ThreadPoolExecutor(max_workers=3) as pool:
|
||||
futures = [
|
||||
pool.submit(cell, survivor, 10.0),
|
||||
pool.submit(cell, victim, 0.2),
|
||||
pool.submit(cell, survivor, 10.0),
|
||||
]
|
||||
first, doomed, second = (future.result() for future in futures)
|
||||
|
||||
time.sleep(1.8)
|
||||
|
||||
assert doomed.state == "forced-kill"
|
||||
assert not victim_sentinel.exists(), "the timed-out cell leaked a descendant"
|
||||
# The siblings were mid-flight when the killpg fired.
|
||||
assert first.ok and second.ok
|
||||
assert "survivor-done" in first.stdout_tail
|
||||
assert "survivor-done" in second.stdout_tail
|
||||
assert survivor_sentinel.exists()
|
||||
|
|
|
|||
|
|
@ -3,11 +3,15 @@
|
|||
import json
|
||||
import os
|
||||
import stat
|
||||
import shlex
|
||||
import subprocess
|
||||
from pathlib import Path, PurePosixPath
|
||||
|
||||
import pytest
|
||||
import yaml
|
||||
|
||||
from workflow_bench import evolve, promotion_apply
|
||||
from workflow_bench import evolution, runner
|
||||
from workflow_bench.evolution import (
|
||||
CANDIDATE_SKILLS,
|
||||
MAX_CANDIDATE_OVERLAY_BYTES,
|
||||
|
|
@ -36,6 +40,120 @@ def _git(repo: Path, *arguments: str) -> str:
|
|||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"arms", [["candidate_review"], ["candidate_workflow"], ["candidate_review", "candidate_workflow"]]
|
||||
)
|
||||
@pytest.mark.parametrize(
|
||||
"tamper",
|
||||
[
|
||||
None,
|
||||
"aborted",
|
||||
"metric",
|
||||
"verdict",
|
||||
"nan",
|
||||
"huge-int",
|
||||
"runs",
|
||||
"task",
|
||||
"duplicate-task",
|
||||
"exclusions",
|
||||
"old-schema",
|
||||
],
|
||||
)
|
||||
def test_produced_promotion_round_trips_to_transactional_apply(tmp_path, arms, tamper):
|
||||
from tests.test_evolve import bound_task_fixture, promotion_fixture
|
||||
|
||||
overlay = tmp_path / "overlay"
|
||||
repo = tmp_path / "repo"
|
||||
for arm in arms:
|
||||
skill = "gitnexus-review" if arm == "candidate_review" else "gitnexus-plan"
|
||||
relative = PurePosixPath(f".claude/skills/{skill}/SKILL.md")
|
||||
(overlay / relative).parent.mkdir(parents=True, exist_ok=True)
|
||||
(overlay / relative).write_text("candidate")
|
||||
for target in mirror_targets(relative):
|
||||
(repo / target).parent.mkdir(parents=True, exist_ok=True)
|
||||
(repo / target).write_text("incumbent")
|
||||
frozen = tmp_path / "frozen"
|
||||
digest = freeze_overlay(overlay, frozen)
|
||||
bases = destination_base_digests(frozen, repo_root=repo)
|
||||
paired = {}
|
||||
for arm in arms:
|
||||
for name, candidate in ((evolution.CANDIDATE_ARMS[arm], False), (arm, True)):
|
||||
paired[name] = runner.aggregate(
|
||||
[
|
||||
{
|
||||
"resolved": True,
|
||||
"cost_usd": 0.8 if candidate else 1.0,
|
||||
"review_weighted_f1": 1.0 if candidate else 0.5,
|
||||
"review_blocker_recall": 1.0,
|
||||
"review_false_positives": 0,
|
||||
"review_verdict_correct": True,
|
||||
"review_clean_control": False,
|
||||
"review_clean_pass": False,
|
||||
}
|
||||
for _ in range(3)
|
||||
]
|
||||
)
|
||||
policy = evolution.promotion_policy(arms)
|
||||
promotion = {
|
||||
**promotion_fixture(),
|
||||
**evolution.promotion_evidence(
|
||||
{"task-a": paired},
|
||||
policy=policy,
|
||||
model="bench-model",
|
||||
complete=True,
|
||||
),
|
||||
"required_candidate_arms": arms,
|
||||
"candidate_overlay_digest": digest,
|
||||
"target_base_digests": bases,
|
||||
"selected_tasks": [bound_task_fixture()],
|
||||
}
|
||||
promotion = json.loads(json.dumps(promotion))
|
||||
decision = promotion["decisions"][0]
|
||||
row = decision["tasks"][0]
|
||||
if tamper == "aborted":
|
||||
promotion["run_status"] = "aborted"
|
||||
elif tamper == "metric":
|
||||
decision["metric"] = "fabricated"
|
||||
elif tamper == "verdict":
|
||||
decision["decision"] = "keep_incumbent"
|
||||
elif tamper == "nan":
|
||||
row["candidate"][policy[arms[0]]["metric"]] = float("nan")
|
||||
elif tamper == "huge-int":
|
||||
row["candidate"][policy[arms[0]]["metric"]] = 10**1000
|
||||
elif tamper == "runs":
|
||||
row["candidate"]["valid_runs"] = 1
|
||||
elif tamper == "task":
|
||||
row["task"] = "unselected"
|
||||
elif tamper == "duplicate-task":
|
||||
decision["tasks"].append(dict(row))
|
||||
elif tamper == "exclusions":
|
||||
row["candidate"]["excluded_runs"] = 1
|
||||
elif tamper == "old-schema":
|
||||
promotion["schema_version"] = 5
|
||||
|
||||
def validate():
|
||||
return evolve.validate_promotion_for_apply(
|
||||
promotion,
|
||||
overlay_digest=digest,
|
||||
benchmark_model="bench-model",
|
||||
proposer_model="proposer-model",
|
||||
effort="xhigh",
|
||||
selected_tasks=[bound_task_fixture()],
|
||||
target_base_digests=bases,
|
||||
required_candidate_arms=arms,
|
||||
policy=policy,
|
||||
)
|
||||
|
||||
if tamper:
|
||||
with pytest.raises(ValueError):
|
||||
validate()
|
||||
assert all((repo / path).read_text() == "incumbent" for path in bases)
|
||||
else:
|
||||
assert all(decision["decision"] == "promote" for decision in validate())
|
||||
written = apply_promoted_overlay(frozen, repo_root=repo, expected_target_bases=bases)
|
||||
assert all((repo / path).read_text() == "candidate" for path in written)
|
||||
|
||||
|
||||
def test_evolve_reexports_public_promotion_helpers():
|
||||
assert evolve.mirror_targets is mirror_targets
|
||||
assert evolve.freeze_overlay is freeze_overlay
|
||||
|
|
@ -52,6 +170,44 @@ def test_mirror_targets_cover_canonical_and_shipped_copies():
|
|||
]
|
||||
|
||||
|
||||
def test_review_mirror_targets_include_cursor_distribution():
|
||||
targets = mirror_targets(PurePosixPath(".claude/skills/gitnexus-review/SKILL.md"))
|
||||
assert targets == [
|
||||
PurePosixPath(".claude/skills/gitnexus-review/SKILL.md"),
|
||||
PurePosixPath("gitnexus/skills/gitnexus-review/SKILL.md"),
|
||||
PurePosixPath("gitnexus-claude-plugin/skills/gitnexus-review/SKILL.md"),
|
||||
PurePosixPath("gitnexus-cursor-integration/skills/gitnexus-review/SKILL.md"),
|
||||
]
|
||||
|
||||
|
||||
def test_workflow_stages_every_review_mirror_and_rejects_unrelated_changes(tmp_path):
|
||||
overlay = tmp_path / "overlay"
|
||||
relative = PurePosixPath(".claude/skills/gitnexus-review/SKILL.md")
|
||||
(overlay / relative).parent.mkdir(parents=True)
|
||||
(overlay / relative).write_text("candidate")
|
||||
repo = tmp_path / "repo"
|
||||
targets = mirror_targets(relative)
|
||||
for path in targets:
|
||||
(repo / path).parent.mkdir(parents=True, exist_ok=True)
|
||||
(repo / path).write_text("incumbent")
|
||||
(repo / "unrelated.txt").write_text("before")
|
||||
_git(repo, "init", "-q")
|
||||
_git(repo, "add", ".")
|
||||
_git(repo, "-c", "user.name=Fixture", "-c", "user.email=fixture@example.test", "commit", "-qm", "base")
|
||||
apply_promoted_overlay(overlay, repo_root=repo)
|
||||
workflow_path = Path(__file__).resolve().parents[2] / ".github/workflows/gitnexus-skill-evolution.yml"
|
||||
steps = yaml.safe_load(workflow_path.read_text())["jobs"]["evolve"]["steps"]
|
||||
publish = next(step["run"] for step in steps if step.get("name") == "Open the promotion PR")
|
||||
staging = next(line for line in publish.splitlines() if line.startswith("git add "))
|
||||
_git(repo, *shlex.split(staging)[1:])
|
||||
assert _git(repo, "diff", "--cached", "--name-only").splitlines() == sorted(map(str, targets))
|
||||
assert {(repo / path).read_text() for path in targets} == {"candidate"}
|
||||
(repo / "unrelated.txt").write_text("after")
|
||||
guard = next(step["run"] for step in steps if step.get("name") == "Detect and bound the applied promotion")
|
||||
guarded = subprocess.run(["bash", "-c", guard], cwd=repo, capture_output=True, text=True)
|
||||
assert guarded.returncode == 1 and "outside the skill trees: unrelated.txt" in guarded.stdout
|
||||
|
||||
|
||||
def test_apply_promoted_overlay_writes_all_mirrors(tmp_path):
|
||||
overlay = tmp_path / "overlay"
|
||||
skill_md = overlay / ".claude" / "skills" / "gitnexus-plan" / "SKILL.md"
|
||||
|
|
@ -492,14 +648,7 @@ def test_committed_destination_bases_ignore_and_reject_live_target_edits(tmp_pat
|
|||
assert dirty.read_text() == "user edit"
|
||||
|
||||
|
||||
def test_mirror_roots_cover_every_candidate_skill_and_omit_none_that_ships_to_cursor():
|
||||
# promotion_apply.mirror_targets writes canonical + MIRROR_SKILL_ROOTS, which
|
||||
# today omits the Cursor tree. That is only safe because no candidate skill is
|
||||
# cursor-shipped. If a future edit adds a cursor-shipped skill (e.g.
|
||||
# gitnexus-review) to CANDIDATE_SKILLS, apply_promoted_overlay would rewrite
|
||||
# the other trees and silently skip Cursor — the PR #2488 asymmetric-sync bug
|
||||
# class. Pin the invariant to the filesystem, the source of truth the TS drift
|
||||
# guard already enforces.
|
||||
def test_mirror_roots_cover_every_candidate_skill_including_cursor_review():
|
||||
repo_root = Path(__file__).resolve().parents[2]
|
||||
cursor_root = repo_root / "gitnexus-cursor-integration" / "skills"
|
||||
for skill in sorted(CANDIDATE_SKILLS):
|
||||
|
|
@ -507,10 +656,9 @@ def test_mirror_roots_cover_every_candidate_skill_and_omit_none_that_ships_to_cu
|
|||
assert canonical.is_dir(), f"candidate skill {skill} has no canonical .claude/skills dir"
|
||||
for target in mirror_targets(PurePosixPath(".claude", "skills", skill, "SKILL.md")):
|
||||
assert (repo_root / target).is_file(), f"candidate skill mirror missing on disk: {target}"
|
||||
assert not (cursor_root / skill).exists(), (
|
||||
f"candidate skill {skill} ships to Cursor, but MIRROR_SKILL_ROOTS does not cover "
|
||||
"gitnexus-cursor-integration/skills — promotion would sync it asymmetrically"
|
||||
)
|
||||
if (cursor_root / skill).exists():
|
||||
expected = PurePosixPath("gitnexus-cursor-integration/skills", skill, "SKILL.md")
|
||||
assert expected in mirror_targets(PurePosixPath(".claude", "skills", skill, "SKILL.md"))
|
||||
|
||||
|
||||
def test_committed_destination_bases_reject_overlay_adding_uncommitted_target(tmp_path):
|
||||
|
|
|
|||
|
|
@ -9,19 +9,28 @@ import stat
|
|||
import subprocess
|
||||
import sys
|
||||
import threading
|
||||
from dataclasses import replace
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
from workflow_bench import runner
|
||||
from workflow_bench import runner, runner_artifacts
|
||||
from workflow_bench import proposer_sandbox
|
||||
|
||||
from workflow_bench.process_control import ManagedProcessResult, run_managed
|
||||
|
||||
|
||||
from workflow_bench.proposer_sandbox import (
|
||||
MAX_BUNDLE_BYTES,
|
||||
MAX_EVIDENCE_FILE_BYTES,
|
||||
SANDBOX_CLAUDE,
|
||||
SANDBOX_NODE,
|
||||
SANDBOX_NODE_PREFIX,
|
||||
SANDBOX_EVIDENCE,
|
||||
SANDBOX_GITNEXUS_CLI,
|
||||
SANDBOX_GIT_EXCLUDES,
|
||||
VITE_TEMP_DIR,
|
||||
SANDBOX_PATH,
|
||||
SANDBOX_PYTHON3,
|
||||
|
|
@ -32,14 +41,77 @@ from workflow_bench.proposer_sandbox import (
|
|||
_runtime_mount_args,
|
||||
build_claude_settings,
|
||||
build_sandbox_environment,
|
||||
_force_rmtree,
|
||||
host_workspace_write_boundary,
|
||||
prepare_sandbox,
|
||||
preflight_bubblewrap,
|
||||
sandbox_workspace_write_boundary,
|
||||
stage_evidence_bundle,
|
||||
stage_task_assets,
|
||||
)
|
||||
from workflow_bench.task_assets import TaskAssetCache, stage_task_assets as stage_immutable_task_assets
|
||||
|
||||
|
||||
@pytest.mark.parametrize("entry", ["file", "directory", "relative-link", "absolute-link"])
|
||||
def test_review_preparation_rejects_existing_output_without_touching_target(tmp_path, entry):
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
sentinel = tmp_path / "sentinel"
|
||||
sentinel.write_text("must survive")
|
||||
output = clone / "review-output.json"
|
||||
if entry == "file":
|
||||
output.write_text("existing result")
|
||||
elif entry == "directory":
|
||||
output.mkdir()
|
||||
else:
|
||||
output.symlink_to(sentinel if entry == "absolute-link" else "../sentinel")
|
||||
with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as sandbox:
|
||||
with pytest.raises(SandboxError, match="already exists"):
|
||||
proposer_sandbox.prepare_review_workspace(sandbox, "review-output.json")
|
||||
assert sentinel.read_text() == "must survive"
|
||||
if entry == "file":
|
||||
assert output.read_text() == "existing result"
|
||||
if "link" in entry:
|
||||
assert output.is_symlink()
|
||||
|
||||
|
||||
def test_review_preparation_creates_a_private_regular_output(tmp_path):
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as sandbox:
|
||||
output = proposer_sandbox.prepare_review_workspace(sandbox, "review-output.json")
|
||||
assert output.read_bytes() == b""
|
||||
assert stat.S_ISREG(output.lstat().st_mode)
|
||||
assert stat.S_IMODE(output.stat().st_mode) == 0o600
|
||||
|
||||
|
||||
def test_review_preparation_preserves_existing_runtime_files_and_tracks_only_created_paths(tmp_path):
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
(clone / "bunfig.toml").write_text("existing configuration\n")
|
||||
with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as sandbox:
|
||||
proposer_sandbox.prepare_review_workspace(replace(sandbox, backend="bwrap"), "review-output.json")
|
||||
created = json.loads((sandbox.private_root / "review-created-paths.json").read_text())
|
||||
assert "bunfig.toml" not in created
|
||||
assert ".npmrc" in created
|
||||
assert ".mcp.json" in created
|
||||
assert json.loads((clone / ".mcp.json").read_text()) == {}
|
||||
assert (clone / "bunfig.toml").read_text() == "existing configuration\n"
|
||||
assert (clone / ".git/commondir").read_text() == ".\n"
|
||||
|
||||
|
||||
def test_review_preparation_rejects_a_runtime_symlink_parent(tmp_path):
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
outside = tmp_path / "outside"
|
||||
outside.mkdir()
|
||||
(clone / ".claude").symlink_to(outside, target_is_directory=True)
|
||||
with prepare_sandbox(clone=clone, claude_bin=sys.executable, backend="host-unsafe") as sandbox:
|
||||
with pytest.raises(SandboxError):
|
||||
proposer_sandbox.prepare_review_workspace(replace(sandbox, backend="bwrap"), "review-output.json")
|
||||
assert list(outside.iterdir()) == []
|
||||
|
||||
|
||||
def test_environment_is_allowlisted_and_shell_children_are_credential_free(monkeypatch) -> None:
|
||||
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "cloud-secret")
|
||||
monkeypatch.setenv("GITHUB_TOKEN", "github-secret")
|
||||
|
|
@ -65,6 +137,11 @@ def test_environment_is_allowlisted_and_shell_children_are_credential_free(monke
|
|||
assert settings["sandbox"]["failIfUnavailable"] is True
|
||||
assert settings["sandbox"]["allowUnsandboxedCommands"] is False
|
||||
assert settings["sandbox"]["network"]["deniedDomains"] == ["*"]
|
||||
assert SANDBOX_EVIDENCE in settings["sandbox"]["filesystem"]["allowRead"]
|
||||
# Headless `claude -p` (2.1.247) never dispatches PreToolUse from any
|
||||
# settings source, so a hook here would be confinement theater: it would
|
||||
# read as a control in review while enforcing nothing at runtime.
|
||||
assert "hooks" not in settings
|
||||
# ENV_SCRUB forces "default" mode; the proposer's tools (Bash writes the
|
||||
# overlay) run headless only because they are explicitly pre-approved.
|
||||
# Requesting a non-default defaultMode would merely warn, so it must be gone.
|
||||
|
|
@ -72,6 +149,160 @@ def test_environment_is_allowlisted_and_shell_children_are_credential_free(monke
|
|||
assert "defaultMode" not in settings["permissions"]
|
||||
|
||||
|
||||
def test_unsafe_host_session_translates_virtual_paths_and_disables_containment(tmp_path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
evidence = tmp_path / "evidence"
|
||||
for directory in (clone, evidence):
|
||||
directory.mkdir()
|
||||
|
||||
with prepare_sandbox(
|
||||
clone=clone,
|
||||
backend="host-unsafe",
|
||||
claude_bin=sys.executable,
|
||||
read_only_mounts=(ReadOnlyMount(evidence, "/evidence"),),
|
||||
) as sandbox:
|
||||
assert sandbox.command_prefix == []
|
||||
assert sandbox.require_pid_namespace is False
|
||||
assert sandbox.host_path("/workspace/review-output.json") == str(clone / "review-output.json")
|
||||
assert sandbox.host_path("/evidence/selected-rows.json") == str(evidence / "selected-rows.json")
|
||||
assert sandbox.host_text("read /evidence and write /workspace/out") == (
|
||||
f"read {evidence} and write {clone}/out"
|
||||
)
|
||||
# Sessions spawn the binary directly, so it must be the host executable
|
||||
# rather than the sandbox-only mount target.
|
||||
assert sandbox.claude_bin != SANDBOX_CLAUDE
|
||||
assert Path(sandbox.claude_bin).exists()
|
||||
assert sandbox.environment()["HOME"] == str(sandbox.home)
|
||||
assert "CLAUDE_CODE_SUBPROCESS_ENV_SCRUB" not in sandbox.environment()
|
||||
assert all(
|
||||
Path(entry).is_dir() for entry in sandbox.environment()["PATH"].split(":")
|
||||
)
|
||||
unsafe_settings = json.loads(sandbox.settings_json)
|
||||
assert unsafe_settings["sandbox"]["enabled"] is False
|
||||
assert unsafe_settings["sandbox"]["failIfUnavailable"] is False
|
||||
assert "disableBypassPermissionsMode" not in unsafe_settings["permissions"]
|
||||
|
||||
|
||||
def test_host_workspace_write_boundary_keeps_only_the_review_artifact_writable(tmp_path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
nested = clone / "src"
|
||||
nested.mkdir(parents=True)
|
||||
source = nested / "source.ts"
|
||||
source.write_text("trusted\n")
|
||||
output = clone / "review-output.json"
|
||||
output.write_text("")
|
||||
original_source_mode = stat.S_IMODE(source.stat().st_mode)
|
||||
original_output_mode = stat.S_IMODE(output.stat().st_mode)
|
||||
|
||||
with host_workspace_write_boundary(clone, writable=(output,)):
|
||||
with pytest.raises(OSError):
|
||||
source.write_text("tampered\n")
|
||||
with pytest.raises(OSError):
|
||||
(clone / "extra.py").write_text("nope\n")
|
||||
output.write_text('{"schema_version":1}\n')
|
||||
|
||||
assert source.read_text() == "trusted\n"
|
||||
assert output.read_text() == '{"schema_version":1}\n'
|
||||
assert not (clone / "extra.py").exists()
|
||||
assert stat.S_IMODE(source.stat().st_mode) == original_source_mode
|
||||
assert stat.S_IMODE(output.stat().st_mode) == original_output_mode
|
||||
|
||||
|
||||
def test_host_workspace_write_boundary_keeps_files_under_an_allowed_directory(tmp_path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
artifacts = clone / "artifacts"
|
||||
artifacts.mkdir(parents=True)
|
||||
existing = artifacts / "review-output.json"
|
||||
existing.write_text("{}\n")
|
||||
(clone / "src").mkdir()
|
||||
locked = clone / "src" / "source.ts"
|
||||
locked.write_text("trusted\n")
|
||||
|
||||
with host_workspace_write_boundary(clone, writable=(artifacts,)):
|
||||
existing.write_text('{"schema_version":1}\n')
|
||||
with pytest.raises(OSError):
|
||||
locked.write_text("tampered\n")
|
||||
|
||||
assert existing.read_text() == '{"schema_version":1}\n'
|
||||
assert locked.read_text() == "trusted\n"
|
||||
|
||||
|
||||
def test_host_workspace_write_boundary_rejects_a_symlinked_writable_artifact(tmp_path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
target = tmp_path / "outside.json"
|
||||
target.write_text("{}\n")
|
||||
output = clone / "review-output.json"
|
||||
output.symlink_to(target)
|
||||
with pytest.raises(SandboxError, match="non-symlink"):
|
||||
with host_workspace_write_boundary(clone, writable=(output,)):
|
||||
pass
|
||||
|
||||
|
||||
def test_sandbox_workspace_write_boundary_is_noop_unless_host_unsafe(tmp_path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
source = clone / "source.ts"
|
||||
source.write_text("trusted\n")
|
||||
bwrap_sandbox = SimpleNamespace(backend="bwrap", clone=clone)
|
||||
with sandbox_workspace_write_boundary(
|
||||
bwrap_sandbox,
|
||||
read_only_workspace=True,
|
||||
writable=(),
|
||||
):
|
||||
source.write_text("still writable under bwrap no-op\n")
|
||||
assert source.read_text() == "still writable under bwrap no-op\n"
|
||||
|
||||
output = clone / "review-output.json"
|
||||
output.write_text("")
|
||||
source.write_text("trusted\n")
|
||||
unsafe = SimpleNamespace(backend="host-unsafe", clone=clone)
|
||||
with sandbox_workspace_write_boundary(
|
||||
unsafe,
|
||||
read_only_workspace=True,
|
||||
writable=(output,),
|
||||
):
|
||||
with pytest.raises(OSError):
|
||||
source.write_text("tampered\n")
|
||||
output.write_text("ok\n")
|
||||
assert source.read_text() == "trusted\n"
|
||||
assert output.read_text() == "ok\n"
|
||||
|
||||
|
||||
def test_force_rmtree_deletes_nonempty_directories_copied_from_a_locked_workspace(tmp_path) -> None:
|
||||
locked = tmp_path / "locked"
|
||||
nested = locked / "gitnexus-shared" / "src"
|
||||
nested.mkdir(parents=True)
|
||||
(nested / "index.ts").write_text("export {}\n")
|
||||
os.chmod(nested, 0o500)
|
||||
os.chmod(locked / "gitnexus-shared", 0o500)
|
||||
os.chmod(locked, 0o500)
|
||||
|
||||
copied = tmp_path / "sandbox-tmp" / "tmp.XXXX" / "gitnexus-shared"
|
||||
copied.parent.mkdir(parents=True)
|
||||
shutil.copytree(locked / "gitnexus-shared", copied)
|
||||
assert stat.S_IMODE(copied.stat().st_mode) & 0o222 == 0
|
||||
|
||||
_force_rmtree(copied.parent)
|
||||
assert not copied.parent.exists()
|
||||
|
||||
|
||||
def test_host_unsafe_sandbox_cleanup_survives_readonly_tmpdir_copies(tmp_path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
leftover = None
|
||||
with prepare_sandbox(clone=clone, backend="host-unsafe", claude_bin=sys.executable) as sandbox:
|
||||
leftover = sandbox.private_root
|
||||
copied = sandbox.temp / "tmp.XXXX" / "gitnexus-shared" / "src"
|
||||
copied.mkdir(parents=True)
|
||||
(copied / "index.ts").write_text("export {}\n")
|
||||
os.chmod(copied, 0o500)
|
||||
os.chmod(copied.parent, 0o500)
|
||||
os.chmod(copied.parent.parent, 0o500)
|
||||
assert leftover is not None
|
||||
assert not leftover.exists()
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"bad_url",
|
||||
["https://user:secret@example.test", "https://example.test/path?token=x", "file:///tmp/model"],
|
||||
|
|
@ -170,6 +401,16 @@ def test_sandbox_command_has_minimal_mounts_and_no_host_root_bind(tmp_path: Path
|
|||
assert probe.returncode == 0, probe.stderr
|
||||
assert probe.stdout == f"/home/agent|{SANDBOX_PATH}"
|
||||
|
||||
gitnexus_index = argv.index(SANDBOX_GITNEXUS_CLI)
|
||||
gitnexus_wrapper = Path(argv[gitnexus_index - 1])
|
||||
assert stat.S_IMODE(gitnexus_wrapper.stat().st_mode) == 0o500
|
||||
assert "/opt/gitnexus/dist/cli/index.js" in gitnexus_wrapper.read_text()
|
||||
|
||||
excludes_index = argv.index(SANDBOX_GIT_EXCLUDES)
|
||||
excludes = Path(argv[excludes_index - 1])
|
||||
assert stat.S_IMODE(excludes.stat().st_mode) == 0o400
|
||||
assert "/.bash_profile" in excludes.read_text().splitlines()
|
||||
|
||||
# The evidence-provenance.mjs plan-writer's PATH-scan trusts a Python 3
|
||||
# candidate only if it (and its directory) is owned by root or by the
|
||||
# current process — real /usr/bin/python3 is root-owned on the host,
|
||||
|
|
@ -617,6 +858,50 @@ finally:
|
|||
assert not (clone / "oracle-leak.txt").exists()
|
||||
|
||||
|
||||
@pytest.mark.skipif(
|
||||
os.environ.get("GITNEXUS_REQUIRE_BWRAP_CANARY") != "1",
|
||||
reason="real Bubblewrap canary is mandatory in the named Ubuntu CI job",
|
||||
)
|
||||
def test_read_only_review_workspace_exposes_only_one_writable_artifact(tmp_path: Path) -> None:
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
source = clone / "source.ts"
|
||||
source.write_text("trusted\n")
|
||||
output = clone / "review-output.json"
|
||||
output.write_text("")
|
||||
script = """
|
||||
from pathlib import Path
|
||||
try:
|
||||
Path('/workspace/source.ts').write_text('tampered')
|
||||
except OSError:
|
||||
pass
|
||||
else:
|
||||
raise SystemExit('review source remained writable')
|
||||
Path('/workspace/review-output.json').write_text('{"schema_version":1}')
|
||||
"""
|
||||
with prepare_sandbox(clone=clone, claude_bin=Path(sys.executable), preflight=True) as sandbox:
|
||||
result = run_managed(
|
||||
[
|
||||
*sandbox.command_prefix_for(
|
||||
read_only_workspace=True,
|
||||
extra_writable_mounts=(
|
||||
ReadOnlyMount(source=output, target="/workspace/review-output.json"),
|
||||
),
|
||||
),
|
||||
"/usr/bin/python3",
|
||||
"-c",
|
||||
script,
|
||||
],
|
||||
timeout=10,
|
||||
env=sandbox.environment(),
|
||||
require_pid_namespace=True,
|
||||
)
|
||||
|
||||
assert result.ok, result.stderr_tail
|
||||
assert source.read_text() == "trusted\n"
|
||||
assert output.read_text() == '{"schema_version":1}'
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="symlink creation may require elevated Windows privileges")
|
||||
@pytest.mark.parametrize("operation", ["stage", "sandbox"])
|
||||
def test_clone_root_symlink_is_rejected_before_host_access(tmp_path: Path, operation: str) -> None:
|
||||
|
|
@ -839,13 +1124,15 @@ def test_clone_controlled_mcp_replacement_is_never_executed_or_credentialed(tmp_
|
|||
os.environ.get("GITNEXUS_REQUIRE_CLAUDE_CANARY") != "1",
|
||||
reason="real Claude/Bash/MCP canary is mandatory in the named Ubuntu CI job",
|
||||
)
|
||||
def test_real_claude_bare_auth_inner_sandbox_and_mcp_permissions(tmp_path: Path) -> None:
|
||||
@pytest.mark.parametrize("review_layout", [False, True])
|
||||
def test_real_claude_auth_inner_sandbox_and_mcp_permissions(tmp_path: Path, review_layout: bool) -> None:
|
||||
"""Exercise the exact CLI boundary without contacting a paid model."""
|
||||
|
||||
claude = Path(os.environ["CLAUDE_CANARY_BIN"]).resolve()
|
||||
assert claude.is_file()
|
||||
clone = tmp_path / "clone"
|
||||
clone.mkdir()
|
||||
(clone / "canary.txt").write_text("hook-readable")
|
||||
fake_mcp = clone / "fake_mcp.py"
|
||||
fake_mcp.write_text(
|
||||
"""import json
|
||||
|
|
@ -872,7 +1159,7 @@ for line in sys.stdin:
|
|||
}]
|
||||
}
|
||||
elif method == "tools/call":
|
||||
Path("/workspace/mcp-called").write_text("ok")
|
||||
Path("/tmp/mcp-called").write_text("ok")
|
||||
result = {"content": [{"type": "text", "text": "repository list ready"}]}
|
||||
else:
|
||||
result = {}
|
||||
|
|
@ -881,6 +1168,31 @@ for line in sys.stdin:
|
|||
)
|
||||
fake_mcp.chmod(0o500)
|
||||
|
||||
review_command = """test -z "${ANTHROPIC_API_KEY:-}" && python3 - <<'PY'
|
||||
import json
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
source = Path('/workspace/canary.txt')
|
||||
assert 'hook-readable' in source.read_text()
|
||||
assert 'changed for review' in subprocess.check_output(['git', 'diff', '--', 'canary.txt'], text=True)
|
||||
for operation in (lambda: source.write_text('forbidden'), lambda: source.rename(source.with_name('renamed')), source.unlink):
|
||||
try:
|
||||
operation()
|
||||
except OSError:
|
||||
pass
|
||||
else:
|
||||
raise AssertionError('source mutation was allowed')
|
||||
Path('/workspace/review-output.json').write_text(json.dumps({'schema_version': 1, 'verdict': 'approve', 'findings': []}))
|
||||
PY"""
|
||||
if review_layout:
|
||||
for command in (
|
||||
["git", "init", "-q"],
|
||||
["git", "add", "canary.txt", "fake_mcp.py"],
|
||||
["git", "-c", "user.name=Canary", "-c", "user.email=canary@example.test", "commit", "-qm", "fixture"],
|
||||
):
|
||||
subprocess.run(command, cwd=clone, check=True, capture_output=True)
|
||||
(clone / "canary.txt").write_text("hook-readable\nchanged for review\n")
|
||||
|
||||
observed_tool_results: dict[str, dict] = {}
|
||||
|
||||
class ModelHandler(BaseHTTPRequestHandler):
|
||||
|
|
@ -910,7 +1222,19 @@ for line in sys.stdin:
|
|||
and isinstance(block.get("tool_use_id"), str)
|
||||
}
|
||||
)
|
||||
if "toolu_mcp_canary" not in tool_result_ids:
|
||||
if "toolu_read_canary" not in tool_result_ids:
|
||||
blocks = [
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "toolu_read_canary",
|
||||
"name": "Read",
|
||||
"input": {
|
||||
"file_path": "/workspace/canary.txt",
|
||||
},
|
||||
}
|
||||
]
|
||||
stop_reason = "tool_use"
|
||||
elif "toolu_mcp_canary" not in tool_result_ids:
|
||||
blocks = [
|
||||
{
|
||||
"type": "tool_use",
|
||||
|
|
@ -927,7 +1251,9 @@ for line in sys.stdin:
|
|||
"id": "toolu_bash_canary",
|
||||
"name": "Bash",
|
||||
"input": {
|
||||
"command": ('test -z "${ANTHROPIC_API_KEY:-}" && printf canary > /workspace/bash-called')
|
||||
"command": review_command
|
||||
if review_layout
|
||||
else ('test -z "${ANTHROPIC_API_KEY:-}" && printf canary > /workspace/bash-called')
|
||||
},
|
||||
}
|
||||
]
|
||||
|
|
@ -1024,6 +1350,18 @@ for line in sys.stdin:
|
|||
}
|
||||
)
|
||||
with prepare_sandbox(clone=clone, claude_bin=claude, preflight=True) as sandbox:
|
||||
output = None
|
||||
before = {}
|
||||
if review_layout:
|
||||
output = proposer_sandbox.prepare_review_workspace(sandbox, "review-output.json")
|
||||
before = runner_artifacts.workspace_snapshot(clone)
|
||||
sandbox = replace(
|
||||
sandbox,
|
||||
command_prefix=sandbox.command_prefix_for(
|
||||
read_only_workspace=True,
|
||||
extra_writable_mounts=(ReadOnlyMount(output, "/workspace/review-output.json"),),
|
||||
),
|
||||
)
|
||||
result = sandbox.run(
|
||||
[
|
||||
sandbox.claude_bin,
|
||||
|
|
@ -1032,7 +1370,6 @@ for line in sys.stdin:
|
|||
"text",
|
||||
"--output-format",
|
||||
"json",
|
||||
"--bare",
|
||||
"--settings",
|
||||
sandbox.settings_json,
|
||||
"--strict-mcp-config",
|
||||
|
|
@ -1044,7 +1381,12 @@ for line in sys.stdin:
|
|||
# authoritative empirical gate for that behavior.
|
||||
"--model",
|
||||
"claude-canary-20260718",
|
||||
"--tools",
|
||||
"Read",
|
||||
"Bash",
|
||||
"mcp__gitnexus__list_repos",
|
||||
"--allowedTools",
|
||||
"Read",
|
||||
"Bash",
|
||||
"mcp__gitnexus__list_repos",
|
||||
],
|
||||
|
|
@ -1053,17 +1395,26 @@ for line in sys.stdin:
|
|||
auth_token="offline-canary-key",
|
||||
base_url=f"http://127.0.0.1:{server.server_port}",
|
||||
),
|
||||
stdin_data=b"Use both available tools, then finish.",
|
||||
stdin_data=b"Use all three available tools, then finish.",
|
||||
)
|
||||
assert result.ok, result.stderr_tail + result.stdout_tail
|
||||
report = json.loads(result.stdout_tail)
|
||||
assert report["subtype"] == "success" and report["is_error"] is False, report
|
||||
read_result = observed_tool_results["toolu_read_canary"]
|
||||
assert read_result.get("is_error") is not True, read_result
|
||||
assert "hook-readable" in json.dumps(read_result)
|
||||
bash_result = observed_tool_results["toolu_bash_canary"]
|
||||
assert bash_result.get("is_error") is not True, bash_result
|
||||
assert (sandbox.temp / "mcp-called").read_text() == "ok"
|
||||
if review_layout:
|
||||
assert output is not None
|
||||
runner_artifacts.enforce_phase_workspace(clone, before, allowed_artifact=output)
|
||||
assert json.loads(output.read_text())["verdict"] == "approve"
|
||||
assert (clone / "canary.txt").read_text() == "hook-readable\nchanged for review\n"
|
||||
finally:
|
||||
server.shutdown()
|
||||
server.server_close()
|
||||
thread.join(timeout=5)
|
||||
|
||||
assert result.ok, result.stderr_tail + result.stdout_tail
|
||||
report = json.loads(result.stdout_tail)
|
||||
assert report["subtype"] == "success" and report["is_error"] is False, report
|
||||
bash_result = observed_tool_results["toolu_bash_canary"]
|
||||
assert bash_result.get("is_error") is not True, bash_result
|
||||
assert (clone / "bash-called").read_text() == "canary"
|
||||
assert (clone / "mcp-called").read_text() == "ok"
|
||||
if not review_layout:
|
||||
assert (clone / "bash-called").read_text() == "canary"
|
||||
|
|
|
|||
55
eval/tests/test_review_corpus.py
Normal file
55
eval/tests/test_review_corpus.py
Normal file
|
|
@ -0,0 +1,55 @@
|
|||
import hashlib
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import yaml
|
||||
|
||||
from workflow_bench.oracle_assets import review_case_setup_command
|
||||
|
||||
|
||||
BENCH_ROOT = Path(__file__).parents[1] / "workflow_bench"
|
||||
|
||||
|
||||
def test_review_corpus_is_immutable_and_task_bound():
|
||||
manifest = json.loads((BENCH_ROOT / "review_cases" / "manifest.json").read_text())
|
||||
tasks = yaml.safe_load((BENCH_ROOT / "tasks.review.scenarios.yaml").read_text())["tasks"]
|
||||
by_id = {task["id"]: task for task in tasks}
|
||||
|
||||
assert len(manifest["cases"]) >= 6
|
||||
assert sum(case["id"].endswith("-defect") for case in manifest["cases"]) >= 4
|
||||
assert sum(case["id"].endswith("-clean") for case in manifest["cases"]) >= 2
|
||||
assert set(by_id) == {case["id"] for case in manifest["cases"]}
|
||||
|
||||
for case in manifest["cases"]:
|
||||
assert len(case["base_sha"]) == len(case["head_sha"]) == 40
|
||||
assert len(case["human_verification_commit"]) == 40
|
||||
patch = BENCH_ROOT / "review_cases" / case["patch"]
|
||||
assert case["patch"]
|
||||
assert "defect" not in case["patch"]
|
||||
assert "clean" not in case["patch"]
|
||||
assert hashlib.sha256(patch.read_bytes()).hexdigest() == case["patch_sha256"]
|
||||
task = by_id[case["id"]]
|
||||
assert task["ref"] == case["base_sha"]
|
||||
assert task["sandbox_copy"] == [f"eval/workflow_bench/review_cases/{patch.name}"]
|
||||
assert task["setup"] == review_case_setup_command(patch.name)
|
||||
|
||||
|
||||
def test_hidden_labels_are_not_recoverable_from_visible_task_input():
|
||||
tasks_path = BENCH_ROOT / "tasks.review.scenarios.yaml"
|
||||
tasks = yaml.safe_load(tasks_path.read_text())["tasks"]
|
||||
|
||||
for task in tasks:
|
||||
visible = json.dumps(
|
||||
{
|
||||
"prompt": task["prompt"],
|
||||
"setup": task["setup"],
|
||||
"sandbox_copy": task["sandbox_copy"],
|
||||
},
|
||||
sort_keys=True,
|
||||
)
|
||||
assert "review-labels.json" not in visible
|
||||
assert "-defect" not in visible
|
||||
assert "-clean" not in visible
|
||||
for oracle_file in task["oracle"]["files"]:
|
||||
assert oracle_file["source"] not in visible
|
||||
assert oracle_file["target"] == "review-labels.json"
|
||||
327
eval/tests/test_review_scoring.py
Normal file
327
eval/tests/test_review_scoring.py
Normal file
|
|
@ -0,0 +1,327 @@
|
|||
import json
|
||||
from dataclasses import replace
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from workflow_bench.oracle_assets import OracleFileSnapshot, TaskOracleSnapshot
|
||||
from workflow_bench.review_scoring import (
|
||||
ExpectedFinding,
|
||||
ReviewFinding,
|
||||
expected_findings,
|
||||
parse_review_output,
|
||||
score_review,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("noise", [False, True])
|
||||
def test_complete_misses_are_measured_zero(noise):
|
||||
actual = (ReviewFinding("noise", "low", "other.py", 1, 1, "style", "s", "e", "r", False),) if noise else ()
|
||||
score = score_review("comment" if noise else "approve", actual, (expected(),))
|
||||
assert score["f1"] == score["weighted_f1"] == 0
|
||||
|
||||
|
||||
def test_downgraded_blocker_loses_weight_and_blocker_credit():
|
||||
actual = ReviewFinding("a", "low", "src/api.ts", 20, 20, "correctness", "s", "e", "r", False)
|
||||
score = score_review("comment", (actual,), (expected(),))
|
||||
assert score["weighted_recall"] == 0.2
|
||||
assert score["blocker_recall"] == 0
|
||||
assert score["verdict_correct"] is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize("size", [2, 17, 100])
|
||||
def test_maximum_matching_at_every_supported_size(size):
|
||||
a = ReviewFinding("a", "high", "src/api.ts", 1, 1, "a", "s", "e", "r", True)
|
||||
actual = [a, replace(a, finding_id="b", line=10, end_line=10, category="b")]
|
||||
labels = [
|
||||
expected(finding_id="broad", line_start=1, line_end=10, category="a"),
|
||||
expected(finding_id="tight", line_start=1, line_end=1, category="b"),
|
||||
]
|
||||
for i in range(2, size):
|
||||
actual.append(replace(a, finding_id=str(i), path=f"{i}.py"))
|
||||
labels.append(expected(finding_id=str(i), path=f"{i}.py", line_start=1, line_end=1))
|
||||
for findings in (actual, list(reversed(actual))):
|
||||
for expected_labels in (labels, list(reversed(labels))):
|
||||
assert score_review("request_changes", findings, expected_labels)["true_positives"] == size
|
||||
|
||||
|
||||
def test_dense_matching_handles_the_full_finding_limit():
|
||||
a = ReviewFinding("a", "high", "src/api.ts", 20, 20, "correctness", "s", "e", "r", True)
|
||||
actual = [replace(a, finding_id=str(i)) for i in range(100)]
|
||||
labels = [expected(finding_id=str(i)) for i in range(100)]
|
||||
assert score_review("request_changes", actual, labels)["true_positives"] == 100
|
||||
|
||||
|
||||
@pytest.mark.parametrize("large_side", ["actual", "expected"])
|
||||
def test_maximum_matching_with_asymmetric_large_inputs(large_side):
|
||||
a = ReviewFinding("a", "high", "src/api.ts", 1, 1, "a", "s", "e", "r", True)
|
||||
actual = [a, replace(a, finding_id="b", line=10, end_line=10, category="b")]
|
||||
labels = [
|
||||
expected(finding_id="broad", line_start=1, line_end=10, category="a"),
|
||||
expected(finding_id="tight", line_start=1, line_end=1, category="b"),
|
||||
]
|
||||
for i in range(15):
|
||||
if large_side == "actual":
|
||||
actual.append(replace(a, finding_id=str(i), path=f"extra-{i}.py"))
|
||||
else:
|
||||
labels.append(expected(finding_id=str(i), path=f"extra-{i}.py"))
|
||||
assert score_review("request_changes", actual, labels)["true_positives"] == 2
|
||||
|
||||
|
||||
def finding(**overrides):
|
||||
values = {
|
||||
"id": "actual-1",
|
||||
"severity": "high",
|
||||
"path": "src/api.ts",
|
||||
"line": 20,
|
||||
"end_line": 24,
|
||||
"category": "correctness",
|
||||
"scenario": "A missing guard lets an invalid request reach the sink.",
|
||||
"evidence": "The changed call at line 20 bypasses validate().",
|
||||
"recommendation": "Restore validation before the call.",
|
||||
"blocking": True,
|
||||
}
|
||||
values.update(overrides)
|
||||
return values
|
||||
|
||||
|
||||
def expected(**overrides):
|
||||
values = {
|
||||
"finding_id": "expected-1",
|
||||
"severity": "high",
|
||||
"path": "src/api.ts",
|
||||
"line_start": 18,
|
||||
"line_end": 22,
|
||||
"category": "correctness",
|
||||
}
|
||||
values.update(overrides)
|
||||
return ExpectedFinding(**values)
|
||||
|
||||
|
||||
def test_parse_review_output_requires_the_strict_schema(tmp_path: Path):
|
||||
output = tmp_path / "review-output.json"
|
||||
output.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"schema_version": 1,
|
||||
"verdict": "request_changes",
|
||||
"findings": [finding()],
|
||||
}
|
||||
)
|
||||
)
|
||||
|
||||
verdict, findings = parse_review_output(output)
|
||||
|
||||
assert verdict == "request_changes"
|
||||
assert findings[0].path == "src/api.ts"
|
||||
assert findings[0].blocking is True
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"document, message",
|
||||
[
|
||||
({"schema_version": 1, "verdict": "approve", "findings": [finding()]}, "approve"),
|
||||
(
|
||||
{
|
||||
"schema_version": 1,
|
||||
"verdict": "request_changes",
|
||||
"findings": [finding(blocking=False)],
|
||||
},
|
||||
"blocking",
|
||||
),
|
||||
(
|
||||
{
|
||||
"schema_version": 1,
|
||||
"verdict": "comment",
|
||||
"findings": [finding(path="../escape.ts")],
|
||||
},
|
||||
"repository-relative",
|
||||
),
|
||||
],
|
||||
)
|
||||
def test_parse_review_output_rejects_incoherent_or_unsafe_documents(tmp_path: Path, document, message):
|
||||
output = tmp_path / "review-output.json"
|
||||
output.write_text(json.dumps(document))
|
||||
with pytest.raises(ValueError, match=message):
|
||||
parse_review_output(output)
|
||||
|
||||
|
||||
def test_expected_findings_are_loaded_from_hidden_snapshot_only():
|
||||
payload = json.dumps(
|
||||
{
|
||||
"schema_version": 1,
|
||||
"findings": [
|
||||
{
|
||||
"id": "hidden-1",
|
||||
"severity": "critical",
|
||||
"path": "src/auth.ts",
|
||||
"line_start": 40,
|
||||
"line_end": 44,
|
||||
"category": "security",
|
||||
}
|
||||
],
|
||||
}
|
||||
).encode()
|
||||
snapshot = TaskOracleSnapshot(
|
||||
command="true",
|
||||
command_digest="command",
|
||||
manifest_digest="manifest",
|
||||
digest="all",
|
||||
files=(
|
||||
OracleFileSnapshot(
|
||||
target="review-labels.json",
|
||||
payload=payload,
|
||||
sha256="payload",
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
assert expected_findings(snapshot)[0].finding_id == "hidden-1"
|
||||
|
||||
|
||||
def test_score_review_matches_by_path_and_overlapping_range():
|
||||
actual = (
|
||||
ReviewFinding(
|
||||
finding_id="actual-1",
|
||||
severity="high",
|
||||
path="src/api.ts",
|
||||
line=20,
|
||||
end_line=24,
|
||||
category="correctness",
|
||||
scenario="scenario",
|
||||
evidence="evidence",
|
||||
recommendation="fix",
|
||||
blocking=True,
|
||||
),
|
||||
ReviewFinding(
|
||||
finding_id="noise",
|
||||
severity="low",
|
||||
path="src/other.ts",
|
||||
line=1,
|
||||
end_line=1,
|
||||
category="style",
|
||||
scenario="noise",
|
||||
evidence="noise",
|
||||
recommendation="noise",
|
||||
blocking=False,
|
||||
),
|
||||
)
|
||||
|
||||
score = score_review("request_changes", actual, (expected(),))
|
||||
|
||||
assert score["true_positives"] == 1
|
||||
assert score["false_positives"] == 1
|
||||
assert score["false_negatives"] == 0
|
||||
assert score["recall"] == 1
|
||||
assert score["precision"] == 0.5
|
||||
assert score["blocker_recall"] == 1
|
||||
assert score["verdict_correct"] is True
|
||||
|
||||
|
||||
def test_score_review_is_independent_of_finding_list_order():
|
||||
expected_labels = (
|
||||
expected(finding_id="broad", line_start=1, line_end=10, category="a"),
|
||||
expected(finding_id="tight", line_start=5, line_end=5, category="b"),
|
||||
)
|
||||
first = ReviewFinding(
|
||||
finding_id="a",
|
||||
severity="high",
|
||||
path="src/api.ts",
|
||||
line=5,
|
||||
end_line=5,
|
||||
category="a",
|
||||
scenario="s",
|
||||
evidence="e",
|
||||
recommendation="r",
|
||||
blocking=True,
|
||||
)
|
||||
second = ReviewFinding(
|
||||
finding_id="b",
|
||||
severity="high",
|
||||
path="src/api.ts",
|
||||
line=1,
|
||||
end_line=1,
|
||||
category="b",
|
||||
scenario="s",
|
||||
evidence="e",
|
||||
recommendation="r",
|
||||
blocking=True,
|
||||
)
|
||||
forward = score_review("request_changes", (first, second), expected_labels)
|
||||
reverse = score_review("request_changes", (second, first), expected_labels)
|
||||
assert forward["true_positives"] == reverse["true_positives"]
|
||||
assert forward["false_positives"] == reverse["false_positives"]
|
||||
assert forward["false_negatives"] == reverse["false_negatives"]
|
||||
assert forward["weighted_f1"] == reverse["weighted_f1"]
|
||||
|
||||
|
||||
def test_score_review_prefers_maximum_cardinality_over_greedy_category_match():
|
||||
expected_labels = (
|
||||
expected(finding_id="broad", line_start=1, line_end=10, category="a"),
|
||||
expected(finding_id="tight", line_start=1, line_end=1, category="b"),
|
||||
)
|
||||
actual = (
|
||||
ReviewFinding(
|
||||
finding_id="actual-1",
|
||||
severity="high",
|
||||
path="src/api.ts",
|
||||
line=1,
|
||||
end_line=1,
|
||||
category="a",
|
||||
scenario="s",
|
||||
evidence="e",
|
||||
recommendation="r",
|
||||
blocking=True,
|
||||
),
|
||||
ReviewFinding(
|
||||
finding_id="actual-2",
|
||||
severity="high",
|
||||
path="src/api.ts",
|
||||
line=10,
|
||||
end_line=10,
|
||||
category="b",
|
||||
scenario="s",
|
||||
evidence="e",
|
||||
recommendation="r",
|
||||
blocking=True,
|
||||
),
|
||||
)
|
||||
|
||||
score = score_review("request_changes", actual, expected_labels)
|
||||
|
||||
assert score["true_positives"] == 2
|
||||
assert score["false_positives"] == 0
|
||||
assert score["false_negatives"] == 0
|
||||
|
||||
|
||||
def test_clean_control_rewards_an_empty_approval_and_penalizes_noise():
|
||||
clean = score_review("approve", (), ())
|
||||
noisy = score_review(
|
||||
"comment",
|
||||
(
|
||||
ReviewFinding(
|
||||
finding_id="noise",
|
||||
severity="medium",
|
||||
path="src/ok.ts",
|
||||
line=1,
|
||||
end_line=1,
|
||||
category="correctness",
|
||||
scenario="noise",
|
||||
evidence="noise",
|
||||
recommendation="noise",
|
||||
blocking=False,
|
||||
),
|
||||
),
|
||||
(),
|
||||
)
|
||||
|
||||
assert clean["weighted_f1"] is None
|
||||
assert clean["precision"] is None
|
||||
assert clean["recall"] is None
|
||||
assert clean["clean_pass"] is True
|
||||
assert clean["verdict_correct"] is True
|
||||
assert noisy["false_positives"] == 1
|
||||
assert noisy["weighted_precision"] == 0
|
||||
assert noisy["recall"] is None
|
||||
assert noisy["clean_pass"] is False
|
||||
assert noisy["verdict_correct"] is False
|
||||
|
|
@ -2,12 +2,18 @@
|
|||
|
||||
import hashlib
|
||||
import json
|
||||
import shutil
|
||||
from contextlib import nullcontext
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
from workflow_bench import runner, runner_artifacts, runner_sessions
|
||||
from workflow_bench.evolution import skill_fingerprint
|
||||
from workflow_bench.oracle_assets import review_case_setup_command
|
||||
from workflow_bench.process_control import ManagedProcessError, ManagedProcessResult
|
||||
from workflow_bench.proposer_sandbox import SandboxError
|
||||
|
||||
|
||||
def _report(**overrides) -> str:
|
||||
|
|
@ -275,6 +281,16 @@ def test_phase_workspace_ignores_claude_sandbox_bootstrap_noise(tmp_path):
|
|||
(tmp_path / ".claude" / "commands").mkdir(parents=True)
|
||||
(tmp_path / ".claude" / ".cc-writes").write_text("{}")
|
||||
(tmp_path / ".env").write_text("")
|
||||
(tmp_path / ".bash_profile").write_text("")
|
||||
(tmp_path / ".bashrc").write_text("")
|
||||
(tmp_path / ".gitconfig").write_text("")
|
||||
(tmp_path / ".idea").mkdir()
|
||||
(tmp_path / ".profile").write_text("")
|
||||
(tmp_path / ".ripgreprc").write_text("")
|
||||
(tmp_path / ".vscode").mkdir()
|
||||
(tmp_path / ".zprofile").write_text("")
|
||||
(tmp_path / ".zshrc").write_text("")
|
||||
(tmp_path / "scripts").write_text("")
|
||||
(tmp_path / ".env.development.local").write_text("")
|
||||
(tmp_path / ".npmrc").write_text("")
|
||||
(tmp_path / "package.json").write_text("{}")
|
||||
|
|
@ -286,6 +302,30 @@ def test_phase_workspace_ignores_claude_sandbox_bootstrap_noise(tmp_path):
|
|||
runner_artifacts.enforce_phase_workspace(tmp_path, before, allowed_artifact=artifact)
|
||||
|
||||
|
||||
def test_phase_workspace_records_a_root_scripts_symlink(tmp_path):
|
||||
target = tmp_path / "helper.py"
|
||||
target.write_text("planted\n")
|
||||
before = runner_artifacts.workspace_snapshot(tmp_path)
|
||||
(tmp_path / "scripts").symlink_to(target)
|
||||
artifact = tmp_path / "review-output.md"
|
||||
artifact.write_text("new review")
|
||||
|
||||
with pytest.raises(ValueError, match="unauthorized workspace path"):
|
||||
runner_artifacts.enforce_phase_workspace(tmp_path, before, allowed_artifact=artifact)
|
||||
|
||||
|
||||
def test_phase_workspace_still_rejects_writes_under_scripts(tmp_path):
|
||||
scripts = tmp_path / "scripts"
|
||||
scripts.mkdir()
|
||||
before = runner_artifacts.workspace_snapshot(tmp_path)
|
||||
(scripts / "helper.py").write_text("planted\n")
|
||||
artifact = tmp_path / "review-output.md"
|
||||
artifact.write_text("new review")
|
||||
|
||||
with pytest.raises(ValueError, match="unauthorized workspace path"):
|
||||
runner_artifacts.enforce_phase_workspace(tmp_path, before, allowed_artifact=artifact)
|
||||
|
||||
|
||||
def test_phase_workspace_still_rejects_a_genuinely_unauthorized_change(tmp_path):
|
||||
# The bootstrap-noise exclusion must stay narrow: an actual source-file
|
||||
# edit outside the allowed artifact still has to be caught.
|
||||
|
|
@ -381,3 +421,584 @@ def test_phase_workspace_still_sees_writes_under_a_pre_existing_nested_claude_di
|
|||
|
||||
with pytest.raises(ValueError, match="unauthorized workspace path"):
|
||||
runner_artifacts.enforce_phase_workspace(tmp_path, before, allowed_artifact=artifact)
|
||||
|
||||
|
||||
def _cell_context(tmp_path, **overrides):
|
||||
"""A TaskCellContext whose per-task inputs are all present and valid."""
|
||||
snapshot = SimpleNamespace(
|
||||
digest="asset-digest",
|
||||
manifest_digest="asset-manifest",
|
||||
dependency_content_digest="dep-content",
|
||||
dependency_manifest_digest="dep-manifest",
|
||||
)
|
||||
graph = SimpleNamespace(
|
||||
digest="graph-digest",
|
||||
manifest_digest="graph-manifest",
|
||||
materialize=lambda *_a, **_k: None,
|
||||
)
|
||||
oracle = SimpleNamespace(
|
||||
digest="oracle-digest",
|
||||
command_digest="oracle-command",
|
||||
manifest_digest="oracle-manifest",
|
||||
)
|
||||
fields = {
|
||||
"task": {"id": "task-a", "class": "demo", "prompt": "do the thing"},
|
||||
"oracle_snapshot": oracle,
|
||||
"repo": tmp_path / "repo",
|
||||
"task_sha": "a" * 40,
|
||||
"graph_snapshot": graph,
|
||||
"graph_snapshot_error": None,
|
||||
"asset_snapshot": snapshot,
|
||||
"asset_snapshot_error": None,
|
||||
"args": SimpleNamespace(
|
||||
claude_bin="claude",
|
||||
model="pinned-model",
|
||||
proposer_model=None,
|
||||
effort="xhigh",
|
||||
auth_token=None,
|
||||
),
|
||||
"out_dir": tmp_path / "out",
|
||||
"ce_plugin_snapshot": None,
|
||||
"trees_dir": tmp_path / "trees",
|
||||
"bwrap_bin": tmp_path / "bwrap",
|
||||
"runtime_mounts": (),
|
||||
"candidate_overlay": None,
|
||||
"overlay_digest": None,
|
||||
}
|
||||
fields.update(overrides)
|
||||
fields["out_dir"].mkdir(parents=True, exist_ok=True)
|
||||
return runner.TaskCellContext(**fields)
|
||||
|
||||
|
||||
def _stub_cell_dependencies(monkeypatch, tmp_path):
|
||||
"""Replace everything a cell shells out to, so only its own logic runs.
|
||||
|
||||
Returns the clone it will hand out and the list its teardown appends to.
|
||||
"""
|
||||
removed: list[Path] = []
|
||||
worktree = tmp_path / "clone"
|
||||
worktree.mkdir()
|
||||
monkeypatch.setattr(runner, "make_worktree", lambda *_a, **_k: worktree)
|
||||
monkeypatch.setattr(runner, "sanitize_clone_for_hidden_oracles", lambda *_a, **_k: "b" * 40)
|
||||
monkeypatch.setattr(runner, "stage_task_assets", lambda *_a, **_k: [])
|
||||
monkeypatch.setattr(runner, "isolated_gitnexus_registry_mount", lambda *_a, **_k: None)
|
||||
monkeypatch.setattr(runner, "ce_plugin_mounts_for_arm", lambda *_a, **_k: [])
|
||||
monkeypatch.setattr(runner, "ce_plugin_dir_for_arm", lambda *_a, **_k: None)
|
||||
monkeypatch.setattr(runner, "prepare_sandbox", lambda **_k: nullcontext(SimpleNamespace(run=None)))
|
||||
monkeypatch.setattr(runner, "skill_fingerprint", lambda *_a, **_k: "skill-digest")
|
||||
monkeypatch.setattr(runner, "require_skill_fingerprint", lambda *_a, **_k: None)
|
||||
monkeypatch.setattr(runner, "_sandbox_git", lambda *_a, **_k: "c" * 40)
|
||||
monkeypatch.setattr(runner, "implementation_diff_digest", lambda *_a, **_k: "")
|
||||
monkeypatch.setattr(runner, "_prepare_untracked_for_diff", lambda *_a, **_k: None)
|
||||
monkeypatch.setattr(runner, "diff_churn", lambda *_a, **_k: {})
|
||||
monkeypatch.setattr(runner, "enforce_work_evidence", lambda *_a, **_k: None)
|
||||
monkeypatch.setattr(runner, "capture_patch", lambda *_a, **_k: b"diff")
|
||||
monkeypatch.setattr(runner, "run_arm", lambda *_a, **_k: {"resolved": True, "ok": True, "error_kind": None})
|
||||
monkeypatch.setattr(runner, "remove_clone", lambda path: removed.append(path))
|
||||
return worktree, removed
|
||||
|
||||
|
||||
def test_run_cell_returns_a_row_bound_to_its_task_and_snapshots(monkeypatch, tmp_path):
|
||||
_, removed = _stub_cell_dependencies(monkeypatch, tmp_path)
|
||||
|
||||
record = runner.run_cell(_cell_context(tmp_path), 2, "workflow")
|
||||
|
||||
assert record["resolved"] is True
|
||||
assert record["error_kind"] is None
|
||||
# The row has to carry its own coordinates: once cells stop running in a
|
||||
# predictable order, position in results.jsonl identifies nothing.
|
||||
assert record["task"] == "task-a"
|
||||
assert record["arm"] == "workflow"
|
||||
assert record["run"] == 2
|
||||
assert record["task_asset_snapshot_digest"] == "asset-digest"
|
||||
assert record["sanitized_graph_snapshot_digest"] == "graph-digest"
|
||||
assert record["oracle_digest"] == "oracle-digest"
|
||||
assert removed == [tmp_path / "clone"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"failure",
|
||||
[
|
||||
ManagedProcessError(
|
||||
["setup"],
|
||||
ManagedProcessResult(
|
||||
state="timeout",
|
||||
returncode=-15,
|
||||
stdout_tail="",
|
||||
stderr_tail="",
|
||||
duration_s=1.0,
|
||||
),
|
||||
),
|
||||
SandboxError("sandbox refused"),
|
||||
OSError("disk went away"),
|
||||
RuntimeError("overlay drifted"),
|
||||
ValueError("bad binding"),
|
||||
],
|
||||
ids=["managed-process", "sandbox", "os", "runtime", "value"],
|
||||
)
|
||||
def test_run_cell_records_an_expected_failure_and_still_removes_its_clone(monkeypatch, tmp_path, failure):
|
||||
_, removed = _stub_cell_dependencies(monkeypatch, tmp_path)
|
||||
|
||||
def explode(*_args, **_kwargs):
|
||||
raise failure
|
||||
|
||||
monkeypatch.setattr(runner, "run_arm", explode)
|
||||
record = runner.run_cell(_cell_context(tmp_path), 0, "workflow")
|
||||
|
||||
assert record["error_kind"] == "infra-error"
|
||||
assert record["resolved"] is False
|
||||
# A cell owns its clone for its whole lifetime; the sweep has no other
|
||||
# chance to reclaim it, so the finally must survive every expected failure.
|
||||
assert removed == [tmp_path / "clone"]
|
||||
|
||||
|
||||
def test_run_cell_redacts_the_auth_token_from_the_failure_it_prints(monkeypatch, tmp_path, capsys):
|
||||
_stub_cell_dependencies(monkeypatch, tmp_path)
|
||||
secret = "sk-ant-not-a-real-key"
|
||||
|
||||
def explode(*_args, **_kwargs):
|
||||
raise ManagedProcessError(
|
||||
["claude"],
|
||||
ManagedProcessResult(
|
||||
state="exited",
|
||||
returncode=1,
|
||||
stdout_tail="",
|
||||
stderr_tail=f"ANTHROPIC_API_KEY={secret}",
|
||||
duration_s=1.0,
|
||||
),
|
||||
)
|
||||
|
||||
monkeypatch.setattr(runner, "run_arm", explode)
|
||||
context = _cell_context(tmp_path)
|
||||
context.args.auth_token = secret
|
||||
# ManagedProcessError stringifies up to 1000 raw bytes of stderr_tail, and
|
||||
# this line streams live into the CI log now that the sweep's stdout is
|
||||
# echoed. results.jsonl already redacts the same field.
|
||||
runner.run_cell(context, 0, "workflow")
|
||||
|
||||
assert secret not in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_run_cell_lets_an_unexpected_failure_escape_rather_than_scoring_it(monkeypatch, tmp_path):
|
||||
_, removed = _stub_cell_dependencies(monkeypatch, tmp_path)
|
||||
|
||||
def explode(*_args, **_kwargs):
|
||||
raise KeyError("harness bug")
|
||||
|
||||
monkeypatch.setattr(runner, "run_arm", explode)
|
||||
# A harness bug recorded as an ordinary infra-error would be averaged into
|
||||
# the evidence and counted toward the outage breaker. It must crash instead.
|
||||
with pytest.raises(KeyError):
|
||||
runner.run_cell(_cell_context(tmp_path), 0, "workflow")
|
||||
assert removed == [tmp_path / "clone"]
|
||||
|
||||
|
||||
def test_run_cell_reports_a_cleanup_failure_over_its_primary_outcome(monkeypatch, tmp_path):
|
||||
_stub_cell_dependencies(monkeypatch, tmp_path)
|
||||
|
||||
def refuse(_path):
|
||||
raise OSError("clone is busy")
|
||||
|
||||
monkeypatch.setattr(runner, "remove_clone", refuse)
|
||||
record = runner.run_cell(_cell_context(tmp_path), 1, "workflow")
|
||||
|
||||
assert record["error_kind"] == "cleanup-failure"
|
||||
assert record["resolved"] is False
|
||||
assert "primary=None" in record["error_detail"]
|
||||
assert "clone is busy" in record["error_detail"]
|
||||
|
||||
|
||||
def test_run_cell_does_not_mask_the_staged_review_patch_before_setup(monkeypatch, tmp_path):
|
||||
"""Review setup applies a patch staged under eval/workflow_bench.
|
||||
|
||||
Overlaying the empty oracle mask on that path is the CI abort:
|
||||
`git apply` dies with `can't open patch`. The staged copy must stay
|
||||
visible to sandboxed setup, then be gone before the model starts.
|
||||
"""
|
||||
|
||||
worktree, _ = _stub_cell_dependencies(monkeypatch, tmp_path)
|
||||
patch = worktree / "eval" / "workflow_bench" / "review_cases" / "pr-2718.patch"
|
||||
patch.parent.mkdir(parents=True)
|
||||
patch.write_text("diff --git a/visible.py b/visible.py\n")
|
||||
captured: dict[str, object] = {}
|
||||
|
||||
def fake_prepare(**kwargs):
|
||||
captured["mounts"] = kwargs.get("read_only_mounts", [])
|
||||
|
||||
def run(_command, **_kwargs):
|
||||
leftover = worktree / "eval" / "workflow_bench"
|
||||
if leftover.exists():
|
||||
shutil.rmtree(leftover)
|
||||
return SimpleNamespace(ok=True)
|
||||
|
||||
return nullcontext(SimpleNamespace(run=run))
|
||||
|
||||
monkeypatch.setattr(runner, "prepare_sandbox", fake_prepare)
|
||||
context = _cell_context(
|
||||
tmp_path,
|
||||
task={
|
||||
"id": "review-pr-2718-defect",
|
||||
"class": "review-defect",
|
||||
"prompt": "review the local diff",
|
||||
"setup": review_case_setup_command("pr-2718.patch"),
|
||||
},
|
||||
)
|
||||
record = runner.run_cell(context, 0, "workflow")
|
||||
|
||||
assert record.get("error_kind") is None
|
||||
targets = [getattr(mount, "target", None) for mount in captured["mounts"] if mount is not None]
|
||||
assert not any(target and "eval/workflow_bench" in str(target) for target in targets)
|
||||
assert not (worktree / "eval" / "workflow_bench").exists()
|
||||
|
||||
|
||||
def test_run_cell_fails_closed_when_setup_leaves_the_hidden_harness(monkeypatch, tmp_path):
|
||||
worktree, _ = _stub_cell_dependencies(monkeypatch, tmp_path)
|
||||
leftover = worktree / "eval" / "workflow_bench" / "review_cases"
|
||||
leftover.mkdir(parents=True)
|
||||
(leftover / "pr-2718.patch").write_text("diff\n")
|
||||
|
||||
def fake_prepare(**kwargs):
|
||||
return nullcontext(SimpleNamespace(run=lambda *_a, **_k: SimpleNamespace(ok=True)))
|
||||
|
||||
monkeypatch.setattr(runner, "prepare_sandbox", fake_prepare)
|
||||
context = _cell_context(
|
||||
tmp_path,
|
||||
task={
|
||||
"id": "review-pr-2718-defect",
|
||||
"class": "review-defect",
|
||||
"prompt": "review the local diff",
|
||||
"setup": review_case_setup_command("pr-2718.patch"),
|
||||
},
|
||||
)
|
||||
record = runner.run_cell(context, 0, "workflow")
|
||||
|
||||
assert record["error_kind"] == "infra-error"
|
||||
assert "hidden harness visible" in str(record["error_detail"])
|
||||
|
||||
|
||||
def test_run_cell_fails_closed_when_a_per_task_snapshot_never_materialized(tmp_path):
|
||||
# The snapshots are prepared once per task, before any cell. If that failed,
|
||||
# every cell of the task has to record it rather than run against nothing.
|
||||
context = _cell_context(tmp_path, asset_snapshot=None, asset_snapshot_error=OSError("no assets"))
|
||||
record = runner.run_cell(context, 0, "workflow")
|
||||
|
||||
assert record["error_kind"] == "infra-error"
|
||||
assert "no assets" in str(record["error_detail"])
|
||||
|
||||
|
||||
def _progress():
|
||||
"""Collector for what a sweep started and kept, readable after it raises."""
|
||||
return SimpleNamespace(started=[], kept=[], streak=0, tripped=False)
|
||||
|
||||
|
||||
def _sweep(cells, *, workers, run, outage_limit=5, streak=0, into=None):
|
||||
"""Drive sweep_task_cells, recording what it started and kept.
|
||||
|
||||
Pass ``into`` a ``_progress()`` when the sweep is expected to raise: the
|
||||
collector survives the exception, the return value does not.
|
||||
"""
|
||||
result = _progress() if into is None else into
|
||||
result.streak, result.tripped = runner.sweep_task_cells(
|
||||
cells,
|
||||
workers=workers,
|
||||
run=run,
|
||||
on_start=lambda run_idx, arm: result.started.append((run_idx, arm)),
|
||||
on_record=lambda run_idx, arm, _record: result.kept.append((run_idx, arm)),
|
||||
outage_streak=streak,
|
||||
outage_limit=outage_limit,
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def _row(error_kind=None):
|
||||
return {"resolved": error_kind is None, "error_kind": error_kind}
|
||||
|
||||
|
||||
CELLS = [(run_idx, arm) for run_idx in range(3) for arm in ("workflow", "candidate_workflow")]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("workers", [1, 3, 8])
|
||||
@pytest.mark.parametrize("primary", ["review-evidence-invalid", "skill-not-invoked", "session-error"])
|
||||
def test_unusable_review_evidence_stops_a_54_cell_sweep(workers, primary):
|
||||
cells = [(i, "review") for i in range(54)]
|
||||
result = _sweep(
|
||||
cells,
|
||||
workers=workers,
|
||||
run=lambda *_: {
|
||||
"resolved": False,
|
||||
"error_kind": primary,
|
||||
"review_evidence_valid": False,
|
||||
},
|
||||
)
|
||||
assert result.tripped
|
||||
assert 5 <= len(result.started) <= 5 + workers - 1
|
||||
assert result.kept == result.started
|
||||
|
||||
|
||||
def test_measured_review_miss_resets_the_outage_streak():
|
||||
result = _sweep(
|
||||
CELLS,
|
||||
workers=1,
|
||||
streak=4,
|
||||
run=lambda *_: {
|
||||
"resolved": False,
|
||||
"error_kind": "oracle-failed",
|
||||
"review_evidence_valid": True,
|
||||
"review_weighted_f1": 0.0,
|
||||
},
|
||||
)
|
||||
assert result.streak == 0 and not result.tripped
|
||||
|
||||
|
||||
def test_invalid_review_with_primary_skill_error_is_excluded_from_aggregation():
|
||||
result = runner.aggregate(
|
||||
[
|
||||
{
|
||||
"resolved": False,
|
||||
"error_kind": "skill-not-invoked",
|
||||
"review_evidence_valid": False,
|
||||
}
|
||||
]
|
||||
)
|
||||
assert result["excluded_runs"] == 1 and result["valid_runs"] == 0
|
||||
|
||||
|
||||
def test_sweep_keeps_rows_in_submission_order_whatever_order_they_finish():
|
||||
# Cells finish in whatever order the machine allows, but a wave is folded
|
||||
# in submission order — the outage streak counts consecutive failures, and
|
||||
# "consecutive" in completion order would make the trip point flaky.
|
||||
import threading
|
||||
|
||||
first_wave = CELLS[:3]
|
||||
rendezvous = threading.Barrier(3, timeout=10)
|
||||
release_first = threading.Event()
|
||||
fast_finished = threading.Event()
|
||||
finished: list[tuple[int, str]] = []
|
||||
result: list[SimpleNamespace] = []
|
||||
|
||||
def run(run_idx, arm):
|
||||
cell = (run_idx, arm)
|
||||
if cell in first_wave:
|
||||
rendezvous.wait()
|
||||
if cell == first_wave[0]:
|
||||
release_first.wait(timeout=10)
|
||||
else:
|
||||
finished.append(cell)
|
||||
if len(finished) == 2:
|
||||
fast_finished.set()
|
||||
return _row()
|
||||
|
||||
sweep = threading.Thread(target=lambda: result.append(_sweep(CELLS, workers=3, run=run)))
|
||||
sweep.start()
|
||||
try:
|
||||
assert fast_finished.wait(timeout=10)
|
||||
assert first_wave[0] not in finished
|
||||
assert set(finished) == set(first_wave[1:])
|
||||
finally:
|
||||
release_first.set()
|
||||
sweep.join(timeout=10)
|
||||
|
||||
assert not sweep.is_alive()
|
||||
assert result[0].kept == CELLS
|
||||
assert result[0].started == CELLS
|
||||
assert result[0].tripped is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize("workers", [1, 2, 3])
|
||||
def test_sweep_trips_the_breaker_within_one_wave_of_the_serial_point(workers):
|
||||
# Serial stops after the 5th consecutive systemic failure. Cells already in
|
||||
# flight when the breaker trips cannot be recalled or erased from the
|
||||
# evidence, so the overrun is bounded by the wave and every completed row
|
||||
# is kept. Ten cells make the bound visible rather than hidden by the end.
|
||||
long_task = [(run_idx, arm) for run_idx in range(5) for arm in ("workflow", "candidate_workflow")]
|
||||
|
||||
result = _sweep(long_task, workers=workers, run=lambda *_: _row("session-error"))
|
||||
|
||||
assert result.tripped is True
|
||||
assert 5 <= len(result.started) <= 5 + workers - 1
|
||||
assert result.kept == result.started
|
||||
assert len(result.started) < len(long_task)
|
||||
|
||||
|
||||
def test_sweep_reads_a_real_failure_as_signal_rather_than_an_outage():
|
||||
# resolved=False with no systemic error_kind is the benchmark working, not
|
||||
# the harness failing; it must reset the streak instead of tripping.
|
||||
result = _sweep(CELLS, workers=3, run=lambda *_: {"resolved": False, "error_kind": None})
|
||||
|
||||
assert result.tripped is False
|
||||
assert result.streak == 0
|
||||
assert result.kept == CELLS
|
||||
|
||||
|
||||
def test_sweep_surfaces_an_unexpected_worker_failure_instead_of_dropping_the_cell():
|
||||
def run(run_idx, arm):
|
||||
if (run_idx, arm) == (0, "candidate_workflow"):
|
||||
raise KeyError("harness bug")
|
||||
return _row()
|
||||
|
||||
# A Future holds its exception until read. Unread, this cell would vanish
|
||||
# from the evidence with no crash and no row — fewer runs in an arm's
|
||||
# aggregate, silently.
|
||||
with pytest.raises(KeyError):
|
||||
_sweep(CELLS, workers=3, run=run)
|
||||
|
||||
|
||||
def test_sweep_runs_cells_of_a_wave_at_the_same_time():
|
||||
import threading
|
||||
|
||||
barrier = threading.Barrier(3, timeout=10)
|
||||
|
||||
def run(run_idx, arm):
|
||||
# Deadlocks unless all three cells of the wave are genuinely in flight
|
||||
# together — a pool that serialised them would time out here.
|
||||
barrier.wait()
|
||||
return _row()
|
||||
|
||||
result = _sweep(CELLS, workers=3, run=run)
|
||||
|
||||
assert result.kept == CELLS
|
||||
|
||||
|
||||
def test_sweep_of_one_worker_never_leaves_the_calling_thread():
|
||||
import threading
|
||||
|
||||
caller = threading.current_thread()
|
||||
seen: list[threading.Thread] = []
|
||||
|
||||
def run(run_idx, arm):
|
||||
seen.append(threading.current_thread())
|
||||
return _row()
|
||||
|
||||
# Ctrl-C reaches only the main thread, so the serial default has to stay on
|
||||
# it: a cell on a worker thread is outside the reach of the cleanup that
|
||||
# kills its sandboxed process tree.
|
||||
_sweep(CELLS, workers=1, run=run)
|
||||
|
||||
assert seen == [caller] * len(CELLS)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("failure", [KeyError("harness bug"), SystemExit(97)])
|
||||
def test_sweep_keeps_the_rows_of_cells_that_finished_beside_a_failing_one(failure):
|
||||
def run(run_idx, arm):
|
||||
if (run_idx, arm) == (0, "candidate_workflow"):
|
||||
raise failure
|
||||
return _row()
|
||||
|
||||
progress = _progress()
|
||||
# The failing cell's two siblings completed and spent their budget before
|
||||
# the harness bug surfaced. Reading the futures in order and raising on the
|
||||
# first failure would drop their rows: money spent, no evidence written.
|
||||
with pytest.raises(type(failure)):
|
||||
_sweep(CELLS, workers=3, run=run, into=progress)
|
||||
|
||||
assert progress.kept == [(0, "workflow"), (1, "workflow")]
|
||||
|
||||
|
||||
def test_sweep_hands_a_ctrl_c_back_without_waiting_for_the_running_cells(monkeypatch):
|
||||
import threading
|
||||
|
||||
in_flight = threading.Barrier(3, timeout=10)
|
||||
release = threading.Event()
|
||||
finished: list[tuple[int, str]] = []
|
||||
|
||||
def run(run_idx, arm):
|
||||
in_flight.wait()
|
||||
release.wait(timeout=10)
|
||||
finished.append((run_idx, arm))
|
||||
return _row()
|
||||
|
||||
def interrupt_once_the_wave_is_running(_futures, *_args, **_kwargs):
|
||||
# Stands in for the Ctrl-C an operator types mid-wave: an async
|
||||
# KeyboardInterrupt is delivered to the main thread, which is the one
|
||||
# blocked here waiting on the wave.
|
||||
in_flight.wait()
|
||||
raise KeyboardInterrupt
|
||||
|
||||
monkeypatch.setattr(runner, "wait", interrupt_once_the_wave_is_running)
|
||||
try:
|
||||
with runner.cancellation_scope(release), pytest.raises(KeyboardInterrupt):
|
||||
_sweep(CELLS, workers=2, run=run)
|
||||
|
||||
# Cancellation releases active work before joining; no worker can
|
||||
# outlive the assets the interrupted sweep is about to clean up.
|
||||
assert set(finished) == set(CELLS[:2])
|
||||
finally:
|
||||
release.set()
|
||||
|
||||
|
||||
def test_workers_is_bounded_at_both_ends_before_the_sweep_starts():
|
||||
base = ["--tasks", "tasks.yaml", "--model", "pinned-model"]
|
||||
|
||||
assert runner.build_parser().parse_args(base).workers == 1
|
||||
at_max = runner.build_parser().parse_args([*base, "--workers", str(runner.MAX_WORKERS)])
|
||||
assert at_max.workers == runner.MAX_WORKERS
|
||||
|
||||
# A mistyped worker count has to fail at the command line: hours later it
|
||||
# only shows up as timed-out sessions, which the promotion gate throws away.
|
||||
for rejected in ("0", "-1", str(runner.MAX_WORKERS + 1)):
|
||||
with pytest.raises(SystemExit):
|
||||
runner.build_parser().parse_args([*base, "--workers", rejected])
|
||||
|
||||
|
||||
def test_partial_wave_submission_preserves_rows_and_original_interruption(monkeypatch):
|
||||
original = runner.ThreadPoolExecutor.submit
|
||||
submissions = 0
|
||||
|
||||
def submit(pool, *args, **kwargs):
|
||||
nonlocal submissions
|
||||
submissions += 1
|
||||
if submissions == 2:
|
||||
raise KeyboardInterrupt("submission interrupted")
|
||||
return original(pool, *args, **kwargs)
|
||||
|
||||
monkeypatch.setattr(runner.ThreadPoolExecutor, "submit", submit)
|
||||
records = []
|
||||
with pytest.raises(KeyboardInterrupt, match="submission interrupted"):
|
||||
runner.sweep_task_cells(
|
||||
CELLS,
|
||||
workers=2,
|
||||
run=lambda *_: _row(),
|
||||
on_start=lambda *_: None,
|
||||
on_record=lambda *record: records.append(record),
|
||||
outage_streak=0,
|
||||
outage_limit=5,
|
||||
)
|
||||
assert len(records) == 2
|
||||
assert records[1][2]["error_kind"] == "cancelled"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("error_kind", ["infra-error", "cleanup-failure"])
|
||||
def test_progress_line_reports_an_unmeasured_failure_as_unmeasured_not_as_free(error_kind):
|
||||
dead = runner.infra_error_record(RuntimeError("bwrap died"))
|
||||
dead["error_kind"] = error_kind
|
||||
|
||||
line = runner.cell_progress_line("task", "workflow", 0, dead)
|
||||
|
||||
# The 0.0s are placeholders for numbers no session ever produced; printed
|
||||
# as numbers they read as a cell that ran instantly for free.
|
||||
assert "cost=n/a" in line
|
||||
assert "took=n/a" in line
|
||||
assert f"error_kind={error_kind}" in line
|
||||
# results.jsonl is promotion evidence — only the display changes.
|
||||
assert dead["cost_usd"] == 0.0
|
||||
assert dead["duration_s"] == 0.0
|
||||
|
||||
|
||||
def test_progress_line_reports_the_numbers_a_real_run_measured():
|
||||
line = runner.cell_progress_line(
|
||||
"task",
|
||||
"workflow",
|
||||
1,
|
||||
{
|
||||
"resolved": True,
|
||||
"input_tokens": 10,
|
||||
"output_tokens": 2,
|
||||
"cost_usd": 0.5,
|
||||
"duration_s": 12.0,
|
||||
"error_kind": None,
|
||||
},
|
||||
)
|
||||
|
||||
assert "cost=$0.5" in line
|
||||
assert "took=12.0s" in line
|
||||
assert "error_kind=none" in line
|
||||
|
|
|
|||
|
|
@ -27,6 +27,32 @@ def test_prebuilt_graph_and_harness_assets_are_rejected(task):
|
|||
sanitized_graph.validate_no_prebuilt_graph_assets(task)
|
||||
|
||||
|
||||
def test_review_case_patches_are_allowed_sandbox_copy():
|
||||
sanitized_graph.validate_no_prebuilt_graph_assets(
|
||||
{"sandbox_copy": ["eval/workflow_bench/review_cases/pr-2718.patch"]}
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"task",
|
||||
[
|
||||
{"sandbox_copy": ["eval/workflow_bench"]},
|
||||
{"sandbox_copy": ["eval/workflow_bench/evolve.py"]},
|
||||
{
|
||||
"sandbox_dependencies": [
|
||||
{
|
||||
"source": "eval/workflow_bench/review_cases/pr-2718.patch",
|
||||
"target": "patch",
|
||||
}
|
||||
]
|
||||
},
|
||||
],
|
||||
)
|
||||
def test_non_corpus_harness_paths_stay_rejected(task):
|
||||
with pytest.raises(SandboxError, match="prebuilt graph or harness"):
|
||||
sanitized_graph.validate_no_prebuilt_graph_assets(task)
|
||||
|
||||
|
||||
def test_graph_environment_is_offline_deterministic_and_ignores_target_gitignore():
|
||||
env = sanitized_graph._graph_environment()
|
||||
|
||||
|
|
@ -34,6 +60,10 @@ def test_graph_environment_is_offline_deterministic_and_ignores_target_gitignore
|
|||
assert env["GITNEXUS_NO_GITIGNORE"] == "1"
|
||||
assert env["GITNEXUS_WORKER_POOL_SIZE"] == "1"
|
||||
assert env["GITNEXUS_PARSE_CHUNK_CONCURRENCY"] == "1"
|
||||
assert env["GITNEXUS_WORKER_READY_TIMEOUT_MS"] == str(
|
||||
sanitized_graph.GRAPH_WORKER_READY_TIMEOUT_MS
|
||||
)
|
||||
assert int(env["GITNEXUS_WORKER_READY_TIMEOUT_MS"]) >= 60_000
|
||||
assert "ANTHROPIC_API_KEY" not in env
|
||||
|
||||
|
||||
|
|
|
|||
327
eval/tests/test_session_progress.py
Normal file
327
eval/tests/test_session_progress.py
Normal file
|
|
@ -0,0 +1,327 @@
|
|||
"""Live progress reporting for long headless sessions.
|
||||
|
||||
Progress includes event metadata and bounded, redacted tool argument/result
|
||||
previews. Model prose and raw event streams are never echoed. These tests pin
|
||||
that boundary along with the signals that distinguish work from a wedged run.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import io
|
||||
import json
|
||||
import time
|
||||
|
||||
from workflow_bench.runner_sessions import SessionProgress
|
||||
|
||||
|
||||
def _drain_lines(stream: io.StringIO) -> list[str]:
|
||||
return [line for line in stream.getvalue().splitlines() if line.strip()]
|
||||
|
||||
|
||||
def _observe(progress: SessionProgress, chunk: bytes) -> None:
|
||||
progress.observe(chunk)
|
||||
progress._emit_pending()
|
||||
|
||||
|
||||
def test_progress_bounds_unanswered_tools_and_undrained_messages() -> None:
|
||||
progress = SessionProgress("bounded", stream=io.StringIO())
|
||||
for index in range(2000):
|
||||
progress.observe(
|
||||
(
|
||||
json.dumps(
|
||||
{
|
||||
"type": "assistant",
|
||||
"message": {
|
||||
"content": [
|
||||
{"type": "tool_use", "id": str(index), "name": "Bash", "input": {"command": "true"}}
|
||||
]
|
||||
},
|
||||
}
|
||||
)
|
||||
+ "\n"
|
||||
).encode()
|
||||
)
|
||||
assert len(progress._pending_tools) <= 256
|
||||
assert len(progress._pending_messages) <= 256
|
||||
assert "1999" in progress._pending_tools
|
||||
assert "0" not in progress._pending_tools
|
||||
progress.observe(
|
||||
(
|
||||
json.dumps(
|
||||
{
|
||||
"type": "user",
|
||||
"message": {
|
||||
"content": [{"type": "tool_result", "tool_use_id": "1999", "content": "recent result"}]
|
||||
},
|
||||
}
|
||||
)
|
||||
+ "\n"
|
||||
).encode()
|
||||
)
|
||||
progress._emit_pending()
|
||||
assert "recent result" in progress._stream.getvalue()
|
||||
assert "1999" not in progress._pending_tools
|
||||
|
||||
|
||||
def test_progress_reports_bounded_redacted_tool_io_but_never_model_prose() -> None:
|
||||
stream = io.StringIO()
|
||||
progress = SessionProgress(
|
||||
"gen 0 proposer",
|
||||
stream=stream,
|
||||
heartbeat_s=3600,
|
||||
secrets=("SECRET-TOKEN-abc123",),
|
||||
)
|
||||
events = [
|
||||
{"type": "system", "subtype": "init"},
|
||||
{
|
||||
"type": "assistant",
|
||||
"message": {
|
||||
"content": [
|
||||
{"type": "text", "text": "SECRET-REASONING-abc123"},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "t1",
|
||||
"name": "Grep",
|
||||
"input": {
|
||||
"pattern": "TODO",
|
||||
"path": "/workspace",
|
||||
"token": "SECRET-TOKEN-abc123",
|
||||
},
|
||||
},
|
||||
]
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "user",
|
||||
"message": {
|
||||
"content": [
|
||||
{
|
||||
"type": "tool_result",
|
||||
"tool_use_id": "t1",
|
||||
"is_error": False,
|
||||
"content": "src/a.py:1: TODO " + "x" * 1000,
|
||||
}
|
||||
]
|
||||
},
|
||||
},
|
||||
{"type": "result", "num_turns": 1, "is_error": False, "total_cost_usd": 1.5},
|
||||
]
|
||||
for event in events:
|
||||
_observe(progress, (json.dumps(event) + "\n").encode())
|
||||
|
||||
output = stream.getvalue()
|
||||
assert "SECRET-REASONING-abc123" not in output
|
||||
assert "SECRET-TOKEN-abc123" not in output
|
||||
assert "[REDACTED]" in output
|
||||
assert "session initialized" in output
|
||||
assert "turn 1 · Grep" in output
|
||||
assert 'tool Grep input={"pattern":"TODO","path":"/workspace","token":"[REDACTED]"}' in output
|
||||
assert "tool Grep result=ok output=" in output
|
||||
assert "truncated" in output
|
||||
assert "finished · 1 turns · ok · $1.50" in output
|
||||
|
||||
|
||||
def test_progress_reports_errors_and_mcp_io_but_skips_other_tool_payloads() -> None:
|
||||
stream = io.StringIO()
|
||||
progress = SessionProgress("flow", stream=stream, heartbeat_s=3600)
|
||||
events = [
|
||||
{
|
||||
"type": "assistant",
|
||||
"message": {
|
||||
"content": [
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "m1",
|
||||
"name": "mcp__gitnexus__query",
|
||||
"input": {"search_query": "call resolution"},
|
||||
},
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "e1",
|
||||
"name": "Edit",
|
||||
"input": {"file_path": "secret.py", "new_string": "do not log"},
|
||||
},
|
||||
]
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "user",
|
||||
"message": {
|
||||
"content": [
|
||||
{
|
||||
"type": "tool_result",
|
||||
"tool_use_id": "m1",
|
||||
"is_error": True,
|
||||
"content": "repository is not indexed",
|
||||
},
|
||||
{
|
||||
"type": "tool_result",
|
||||
"tool_use_id": "e1",
|
||||
"content": "edited secret.py",
|
||||
},
|
||||
]
|
||||
},
|
||||
},
|
||||
]
|
||||
for event in events:
|
||||
_observe(progress, (json.dumps(event) + "\n").encode())
|
||||
|
||||
output = stream.getvalue()
|
||||
assert 'tool mcp__gitnexus__query input={"search_query":"call resolution"}' in output
|
||||
assert 'tool mcp__gitnexus__query result=error output="repository is not indexed"' in output
|
||||
assert "do not log" not in output
|
||||
assert "edited secret.py" not in output
|
||||
|
||||
|
||||
def test_progress_distinguishes_mcp_semantic_errors_from_transport_success() -> None:
|
||||
stream = io.StringIO()
|
||||
progress = SessionProgress("flow", stream=stream, heartbeat_s=3600)
|
||||
events = [
|
||||
{
|
||||
"type": "assistant",
|
||||
"message": {
|
||||
"content": [
|
||||
{
|
||||
"type": "tool_use",
|
||||
"id": "m1",
|
||||
"name": "mcp__gitnexus__impact",
|
||||
"input": {"target": "missing", "direction": "upstream"},
|
||||
}
|
||||
]
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "user",
|
||||
"message": {
|
||||
"content": [
|
||||
{
|
||||
"type": "tool_result",
|
||||
"tool_use_id": "m1",
|
||||
"is_error": False,
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": '{"error":"Target missing not found"}\n\n---\n**Next:** retry',
|
||||
}
|
||||
],
|
||||
}
|
||||
]
|
||||
},
|
||||
},
|
||||
]
|
||||
for event in events:
|
||||
_observe(progress, (json.dumps(event) + "\n").encode())
|
||||
|
||||
assert "tool mcp__gitnexus__impact result=semantic-error" in stream.getvalue()
|
||||
|
||||
|
||||
def test_progress_calls_out_api_retries_because_that_is_the_stuck_signature() -> None:
|
||||
stream = io.StringIO()
|
||||
progress = SessionProgress("proposer", stream=stream, heartbeat_s=3600)
|
||||
event = {
|
||||
"type": "system",
|
||||
"subtype": "api_retry",
|
||||
"attempt": 7,
|
||||
"max_retries": 10,
|
||||
"retry_delay_ms": 34199.87,
|
||||
"error": "unknown",
|
||||
}
|
||||
_observe(progress, (json.dumps(event) + "\n").encode())
|
||||
|
||||
line = _drain_lines(stream)[-1]
|
||||
assert "API retry 7/10 in 34s" in line
|
||||
assert "no response from the model endpoint" in line
|
||||
|
||||
|
||||
def test_progress_speaks_up_while_a_session_is_silent() -> None:
|
||||
stream = io.StringIO()
|
||||
with SessionProgress("proposer", stream=stream, heartbeat_s=0.05):
|
||||
time.sleep(0.35)
|
||||
|
||||
heartbeats = [line for line in _drain_lines(stream) if "still running" in line]
|
||||
assert heartbeats, "a silent session must still report that it is alive"
|
||||
assert "0 turns" in heartbeats[0]
|
||||
|
||||
|
||||
def test_progress_survives_partial_chunks_garbage_and_unbounded_lines() -> None:
|
||||
stream = io.StringIO()
|
||||
progress = SessionProgress("proposer", stream=stream, heartbeat_s=3600)
|
||||
payload = json.dumps(
|
||||
{"type": "assistant", "message": {"content": [{"type": "tool_use", "id": "t1", "name": "Bash"}]}}
|
||||
).encode()
|
||||
# An event split across reads, non-JSON noise, and a huge newline-free run.
|
||||
_observe(progress, payload[:10])
|
||||
_observe(progress, payload[10:] + b"\nnot json at all\n")
|
||||
_observe(progress, b"x" * (4 * 1024 * 1024))
|
||||
_observe(progress, b'\n{"type":"result","num_turns":2,"is_error":true}\n')
|
||||
|
||||
output = stream.getvalue()
|
||||
assert "turn 1 · Bash" in output
|
||||
assert "finished · 2 turns · error" in output
|
||||
|
||||
|
||||
def test_progress_sanitizes_a_hostile_tool_name() -> None:
|
||||
stream = io.StringIO()
|
||||
progress = SessionProgress("proposer", stream=stream, heartbeat_s=3600)
|
||||
event = {
|
||||
"type": "assistant",
|
||||
"message": {"content": [{"type": "tool_use", "id": "t1", "name": "Bash\nFAKE-LOG-LINE injected"}]},
|
||||
}
|
||||
_observe(progress, (json.dumps(event) + "\n").encode())
|
||||
|
||||
assert "FAKE-LOG-LINE" not in stream.getvalue()
|
||||
assert len(_drain_lines(stream)) == 1
|
||||
|
||||
|
||||
def test_progress_redacts_non_ascii_secrets_before_json_escaping() -> None:
|
||||
stream = io.StringIO()
|
||||
secret = "tokén-密码"
|
||||
progress = SessionProgress("flow", stream=stream, heartbeat_s=3600, secrets=(secret,))
|
||||
event = {
|
||||
"type": "assistant",
|
||||
"message": {
|
||||
"content": [
|
||||
{"type": "tool_use", "id": "t1", "name": "Grep", "input": {"token": secret}},
|
||||
]
|
||||
},
|
||||
}
|
||||
_observe(progress, (json.dumps(event, ensure_ascii=False) + "\n").encode())
|
||||
|
||||
output = stream.getvalue()
|
||||
assert secret not in output
|
||||
assert json.dumps(secret)[1:-1] not in output
|
||||
assert "[REDACTED]" in output
|
||||
|
||||
|
||||
def test_cell_failure_detail_line_explains_why_a_cell_failed() -> None:
|
||||
from workflow_bench.runner import cell_failure_detail_line
|
||||
|
||||
assert cell_failure_detail_line("t", "workflow", 0, {"error_kind": None}) is None
|
||||
assert cell_failure_detail_line("t", "workflow", 0, {"error_kind": "x"}) is None
|
||||
|
||||
line = cell_failure_detail_line(
|
||||
"trivial-status-json-alias",
|
||||
"candidate_workflow",
|
||||
1,
|
||||
{
|
||||
"error_kind": "plan-evidence-invalid",
|
||||
"error_detail": "unauthorized workspace path\ntoken=sk-secret-value",
|
||||
},
|
||||
("sk-secret-value",),
|
||||
)
|
||||
assert line is not None
|
||||
assert line.startswith("[trivial-status-json-alias][candidate_workflow][run 1] detail: ")
|
||||
assert "unauthorized workspace path" in line
|
||||
assert "sk-secret-value" not in line
|
||||
assert "\n" not in line
|
||||
|
||||
|
||||
def test_cell_failure_detail_line_bounds_a_huge_detail() -> None:
|
||||
from workflow_bench.runner import MAX_CELL_DETAIL_CHARS, cell_failure_detail_line
|
||||
|
||||
line = cell_failure_detail_line(
|
||||
"t", "workflow", 0, {"error_kind": "session-error", "error_detail": {"stdout_tail": "y" * 50_000}}
|
||||
)
|
||||
assert line is not None
|
||||
assert "truncated" in line
|
||||
assert len(line) < MAX_CELL_DETAIL_CHARS + 200
|
||||
|
|
@ -5,15 +5,15 @@ from __future__ import annotations
|
|||
import os
|
||||
import stat
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
from pathlib import Path, PurePosixPath
|
||||
|
||||
import pytest
|
||||
|
||||
from workflow_bench.proposer_sandbox import VITE_TEMP_DIR, SandboxError
|
||||
from workflow_bench.oracle_assets import TaskOracleSnapshot
|
||||
from workflow_bench.runner_tasks import resolve_task_bindings
|
||||
from workflow_bench.task_assets import TaskAssetCache, stage_task_assets
|
||||
from workflow_bench import task_assets
|
||||
from workflow_bench.task_assets import TaskAssetCache, _is_harness_sandbox_copy, stage_task_assets
|
||||
from workflow_bench import runtime_mounts, task_assets
|
||||
|
||||
|
||||
SHA = "a" * 40
|
||||
|
|
@ -443,3 +443,53 @@ def test_non_node_modules_dependency_snapshot_has_no_vite_temp(tmp_path: Path) -
|
|||
snapshot = cache.prepare(task, repo=repo, resolved_sha=SHA)
|
||||
captured = {entry.path.as_posix() for entry in snapshot.dependencies[0].entries}
|
||||
assert not any(path.endswith(VITE_TEMP_DIR) for path in captured)
|
||||
|
||||
|
||||
def test_review_case_sandbox_copy_is_read_from_the_harness_not_the_task_repo(
|
||||
monkeypatch, tmp_path: Path
|
||||
) -> None:
|
||||
repo = tmp_path / "task-repo"
|
||||
repo.mkdir()
|
||||
(repo / "eval" / "workflow_bench").mkdir(parents=True)
|
||||
harness = tmp_path / "harness"
|
||||
patch = harness / "eval" / "workflow_bench" / "review_cases" / "pr-2718.patch"
|
||||
patch.parent.mkdir(parents=True)
|
||||
patch.write_bytes(b"diff --git a/a b/a\n")
|
||||
monkeypatch.setattr(runtime_mounts, "HARNESS_ROOT", harness)
|
||||
|
||||
task = {
|
||||
"sandbox_copy": ["eval/workflow_bench/review_cases/pr-2718.patch"],
|
||||
"sandbox_dependencies": [],
|
||||
}
|
||||
with TaskAssetCache(tmp_path / "cache") as cache:
|
||||
snapshot = cache.prepare(task, repo=repo, resolved_sha=SHA)
|
||||
copied = snapshot.root / "sandbox-copy" / "eval" / "workflow_bench" / "review_cases" / "pr-2718.patch"
|
||||
assert copied.read_bytes() == b"diff --git a/a b/a\n"
|
||||
|
||||
|
||||
def test_review_case_sandbox_copy_does_not_fall_back_to_the_task_repo(
|
||||
monkeypatch, tmp_path: Path
|
||||
) -> None:
|
||||
repo = tmp_path / "task-repo"
|
||||
planted = repo / "eval" / "workflow_bench" / "review_cases" / "pr-2718.patch"
|
||||
planted.parent.mkdir(parents=True)
|
||||
planted.write_bytes(b"from-task-repo")
|
||||
harness = tmp_path / "harness"
|
||||
harness.mkdir()
|
||||
monkeypatch.setattr(runtime_mounts, "HARNESS_ROOT", harness)
|
||||
|
||||
task = {
|
||||
"sandbox_copy": ["eval/workflow_bench/review_cases/pr-2718.patch"],
|
||||
"sandbox_dependencies": [],
|
||||
}
|
||||
with TaskAssetCache(tmp_path / "cache") as cache:
|
||||
with pytest.raises(SandboxError, match="unavailable"):
|
||||
cache.prepare(task, repo=repo, resolved_sha=SHA)
|
||||
|
||||
|
||||
def test_harness_sandbox_copy_does_not_treat_parent_escapes_as_corpus() -> None:
|
||||
assert _is_harness_sandbox_copy(PurePosixPath("eval/workflow_bench/review_cases/pr.patch"))
|
||||
assert not _is_harness_sandbox_copy(
|
||||
PurePosixPath("eval/workflow_bench/review_cases/../oracles/hidden.json")
|
||||
)
|
||||
assert not _is_harness_sandbox_copy(PurePosixPath("eval/workflow_bench/oracles"))
|
||||
|
|
|
|||
|
|
@ -1,7 +1,9 @@
|
|||
"""Unit tests for workflow benchmark aggregation, reporting, task, and CI contracts."""
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import shlex
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
|
|
@ -89,7 +91,7 @@ def task_row(task_id: str, **overrides):
|
|||
"command": "true",
|
||||
"files": [
|
||||
{
|
||||
"source": "trivial-version-alias.oracle.test.ts",
|
||||
"source": "trivial-status-json-alias.oracle.test.ts",
|
||||
"target": "oracle.test.ts",
|
||||
}
|
||||
],
|
||||
|
|
@ -189,11 +191,13 @@ def test_eval_ci_uses_locked_uv_and_blocking_native_containment_jobs():
|
|||
assert claude_lock["packages"]["node_modules/@anthropic-ai/claude-code"]["integrity"].startswith("sha512-")
|
||||
assert "if(p.version!=='2.1.214') process.exit(1)" in workflow
|
||||
assert "'2.1.214 (Claude Code)'" in workflow
|
||||
assert containment_steps["Build pinned shared runtime"]["working-directory"] == "gitnexus-shared"
|
||||
assert containment_steps["Build pinned shared runtime"]["run"].splitlines() == [
|
||||
"npm ci",
|
||||
"npm run build",
|
||||
]
|
||||
# Shared is compiled by gitnexus `npm run build` (scripts/build.js runTsc).
|
||||
# A dedicated npm ci in gitnexus-shared pulls TypeScript 7 and stalls CI.
|
||||
assert "Build pinned shared runtime" not in containment_steps
|
||||
assert not any(
|
||||
step.get("working-directory") == "gitnexus-shared" and "npm ci" in str(step.get("run", ""))
|
||||
for step in containment["steps"]
|
||||
)
|
||||
assert containment_steps["Install and build pinned GitNexus runtime"]["working-directory"] == "gitnexus"
|
||||
assert containment_steps["Install and build pinned GitNexus runtime"]["run"].splitlines() == [
|
||||
"npm ci",
|
||||
|
|
@ -234,8 +238,8 @@ def test_shipped_scenarios_opt_out_the_cross_module_cell_and_rebuild_graph_asset
|
|||
tasks = yaml.safe_load(task_file.read_text())["tasks"]
|
||||
selected, skipped = select_tasks(tasks, include_expensive=False)
|
||||
assert [task["id"] for task in selected] == [
|
||||
"trivial-version-alias",
|
||||
"inv-bug-pdg-note",
|
||||
"trivial-status-json-alias",
|
||||
"inv-bug-c-system-include",
|
||||
"inv-feature-list-repos-filter",
|
||||
]
|
||||
assert skipped == ["cross-module-parse-retry"]
|
||||
|
|
@ -318,6 +322,36 @@ def test_aggregate_excludes_unverified_transcript_evidence():
|
|||
assert agg["excluded_runs"] == 1
|
||||
|
||||
|
||||
def test_aggregate_excludes_invalid_review_artifacts_from_quality_metrics():
|
||||
scored = record(
|
||||
cost_usd=1.0,
|
||||
review_weighted_f1=0.8,
|
||||
review_true_positives=2,
|
||||
review_false_positives=0,
|
||||
review_false_negatives=1,
|
||||
review_precision=1.0,
|
||||
review_recall=0.67,
|
||||
review_f1=0.8,
|
||||
review_weighted_precision=0.8,
|
||||
review_weighted_recall=0.8,
|
||||
review_blocker_recall=1.0,
|
||||
review_severity_accuracy=1.0,
|
||||
review_category_accuracy=1.0,
|
||||
review_grounded_evidence=1.0,
|
||||
review_clean_control=False,
|
||||
)
|
||||
agg = aggregate(
|
||||
[
|
||||
scored,
|
||||
record(cost_usd=2.0, resolved=False, error_kind="review-evidence-invalid"),
|
||||
]
|
||||
)
|
||||
assert agg["valid_runs"] == 1
|
||||
assert agg["excluded_runs"] == 1
|
||||
assert agg["review_weighted_f1"] == 0.8
|
||||
assert agg["review_true_positives"] == 2
|
||||
|
||||
|
||||
def test_render_report_surfaces_excluded_and_unverified_runs():
|
||||
results = {
|
||||
"t": {
|
||||
|
|
@ -425,3 +459,56 @@ def test_outage_streak_flag_defaults_and_disables():
|
|||
base = ["--tasks", "tasks.yaml", "--model", "claude-sonnet-4-20250514"]
|
||||
assert build_parser().parse_args(base).outage_streak == 5
|
||||
assert build_parser().parse_args([*base, "--outage-streak", "0"]).outage_streak == 0
|
||||
|
||||
|
||||
def test_run_evolution_script_is_the_shared_ci_and_local_entrypoint():
|
||||
eval_dir = Path(__file__).resolve().parents[1]
|
||||
script = eval_dir / "workflow_bench" / "run-evolution.sh"
|
||||
workflow = eval_dir.parent / ".github" / "workflows" / "gitnexus-skill-evolution.yml"
|
||||
assert script.is_file()
|
||||
assert script.stat().st_mode & 0o111
|
||||
workflow_text = workflow.read_text()
|
||||
assert "./workflow_bench/run-evolution.sh --apply" in workflow_text
|
||||
assert "python -m workflow_bench.evolve" not in workflow_text
|
||||
|
||||
env = {
|
||||
"PATH": os.environ.get("PATH", "/usr/bin"),
|
||||
"MODEL": "claude-sonnet-5",
|
||||
"PROPOSER_MODEL": "claude-opus-4-8",
|
||||
"EFFORT": "xhigh",
|
||||
"GENERATIONS": "1",
|
||||
"RUNS": "3",
|
||||
"WORKERS": "2",
|
||||
"PROVIDER": "openai",
|
||||
"INCLUDE_EXPENSIVE": "1",
|
||||
"SEED_RESULTS": "/tmp/seed-bench",
|
||||
"CLAUDE_BIN": "/opt/claude",
|
||||
"OUT_ROOT": "/tmp/wfevolve",
|
||||
"CE_PLUGIN_DIR": "/tmp/ce-plugin",
|
||||
"CE_PLUGIN_VERSION": "3.24.0",
|
||||
"HOME": os.environ.get("HOME", "/tmp"),
|
||||
}
|
||||
printed = subprocess.run(
|
||||
[str(script), "--dry-run", "--apply"],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
env=env,
|
||||
)
|
||||
argv = shlex.split(printed.stdout)
|
||||
assert argv[:7] == ["uv", "run", "--locked", "--extra", "dev", "python", "-m"]
|
||||
assert argv[7] == "workflow_bench.evolve"
|
||||
assert argv[argv.index("--tasks") + 1] == "workflow_bench/tasks.review.scenarios.yaml"
|
||||
assert argv[argv.index("--arms") + 1] == "review"
|
||||
assert argv[argv.index("--ce-plugin-version") + 1] == "3.24.0"
|
||||
assert argv[argv.index("--model") + 1] == "gpt-5.6-sol"
|
||||
assert argv[argv.index("--proposer-model") + 1] == "gpt-5.6-sol"
|
||||
assert argv[argv.index("--effort") + 1] == "xhigh"
|
||||
assert argv[argv.index("--workers") + 1] == "2"
|
||||
assert argv[argv.index("--claude-bin") + 1] == "/opt/claude"
|
||||
assert argv[argv.index("--out-root") + 1] == "/tmp/wfevolve"
|
||||
assert argv[argv.index("--seed-results") + 1] == "/tmp/seed-bench"
|
||||
assert "--apply" in argv
|
||||
assert "--include-expensive" in argv
|
||||
assert "claude-sonnet-5" not in argv
|
||||
assert printed.stderr # rewrite notice goes to stderr
|
||||
|
|
|
|||
|
|
@ -8,14 +8,18 @@ from types import SimpleNamespace
|
|||
import pytest
|
||||
|
||||
from workflow_bench.evolution import (
|
||||
CANDIDATE_SKILLS,
|
||||
MAX_CANDIDATE_ENTRIES,
|
||||
apply_candidate_overlay,
|
||||
candidate_overlay_digest,
|
||||
evaluate_candidate,
|
||||
evaluate_review_candidate,
|
||||
required_candidate_arms,
|
||||
seed_evaluated_skills,
|
||||
skill_fingerprint,
|
||||
unexercised_overlay_skills,
|
||||
)
|
||||
from workflow_bench.promotion_apply import mirror_targets
|
||||
from workflow_bench.process_control import ManagedProcessResult
|
||||
from workflow_bench.runner import aggregate, build_parser
|
||||
|
||||
|
|
@ -185,8 +189,7 @@ def test_candidate_overlay_is_skill_only_and_content_addressed(tmp_path):
|
|||
review_skill = review_overlay / ".claude" / "skills" / "gitnexus-review" / "SKILL.md"
|
||||
review_skill.parent.mkdir(parents=True)
|
||||
review_skill.write_text("review candidate\n")
|
||||
with pytest.raises(ValueError, match="plan,work"):
|
||||
candidate_overlay_digest(review_overlay)
|
||||
assert candidate_overlay_digest(review_overlay)
|
||||
|
||||
invalid = tmp_path / "invalid"
|
||||
source = invalid / "gitnexus" / "src" / "cli" / "index.ts"
|
||||
|
|
@ -238,6 +241,116 @@ def test_required_candidate_arms_are_minimal_for_touched_skills(tmp_path):
|
|||
"candidate_workflow_direct",
|
||||
]
|
||||
|
||||
review = tmp_path / "review"
|
||||
write_overlay_skill(review, "gitnexus-review")
|
||||
assert required_candidate_arms(review) == ["candidate_review"]
|
||||
|
||||
|
||||
def test_review_gate_is_quality_first_and_requires_repeated_evidence():
|
||||
def arm(score, blocker=1.0, false_positives=0, runs=3):
|
||||
return {
|
||||
"runs": runs,
|
||||
"valid_runs": runs,
|
||||
"excluded_runs": 0,
|
||||
"class": "review-defect",
|
||||
"review_weighted_f1": score,
|
||||
"review_blocker_recall": blocker,
|
||||
"review_false_positives": false_positives,
|
||||
"review_clean_control": False,
|
||||
"review_verdict_correct": True,
|
||||
}
|
||||
|
||||
decision = evaluate_review_candidate(
|
||||
{
|
||||
"case-a": {
|
||||
"review": arm(0.6),
|
||||
"candidate_review": arm(0.8),
|
||||
}
|
||||
},
|
||||
incumbent_arm="review",
|
||||
candidate_arm="candidate_review",
|
||||
model="pinned-model",
|
||||
)
|
||||
assert decision["decision"] == "promote"
|
||||
|
||||
regression = evaluate_review_candidate(
|
||||
{
|
||||
"case-a": {
|
||||
"review": arm(0.6, blocker=1.0),
|
||||
"candidate_review": arm(0.8, blocker=0.0),
|
||||
}
|
||||
},
|
||||
incumbent_arm="review",
|
||||
candidate_arm="candidate_review",
|
||||
model="pinned-model",
|
||||
)
|
||||
assert regression["decision"] == "keep_incumbent"
|
||||
assert any("blocker recall" in reason for reason in regression["reasons"])
|
||||
|
||||
|
||||
def test_review_gate_rejects_added_false_positives_on_clean_controls():
|
||||
base = {
|
||||
"runs": 3,
|
||||
"valid_runs": 3,
|
||||
"excluded_runs": 0,
|
||||
"class": "review-clean",
|
||||
"review_weighted_f1": 1.0,
|
||||
"review_blocker_recall": 1.0,
|
||||
"review_clean_control": True,
|
||||
"review_clean_pass": True,
|
||||
"review_verdict_correct": True,
|
||||
}
|
||||
decision = evaluate_review_candidate(
|
||||
{
|
||||
"clean": {
|
||||
"review": {**base, "review_false_positives": 0},
|
||||
"candidate_review": {**base, "review_false_positives": 1, "review_clean_pass": False},
|
||||
}
|
||||
},
|
||||
incumbent_arm="review",
|
||||
candidate_arm="candidate_review",
|
||||
model="pinned-model",
|
||||
)
|
||||
assert decision["decision"] == "keep_incumbent"
|
||||
assert any("clean control" in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
@pytest.mark.parametrize("wrong_verdict,bad_blocker", [(True, False), (False, True), (False, False)])
|
||||
def test_review_gate_preserves_every_repeat_safeguard(wrong_verdict, bad_blocker):
|
||||
from workflow_bench.runner import aggregate
|
||||
|
||||
def row(score, verdict=True, blocker=1.0):
|
||||
return {
|
||||
"resolved": True,
|
||||
"review_weighted_f1": score,
|
||||
"review_blocker_recall": blocker,
|
||||
"review_false_positives": 0,
|
||||
"review_verdict_correct": verdict,
|
||||
"review_clean_control": False,
|
||||
"review_clean_pass": False,
|
||||
}
|
||||
|
||||
incumbent = aggregate([row(0.5) for _ in range(3)])
|
||||
candidate = aggregate([row(0.8), row(0.8), row(0.8, not wrong_verdict, 0.0 if bad_blocker else 1.0)])
|
||||
decision = evaluate_review_candidate(
|
||||
{"case": {"review": incumbent, "candidate_review": candidate}},
|
||||
incumbent_arm="review",
|
||||
candidate_arm="candidate_review",
|
||||
model="pinned-model",
|
||||
)
|
||||
assert decision["decision"] == ("keep_incumbent" if wrong_verdict or bad_blocker else "promote")
|
||||
|
||||
|
||||
def test_review_gate_treats_an_empty_corpus_as_insufficient_evidence():
|
||||
decision = evaluate_review_candidate(
|
||||
{},
|
||||
incumbent_arm="review",
|
||||
candidate_arm="candidate_review",
|
||||
model="pinned-model",
|
||||
)
|
||||
assert decision["decision"] == "insufficient_evidence"
|
||||
assert any("no paired review task results" in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="candidate overlays require the Linux outer sandbox")
|
||||
def test_apply_candidate_overlay_creates_a_clean_ephemeral_commit(tmp_path):
|
||||
|
|
@ -314,6 +427,12 @@ def test_apply_candidate_overlay_creates_a_clean_ephemeral_commit(tmp_path):
|
|||
) == candidate_overlay_digest(overlay)
|
||||
assert incumbent.read_text() == "candidate\n"
|
||||
git_commands = [command for command in sandbox.commands if command[0] == "/usr/bin/git"]
|
||||
assert git_commands[0][-4:] == [
|
||||
"add",
|
||||
"-f",
|
||||
"--",
|
||||
".claude/skills/gitnexus-work/SKILL.md",
|
||||
]
|
||||
assert [command[-1] for command in git_commands[:2]] == [
|
||||
".claude/skills/gitnexus-work/SKILL.md",
|
||||
"--",
|
||||
|
|
@ -349,6 +468,181 @@ def test_candidate_overlay_rejects_linked_destination_parents(tmp_path):
|
|||
apply_candidate_overlay(overlay, repo, sandbox=sandbox)
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="candidate overlays require the Linux outer sandbox")
|
||||
def test_apply_candidate_overlay_force_adds_historically_ignored_skill(tmp_path):
|
||||
repo = tmp_path / "repo"
|
||||
repo.mkdir()
|
||||
subprocess.run(["git", "init", "--quiet", str(repo)], check=True)
|
||||
(repo / ".gitignore").write_text(".claude/skills/*\n")
|
||||
(repo / "README").write_text("subject\n")
|
||||
subprocess.run(["git", "-C", str(repo), "add", "."], check=True)
|
||||
subprocess.run(
|
||||
[
|
||||
"git",
|
||||
"-C",
|
||||
str(repo),
|
||||
"-c",
|
||||
"user.name=test",
|
||||
"-c",
|
||||
"user.email=test@invalid",
|
||||
"commit",
|
||||
"--quiet",
|
||||
"-m",
|
||||
"historical checkout that ignores skills",
|
||||
],
|
||||
check=True,
|
||||
)
|
||||
|
||||
overlay = tmp_path / "candidate"
|
||||
write_overlay_skill(overlay, "gitnexus-review")
|
||||
|
||||
class LocalSandbox:
|
||||
def __init__(self):
|
||||
self.clone = repo
|
||||
|
||||
def run(self, command, **kwargs):
|
||||
if command[0] == "/bin/mkdir":
|
||||
return ManagedProcessResult(
|
||||
state="exited",
|
||||
returncode=0,
|
||||
stdout_tail="",
|
||||
stderr_tail="",
|
||||
duration_s=0.0,
|
||||
)
|
||||
translated = [str(repo) if item == "/workspace" else item for item in command]
|
||||
completed = subprocess.run(
|
||||
translated,
|
||||
cwd=repo,
|
||||
env=dict(kwargs["env"]),
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
)
|
||||
return ManagedProcessResult(
|
||||
state="exited",
|
||||
returncode=completed.returncode,
|
||||
stdout_tail=completed.stdout,
|
||||
stderr_tail=completed.stderr,
|
||||
duration_s=0.0,
|
||||
)
|
||||
|
||||
apply_candidate_overlay(overlay, repo, sandbox=LocalSandbox())
|
||||
assert (repo / ".claude" / "skills" / "gitnexus-review" / "SKILL.md").read_text() == (
|
||||
"gitnexus-review candidate\n"
|
||||
)
|
||||
status = subprocess.run(
|
||||
["git", "-C", str(repo), "status", "--porcelain"],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
)
|
||||
assert status.stdout == ""
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="skill seeds require the Linux outer sandbox")
|
||||
def test_seed_evaluated_skills_installs_missing_review_skill_and_is_idempotent(tmp_path):
|
||||
repo = tmp_path / "clone"
|
||||
repo.mkdir()
|
||||
subprocess.run(["git", "init", "--quiet", str(repo)], check=True)
|
||||
# Historical review SHAs ignore the whole skill tree and lack today's
|
||||
# `!.claude/skills/gitnexus-review/` allowlist. Seeding must still commit.
|
||||
(repo / ".gitignore").write_text(".claude/skills/*\n")
|
||||
(repo / "README").write_text("subject\n")
|
||||
subprocess.run(["git", "-C", str(repo), "add", "."], check=True)
|
||||
subprocess.run(
|
||||
[
|
||||
"git",
|
||||
"-C",
|
||||
str(repo),
|
||||
"-c",
|
||||
"user.name=test",
|
||||
"-c",
|
||||
"user.email=test@invalid",
|
||||
"commit",
|
||||
"--quiet",
|
||||
"-m",
|
||||
"historical checkout without review skill",
|
||||
],
|
||||
check=True,
|
||||
)
|
||||
before = subprocess.run(
|
||||
["git", "-C", str(repo), "rev-parse", "HEAD"],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
).stdout.strip()
|
||||
|
||||
source = tmp_path / "harness"
|
||||
skill = source / ".claude" / "skills" / "gitnexus-review" / "SKILL.md"
|
||||
persona = source / ".claude" / "skills" / "gitnexus-review" / "ci-personas" / "lens.md"
|
||||
persona.parent.mkdir(parents=True)
|
||||
skill.write_text("current incumbent review skill\n")
|
||||
persona.write_text("persona\n")
|
||||
|
||||
class LocalSandbox:
|
||||
def __init__(self):
|
||||
self.clone = repo
|
||||
|
||||
def run(self, command, **kwargs):
|
||||
if command[0] == "/bin/mkdir":
|
||||
return ManagedProcessResult(
|
||||
state="exited",
|
||||
returncode=0,
|
||||
stdout_tail="",
|
||||
stderr_tail="",
|
||||
duration_s=0.0,
|
||||
)
|
||||
translated = [str(repo) if item == "/workspace" else item for item in command]
|
||||
completed = subprocess.run(
|
||||
translated,
|
||||
cwd=repo,
|
||||
env=dict(kwargs["env"]),
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
)
|
||||
return ManagedProcessResult(
|
||||
state="exited",
|
||||
returncode=completed.returncode,
|
||||
stdout_tail=completed.stdout,
|
||||
stderr_tail=completed.stderr,
|
||||
duration_s=0.0,
|
||||
)
|
||||
|
||||
sandbox = LocalSandbox()
|
||||
seed_evaluated_skills(source, repo, sandbox=sandbox, arm="review")
|
||||
assert skill_fingerprint(repo, "review") is not None
|
||||
assert (repo / ".claude" / "skills" / "gitnexus-review" / "SKILL.md").read_text() == (
|
||||
"current incumbent review skill\n"
|
||||
)
|
||||
assert (
|
||||
repo / ".claude" / "skills" / "gitnexus-review" / "ci-personas" / "lens.md"
|
||||
).read_text() == "persona\n"
|
||||
status = subprocess.run(
|
||||
["git", "-C", str(repo), "status", "--porcelain"],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
)
|
||||
assert status.stdout == ""
|
||||
after = subprocess.run(
|
||||
["git", "-C", str(repo), "rev-parse", "HEAD"],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
).stdout.strip()
|
||||
assert after != before
|
||||
|
||||
seed_evaluated_skills(source, repo, sandbox=sandbox, arm="review")
|
||||
again = subprocess.run(
|
||||
["git", "-C", str(repo), "rev-parse", "HEAD"],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
).stdout.strip()
|
||||
assert again == after
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name == "nt", reason="skill links are rejected by the Linux sandbox harness")
|
||||
def test_skill_fingerprint_rejects_linked_skill_roots(tmp_path):
|
||||
outside = tmp_path / "outside"
|
||||
|
|
@ -424,9 +718,7 @@ def test_candidate_gate_refuses_promotion_on_unmeasured_cost():
|
|||
results = {
|
||||
"task-a": {
|
||||
"workflow_direct": aggregate([record(cost_usd=1.0) for _ in range(3)]),
|
||||
"candidate_workflow_direct": aggregate(
|
||||
[record(cost_usd=0.1), record(cost_usd=None), record(cost_usd=0.1)]
|
||||
),
|
||||
"candidate_workflow_direct": aggregate([record(cost_usd=0.1), record(cost_usd=None), record(cost_usd=0.1)]),
|
||||
}
|
||||
}
|
||||
decision = evaluate_candidate(
|
||||
|
|
@ -510,8 +802,23 @@ def test_candidate_gate_rejects_a_partial_candidate_even_with_a_resolution_edge(
|
|||
assert any("oracle-backed quality floor" in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
@pytest.mark.parametrize("resolved", [0, 2])
|
||||
def test_candidate_gate_never_promotes_zero_or_partial_success_for_efficiency(resolved):
|
||||
@pytest.mark.parametrize(
|
||||
("resolved", "expected_decision", "expected_reason"),
|
||||
[
|
||||
# Nothing resolved anywhere: the task is ungated, which leaves the
|
||||
# generation with no quality signal at all — refuse outright rather
|
||||
# than rank a 100x cost win across runs that all failed the oracle.
|
||||
(0, "insufficient_evidence", "no task supplied quality signal"),
|
||||
# Partial success on a task the incumbent also partly resolves stays a
|
||||
# quality-floor rejection: the candidate has to be reliable, not lucky.
|
||||
(2, "keep_incumbent", "oracle-backed quality floor"),
|
||||
],
|
||||
)
|
||||
def test_candidate_gate_never_promotes_zero_or_partial_success_for_efficiency(
|
||||
resolved,
|
||||
expected_decision,
|
||||
expected_reason,
|
||||
):
|
||||
incumbent_records = [record(cost_usd=1.0, resolved=index < resolved) for index in range(3)]
|
||||
candidate_records = [record(cost_usd=0.01, resolved=index < resolved) for index in range(3)]
|
||||
decision = evaluate_candidate(
|
||||
|
|
@ -526,9 +833,184 @@ def test_candidate_gate_never_promotes_zero_or_partial_success_for_efficiency(re
|
|||
model="pinned-model",
|
||||
)
|
||||
|
||||
assert decision["decision"] == "keep_incumbent"
|
||||
assert decision["decision"] == expected_decision
|
||||
assert decision["tasks"][0]["candidate_quality_floor_met"] is False
|
||||
assert any("oracle-backed quality floor" in reason for reason in decision["reasons"])
|
||||
assert any(expected_reason in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
def test_a_task_no_arm_can_resolve_is_reported_but_does_not_veto_promotion():
|
||||
# inv-feature-list-repos-filter fails its hidden oracle on every run of
|
||||
# both arms. Gating on it made promotion unreachable for as long as it
|
||||
# stayed in the set, while saying nothing about the candidate.
|
||||
solvable = {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=index > 1) for index in range(3)]),
|
||||
"candidate_workflow": aggregate([record(cost_usd=1.0) for _ in range(3)]),
|
||||
}
|
||||
unsolvable = {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=False) for _ in range(3)]),
|
||||
"candidate_workflow": aggregate([record(cost_usd=1.3, resolved=False) for _ in range(3)]),
|
||||
}
|
||||
decision = evaluate_candidate(
|
||||
{"task-a": solvable, "task-impossible": unsolvable},
|
||||
incumbent_arm="workflow",
|
||||
candidate_arm="candidate_workflow",
|
||||
model="pinned-model",
|
||||
)
|
||||
|
||||
assert decision["decision"] == "promote"
|
||||
assert decision["ungated_tasks"] == ["task-impossible"]
|
||||
assert decision["gated_tasks"] == ["task-a"]
|
||||
assert [row["gated"] for row in decision["tasks"]] == [True, False]
|
||||
# The ungated task's 30% cost regression stays under the failed-task cap
|
||||
# but must not reach the median or the (tighter) gated per-task cap.
|
||||
assert decision["median_improvement_pct"] == 0.0
|
||||
assert not any("above the" in reason for reason in decision["reasons"])
|
||||
# One aggregate line, so a growing set of unsolvable tasks cannot crowd the
|
||||
# real verdict out of the three reasons the proposer is shown — and it
|
||||
# discloses how much of the set the verdict actually rests on.
|
||||
assert [reason for reason in decision["reasons"] if "not gated on" in reason] == [
|
||||
"not gated on 1 task(s) neither arm resolved: task-impossible (evidence base: 1/2 paired tasks gated)"
|
||||
]
|
||||
|
||||
|
||||
def test_an_ungated_task_still_ranks_against_the_failed_task_cost_cap():
|
||||
# Leaving the quality gate is not leaving the spend gate: burning 9x the
|
||||
# incumbent's cost to fail the same oracle is a regression the gate has to
|
||||
# see, or a candidate can hide unbounded waste inside "task health".
|
||||
solvable = {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=index > 1) for index in range(3)]),
|
||||
"candidate_workflow": aggregate([record(cost_usd=1.0) for _ in range(3)]),
|
||||
}
|
||||
unsolvable = {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=False) for _ in range(3)]),
|
||||
"candidate_workflow": aggregate([record(cost_usd=9.0, resolved=False) for _ in range(3)]),
|
||||
}
|
||||
|
||||
decision = evaluate_candidate(
|
||||
{"task-a": solvable, "task-impossible": unsolvable},
|
||||
incumbent_arm="workflow",
|
||||
candidate_arm="candidate_workflow",
|
||||
model="pinned-model",
|
||||
)
|
||||
|
||||
assert decision["decision"] == "keep_incumbent"
|
||||
assert decision["ungated_tasks"] == ["task-impossible"]
|
||||
assert any("failed-task cap" in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
def test_a_mutually_failed_task_stays_gated_when_the_skill_never_loaded():
|
||||
# skill-not-invoked is prompt evidence, not task health: the skill under
|
||||
# test never ran, so the task cannot be written off as beyond both arms.
|
||||
results = {
|
||||
"task-a": {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=False) for _ in range(3)]),
|
||||
"candidate_workflow": aggregate(
|
||||
[record(cost_usd=0.01, resolved=False, error_kind="skill-not-invoked") for _ in range(3)]
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
decision = evaluate_candidate(
|
||||
results,
|
||||
incumbent_arm="workflow",
|
||||
candidate_arm="candidate_workflow",
|
||||
model="pinned-model",
|
||||
)
|
||||
|
||||
assert decision["ungated_tasks"] == []
|
||||
assert decision["tasks"][0]["gated"] is True
|
||||
assert decision["tasks"][0]["skill_attributable_failure"] is True
|
||||
# Gated with teeth: the 99% cost "win" must not carry a candidate whose
|
||||
# skill never loaded.
|
||||
assert decision["decision"] == "keep_incumbent"
|
||||
assert any("never invoked the skill under test" in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
def test_a_mutually_failed_task_stays_gated_when_its_metric_was_never_measured():
|
||||
# Ungating is a claim about spend as well as quality. With no measured
|
||||
# cost there is nothing to claim, so the task stays in the gate and the
|
||||
# missing measurement is named instead of silently skipped.
|
||||
results = {
|
||||
"task-a": {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=False) for _ in range(3)]),
|
||||
"candidate_workflow": aggregate(
|
||||
[record(cost_usd=None, resolved=False), *(record(cost_usd=0.1, resolved=False) for _ in range(2))]
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
decision = evaluate_candidate(
|
||||
results,
|
||||
incumbent_arm="workflow",
|
||||
candidate_arm="candidate_workflow",
|
||||
model="pinned-model",
|
||||
)
|
||||
|
||||
assert decision["ungated_tasks"] == []
|
||||
assert decision["decision"] == "insufficient_evidence"
|
||||
assert any("was not measured on every run" in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
def test_partial_progress_on_a_task_the_incumbent_never_resolves_is_not_punished():
|
||||
# Resolving 1 of 3 runs where the incumbent resolves none is strictly
|
||||
# better than resolving none — which the gate ungates and forgives. Holding
|
||||
# the partial run to the quality floor made improvement score worse than
|
||||
# inaction.
|
||||
def outcome(candidate_resolved: int) -> dict[str, object]:
|
||||
return evaluate_candidate(
|
||||
{
|
||||
"task-a": {
|
||||
"workflow": aggregate([record(cost_usd=1.0) for _ in range(3)]),
|
||||
"candidate_workflow": aggregate([record(cost_usd=0.5) for _ in range(3)]),
|
||||
},
|
||||
"task-hard": {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=False) for _ in range(3)]),
|
||||
"candidate_workflow": aggregate(
|
||||
[record(cost_usd=1.0, resolved=index < candidate_resolved) for index in range(3)]
|
||||
),
|
||||
},
|
||||
},
|
||||
incumbent_arm="workflow",
|
||||
candidate_arm="candidate_workflow",
|
||||
model="pinned-model",
|
||||
)
|
||||
|
||||
no_progress = outcome(0)
|
||||
some_progress = outcome(1)
|
||||
|
||||
assert no_progress["decision"] == "promote"
|
||||
assert no_progress["ungated_tasks"] == ["task-hard"]
|
||||
# The partial run gives the task quality signal, so it is gated — but as
|
||||
# improvement, not as a floor failure the zero-progress candidate escapes.
|
||||
assert some_progress["decision"] == "promote"
|
||||
assert some_progress["ungated_tasks"] == []
|
||||
assert some_progress["tasks"][1]["quality_floor_enforced"] is False
|
||||
assert not any("quality floor" in reason for reason in some_progress["reasons"])
|
||||
|
||||
|
||||
def test_promotion_requires_a_gated_majority_of_the_paired_tasks():
|
||||
# Two of three tasks written off as task health leaves one task deciding
|
||||
# the whole promotion. Ungating keeps promotion reachable; it must not
|
||||
# hollow out the evidence base that makes a promotion mean anything.
|
||||
solvable = {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=index > 1) for index in range(3)]),
|
||||
"candidate_workflow": aggregate([record(cost_usd=0.1) for _ in range(3)]),
|
||||
}
|
||||
unsolvable = {
|
||||
"workflow": aggregate([record(cost_usd=1.0, resolved=False) for _ in range(3)]),
|
||||
"candidate_workflow": aggregate([record(cost_usd=1.0, resolved=False) for _ in range(3)]),
|
||||
}
|
||||
|
||||
decision = evaluate_candidate(
|
||||
{"task-a": solvable, "task-impossible": unsolvable, "task-impossible-2": dict(unsolvable)},
|
||||
incumbent_arm="workflow",
|
||||
candidate_arm="candidate_workflow",
|
||||
model="pinned-model",
|
||||
)
|
||||
|
||||
assert decision["decision"] == "insufficient_evidence"
|
||||
assert decision["gated_tasks"] == ["task-a"]
|
||||
assert any("evidence base is too thin" in reason for reason in decision["reasons"])
|
||||
|
||||
|
||||
def test_candidate_gate_promotes_on_a_two_run_resolution_margin():
|
||||
|
|
@ -553,3 +1035,34 @@ def test_overlay_skills_must_be_exercised_by_selected_candidate_arms(tmp_path):
|
|||
write_overlay_skill(plan_overlay, "gitnexus-plan")
|
||||
assert unexercised_overlay_skills(plan_overlay, ["candidate_workflow_direct"]) == ["gitnexus-plan"]
|
||||
assert unexercised_overlay_skills(plan_overlay, ["candidate_workflow"]) == []
|
||||
|
||||
|
||||
@pytest.mark.parametrize("skill", sorted(CANDIDATE_SKILLS))
|
||||
def test_a_promoted_skill_is_visible_to_git_status_in_every_shipped_tree(skill):
|
||||
"""A promotion the repository cannot see is a promotion that never happens.
|
||||
|
||||
The workflow detects an applied promotion with `git status --porcelain`,
|
||||
which is blind to ignored paths, and `.claude/skills/*` is ignored with a
|
||||
hand-maintained per-skill allowlist. A candidate skill missing from that
|
||||
allowlist would leave the run reporting "No promotion this run" after the
|
||||
gate had already said promote — silently, and only after a full generation
|
||||
of benchmark spend.
|
||||
"""
|
||||
repo_root = Path(__file__).resolve().parents[2]
|
||||
from pathlib import PurePosixPath
|
||||
|
||||
targets = [
|
||||
str(path.parent)
|
||||
for path in mirror_targets(PurePosixPath(".claude/skills") / skill / "SKILL.md")
|
||||
]
|
||||
ignored = [
|
||||
target
|
||||
for target in targets
|
||||
if subprocess.run(
|
||||
["git", "check-ignore", "-q", f"{target}/SKILL.md"],
|
||||
cwd=repo_root,
|
||||
check=False,
|
||||
).returncode
|
||||
== 0
|
||||
]
|
||||
assert ignored == []
|
||||
|
|
|
|||
|
|
@ -4,6 +4,7 @@ import argparse
|
|||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
|
@ -14,6 +15,7 @@ import pytest
|
|||
from workflow_bench import evolve, runner, runner_sessions, runtime_mounts
|
||||
from workflow_bench.evolution import skill_fingerprint
|
||||
from workflow_bench.process_control import ManagedProcessResult
|
||||
from workflow_bench.proposer_sandbox import SandboxError
|
||||
from workflow_bench.runner import snapshot_plan_docs
|
||||
|
||||
|
||||
|
|
@ -71,6 +73,7 @@ def bench_args(**overrides):
|
|||
"claude_bin": "claude",
|
||||
"timeout": 5,
|
||||
"model": None,
|
||||
"effort": "xhigh",
|
||||
"base_url": None,
|
||||
"auth_token": None,
|
||||
"permission_mode": None,
|
||||
|
|
@ -118,6 +121,7 @@ def skill_events(skill_input: dict, *, tool_id: str = "skill-1", is_error: bool
|
|||
|
||||
def fake_sandbox(root: Path) -> SimpleNamespace:
|
||||
return SimpleNamespace(
|
||||
backend="test-double",
|
||||
claude_bin="claude",
|
||||
clone=root,
|
||||
private_root=root,
|
||||
|
|
@ -167,6 +171,25 @@ def test_run_claude_forwards_the_named_model_to_every_session(monkeypatch, tmp_p
|
|||
assert captured[captured.index("--model") + 1] == "claude-sonnet-4-20250514"
|
||||
|
||||
|
||||
def test_run_claude_forwards_xhigh_effort_to_every_session(monkeypatch, tmp_path):
|
||||
captured: list[str] = []
|
||||
|
||||
def fake_run(command, **kwargs):
|
||||
captured.extend(command)
|
||||
return fake_cli_result(VALID_REPORT)
|
||||
|
||||
monkeypatch.setattr(runner_sessions, "run_managed", fake_run)
|
||||
runner.run_claude(
|
||||
"task",
|
||||
tmp_path,
|
||||
claude_bin="claude",
|
||||
timeout=5,
|
||||
model="gpt-5.6-sol",
|
||||
effort="xhigh",
|
||||
)
|
||||
assert captured[captured.index("--effort") + 1] == "xhigh"
|
||||
|
||||
|
||||
def test_run_claude_restricts_tools_via_tools_flag_outside_bare(monkeypatch, tmp_path):
|
||||
# Outside --bare, the built-in toolset defaults to everything (subagents,
|
||||
# WebFetch, Task, ...) and --allowedTools only pre-approves within that —
|
||||
|
|
@ -295,10 +318,18 @@ def test_run_arm_keeps_session_error_kind_over_verify(monkeypatch, tmp_path):
|
|||
|
||||
def test_agent_tool_grants_are_exact_and_nomcp_has_no_graph_tools(monkeypatch, tmp_path):
|
||||
read_only = runner.allowed_agent_tools(implementation=False)
|
||||
review_tools = runner.allowed_agent_tools(implementation=False, allow_edit=False)
|
||||
implementation = runner.allowed_agent_tools(implementation=True)
|
||||
no_mcp = runner.allowed_agent_tools(implementation=True, include_mcp=False)
|
||||
|
||||
assert read_only == [*runner.BUILTIN_AGENT_TOOLS, *runner.GITNEXUS_READ_ONLY_TOOLS]
|
||||
assert review_tools == [
|
||||
tool
|
||||
for tool in [*runner.BUILTIN_AGENT_TOOLS, *runner.GITNEXUS_READ_ONLY_TOOLS]
|
||||
if tool != "Edit"
|
||||
]
|
||||
assert "Write" in review_tools
|
||||
assert "Edit" not in review_tools
|
||||
assert implementation == [
|
||||
*runner.BUILTIN_AGENT_TOOLS,
|
||||
*runner.GITNEXUS_READ_ONLY_TOOLS,
|
||||
|
|
@ -326,7 +357,7 @@ def test_agent_tool_grants_are_exact_and_nomcp_has_no_graph_tools(monkeypatch, t
|
|||
)
|
||||
|
||||
assert captured[0]["allowed_tools"] == read_only # planning
|
||||
assert captured[1]["allowed_tools"] == read_only # review
|
||||
assert captured[1]["allowed_tools"] == review_tools # review
|
||||
assert captured[2]["allowed_tools"] == implementation
|
||||
assert captured[3]["allowed_tools"] == list(runner.BUILTIN_AGENT_TOOLS)
|
||||
assert captured[3]["mcp_config_json"] == '{"mcpServers":{}}'
|
||||
|
|
@ -355,7 +386,7 @@ def test_mcp_config_uses_only_the_minimal_pinned_harness_runtime(monkeypatch, tm
|
|||
directory.mkdir(parents=True)
|
||||
(runtime / "dist" / "cli" / "index.js").write_text("")
|
||||
(runtime / "hooks" / "claude" / "resolve-analyze-cmd.cjs").write_text("")
|
||||
(runtime / "package.json").write_text(json.dumps({"version": runner.PINNED_GITNEXUS_VERSION}))
|
||||
(runtime / "package.json").write_text(json.dumps({"version": "9.9.9-test"}))
|
||||
(runtime / "node_modules" / "gitnexus-shared").symlink_to(shared, target_is_directory=True)
|
||||
(shared / "package.json").write_text(json.dumps({"name": "gitnexus-shared"}))
|
||||
monkeypatch.setattr(runtime_mounts, "HARNESS_ROOT", tmp_path)
|
||||
|
|
@ -379,9 +410,6 @@ def test_mcp_config_uses_only_the_minimal_pinned_harness_runtime(monkeypatch, tm
|
|||
(shared / "package.json", f"{runner.SANDBOX_GITNEXUS_SHARED}/package.json"),
|
||||
(runtime / "hooks" / "claude", f"{runner.SANDBOX_GITNEXUS}/hooks/claude"),
|
||||
]
|
||||
package = json.loads((runtime / "package.json").read_text())
|
||||
assert package["version"] == runner.PINNED_GITNEXUS_VERSION
|
||||
|
||||
mounted_sources = {mount.source for mount in mounts}
|
||||
mounted_targets = {mount.target for mount in mounts}
|
||||
assert runtime not in mounted_sources
|
||||
|
|
@ -399,6 +427,118 @@ def test_mcp_config_uses_only_the_minimal_pinned_harness_runtime(monkeypatch, tm
|
|||
assert f"{runner.SANDBOX_GITNEXUS}/hooks" not in mounted_targets
|
||||
|
||||
|
||||
def _install_pinned_runtime(root: Path) -> None:
|
||||
runtime = root / "gitnexus"
|
||||
shared = root / "gitnexus-shared"
|
||||
for directory in (
|
||||
runtime / "dist" / "cli",
|
||||
runtime / "node_modules",
|
||||
runtime / "vendor",
|
||||
runtime / "hooks" / "claude",
|
||||
shared / "dist",
|
||||
):
|
||||
directory.mkdir(parents=True)
|
||||
(runtime / "dist" / "cli" / "index.js").write_text("")
|
||||
(runtime / "hooks" / "claude" / "resolve-analyze-cmd.cjs").write_text("")
|
||||
(runtime / "package.json").write_text(json.dumps({"version": "9.9.9-test"}))
|
||||
(runtime / "node_modules" / "gitnexus-shared").symlink_to(shared, target_is_directory=True)
|
||||
(shared / "package.json").write_text(json.dumps({"name": "gitnexus-shared"}))
|
||||
|
||||
|
||||
def test_runtime_mounts_reuse_primary_checkout_node_modules_from_a_worktree(
|
||||
monkeypatch, tmp_path
|
||||
) -> None:
|
||||
primary = tmp_path / "primary"
|
||||
worktree = tmp_path / "worktree"
|
||||
_install_pinned_runtime(primary)
|
||||
(primary / ".git" / "worktrees" / "wt").mkdir(parents=True)
|
||||
_install_pinned_runtime(worktree)
|
||||
shutil.rmtree(worktree / "gitnexus" / "node_modules")
|
||||
(worktree / "gitnexus" / "node_modules").symlink_to(
|
||||
primary / "gitnexus" / "node_modules",
|
||||
target_is_directory=True,
|
||||
)
|
||||
(worktree / ".git").write_text(f"gitdir: {primary / '.git' / 'worktrees' / 'wt'}\n")
|
||||
monkeypatch.setattr(runtime_mounts, "HARNESS_ROOT", worktree)
|
||||
|
||||
mounts = runner.trusted_gitnexus_runtime_mounts()
|
||||
by_target = {mount.target: mount.source for mount in mounts}
|
||||
|
||||
assert by_target[f"{runner.SANDBOX_GITNEXUS}/node_modules"] == (
|
||||
primary / "gitnexus" / "node_modules"
|
||||
)
|
||||
assert by_target[f"{runner.SANDBOX_GITNEXUS_SHARED}/package.json"] == (
|
||||
primary / "gitnexus-shared" / "package.json"
|
||||
)
|
||||
assert by_target[f"{runner.SANDBOX_GITNEXUS}/dist"] == worktree / "gitnexus" / "dist"
|
||||
|
||||
|
||||
def test_runtime_mounts_reuse_primary_shared_when_only_the_inner_link_points_there(
|
||||
monkeypatch, tmp_path
|
||||
) -> None:
|
||||
primary = tmp_path / "primary"
|
||||
worktree = tmp_path / "worktree"
|
||||
_install_pinned_runtime(primary)
|
||||
(primary / ".git" / "worktrees" / "wt").mkdir(parents=True)
|
||||
_install_pinned_runtime(worktree)
|
||||
linked = worktree / "gitnexus" / "node_modules" / "gitnexus-shared"
|
||||
linked.unlink()
|
||||
linked.symlink_to(primary / "gitnexus-shared", target_is_directory=True)
|
||||
(worktree / ".git").write_text(f"gitdir: {primary / '.git' / 'worktrees' / 'wt'}\n")
|
||||
monkeypatch.setattr(runtime_mounts, "HARNESS_ROOT", worktree)
|
||||
|
||||
mounts = runner.trusted_gitnexus_runtime_mounts()
|
||||
by_target = {mount.target: mount.source for mount in mounts}
|
||||
|
||||
assert by_target[f"{runner.SANDBOX_GITNEXUS}/node_modules"] == (
|
||||
worktree / "gitnexus" / "node_modules"
|
||||
)
|
||||
assert by_target[f"{runner.SANDBOX_GITNEXUS_SHARED}/package.json"] == (
|
||||
primary / "gitnexus-shared" / "package.json"
|
||||
)
|
||||
|
||||
|
||||
def test_runtime_mounts_reject_a_node_modules_symlink_outside_the_primary_checkout(
|
||||
monkeypatch, tmp_path
|
||||
) -> None:
|
||||
primary = tmp_path / "primary"
|
||||
worktree = tmp_path / "worktree"
|
||||
outsider = tmp_path / "outsider"
|
||||
_install_pinned_runtime(primary)
|
||||
_install_pinned_runtime(outsider)
|
||||
(primary / ".git" / "worktrees" / "wt").mkdir(parents=True)
|
||||
_install_pinned_runtime(worktree)
|
||||
shutil.rmtree(worktree / "gitnexus" / "node_modules")
|
||||
(worktree / "gitnexus" / "node_modules").symlink_to(
|
||||
outsider / "gitnexus" / "node_modules",
|
||||
target_is_directory=True,
|
||||
)
|
||||
(worktree / ".git").write_text(f"gitdir: {primary / '.git' / 'worktrees' / 'wt'}\n")
|
||||
monkeypatch.setattr(runtime_mounts, "HARNESS_ROOT", worktree)
|
||||
|
||||
with pytest.raises(SandboxError, match="primary checkout"):
|
||||
runner.trusted_gitnexus_runtime_mounts()
|
||||
|
||||
|
||||
def test_runtime_mounts_reject_a_node_modules_symlink_in_a_regular_checkout(
|
||||
monkeypatch, tmp_path
|
||||
) -> None:
|
||||
checkout = tmp_path / "checkout"
|
||||
other = tmp_path / "other"
|
||||
_install_pinned_runtime(checkout)
|
||||
_install_pinned_runtime(other)
|
||||
(checkout / ".git").mkdir()
|
||||
shutil.rmtree(checkout / "gitnexus" / "node_modules")
|
||||
(checkout / "gitnexus" / "node_modules").symlink_to(
|
||||
other / "gitnexus" / "node_modules",
|
||||
target_is_directory=True,
|
||||
)
|
||||
monkeypatch.setattr(runtime_mounts, "HARNESS_ROOT", checkout)
|
||||
|
||||
with pytest.raises(SandboxError, match="must be a real directory"):
|
||||
runner.trusted_gitnexus_runtime_mounts()
|
||||
|
||||
|
||||
@pytest.mark.skipif(
|
||||
os.environ.get("GITNEXUS_REQUIRE_BWRAP_CANARY") != "1",
|
||||
reason="real Bubblewrap canary is mandatory in the named Ubuntu CI job",
|
||||
|
|
@ -458,7 +598,11 @@ def test_real_bubblewrap_runtime_mount_imports_cli_without_exposing_checkout(tmp
|
|||
assert visibility.ok, visibility.stderr_tail
|
||||
assert imported.ok, imported.stderr_tail
|
||||
assert analyze_imported.ok, analyze_imported.stderr_tail
|
||||
assert imported.stdout_tail.strip() == runner.PINNED_GITNEXUS_VERSION
|
||||
# The runtime the sandbox sees must be the one this checkout built —
|
||||
# compared against the checkout itself rather than a constant, so a release
|
||||
# bump cannot fail a benchmark that is running exactly what it should.
|
||||
built = json.loads((runtime_mounts.HARNESS_ROOT / "gitnexus" / "package.json").read_text())
|
||||
assert imported.stdout_tail.strip() == built["version"]
|
||||
|
||||
|
||||
def test_isolated_mcp_registry_contains_only_the_sandbox_clone(tmp_path):
|
||||
|
|
@ -641,6 +785,23 @@ def test_skill_invocation_parses_supported_exact_identifier_fields(skill_input):
|
|||
)
|
||||
|
||||
|
||||
def test_skill_invocation_accepts_plugin_qualified_identifier():
|
||||
assert (
|
||||
runner_sessions.skill_was_invoked_events(
|
||||
skill_events({"skill": "compound-engineering:ce-code-review"}),
|
||||
"ce-code-review",
|
||||
)
|
||||
is True
|
||||
)
|
||||
assert (
|
||||
runner_sessions.skill_was_invoked_events(
|
||||
skill_events({"skill": "compound-engineering:ce-plan"}),
|
||||
"ce-code-review",
|
||||
)
|
||||
is False
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"skill_input",
|
||||
[
|
||||
|
|
@ -967,6 +1128,35 @@ def test_final_result_event_must_be_last(monkeypatch, tmp_path):
|
|||
assert "not the last event" in rec["error_detail"]["event_stream_error"]
|
||||
|
||||
|
||||
def test_background_task_teardown_after_the_result_stays_valid_evidence(monkeypatch, tmp_path):
|
||||
# Claude Code drains background-task bookkeeping after the final result
|
||||
# event. Those `system` events carry no tool or usage payload, so they must
|
||||
# not invalidate an otherwise complete session (run 29907431284 lost three
|
||||
# runs this way, and the gate demands zero excluded runs).
|
||||
teardown = [
|
||||
{"type": "system", "subtype": "background_tasks_changed", "tasks": []},
|
||||
{"type": "system", "subtype": "task_updated", "task_id": "bdw43oy7j", "patch": {"status": "killed"}},
|
||||
{"type": "system", "subtype": "task_notification", "task_id": "bdw43oy7j", "status": "stopped"},
|
||||
]
|
||||
stream = event_stream(*skill_events({"skill": "gitnexus-work"})) + "".join(
|
||||
json.dumps(event) + "\n" for event in teardown
|
||||
)
|
||||
monkeypatch.setattr(runner_sessions, "run_managed", lambda *a, **k: fake_cli_result(stream))
|
||||
rec = runner.run_claude(
|
||||
"task",
|
||||
tmp_path,
|
||||
claude_bin="claude",
|
||||
timeout=5,
|
||||
expected_skill="gitnexus-work",
|
||||
)
|
||||
|
||||
assert rec["ok"] is True
|
||||
assert rec["error_kind"] is None
|
||||
assert rec["skill_invoked"] is True
|
||||
assert rec["transcript_missing"] is False
|
||||
assert "evidence_diagnostics" not in rec
|
||||
|
||||
|
||||
def test_snapshot_plan_docs_detects_one_modified_plan_and_rejects_ambiguous_output(tmp_path):
|
||||
plans = tmp_path / "docs" / "plans"
|
||||
plans.mkdir(parents=True)
|
||||
|
|
@ -1075,7 +1265,9 @@ def test_review_phase_rejects_workspace_or_skill_mutation(
|
|||
expected_skill_digest = "expected-skill-fingerprint"
|
||||
|
||||
def adversarial_review(prompt, *args, **kwargs):
|
||||
(tmp_path / "review-output.md").write_text("review findings")
|
||||
(tmp_path / "review-output.json").write_text(
|
||||
'{"schema_version":1,"verdict":"approve","findings":[]}'
|
||||
)
|
||||
if attack == "workspace":
|
||||
source.write_text("review silently changed source")
|
||||
return session_record()
|
||||
|
|
|
|||
1119
eval/uv.lock
generated
1119
eval/uv.lock
generated
File diff suppressed because it is too large
Load diff
|
|
@ -1,4 +1,4 @@
|
|||
# Workflow benchmark — observe the token savings
|
||||
# Skill benchmark — evolve review quality, measure workflow cost
|
||||
|
||||
Measures whether the `gitnexus-plan` → `gitnexus-work` engineering workflow
|
||||
actually saves tokens versus a baseline agent on the same tasks, using real
|
||||
|
|
@ -8,18 +8,19 @@ report.
|
|||
|
||||
## What it compares
|
||||
|
||||
| Arm | Sessions | Notes |
|
||||
| --- | --- | --- |
|
||||
| `workflow` | `gitnexus-plan` on the task, then `gitnexus-work` on the produced plan | The skills must be installed (`gitnexus setup`, or repo-local `.claude/skills/`) |
|
||||
| `candidate_workflow` | same sessions as `workflow`, with a candidate skill overlay | Paired with `workflow` on the same task/ref/model |
|
||||
| `workflow_direct` | one `gitnexus-work` direct-mode session | The middle option — execution discipline without a planning pass |
|
||||
| `candidate_workflow_direct` | same session as `workflow_direct`, with a candidate skill overlay | Paired with `workflow_direct` on the same task/ref/model |
|
||||
| `ce_workflow` | `ce-plan` on the task, then `ce-work` on the produced plan | External comparator: the explicitly supplied, pinned compound-engineering plugin's plan→work family |
|
||||
| `ce_workflow_direct` | one `ce-work` direct-mode session | External comparator paired with `workflow_direct` |
|
||||
| `review` | one `gitnexus-review` session over local uncommitted changes | The task's `setup` applies the diff under review; the review is written to `review-output.md` so `verify` can gate on it |
|
||||
| `ce_review` | one `ce-code-review` session over the same changes | External comparator paired with `review` |
|
||||
| `baseline` | one session with the identical task text | `--disallowedTools Skill` so it cannot borrow the workflow; same repo, same MCP tools |
|
||||
| `baseline_nomcp` | like baseline, graph tools also disallowed | Separates the workflow-discipline question from the GitNexus-tools question (off by default) |
|
||||
| Arm | Sessions | Notes |
|
||||
| --------------------------- | ---------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
|
||||
| `workflow` | `gitnexus-plan` on the task, then `gitnexus-work` on the produced plan | The skills must be installed (`gitnexus setup`, or repo-local `.claude/skills/`) |
|
||||
| `candidate_workflow` | same sessions as `workflow`, with a candidate skill overlay | Paired with `workflow` on the same task/ref/model |
|
||||
| `workflow_direct` | one `gitnexus-work` direct-mode session | The middle option — execution discipline without a planning pass |
|
||||
| `candidate_workflow_direct` | same session as `workflow_direct`, with a candidate skill overlay | Paired with `workflow_direct` on the same task/ref/model |
|
||||
| `ce_workflow` | `ce-plan` on the task, then `ce-work` on the produced plan | External comparator: the explicitly supplied, pinned compound-engineering plugin's plan→work family |
|
||||
| `ce_workflow_direct` | one `ce-work` direct-mode session | External comparator paired with `workflow_direct` |
|
||||
| `review` | one `gitnexus-review` session over an immutable historical PR snapshot | Emits strict `review-output.json`; hidden human labels score quality after the session |
|
||||
| `candidate_review` | the same review with a `gitnexus-review` candidate overlay | Paired with `review` on the same case/ref/model/runtime |
|
||||
| `ce_review` | one pinned `ce-code-review` session over the same changes | External comparator paired with both review arms |
|
||||
| `baseline` | one session with the identical task text | `--disallowedTools Skill` so it cannot borrow the workflow; same repo, same MCP tools |
|
||||
| `baseline_nomcp` | like baseline, graph tools also disallowed | Separates the workflow-discipline question from the GitNexus-tools question (off by default) |
|
||||
|
||||
Every arm runs in a fresh detached git worktree of the task's `ref`, once per
|
||||
`--runs`. The model-visible `verify` command is recorded as
|
||||
|
|
@ -36,7 +37,7 @@ lfg's gate and work's direct-mode triage should encode.
|
|||
|
||||
```bash
|
||||
cd eval
|
||||
export GITNEXUS_BENCH_AUTH_TOKEN="$ANTHROPIC_API_KEY"
|
||||
export GITNEXUS_BENCH_ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY"
|
||||
uv run --locked --extra dev python -m workflow_bench.runner \
|
||||
--tasks workflow_bench/tasks.scenarios.yaml --runs 3 \
|
||||
--model claude-sonnet-4-20250514
|
||||
|
|
@ -120,10 +121,15 @@ digest. Files written beneath the agent's `$HOME` are never trusted as
|
|||
evidence.
|
||||
|
||||
Bare mode is deliberately non-interactive: it does not consult a stored
|
||||
Claude login/keychain or `ANTHROPIC_AUTH_TOKEN`. Supply one explicit API or
|
||||
proxy key through `GITNEXUS_BENCH_AUTH_TOKEN` (preferred) or `--auth-token`;
|
||||
Claude login/keychain or `ANTHROPIC_AUTH_TOKEN`. Supply an Anthropic API key
|
||||
through `GITNEXUS_BENCH_ANTHROPIC_API_KEY` (preferred) or `--anthropic-api-key`;
|
||||
the harness maps it to `ANTHROPIC_API_KEY` only for the trusted Claude parent
|
||||
and scrubs it from agent-launched tools.
|
||||
and scrubs it from agent-launched tools. `GITNEXUS_BENCH_AUTH_TOKEN` and
|
||||
`--auth-token` remain as aliases. OpenAI keys are not a drop-in
|
||||
replacement: pass `--openai-api-key` / `GITNEXUS_BENCH_OPENAI_API_KEY` with
|
||||
`gpt-*` / `o*` / `openai/*` model ids and the harness starts a loopback
|
||||
LiteLLM proxy. The OpenAI key stays on that host process; Claude still sees
|
||||
only a minted `ANTHROPIC_API_KEY` plus `ANTHROPIC_BASE_URL`.
|
||||
|
||||
The trusted Claude CLI still needs outbound access to the explicitly supplied
|
||||
model endpoint. This is not a network broker, so the CLI itself retains that
|
||||
|
|
@ -133,6 +139,28 @@ Native benchmark execution is therefore Linux/WSL2-only. Evidence assembly
|
|||
and hand-authored overlay preparation can happen elsewhere, but
|
||||
`--initial-overlay` does not bypass containment.
|
||||
|
||||
For a local diagnostic inside a container that blocks user namespaces, an
|
||||
operator may explicitly choose the non-containment host backend:
|
||||
|
||||
```bash
|
||||
UNSAFE_NO_BWRAP=1 RUNS=1 ./workflow_bench/run-evolution.sh
|
||||
```
|
||||
|
||||
This mode runs review sessions directly in disposable host worktrees and is
|
||||
**not** a security boundary: it does not isolate the network or create a PID
|
||||
namespace, and a session that can `chmod` can undo the workspace lock. The
|
||||
harness still drops write bits on the clone except `review-output.json` so
|
||||
accidental `npm install` / analyze writes cannot invalidate review evidence.
|
||||
Sandbox cleanup restores owner write bits before deleting the private TMPDIR,
|
||||
because a session that `copytree`s the locked clone would otherwise leave
|
||||
non-empty 0555 directories that `rmtree` cannot remove. Historical review
|
||||
SHAs that gitignore `.claude/skills/*` are force-added when the harness seeds
|
||||
or overlays the evaluated `gitnexus-review` skill.
|
||||
Treat model and verifier processes as able to access host files and
|
||||
credentials available to the invoking user. It is restricted to the review
|
||||
benchmark, forbidden with `--apply` and whenever `CI` is set;
|
||||
promotion-capable and CI runs must use Bubblewrap.
|
||||
|
||||
## Prompt and skill evolution loop
|
||||
|
||||
Prompts age as models and tool harnesses change. Treat the current skills and
|
||||
|
|
@ -161,7 +189,7 @@ paid work. For a work overlay:
|
|||
cd eval
|
||||
uv run --locked --extra dev python -m workflow_bench.runner \
|
||||
--tasks workflow_bench/tasks.scenarios.yaml \
|
||||
--runs 3 --model claude-sonnet-4-20250514 \
|
||||
--runs 3 --workers 1 --model claude-sonnet-4-20250514 \
|
||||
--arms workflow candidate_workflow \
|
||||
workflow_direct candidate_workflow_direct \
|
||||
--candidate-overlay /tmp/gn-skill-candidate
|
||||
|
|
@ -177,7 +205,7 @@ artifacts. Those artifacts are the trajectory evidence: cluster failures and
|
|||
expensive detours, propose one bounded prompt change, and feed it back as the
|
||||
next overlay.
|
||||
|
||||
When candidate arms are present the runner also writes schema-3
|
||||
When candidate arms are present the runner also writes schema-6
|
||||
`promotion.json`. It
|
||||
binds the immutable overlay digest, benchmark model, truthful candidate origin
|
||||
(a named proposer model or `manual-initial-overlay`), selected
|
||||
|
|
@ -186,9 +214,36 @@ immutable dependency bytes, committed base digest of every apply
|
|||
destination, exact required arms, thresholds, and evidence expiry. Its default
|
||||
deterministic gate is deliberately conservative:
|
||||
|
||||
Schema 6 binds a separate policy to each required candidate arm and records
|
||||
whether the sweep completed. Apply validates the paired metrics and recomputes
|
||||
each decision. Historical schema 5 reports remain readable; regenerate their
|
||||
benchmark evidence before applying an overlay. Editing a schema number does
|
||||
not supply the missing evidence.
|
||||
|
||||
Review candidates optimize weighted F1 with a minimum improvement of 0.01,
|
||||
complete paired evidence on every selected task, and no per-task quality
|
||||
regression. Complete misses score zero. Matching uses maximum cardinality
|
||||
throughout the 100-finding limit. Downgraded findings receive at most their
|
||||
reported severity's weight; only blocking-severity matches count toward blocker
|
||||
recall. Every valid candidate repeat must have the correct verdict, and the
|
||||
minimum blocker recall across repeats must not regress. Clean controls retain
|
||||
their false-positive and verdict safeguards. Implementation candidates retain
|
||||
the efficiency policy below:
|
||||
|
||||
- at least 3 paired VALID runs per task, zero excluded runs in either arm
|
||||
(session/infra-error rows therefore block promotion), and a named model;
|
||||
- the candidate must pass the hidden oracle on every valid run for every task;
|
||||
- a fully measured task that neither arm ever resolves remains reported but is
|
||||
ungated from the quality comparison — only if its metric was measured in both
|
||||
arms and no run hit `skill-not-invoked` (a skill that never loaded is prompt
|
||||
evidence, not task health). An ungated task still ranks against a looser 100%
|
||||
failed-task regression cap on the promotion metric;
|
||||
- at least half the paired tasks must stay gated, and `promotion.json` discloses
|
||||
the gated/ungated split per decision; a set with no gated task at all is
|
||||
`insufficient_evidence`;
|
||||
- the candidate must pass the hidden oracle on every valid run of every gated
|
||||
task the incumbent resolves at least once — on a task the incumbent never
|
||||
resolves, partial candidate progress counts as improvement instead of failing
|
||||
the floor, so making some progress is never scored worse than making none;
|
||||
- no per-task resolution-rate regression (quality is lexicographically first);
|
||||
- promotion by resolution needs a margin of at least 2 resolved runs —
|
||||
a 1-run difference is noise at this run count and falls through to the
|
||||
|
|
@ -219,28 +274,82 @@ without weakening today's deterministic promotion boundary.
|
|||
|
||||
### Closing the loop automatically (`evolve.py`)
|
||||
|
||||
The evolution workflow runs an offline containment preflight with the pinned
|
||||
Claude Code 2.1.214 binary before starting a paid proposer or benchmark. The
|
||||
review canary seals the workspace read-only and exposes only the pre-created
|
||||
`review-output.json` as writable. Runtime mount placeholders are prepared in
|
||||
the disposable clone before sealing it; existing config bytes are preserved.
|
||||
Any pre-existing result entry, including a symlink, is rejected. Required
|
||||
canaries fail when their runtime or Bubblewrap is unavailable.
|
||||
|
||||
The default outage limit is five consecutive unusable results, across task
|
||||
boundaries. Invalid review JSON advances this limit even when a skill or session
|
||||
error was recorded first. A valid zero-quality review resets it. Concurrent
|
||||
waves can exceed the limit by at most `workers - 1` completed cells; no further
|
||||
wave starts after a trip. Completed rows and redacted diagnostics remain in the
|
||||
partial report, the runner exits nonzero, and the evolution driver stops without
|
||||
applying or starting another generation.
|
||||
|
||||
SIGINT and SIGTERM propagate one cancellation event through managed commands,
|
||||
including clone, setup, Claude, and verification. Executor submissions copy
|
||||
the run context so indirect subprocess helpers receive the same event. Active
|
||||
process groups or Windows Job Objects are terminated and workers joined before
|
||||
shared assets or the gateway are released. Controlled cancellation tests require
|
||||
cleanup within 15 seconds. Cancellation remains distinct from timeout and
|
||||
quality failure in recorded evidence.
|
||||
|
||||
The gateway runs under a private supervisor watching a pipe owned only by the
|
||||
harness. Parent exit, including SIGKILL, closes that pipe and stops the proxy
|
||||
group; Windows also retains kill-on-close Job Object ownership. Keep completed
|
||||
JSONL rows, transcripts, the partial report, and gateway diagnostics when
|
||||
investigating an interrupted run. A subsequent paid comparison needs fresh
|
||||
evidence from all arms under the same dependency lock. LiteLLM pricing comes
|
||||
from that locked release's local cost map; compare no old/new-lock costs as
|
||||
quality evidence.
|
||||
|
||||
`workflow_bench.evolve` automates the three manual arrows — propose,
|
||||
benchmark, apply — without moving the trust boundary:
|
||||
|
||||
```bash
|
||||
cd eval
|
||||
uv run --locked --extra dev python -m workflow_bench.evolve \
|
||||
--tasks workflow_bench/tasks.scenarios.yaml \
|
||||
--model claude-sonnet-4-20250514 --generations 2 \
|
||||
--seed-results results/wfbench-<prior-run> # optional gen-0 evidence
|
||||
./workflow_bench/run-evolution.sh # local; no working-tree apply
|
||||
./workflow_bench/run-evolution.sh --apply # CI; same argv the workflow uses
|
||||
./workflow_bench/run-evolution.sh --dry-run # print the evolve command
|
||||
```
|
||||
|
||||
Each generation: a confined **proposer** session reads the incumbent plan/work
|
||||
skills, the prior generation's `results.jsonl`
|
||||
loser rows, their session transcripts and patches, and the learning queue,
|
||||
The GitHub skill-evolution job calls this script. Do not invoke
|
||||
`python -m workflow_bench.evolve` directly for a full loop. Environment knobs
|
||||
match the workflow: `MODEL`, `PROPOSER_MODEL`, `GENERATIONS`, `RUNS`,
|
||||
`WORKERS`, `PROVIDER`, `EFFORT`, `SEED_RESULTS`, `INCLUDE_EXPENSIVE`. The
|
||||
checked-in production defaults are `PROVIDER=openai`, `MODEL=gpt-5.6-sol`,
|
||||
`PROPOSER_MODEL=gpt-5.6-sol`, and `EFFORT=xhigh`.
|
||||
|
||||
The scheduled/default profile is read-only review evolution. Set
|
||||
`EVOLUTION_PROFILE=implementation` explicitly to run the legacy plan/work
|
||||
benchmark. Review mode requires `CE_PLUGIN_DIR` and `CE_PLUGIN_VERSION`.
|
||||
|
||||
Each review generation: a confined **proposer** session reads only the incumbent
|
||||
`gitnexus-review` skill, normalized CE/incumbent/candidate result rows, bounded
|
||||
review artifacts and session transcripts, and the rejected
|
||||
`proposal.md` when available (including a workflow seed from a prior run), and
|
||||
the learning queue,
|
||||
then writes ONE bounded candidate overlay plus a reviewer-facing
|
||||
`proposal.md`. The overlay is re-validated by `candidate_overlay_files`
|
||||
(same boundary: Markdown under the plan/work trees, nothing else), frozen,
|
||||
`proposal.md`. The proposer's clone is sanitized exactly like an arm's before
|
||||
its session starts: it authors the artifact the arms are scored with, so
|
||||
letting it read `eval/workflow_bench` would hand it the task prompts and the
|
||||
hidden oracles it is about to be graded against, and a proposal could win the
|
||||
gate by encoding the expected behavior into a skill rather than by being a
|
||||
better skill. The overlay is re-validated by `candidate_overlay_files`
|
||||
(same boundary: Markdown under `gitnexus-review`, including exercised
|
||||
`ci-personas/`, nothing else), frozen,
|
||||
and exercised only by its exact required pairs. Task refs are resolved once
|
||||
before generation zero and the immutable task bindings are forwarded to every
|
||||
generated runner invocation, so a moving branch cannot change later evidence.
|
||||
The deterministic gate then decides. Promotion application rejects older
|
||||
pre-oracle evidence schemas. `promote` stops the loop; with `--apply`
|
||||
The deterministic quality-first gate rejects blocker-recall regressions,
|
||||
new false positives on clean controls, and any weighted-score regression.
|
||||
Repeated evidence (`RUNS>=3`) is required for promotion; `RUNS=1` is
|
||||
diagnostic-only. CE is the external comparator. Cost and latency are
|
||||
tiebreakers and never compensate for quality loss. `promote` stops the loop; with `--apply`
|
||||
the authorized frozen bytes
|
||||
are transactionally applied to the canonical
|
||||
`.claude/skills/` trees and their shipped mirrors as an ordinary
|
||||
|
|
@ -250,17 +359,19 @@ generation's trajectories to the next proposer. `--initial-overlay` skips
|
|||
the generation-0 proposer to benchmark a hand-written candidate;
|
||||
`--proposer-model` upgrades only the diagnosis session.
|
||||
|
||||
**Learning queue.** Live plan/work skill runs never self-edit (see each
|
||||
**Learning queue.** Live skill runs never self-edit (see each
|
||||
skill's "Skill feedback" section) — instead they may append one-line JSON notes to
|
||||
`workflow_bench/learnings.jsonl` (gitignored, machine-local like the
|
||||
transcripts they complement). The proposer reads the queue as hints, not
|
||||
ground truth: a learning only reaches a shipped skill by surviving the same
|
||||
paired benchmark as any other candidate. Legacy review/LFG rows are ignored;
|
||||
those skills do not yet have honest candidate lanes or promotion gates.
|
||||
paired benchmark as any other candidate.
|
||||
|
||||
Run the driver on the existing re-evaluation triggers (model/harness change,
|
||||
90-day staleness), not on a tight schedule — every generation costs ≥3 paired
|
||||
runs per task, and `--generations` is the only loop bound.
|
||||
For ad-hoc use, run the driver on the existing re-evaluation triggers
|
||||
(model/harness change or 90-day staleness). The repository workflow runs a
|
||||
deliberate weekly drift check: scheduled concurrency stays serial unless
|
||||
`GITNEXUS_EVOLUTION_WORKERS` is raised after a funded host-sized proof, and
|
||||
`--workers` is bounded to 1–8 before paid work starts. `--generations` remains
|
||||
the only loop bound.
|
||||
|
||||
## Free-model setup (no paid tokens)
|
||||
|
||||
|
|
@ -279,12 +390,30 @@ uv run --locked --with 'litellm[proxy]' litellm --config workflow_bench/free-mod
|
|||
# 2. Point the benchmark at it
|
||||
uv run --locked --extra dev python -m workflow_bench.runner \
|
||||
--tasks workflow_bench/tasks.scenarios.yaml --runs 3 \
|
||||
--base-url http://localhost:4000 --auth-token "$LITELLM_MASTER_KEY" --model free-coder
|
||||
--base-url http://localhost:4000 --anthropic-api-key "$LITELLM_MASTER_KEY" --model free-coder
|
||||
```
|
||||
|
||||
## OpenAI API keys
|
||||
|
||||
Claude Code still speaks Anthropic `/v1/messages`. For a paid OpenAI backend,
|
||||
do not point `--anthropic-api-key` at an `sk-...` OpenAI key. Export the OpenAI key
|
||||
and use OpenAI model ids; the driver starts the proxy itself:
|
||||
|
||||
```bash
|
||||
export GITNEXUS_BENCH_OPENAI_API_KEY="$OPENAI_API_KEY"
|
||||
PROVIDER=openai ./workflow_bench/run-evolution.sh
|
||||
```
|
||||
|
||||
The GitHub skill-evolution workflow accepts `GITNEXUS_BENCH_OPENAI_API_KEY` on
|
||||
the `gitnexus-evolution` environment. Dispatch with `provider=openai` to force
|
||||
that backend even when an Anthropic token is also configured (otherwise `auto`
|
||||
keeps using Anthropic whenever that secret exists). Claude default model
|
||||
inputs are then rewritten to `gpt-5.6-sol`; every proposer and benchmark
|
||||
session receives `--effort xhigh`.
|
||||
|
||||
Caveats, honestly:
|
||||
|
||||
- Both arms run on the same model, so the *comparison* stays fair at any
|
||||
- Both arms run on the same model, so the _comparison_ stays fair at any
|
||||
quality level — but small free models follow skills less reliably, so
|
||||
expect lower resolve rates and noisier savings than on frontier models.
|
||||
Treat free-model runs as directional; confirm headline numbers with a
|
||||
|
|
@ -311,16 +440,16 @@ Three task classes × three arms, single-repo (GitNexus itself). **Every arm
|
|||
resolved every task** — at this difficulty, pass/fail quality is saturated
|
||||
and the comparison is pure cost:
|
||||
|
||||
| task (class) | arm | resolved | cost $ | wall | turns | vs baseline cost |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| trivial-version-alias | workflow | 1/1 | 9.16 | 16m | 63 | −333% |
|
||||
| trivial-version-alias | baseline | 1/1 | 2.11 | 2.8m | 16 | — |
|
||||
| inv-bug-pdg-note | workflow | 1/1 | 14.56 | 21m | 83 | −331% |
|
||||
| inv-bug-pdg-note | workflow_direct | 1/1 | 5.23 | 7.5m | 32 | −55% |
|
||||
| inv-bug-pdg-note | baseline | 1/1 | 3.38 | 4.7m | 22 | — |
|
||||
| inv-feature-list-repos-filter | workflow | 1/1 | 13.22 | 19m | 84 | −211% |
|
||||
| inv-feature-list-repos-filter | workflow_direct | 1/1 | 4.87 | 4.8m | 38 | −15% (wall +14% faster) |
|
||||
| inv-feature-list-repos-filter | baseline | 1/1 | 4.25 | 5.5m | 32 | — |
|
||||
| task (class) | arm | resolved | cost $ | wall | turns | vs baseline cost |
|
||||
| ----------------------------- | --------------- | -------- | ------ | ---- | ----- | ----------------------- |
|
||||
| trivial-version-alias | workflow | 1/1 | 9.16 | 16m | 63 | −333% |
|
||||
| trivial-version-alias | baseline | 1/1 | 2.11 | 2.8m | 16 | — |
|
||||
| inv-bug-pdg-note | workflow | 1/1 | 14.56 | 21m | 83 | −331% |
|
||||
| inv-bug-pdg-note | workflow_direct | 1/1 | 5.23 | 7.5m | 32 | −55% |
|
||||
| inv-bug-pdg-note | baseline | 1/1 | 3.38 | 4.7m | 22 | — |
|
||||
| inv-feature-list-repos-filter | workflow | 1/1 | 13.22 | 19m | 84 | −211% |
|
||||
| inv-feature-list-repos-filter | workflow_direct | 1/1 | 4.87 | 4.8m | 38 | −15% (wall +14% faster) |
|
||||
| inv-feature-list-repos-filter | baseline | 1/1 | 4.25 | 5.5m | 32 | — |
|
||||
|
||||
What the ground base says, honestly:
|
||||
|
||||
|
|
@ -335,7 +464,7 @@ What the ground base says, honestly:
|
|||
detect_changes-before-commit) is cheap. It produced noticeably more test
|
||||
coverage than baseline for near-equal cost on the feature task.
|
||||
- **Quality didn't differentiate because nothing failed.** The regime where
|
||||
the workflow should win on *resolve rate* — cross-module tasks where
|
||||
the workflow should win on _resolve rate_ — cross-module tasks where
|
||||
baselines flail — is the unmeasured cell (`cross-module-parse-retry`), and
|
||||
the next thing to measure, ideally with `--runs 3+` on a free backend.
|
||||
- Caveats: n=1 per cell, one repo, one model; churn numbers from this run
|
||||
|
|
@ -353,11 +482,11 @@ If a future run shows the workflow flattering itself here, distrust the run.
|
|||
The hardest class — retry-with-backoff across the worker-pool/pipeline
|
||||
seams, transient-vs-deterministic classification:
|
||||
|
||||
| arm | resolved | cost $ | wall | turns | churn |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| workflow | 1/1 | 18.32 | 37m | 107 | 4/+373/−17 |
|
||||
| **workflow_direct** | 1/1 | **9.53** | **15m** | **52** | 11/+244/−66 |
|
||||
| baseline | 1/1 | 18.03 | 34m | 98 | 6/+345/−69 |
|
||||
| arm | resolved | cost $ | wall | turns | churn |
|
||||
| ------------------- | -------- | -------- | ------- | ------ | ----------- |
|
||||
| workflow | 1/1 | 18.32 | 37m | 107 | 4/+373/−17 |
|
||||
| **workflow_direct** | 1/1 | **9.53** | **15m** | **52** | 11/+244/−66 |
|
||||
| baseline | 1/1 | 18.03 | 34m | 98 | 6/+345/−69 |
|
||||
|
||||
(The workflow_direct row is the clean re-run under clone isolation — the
|
||||
original was contaminated, see the integrity note below.)
|
||||
|
|
@ -391,14 +520,14 @@ category-priced freshness (`accept` for compact classes), per-category turn
|
|||
budgets, and the work-phase HEAD==pin fast path, the same
|
||||
`inv-bug-pdg-note` workflow cell re-measured (n=1):
|
||||
|
||||
| | ground base | optimized | delta |
|
||||
| --- | --- | --- | --- |
|
||||
| resolved | ✅ | ✅ | — |
|
||||
| cost $ | 14.56 | 11.70 | **−20%** |
|
||||
| turns | 83 | 72 | −13% |
|
||||
| output tokens | 59,789 | 53,345 | −11% |
|
||||
| cache_read | 6.64M | 5.07M | −24% |
|
||||
| wall | 21m | 25m | +15% |
|
||||
| | ground base | optimized | delta |
|
||||
| ------------- | ----------- | --------- | -------- |
|
||||
| resolved | ✅ | ✅ | — |
|
||||
| cost $ | 14.56 | 11.70 | **−20%** |
|
||||
| turns | 83 | 72 | −13% |
|
||||
| output tokens | 59,789 | 53,345 | −11% |
|
||||
| cache_read | 6.64M | 5.07M | −24% |
|
||||
| wall | 21m | 25m | +15% |
|
||||
|
||||
Verified in-transcript: the compact form fired (115-line plan vs 209 for a
|
||||
simpler task pre-optimization), the plan session dropped 72→49 turns, and
|
||||
|
|
@ -411,8 +540,8 @@ this task class, so the routing rule above stands unchanged.
|
|||
## Writing good tasks
|
||||
|
||||
See `tasks.scenarios.yaml`. Small enough to finish headless, real enough to
|
||||
require investigation — the workflow's savings come from *not re-reading and
|
||||
not re-investigating*, which trivial tasks never exercise. Keep `verify` as a
|
||||
require investigation — the workflow's savings come from _not re-reading and
|
||||
not re-investigating_, which trivial tasks never exercise. Keep `verify` as a
|
||||
model-visible authored-test quality signal, and add an independent `oracle`
|
||||
whose source files live under `workflow_bench/oracles/`. Oracle commands must
|
||||
run only files staged beneath `$GITNEXUS_BENCH_ORACLE_ROOT`; for Vitest, include
|
||||
|
|
@ -422,7 +551,7 @@ carry build pre-hooks).
|
|||
|
||||
## Relation to the SWE-bench harness
|
||||
|
||||
The rest of `eval/` benchmarks GitNexus *tools* inside a litellm agent loop
|
||||
(baseline vs graph-enhanced). This module benchmarks the *skill workflow*
|
||||
The rest of `eval/` benchmarks GitNexus _tools_ inside a litellm agent loop
|
||||
(baseline vs graph-enhanced). This module benchmarks the _skill workflow_
|
||||
inside the real CLI harness those skills ship for. Different question, same
|
||||
spirit: measure, don't assume.
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@ import os
|
|||
import secrets
|
||||
import stat
|
||||
import statistics
|
||||
from collections.abc import Sequence
|
||||
from pathlib import Path, PurePosixPath
|
||||
from typing import Any
|
||||
|
||||
|
|
@ -21,9 +22,94 @@ from .proposer_sandbox import (
|
|||
CANDIDATE_ARMS = {
|
||||
"candidate_workflow": "workflow",
|
||||
"candidate_workflow_direct": "workflow_direct",
|
||||
"candidate_review": "review",
|
||||
}
|
||||
PROMOTION_SCHEMA_VERSION = 6
|
||||
|
||||
|
||||
def promotion_policy(
|
||||
candidate_arms: Sequence[str],
|
||||
*,
|
||||
metric: str = "cost_usd",
|
||||
min_runs: int = 3,
|
||||
min_improvement_pct: float = 5.0,
|
||||
max_task_regression_pct: float = 20.0,
|
||||
) -> dict[str, dict[str, Any]]:
|
||||
"""The exact per-arm policy shared by evidence production and application."""
|
||||
if not candidate_arms or len(set(candidate_arms)) != len(candidate_arms):
|
||||
raise ValueError("promotion policy requires unique candidate arms")
|
||||
policies = {}
|
||||
for arm in candidate_arms:
|
||||
if arm not in CANDIDATE_ARMS:
|
||||
raise ValueError(f"unsupported candidate arm: {arm}")
|
||||
policies[arm] = (
|
||||
{
|
||||
"metric": "review_weighted_f1",
|
||||
"min_runs": min_runs,
|
||||
"min_improvement": 0.01,
|
||||
"quality_rule": "correct verdict on every repeat; minimum blocker recall; no clean-control regression",
|
||||
}
|
||||
if arm == "candidate_review"
|
||||
else {
|
||||
"metric": metric,
|
||||
"min_runs": min_runs,
|
||||
"min_improvement_pct": min_improvement_pct,
|
||||
"max_task_regression_pct": max_task_regression_pct,
|
||||
"max_failed_task_regression_pct": MAX_FAILED_TASK_REGRESSION_PCT,
|
||||
"min_gated_task_ratio": MIN_GATED_TASK_RATIO,
|
||||
"quality_rule": "no per-task resolution-rate regression",
|
||||
}
|
||||
)
|
||||
return policies
|
||||
|
||||
|
||||
def promotion_evidence(
|
||||
results: dict[str, dict[str, dict[str, Any]]],
|
||||
*,
|
||||
policy: dict[str, dict[str, Any]],
|
||||
model: str | None,
|
||||
complete: bool,
|
||||
) -> dict[str, Any]:
|
||||
"""Produce discriminated, recomputable decisions, including partial reports."""
|
||||
decisions = []
|
||||
for candidate, rules in policy.items():
|
||||
common = {
|
||||
"incumbent_arm": CANDIDATE_ARMS[candidate],
|
||||
"candidate_arm": candidate,
|
||||
"model": model,
|
||||
"min_runs": rules["min_runs"],
|
||||
}
|
||||
if candidate == "candidate_review":
|
||||
decision = evaluate_review_candidate(results, **common, min_improvement=rules["min_improvement"])
|
||||
else:
|
||||
decision = evaluate_candidate(
|
||||
results,
|
||||
**common,
|
||||
**{
|
||||
key: rules[key]
|
||||
for key in (
|
||||
"metric",
|
||||
"min_improvement_pct",
|
||||
"max_task_regression_pct",
|
||||
"max_failed_task_regression_pct",
|
||||
)
|
||||
},
|
||||
)
|
||||
if not complete:
|
||||
decision["decision"] = "insufficient_evidence"
|
||||
decision["reasons"].append("sweep aborted; partial evidence cannot promote")
|
||||
decisions.append(decision)
|
||||
return {
|
||||
"schema_version": PROMOTION_SCHEMA_VERSION,
|
||||
"run_status": "complete" if complete else "aborted",
|
||||
"policy": policy,
|
||||
"decisions": decisions,
|
||||
}
|
||||
|
||||
|
||||
CANDIDATE_SKILLS = {
|
||||
"gitnexus-plan",
|
||||
"gitnexus-review",
|
||||
"gitnexus-work",
|
||||
}
|
||||
# Skills each incumbent arm actually loads in its sessions. An overlay that
|
||||
|
|
@ -32,6 +118,7 @@ CANDIDATE_SKILLS = {
|
|||
ARM_SKILLS = {
|
||||
"workflow": ("gitnexus-plan", "gitnexus-work"),
|
||||
"workflow_direct": ("gitnexus-work",),
|
||||
"review": ("gitnexus-review",),
|
||||
}
|
||||
# Repo-local prompts whose bytes are evidence for each executed arm. Keep this
|
||||
# distinct from ``ARM_SKILLS``: that mapping defines which skills a promotable
|
||||
|
|
@ -54,6 +141,19 @@ MAIN_LOOP_ONLY_WARNING = (
|
|||
"each run output, deduplicating events "
|
||||
"that share one message.id."
|
||||
)
|
||||
# Failure kinds the prompts under test cause, not the task: the skill never
|
||||
# ran at all. Both arms failing a task this way is evidence about the skills,
|
||||
# so such a task stays inside the gate however unresolvable it looks.
|
||||
SKILL_ATTRIBUTABLE_ERROR_KINDS = frozenset({"skill-not-invoked"})
|
||||
# Leaving the quality gate is not leaving the spend gate. A candidate may fail
|
||||
# the same oracle the incumbent fails, but not at a multiple of its cost — an
|
||||
# ungated task is still real money and still ranks on the metric.
|
||||
MAX_FAILED_TASK_REGRESSION_PCT = 100.0
|
||||
# Promotion must rest on a real evidence base. Half the paired tasks is the
|
||||
# loosest rule the three-task production set can carry: it tolerates the one
|
||||
# scenario neither arm resolves and refuses a generation that has quietly
|
||||
# decayed to a single gated task deciding everything.
|
||||
MIN_GATED_TASK_RATIO = 0.5
|
||||
EVIDENCE_MAX_AGE_DAYS = 90
|
||||
MAX_CANDIDATE_OVERLAY_BYTES = 4 * 1024 * 1024
|
||||
MAX_SKILL_FINGERPRINT_BYTES = 4 * 1024 * 1024
|
||||
|
|
@ -325,7 +425,7 @@ def candidate_overlay_files(overlay: Path) -> list[Path]:
|
|||
):
|
||||
raise ValueError(
|
||||
"candidate overlays may only contain Markdown files under "
|
||||
".claude/skills/gitnexus-{plan,work}: "
|
||||
".claude/skills/gitnexus-{plan,review,work}: "
|
||||
f"{relative}"
|
||||
)
|
||||
return entries
|
||||
|
|
@ -345,6 +445,8 @@ def required_candidate_arms(overlay: Path) -> list[str]:
|
|||
required.append("candidate_workflow")
|
||||
if "gitnexus-work" in touched:
|
||||
required.append("candidate_workflow_direct")
|
||||
if "gitnexus-review" in touched:
|
||||
required.append("candidate_review")
|
||||
return required
|
||||
|
||||
|
||||
|
|
@ -365,6 +467,69 @@ def candidate_overlay_digest(overlay: Path) -> str:
|
|||
return digest
|
||||
|
||||
|
||||
def _commit_sandbox_paths(
|
||||
sandbox: SandboxSession,
|
||||
relative_paths: Sequence[str],
|
||||
*,
|
||||
message: str,
|
||||
require_change: bool,
|
||||
) -> bool:
|
||||
"""Stage and commit paths inside the outer sandbox.
|
||||
|
||||
Returns True when a commit was created. ``require_change`` keeps the
|
||||
candidate-overlay contract: a no-op overlay is an error, while an
|
||||
incumbent skill seed may already match the historical tree.
|
||||
"""
|
||||
|
||||
mkdir_command = ["/bin/mkdir", "-p", f"{SANDBOX_TMP}/wfbench-empty-hooks"]
|
||||
mkdir_result = sandbox.run(
|
||||
mkdir_command,
|
||||
timeout=60,
|
||||
env=build_sandbox_environment(),
|
||||
)
|
||||
if not mkdir_result.ok:
|
||||
raise ManagedProcessError(mkdir_command, mkdir_result)
|
||||
if not relative_paths:
|
||||
if require_change:
|
||||
raise ValueError("candidate overlay is byte-identical to the incumbent skills")
|
||||
return False
|
||||
|
||||
# Historical review SHAs gitignore `.claude/skills/*` and lack the current
|
||||
# per-skill allowlist. Force-add so a seed or overlay of harness-owned
|
||||
# skill bytes is not rejected as an ignored path.
|
||||
command, added = _sandbox_overlay_git(sandbox, ["add", "-f", "--", *relative_paths])
|
||||
if not added.ok:
|
||||
raise ManagedProcessError(command, added)
|
||||
command, changed = _sandbox_overlay_git(
|
||||
sandbox,
|
||||
["diff", "--cached", "--quiet", "--no-ext-diff", "--no-textconv", "--"],
|
||||
)
|
||||
if changed.returncode == 0:
|
||||
if require_change:
|
||||
raise ValueError("candidate overlay is byte-identical to the incumbent skills")
|
||||
return False
|
||||
if changed.returncode != 1:
|
||||
raise ManagedProcessError(command, changed)
|
||||
|
||||
command, committed = _sandbox_overlay_git(
|
||||
sandbox,
|
||||
[
|
||||
"commit",
|
||||
"--quiet",
|
||||
"--no-verify",
|
||||
"-m",
|
||||
message,
|
||||
],
|
||||
extra_config=(
|
||||
"user.name=workflow-bench",
|
||||
"user.email=workflow-bench@invalid",
|
||||
),
|
||||
)
|
||||
if not committed.ok:
|
||||
raise ManagedProcessError(command, committed)
|
||||
return True
|
||||
|
||||
|
||||
def apply_candidate_overlay(
|
||||
overlay: Path,
|
||||
worktree: Path,
|
||||
|
|
@ -383,47 +548,96 @@ def apply_candidate_overlay(
|
|||
for relative, content in payload:
|
||||
_replace_regular_file(worktree, relative, content)
|
||||
relative_paths.append(relative.as_posix())
|
||||
|
||||
mkdir_command = ["/bin/mkdir", "-p", f"{SANDBOX_TMP}/wfbench-empty-hooks"]
|
||||
mkdir_result = sandbox.run(
|
||||
mkdir_command,
|
||||
timeout=60,
|
||||
env=build_sandbox_environment(),
|
||||
)
|
||||
if not mkdir_result.ok:
|
||||
raise ManagedProcessError(mkdir_command, mkdir_result)
|
||||
|
||||
command, added = _sandbox_overlay_git(sandbox, ["add", "--", *relative_paths])
|
||||
if not added.ok:
|
||||
raise ManagedProcessError(command, added)
|
||||
command, changed = _sandbox_overlay_git(
|
||||
_commit_sandbox_paths(
|
||||
sandbox,
|
||||
["diff", "--cached", "--quiet", "--no-ext-diff", "--no-textconv", "--"],
|
||||
relative_paths,
|
||||
message="benchmark candidate skill overlay",
|
||||
require_change=True,
|
||||
)
|
||||
if changed.returncode == 0:
|
||||
raise ValueError("candidate overlay is byte-identical to the incumbent skills")
|
||||
if changed.returncode != 1:
|
||||
raise ManagedProcessError(command, changed)
|
||||
|
||||
command, committed = _sandbox_overlay_git(
|
||||
sandbox,
|
||||
[
|
||||
"commit",
|
||||
"--quiet",
|
||||
"--no-verify",
|
||||
"-m",
|
||||
"benchmark candidate skill overlay",
|
||||
],
|
||||
extra_config=(
|
||||
"user.name=workflow-bench",
|
||||
"user.email=workflow-bench@invalid",
|
||||
),
|
||||
)
|
||||
if not committed.ok:
|
||||
raise ManagedProcessError(command, committed)
|
||||
return digest
|
||||
|
||||
|
||||
def seed_evaluated_skills(
|
||||
source_repo: Path,
|
||||
worktree: Path,
|
||||
*,
|
||||
sandbox: SandboxSession,
|
||||
arm: str,
|
||||
) -> None:
|
||||
"""Install the current evaluated skill tree into a historical clone.
|
||||
|
||||
Review evolution scores the current (or overlay) ``gitnexus-review`` skill
|
||||
against a historical PR checkout. Older SHAs predate that skill, and
|
||||
using whatever prose happened to exist at the reviewed commit would make
|
||||
the incumbent arm a moving target. Copy the harness checkout's skill
|
||||
bytes and commit them before setup so ``git status`` still shows only
|
||||
the task patch.
|
||||
"""
|
||||
|
||||
skill_names = EVALUATED_ARM_SKILLS.get(arm)
|
||||
if not skill_names:
|
||||
return
|
||||
|
||||
source_repo = source_repo.expanduser().absolute()
|
||||
expected_clone = Path(os.path.abspath(worktree.expanduser()))
|
||||
sandbox_clone = Path(os.path.abspath(sandbox.clone.expanduser()))
|
||||
if sandbox_clone != expected_clone:
|
||||
raise ValueError("skill seed sandbox does not bind the requested clone")
|
||||
_require_real_directory(source_repo, label="incumbent skill repository")
|
||||
if source_repo.resolve(strict=True) != source_repo:
|
||||
raise ValueError(f"incumbent skill repository cannot traverse symlinks: {source_repo}")
|
||||
|
||||
relative_paths: list[str] = []
|
||||
total = 0
|
||||
for skill_name in skill_names:
|
||||
_require_directory_chain(
|
||||
source_repo,
|
||||
Path(".claude") / "skills" / skill_name,
|
||||
label="incumbent skill root",
|
||||
)
|
||||
skill_root = source_repo / ".claude" / "skills" / skill_name
|
||||
pending = [skill_root]
|
||||
while pending:
|
||||
directory = pending.pop()
|
||||
try:
|
||||
children = list(os.scandir(directory))
|
||||
except OSError as exc:
|
||||
raise ValueError(f"incumbent skill directory is unreadable: {directory}: {exc}") from exc
|
||||
for item in children:
|
||||
path = Path(item.path)
|
||||
if item.is_symlink():
|
||||
raise ValueError(
|
||||
"incumbent skill seed cannot contain symlinks: "
|
||||
f"{path.relative_to(source_repo)}"
|
||||
)
|
||||
if item.is_dir(follow_symlinks=False):
|
||||
pending.append(path)
|
||||
continue
|
||||
if not item.is_file(follow_symlinks=False):
|
||||
raise ValueError(
|
||||
"incumbent skill seed entries must be regular files: "
|
||||
f"{path.relative_to(source_repo)}"
|
||||
)
|
||||
total += item.stat(follow_symlinks=False).st_size
|
||||
if total > MAX_SKILL_FINGERPRINT_BYTES:
|
||||
raise ValueError("incumbent skill seed exceeds the bounded evidence limit")
|
||||
relative = Path(".claude") / "skills" / skill_name / path.relative_to(skill_root)
|
||||
content = _bounded_regular_bytes(
|
||||
path,
|
||||
limit=MAX_SKILL_FINGERPRINT_BYTES,
|
||||
label="incumbent skill file",
|
||||
)
|
||||
_replace_regular_file(worktree, relative, content)
|
||||
relative_paths.append(PurePosixPath(relative.as_posix()).as_posix())
|
||||
|
||||
_commit_sandbox_paths(
|
||||
sandbox,
|
||||
relative_paths,
|
||||
message="benchmark incumbent review skill",
|
||||
require_change=False,
|
||||
)
|
||||
|
||||
|
||||
def unexercised_overlay_skills(overlay: Path, candidate_arms: list[str]) -> list[str]:
|
||||
"""Overlay skills that no selected candidate arm would ever load.
|
||||
|
||||
|
|
@ -475,6 +689,133 @@ def skill_fingerprint(worktree: Path, arm: str) -> str | None:
|
|||
return fingerprint_files(worktree, files)
|
||||
|
||||
|
||||
def evaluate_review_candidate(
|
||||
results: dict[str, dict[str, dict[str, Any]]],
|
||||
*,
|
||||
incumbent_arm: str,
|
||||
candidate_arm: str,
|
||||
model: str | None,
|
||||
min_runs: int = 3,
|
||||
min_improvement: float = 0.01,
|
||||
) -> dict[str, Any]:
|
||||
"""Quality-first promotion gate for paired read-only review arms."""
|
||||
|
||||
reasons: list[str] = []
|
||||
task_rows: list[dict[str, Any]] = []
|
||||
insufficient = not model
|
||||
regression = False
|
||||
improvement = False
|
||||
if not model:
|
||||
reasons.append("a named --model is required so review evidence cannot drift")
|
||||
|
||||
for task_id, arms in sorted(results.items()):
|
||||
if incumbent_arm not in arms or candidate_arm not in arms:
|
||||
insufficient = True
|
||||
reasons.append(f"{task_id}: both {incumbent_arm} and {candidate_arm} are required")
|
||||
continue
|
||||
incumbent = arms[incumbent_arm]
|
||||
candidate = arms[candidate_arm]
|
||||
incumbent_runs = int(incumbent.get("valid_runs", 0))
|
||||
candidate_runs = int(candidate.get("valid_runs", 0))
|
||||
incumbent_score = incumbent.get("review_weighted_f1")
|
||||
candidate_score = candidate.get("review_weighted_f1")
|
||||
incumbent_blockers = incumbent.get("review_blocker_recall")
|
||||
candidate_blockers = candidate.get("review_blocker_recall")
|
||||
incumbent_fp = incumbent.get("review_false_positives")
|
||||
candidate_fp = candidate.get("review_false_positives")
|
||||
clean = bool(incumbent.get("review_clean_control", candidate.get("review_clean_control", False)))
|
||||
incumbent_clean_pass = incumbent.get("review_clean_pass")
|
||||
candidate_clean_pass = candidate.get("review_clean_pass")
|
||||
task_rows.append(
|
||||
{
|
||||
"task": task_id,
|
||||
"class": incumbent.get("class", ""),
|
||||
"incumbent_weighted_f1": incumbent_score,
|
||||
"incumbent": dict(incumbent),
|
||||
"candidate": dict(candidate),
|
||||
"gated": True,
|
||||
"candidate_weighted_f1": candidate_score,
|
||||
"incumbent_blocker_recall": incumbent_blockers,
|
||||
"candidate_blocker_recall": candidate_blockers,
|
||||
"incumbent_false_positives": incumbent_fp,
|
||||
"candidate_false_positives": candidate_fp,
|
||||
"clean_control": clean,
|
||||
"incumbent_clean_pass": incumbent_clean_pass,
|
||||
"candidate_clean_pass": candidate_clean_pass,
|
||||
}
|
||||
)
|
||||
if (
|
||||
incumbent_runs < min_runs
|
||||
or candidate_runs < min_runs
|
||||
or incumbent_runs != candidate_runs
|
||||
or incumbent.get("excluded_runs")
|
||||
or candidate.get("excluded_runs")
|
||||
):
|
||||
insufficient = True
|
||||
reasons.append(
|
||||
f"{task_id}: needs {min_runs} valid paired runs with zero exclusions "
|
||||
f"(got {incumbent_runs}/{candidate_runs})"
|
||||
)
|
||||
required_values = (
|
||||
(incumbent_fp, candidate_fp, incumbent_clean_pass, candidate_clean_pass)
|
||||
if clean
|
||||
else (incumbent_score, candidate_score, incumbent_fp, candidate_fp)
|
||||
)
|
||||
if any(value is None for value in required_values):
|
||||
insufficient = True
|
||||
reasons.append(f"{task_id}: structured review quality metrics are incomplete")
|
||||
continue
|
||||
if candidate.get("review_verdict_correct") is None:
|
||||
insufficient = True
|
||||
reasons.append(f"{task_id}: candidate verdict evidence is incomplete")
|
||||
elif candidate["review_verdict_correct"] is not True:
|
||||
regression = True
|
||||
reasons.append(f"{task_id}: candidate verdict was incorrect on a valid repeat")
|
||||
if (incumbent_blockers is None) != (candidate_blockers is None):
|
||||
insufficient = True
|
||||
reasons.append(f"{task_id}: blocker recall evidence is incomplete")
|
||||
if incumbent_blockers is not None and candidate_blockers is not None and float(candidate_blockers) < float(
|
||||
incumbent_blockers
|
||||
):
|
||||
regression = True
|
||||
reasons.append(f"{task_id}: blocker recall regressed")
|
||||
if clean and float(candidate_fp) > float(incumbent_fp):
|
||||
regression = True
|
||||
reasons.append(f"{task_id}: false positives increased on a clean control")
|
||||
if clean and bool(incumbent_clean_pass) and not bool(candidate_clean_pass):
|
||||
regression = True
|
||||
reasons.append(f"{task_id}: clean-control verdict regressed")
|
||||
if not clean and float(candidate_score) + 1e-9 < float(incumbent_score):
|
||||
regression = True
|
||||
reasons.append(f"{task_id}: weighted review score regressed")
|
||||
if not clean and float(candidate_score) >= float(incumbent_score) + min_improvement:
|
||||
improvement = True
|
||||
|
||||
if not task_rows:
|
||||
insufficient = True
|
||||
reasons.append("no paired review task results were found")
|
||||
if insufficient:
|
||||
decision = "insufficient_evidence"
|
||||
elif regression:
|
||||
decision = "keep_incumbent"
|
||||
elif not improvement:
|
||||
decision = "keep_incumbent"
|
||||
reasons.append("candidate did not improve weighted review quality on any corpus case")
|
||||
else:
|
||||
decision = "promote"
|
||||
reasons.append("candidate improved weighted review quality without blocker or clean-control regression")
|
||||
return {
|
||||
"candidate_arm": candidate_arm,
|
||||
"incumbent_arm": incumbent_arm,
|
||||
"decision": decision,
|
||||
"metric": "review_weighted_f1",
|
||||
"model": model,
|
||||
"tasks": task_rows,
|
||||
"ungated_tasks": [],
|
||||
"reasons": reasons,
|
||||
}
|
||||
|
||||
|
||||
def evaluate_candidate(
|
||||
results: dict[str, dict[str, dict[str, Any]]],
|
||||
*,
|
||||
|
|
@ -485,12 +826,18 @@ def evaluate_candidate(
|
|||
min_runs: int = 3,
|
||||
min_improvement_pct: float = 5.0,
|
||||
max_task_regression_pct: float = 20.0,
|
||||
max_failed_task_regression_pct: float = MAX_FAILED_TASK_REGRESSION_PCT,
|
||||
) -> dict[str, Any]:
|
||||
"""Deterministically decide whether a prompt candidate is promotable.
|
||||
|
||||
Resolution is lexicographically primary: a cheaper candidate that fails
|
||||
more tasks never wins. With equal quality, the candidate must clear the
|
||||
configured median efficiency gain without a large per-task regression.
|
||||
|
||||
A task neither arm can resolve leaves the quality gate, but only on
|
||||
evidence: a comparable metric, no skill-not-invoked run, and enough tasks
|
||||
left inside the gate to decide anything. It still ranks against the
|
||||
failed-task spend cap.
|
||||
"""
|
||||
if metric not in PROMOTION_METRICS:
|
||||
raise ValueError(f"unsupported promotion metric: {metric}")
|
||||
|
|
@ -536,20 +883,66 @@ def evaluate_candidate(
|
|||
if (not metric_unavailable and incumbent_metric)
|
||||
else None
|
||||
)
|
||||
# A task that both arms measured cleanly and neither ever resolved sits
|
||||
# outside both arms' current capability. It carries no quality signal
|
||||
# about the candidate, and its metric compares who spent more while
|
||||
# failing the same oracle — so gating on it measures the task, not the
|
||||
# candidate, and one such task vetoes every future promotion for as
|
||||
# long as it stays in the set. Keep it in the evidence, out of the gate,
|
||||
# and name it as task health instead.
|
||||
#
|
||||
# Ungating is itself a claim, so it needs evidence: the failures must
|
||||
# be the task's (not a skill that never loaded) and the metric must be
|
||||
# comparable, otherwise the task stays gated and the checks below name
|
||||
# what is missing.
|
||||
fully_measured = (
|
||||
incumbent_runs >= min_runs
|
||||
and candidate_runs >= min_runs
|
||||
and incumbent_runs == candidate_runs
|
||||
and not incumbent_excluded
|
||||
and not candidate_excluded
|
||||
)
|
||||
skill_attributable = bool(
|
||||
(set(incumbent.get("error_kinds", {})) | set(candidate.get("error_kinds", {})))
|
||||
& SKILL_ATTRIBUTABLE_ERROR_KINDS
|
||||
)
|
||||
mutually_unresolved = fully_measured and not incumbent["resolved"] and not candidate["resolved"]
|
||||
gated = not (mutually_unresolved and not skill_attributable and improvement is not None)
|
||||
# The floor asks the candidate to be reliable where the incumbent is.
|
||||
# On a task the incumbent never resolves there is no reliability to
|
||||
# match, and holding partial candidate progress to it punished a
|
||||
# candidate for resolving 1 of 3 runs while excusing it for resolving
|
||||
# none — the strictly worse result. A skill that never loaded is the
|
||||
# exception: those failures belong to the prompts, so the floor applies
|
||||
# even with nothing on the incumbent's side to match.
|
||||
quality_floor_enforced = bool(incumbent["resolved"]) or skill_attributable
|
||||
task_rows.append(
|
||||
{
|
||||
"task": task_id,
|
||||
"class": incumbent.get("class", ""),
|
||||
"incumbent_resolved": f"{incumbent['resolved']}/{incumbent_runs}",
|
||||
"incumbent": dict(incumbent),
|
||||
"candidate": dict(candidate),
|
||||
"candidate_resolved": f"{candidate['resolved']}/{candidate_runs}",
|
||||
"incumbent_excluded_runs": incumbent_excluded,
|
||||
"candidate_excluded_runs": candidate_excluded,
|
||||
"candidate_quality_floor_met": candidate_runs > 0 and candidate["resolved"] == candidate_runs,
|
||||
"quality_floor_enforced": quality_floor_enforced,
|
||||
"incumbent_metric": incumbent_metric,
|
||||
"candidate_metric": candidate_metric,
|
||||
"improvement_pct": improvement,
|
||||
"gated": gated,
|
||||
"skill_attributable_failure": skill_attributable,
|
||||
}
|
||||
)
|
||||
if not gated:
|
||||
if improvement < -max_failed_task_regression_pct:
|
||||
efficiency_regression = True
|
||||
reasons.append(
|
||||
f"{task_id}: {metric} regressed {-improvement:.1f}% on a task neither arm resolved, "
|
||||
f"above the {max_failed_task_regression_pct:.1f}% failed-task cap"
|
||||
)
|
||||
continue
|
||||
|
||||
if incumbent_runs < min_runs or candidate_runs < min_runs:
|
||||
insufficient = True
|
||||
|
|
@ -571,11 +964,16 @@ def evaluate_candidate(
|
|||
if candidate_rate < incumbent_rate:
|
||||
quality_regression = True
|
||||
reasons.append(f"{task_id}: resolution regressed from {incumbent_rate:.0%} to {candidate_rate:.0%}")
|
||||
if candidate_runs > 0 and candidate["resolved"] != candidate_runs:
|
||||
if quality_floor_enforced and candidate_runs > 0 and candidate["resolved"] != candidate_runs:
|
||||
quality_floor_failed = True
|
||||
floor_trigger = (
|
||||
"a run never invoked the skill under test"
|
||||
if skill_attributable
|
||||
else f"the incumbent resolves {incumbent['resolved']}/{incumbent_runs}"
|
||||
)
|
||||
reasons.append(
|
||||
f"{task_id}: candidate must resolve every valid run for the oracle-backed quality floor "
|
||||
f"(got {candidate['resolved']}/{candidate_runs})"
|
||||
f"({floor_trigger}; got {candidate['resolved']}/{candidate_runs})"
|
||||
)
|
||||
if metric_unavailable:
|
||||
insufficient = True
|
||||
|
|
@ -592,11 +990,36 @@ def evaluate_candidate(
|
|||
f"{task_id}: {metric} regressed {-improvement:.1f}%, above the {max_task_regression_pct:.1f}% task cap"
|
||||
)
|
||||
|
||||
ungated_tasks = [row["task"] for row in task_rows if not row["gated"]]
|
||||
gated_tasks = [row["task"] for row in task_rows if row["gated"]]
|
||||
if ungated_tasks:
|
||||
# One line, not one per task: `reasons` is truncated to three entries
|
||||
# when it is fed back to the proposer (evolve.summarize_gate), and a
|
||||
# growing set of unsolvable tasks must not crowd out the reason the
|
||||
# candidate actually won or lost. The full list ships structurally.
|
||||
reasons.append(
|
||||
f"not gated on {len(ungated_tasks)} task(s) neither arm resolved: {', '.join(ungated_tasks)} "
|
||||
f"(evidence base: {len(gated_tasks)}/{len(task_rows)} paired tasks gated)"
|
||||
)
|
||||
if not task_rows:
|
||||
insufficient = True
|
||||
reasons.append("no paired task results were found")
|
||||
elif not gated_tasks:
|
||||
# Every paired task was ungated, so nothing in this generation says
|
||||
# anything about candidate quality. Refuse rather than fall through to
|
||||
# an efficiency-only verdict on runs that all failed their oracle.
|
||||
insufficient = True
|
||||
reasons.append("no task supplied quality signal: neither arm resolved a run anywhere in the set")
|
||||
elif len(gated_tasks) < MIN_GATED_TASK_RATIO * len(task_rows):
|
||||
# Ungating one unsolvable task keeps promotion reachable; ungating most
|
||||
# of the set turns "promote" into a verdict from whatever is left.
|
||||
insufficient = True
|
||||
reasons.append(
|
||||
f"promotion evidence base is too thin: {len(gated_tasks)}/{len(task_rows)} paired tasks are gated "
|
||||
f"(at least {MIN_GATED_TASK_RATIO:.0%} required)"
|
||||
)
|
||||
|
||||
improvements = [row["improvement_pct"] for row in task_rows if row["improvement_pct"] is not None]
|
||||
improvements = [row["improvement_pct"] for row in task_rows if row["gated"] and row["improvement_pct"] is not None]
|
||||
median_improvement = round(statistics.median(improvements), 1) if improvements else None
|
||||
incumbent_resolved = sum(
|
||||
arms[incumbent_arm]["resolved"] for arms in results.values() if incumbent_arm in arms and candidate_arm in arms
|
||||
|
|
@ -640,6 +1063,8 @@ def evaluate_candidate(
|
|||
"metric": metric,
|
||||
"metric_warning": (MAIN_LOOP_ONLY_WARNING if metric in MAIN_LOOP_ONLY_METRICS else None),
|
||||
"median_improvement_pct": median_improvement,
|
||||
"ungated_tasks": ungated_tasks,
|
||||
"gated_tasks": gated_tasks,
|
||||
"reasons": reasons,
|
||||
"tasks": task_rows,
|
||||
}
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -9,7 +9,7 @@
|
|||
#
|
||||
# uv run python -m workflow_bench.runner \
|
||||
# --tasks workflow_bench/tasks.scenarios.yaml \
|
||||
# --base-url http://localhost:4000 --auth-token "$LITELLM_MASTER_KEY" \
|
||||
# --base-url http://localhost:4000 --anthropic-api-key "$LITELLM_MASTER_KEY" \
|
||||
# --model free-coder
|
||||
#
|
||||
# Keep the proxy on loopback (litellm's default host). Anyone who can reach
|
||||
|
|
@ -39,5 +39,5 @@ model_list:
|
|||
|
||||
general_settings:
|
||||
# No static default — export LITELLM_MASTER_KEY before starting the proxy
|
||||
# and pass the same value as --auth-token (see header).
|
||||
# and pass the same value as --anthropic-api-key (see header).
|
||||
master_key: os.environ/LITELLM_MASTER_KEY
|
||||
|
|
|
|||
45
eval/workflow_bench/gateway_supervisor.py
Normal file
45
eval/workflow_bench/gateway_supervisor.py
Normal file
|
|
@ -0,0 +1,45 @@
|
|||
"""Private gateway owner. EOF on stdin means the harness no longer exists.
|
||||
|
||||
Executed by absolute script path so the isolated gateway environment needs no
|
||||
PYTHONPATH. The gateway receives DEVNULL, never the owner-liveness descriptor.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import signal
|
||||
import sys
|
||||
import threading
|
||||
|
||||
from process_control import run_managed
|
||||
|
||||
|
||||
def main() -> int:
|
||||
cancelled = threading.Event()
|
||||
|
||||
def watch_owner() -> None:
|
||||
try:
|
||||
# A buffered stdin lock held by a daemon aborts CPython shutdown
|
||||
# when the proxy exits while the owner is still alive.
|
||||
os.read(sys.stdin.fileno(), 1)
|
||||
finally:
|
||||
cancelled.set()
|
||||
|
||||
threading.Thread(target=watch_owner, daemon=True).start()
|
||||
for signum in (signal.SIGTERM, signal.SIGINT):
|
||||
signal.signal(signum, lambda *_: cancelled.set())
|
||||
result = run_managed(
|
||||
sys.argv[1:],
|
||||
timeout=7 * 24 * 60 * 60,
|
||||
cancel_event=cancelled,
|
||||
echo_stdout=True,
|
||||
)
|
||||
if result.stderr_tail:
|
||||
print(result.stderr_tail, file=sys.stderr, flush=True)
|
||||
if result.detail:
|
||||
print(result.detail, file=sys.stderr, flush=True)
|
||||
return 0 if result.ok or result.state == "cancelled" else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
458
eval/workflow_bench/model_gateway.py
Normal file
458
eval/workflow_bench/model_gateway.py
Normal file
|
|
@ -0,0 +1,458 @@
|
|||
"""Route Claude Code sessions at a non-Anthropic backend without leaking it.
|
||||
|
||||
Headless Claude Code speaks the Anthropic Messages API and authenticates with
|
||||
``ANTHROPIC_API_KEY``. OpenAI keys are not a drop-in replacement. The documented
|
||||
escape hatch is already in this package: an Anthropic-compatible loopback proxy
|
||||
(LiteLLM) plus ``ANTHROPIC_BASE_URL``. This module starts that proxy for OpenAI,
|
||||
mints a random master key for Claude, and keeps ``OPENAI_API_KEY`` on the host
|
||||
proxy process — never in the sandboxed agent environment.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
import secrets
|
||||
import shutil
|
||||
import socket
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from collections.abc import Sequence
|
||||
from contextlib import AbstractContextManager
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import yaml
|
||||
|
||||
ANTHROPIC_API_KEY_ENV = "GITNEXUS_BENCH_ANTHROPIC_API_KEY"
|
||||
LEGACY_ANTHROPIC_API_KEY_ENV = "GITNEXUS_BENCH_AUTH_TOKEN"
|
||||
OPENAI_API_KEY_ENV = "GITNEXUS_BENCH_OPENAI_API_KEY"
|
||||
_OPENAI_MODEL = re.compile(
|
||||
r"^(?:openai/)?(?:gpt-|chatgpt-|o[0-9])",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# High reasoning effort on a full context window can leave a request without a
|
||||
# first token for many minutes. Claude Code's default client timeout is far
|
||||
# shorter than that, so both ends of the loopback hop get the same generous
|
||||
# budget and the session fails on real errors instead of on the clock.
|
||||
GATEWAY_REQUEST_TIMEOUT_S = 1800
|
||||
# Importing LiteLLM alone costs ~17s on a cold container filesystem, and the
|
||||
# proxy only binds its port after that. A budget tight enough to lose that race
|
||||
# reads as "connection refused", which looks like a dead proxy rather than a
|
||||
# slow import.
|
||||
GATEWAY_READY_TIMEOUT_ENV = "GITNEXUS_BENCH_GATEWAY_READY_TIMEOUT_S"
|
||||
DEFAULT_GATEWAY_READY_TIMEOUT_S = 180.0
|
||||
|
||||
|
||||
def gateway_ready_timeout_s() -> float:
|
||||
"""Startup budget for the loopback proxy, overridable for slow hosts."""
|
||||
|
||||
raw = (os.environ.get(GATEWAY_READY_TIMEOUT_ENV) or "").strip()
|
||||
if not raw:
|
||||
return DEFAULT_GATEWAY_READY_TIMEOUT_S
|
||||
try:
|
||||
value = float(raw)
|
||||
except ValueError as exc:
|
||||
raise ValueError(f"{GATEWAY_READY_TIMEOUT_ENV} must be a number of seconds, not {raw!r}") from exc
|
||||
if not math.isfinite(value) or value <= 0:
|
||||
raise ValueError(f"{GATEWAY_READY_TIMEOUT_ENV} must be finite and positive, not {raw!r}")
|
||||
return value
|
||||
|
||||
|
||||
def is_openai_model(model: str) -> bool:
|
||||
return bool(_OPENAI_MODEL.match((model or "").strip()))
|
||||
|
||||
|
||||
def openai_backend_model(model: str) -> str:
|
||||
name = model.strip()
|
||||
if name.lower().startswith("openai/"):
|
||||
return f"openai/{name.split('/', 1)[1]}"
|
||||
return f"openai/{name}"
|
||||
|
||||
|
||||
def claude_gateway_model_env(model: str) -> dict[str, str]:
|
||||
"""Stop Claude Code from spawning unpaid Anthropic-named subagent models."""
|
||||
|
||||
return {
|
||||
"ANTHROPIC_MODEL": model,
|
||||
"ANTHROPIC_DEFAULT_OPUS_MODEL": model,
|
||||
"ANTHROPIC_DEFAULT_SONNET_MODEL": model,
|
||||
"ANTHROPIC_DEFAULT_HAIKU_MODEL": model,
|
||||
"CLAUDE_CODE_SUBAGENT_MODEL": model,
|
||||
"API_TIMEOUT_MS": str(GATEWAY_REQUEST_TIMEOUT_S * 1000),
|
||||
}
|
||||
|
||||
|
||||
def model_session_environment(
|
||||
*,
|
||||
auth_token: str | None,
|
||||
base_url: str | None,
|
||||
model: str,
|
||||
build_sandbox_environment: Any,
|
||||
) -> dict[str, str]:
|
||||
env = build_sandbox_environment(auth_token=auth_token, base_url=base_url)
|
||||
if base_url:
|
||||
env.update(claude_gateway_model_env(model))
|
||||
return env
|
||||
|
||||
|
||||
def credential_secrets(args: argparse.Namespace) -> list[str]:
|
||||
return [
|
||||
secret
|
||||
for secret in (
|
||||
getattr(args, "auth_token", None),
|
||||
getattr(args, "openai_api_key", None),
|
||||
)
|
||||
if secret
|
||||
]
|
||||
|
||||
|
||||
def _first_nonblank_env(*names: str) -> str | None:
|
||||
for name in names:
|
||||
value = (os.environ.get(name) or "").strip()
|
||||
if value:
|
||||
return value
|
||||
return None
|
||||
|
||||
|
||||
def anthropic_api_key_from_environ() -> str | None:
|
||||
return _first_nonblank_env(ANTHROPIC_API_KEY_ENV, LEGACY_ANTHROPIC_API_KEY_ENV)
|
||||
|
||||
|
||||
def openai_api_key_from_environ() -> str | None:
|
||||
return _first_nonblank_env(OPENAI_API_KEY_ENV, "OPENAI_API_KEY")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ModelAccess:
|
||||
start_proxy: bool
|
||||
openai_api_key: str | None = None
|
||||
|
||||
|
||||
def resolve_model_access(
|
||||
*,
|
||||
auth_token: str | None,
|
||||
openai_api_key: str | None,
|
||||
base_url: str | None,
|
||||
models: Sequence[str],
|
||||
) -> ModelAccess:
|
||||
token = (auth_token or "").strip() or None
|
||||
openai_key = (openai_api_key or "").strip() or None
|
||||
gateway = (base_url or "").strip() or None
|
||||
names = [name.strip() for name in models if name and name.strip()]
|
||||
openai_models = [name for name in names if is_openai_model(name)]
|
||||
other_models = [name for name in names if not is_openai_model(name)]
|
||||
|
||||
if gateway:
|
||||
if not token:
|
||||
raise ValueError(
|
||||
"--base-url requires --anthropic-api-key / GITNEXUS_BENCH_ANTHROPIC_API_KEY "
|
||||
"(legacy --auth-token / GITNEXUS_BENCH_AUTH_TOKEN is still accepted)"
|
||||
)
|
||||
return ModelAccess(start_proxy=False)
|
||||
if openai_models and other_models:
|
||||
raise ValueError(
|
||||
"do not mix OpenAI model ids with Anthropic/other ids in one run; "
|
||||
f"openai={openai_models!r} other={other_models!r}"
|
||||
)
|
||||
if openai_models:
|
||||
if not openai_key:
|
||||
raise ValueError(
|
||||
"OpenAI model ids require GITNEXUS_BENCH_OPENAI_API_KEY "
|
||||
"(or OPENAI_API_KEY). Claude Code still speaks Anthropic "
|
||||
"protocol; the harness starts a loopback LiteLLM proxy."
|
||||
)
|
||||
return ModelAccess(start_proxy=True, openai_api_key=openai_key)
|
||||
return ModelAccess(start_proxy=False)
|
||||
|
||||
|
||||
def openai_litellm_config(model_names: Sequence[str]) -> dict[str, Any]:
|
||||
seen: list[str] = []
|
||||
for name in model_names:
|
||||
trimmed = name.strip()
|
||||
if trimmed and trimmed not in seen:
|
||||
seen.append(trimmed)
|
||||
if not seen:
|
||||
raise ValueError("OpenAI gateway requires at least one model name")
|
||||
return {
|
||||
"model_list": [
|
||||
{
|
||||
"model_name": name,
|
||||
"litellm_params": {
|
||||
"model": openai_backend_model(name),
|
||||
"api_key": "os.environ/OPENAI_API_KEY",
|
||||
"timeout": GATEWAY_REQUEST_TIMEOUT_S,
|
||||
},
|
||||
# GPT-5.6 tool use with active reasoning is supported through
|
||||
# OpenAI's Responses API, not Chat Completions.
|
||||
"model_info": {"mode": "responses"},
|
||||
}
|
||||
for name in seen
|
||||
],
|
||||
"litellm_settings": {"request_timeout": GATEWAY_REQUEST_TIMEOUT_S},
|
||||
"general_settings": {"master_key": "os.environ/LITELLM_MASTER_KEY"},
|
||||
}
|
||||
|
||||
|
||||
def write_openai_litellm_config(path: Path, model_names: Sequence[str]) -> Path:
|
||||
path.write_text(yaml.safe_dump(openai_litellm_config(model_names), sort_keys=False))
|
||||
path.chmod(0o600)
|
||||
return path
|
||||
|
||||
|
||||
def _free_loopback_port() -> int:
|
||||
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as sock:
|
||||
sock.bind(("127.0.0.1", 0))
|
||||
return int(sock.getsockname()[1])
|
||||
|
||||
|
||||
def litellm_proxy_argv(
|
||||
*,
|
||||
config: Path,
|
||||
host: str,
|
||||
port: int,
|
||||
python_executable: str | None = None,
|
||||
) -> list[str]:
|
||||
"""Build the LiteLLM proxy argv for this interpreter.
|
||||
|
||||
``python -m litellm`` fails on current releases (no ``litellm.__main__``).
|
||||
Prefer the console script next to ``sys.executable``; under ``uv run`` that
|
||||
path is the base CPython, so also honor ``VIRTUAL_ENV`` and ``PATH``.
|
||||
"""
|
||||
|
||||
python = Path(python_executable or sys.executable).resolve()
|
||||
candidates: list[Path] = [python.with_name("litellm")]
|
||||
virtual_env = (os.environ.get("VIRTUAL_ENV") or "").strip()
|
||||
if virtual_env:
|
||||
candidates.append(Path(virtual_env) / "bin" / "litellm")
|
||||
which = shutil.which("litellm")
|
||||
if which:
|
||||
candidates.append(Path(which))
|
||||
litellm_bin: Path | None = None
|
||||
for candidate in candidates:
|
||||
if candidate.is_file() and os.access(candidate, os.X_OK):
|
||||
litellm_bin = candidate.resolve()
|
||||
break
|
||||
if litellm_bin is None:
|
||||
raise RuntimeError(
|
||||
f"LiteLLM console script missing next to {python} "
|
||||
"(install litellm[proxy]; do not use python -m litellm)"
|
||||
)
|
||||
return [
|
||||
str(litellm_bin),
|
||||
"--config",
|
||||
str(config),
|
||||
"--host",
|
||||
host,
|
||||
"--port",
|
||||
str(port),
|
||||
]
|
||||
|
||||
|
||||
class OpenAIGateway(AbstractContextManager["OpenAIGateway"]):
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
openai_api_key: str,
|
||||
model_names: Sequence[str],
|
||||
work_dir: Path,
|
||||
ready_timeout_s: float | None = None,
|
||||
) -> None:
|
||||
self.openai_api_key = openai_api_key
|
||||
self.model_names = tuple(model_names)
|
||||
self.work_dir = work_dir
|
||||
self.ready_timeout_s = gateway_ready_timeout_s() if ready_timeout_s is None else ready_timeout_s
|
||||
if not math.isfinite(self.ready_timeout_s) or self.ready_timeout_s <= 0:
|
||||
raise ValueError("gateway readiness timeout must be finite and positive")
|
||||
self.auth_token = secrets.token_hex(16)
|
||||
self.port = _free_loopback_port()
|
||||
self.base_url = f"http://127.0.0.1:{self.port}"
|
||||
self.log_path = work_dir / "litellm.log"
|
||||
self._process: subprocess.Popen[bytes] | None = None
|
||||
self._log: Any = None
|
||||
self._job: Any = None
|
||||
|
||||
def log_tail(self, limit: int = 1000) -> str:
|
||||
try:
|
||||
text = self.log_path.read_bytes()[-limit:].decode(errors="replace")
|
||||
for secret in (self.openai_api_key, self.auth_token):
|
||||
if secret:
|
||||
text = text.replace(secret, "[REDACTED]")
|
||||
return text
|
||||
except OSError:
|
||||
return ""
|
||||
|
||||
def __enter__(self) -> OpenAIGateway:
|
||||
self.work_dir.mkdir(parents=True, exist_ok=True)
|
||||
config = write_openai_litellm_config(self.work_dir / "litellm.yaml", self.model_names)
|
||||
env = {
|
||||
"HOME": str(self.work_dir),
|
||||
"PATH": os.environ.get("PATH", "/usr/local/bin:/usr/bin:/bin"),
|
||||
"LANG": "C.UTF-8",
|
||||
"LC_ALL": "C.UTF-8",
|
||||
"PYTHONUNBUFFERED": "1",
|
||||
"LITELLM_LOCAL_MODEL_COST_MAP": "True",
|
||||
"OPENAI_API_KEY": self.openai_api_key,
|
||||
"LITELLM_MASTER_KEY": self.auth_token,
|
||||
}
|
||||
if os.name == "nt":
|
||||
# Windows subprocess DLL/socket initialization needs SystemRoot.
|
||||
# Keep the rest of the gateway's credential boundary explicit.
|
||||
env["SystemRoot"] = os.environ["SystemRoot"]
|
||||
# Never hand the proxy a pipe: nothing drains it after startup, so the
|
||||
# proxy would block forever once its request logs fill the 64 KiB pipe
|
||||
# buffer, and every later session request would hang without a status.
|
||||
try:
|
||||
self._log = self.log_path.open("wb")
|
||||
self.log_path.chmod(0o600)
|
||||
self._process = subprocess.Popen(
|
||||
[
|
||||
sys.executable,
|
||||
str(Path(__file__).with_name("gateway_supervisor.py")),
|
||||
*litellm_proxy_argv(
|
||||
config=config,
|
||||
host="127.0.0.1",
|
||||
port=self.port,
|
||||
),
|
||||
],
|
||||
cwd=str(self.work_dir),
|
||||
env=env,
|
||||
# This pipe's non-inheritable write end belongs only to this
|
||||
# parent. EOF reaches the supervisor even after SIGKILL.
|
||||
stdin=subprocess.PIPE,
|
||||
stdout=self._log,
|
||||
stderr=subprocess.STDOUT,
|
||||
start_new_session=os.name != "nt",
|
||||
creationflags=0x00000004 | 0x00000200 if os.name == "nt" else 0,
|
||||
)
|
||||
if os.name == "nt":
|
||||
from .process_control import _WindowsJob
|
||||
|
||||
self._job = _WindowsJob(self._process, [(self._process, None, None)])
|
||||
except BaseException as exc:
|
||||
if os.name == "nt" and self._job is None and self._process is not None:
|
||||
# A failed Job Object creation must not leave the supervisor
|
||||
# suspended before it can observe the liveness pipe.
|
||||
self._process.kill()
|
||||
self._process.wait(timeout=5)
|
||||
self.close()
|
||||
if not isinstance(exc, OSError):
|
||||
raise
|
||||
raise RuntimeError(f"failed to start the OpenAI LiteLLM gateway: {exc}") from exc
|
||||
try:
|
||||
self._wait_until_ready()
|
||||
except BaseException:
|
||||
self.close()
|
||||
raise
|
||||
return self
|
||||
|
||||
def __exit__(self, *exc: object) -> None:
|
||||
self.close()
|
||||
|
||||
def close(self) -> None:
|
||||
process = self._process
|
||||
try:
|
||||
if process is not None:
|
||||
if process.stdin is not None and not process.stdin.closed:
|
||||
process.stdin.close()
|
||||
# Keep the supervisor alive until its process-tree cleanup
|
||||
# completes. Killing it early would discard the ownership.
|
||||
process.wait(timeout=15)
|
||||
self._process = None
|
||||
finally:
|
||||
job, self._job = self._job, None
|
||||
if job is not None:
|
||||
job.close()
|
||||
log, self._log = self._log, None
|
||||
if log is not None:
|
||||
log.close()
|
||||
|
||||
def _wait_until_ready(self) -> None:
|
||||
deadline = time.monotonic() + self.ready_timeout_s
|
||||
request = urllib.request.Request(
|
||||
f"{self.base_url}/health/liveliness",
|
||||
method="GET",
|
||||
)
|
||||
last_error = "gateway did not become ready"
|
||||
while time.monotonic() < deadline:
|
||||
process = self._process
|
||||
if process is not None and process.poll() is not None:
|
||||
detail = self.log_tail()
|
||||
raise RuntimeError(
|
||||
"OpenAI LiteLLM gateway exited before becoming ready"
|
||||
+ (f": {detail}" if detail else "")
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(request, timeout=1) as response:
|
||||
if 200 <= int(response.status) < 300:
|
||||
return
|
||||
except (urllib.error.URLError, TimeoutError, ConnectionError, OSError) as exc:
|
||||
last_error = str(exc)
|
||||
time.sleep(0.1)
|
||||
detail = self.log_tail()
|
||||
raise RuntimeError(
|
||||
f"OpenAI LiteLLM gateway was not ready on {self.base_url} after "
|
||||
f"{self.ready_timeout_s:.0f}s: {last_error}"
|
||||
+ (f"; proxy log: {detail}" if detail else "; proxy wrote no output yet")
|
||||
+ f" (raise {GATEWAY_READY_TIMEOUT_ENV} if this host is simply slow to import LiteLLM)"
|
||||
)
|
||||
|
||||
|
||||
class attach_openai_gateway(AbstractContextManager[argparse.Namespace]):
|
||||
"""Start a loopback OpenAI gateway when the selected models need one."""
|
||||
|
||||
def __init__(self, args: argparse.Namespace) -> None:
|
||||
self.args = args
|
||||
self._gateway: OpenAIGateway | None = None
|
||||
self._work_dir: Path | None = None
|
||||
|
||||
def __enter__(self) -> argparse.Namespace:
|
||||
models = [self.args.model]
|
||||
proposer = getattr(self.args, "proposer_model", None)
|
||||
if proposer:
|
||||
models.append(proposer)
|
||||
access = resolve_model_access(
|
||||
auth_token=getattr(self.args, "auth_token", None),
|
||||
openai_api_key=getattr(self.args, "openai_api_key", None),
|
||||
base_url=getattr(self.args, "base_url", None),
|
||||
models=models,
|
||||
)
|
||||
if not access.start_proxy:
|
||||
return self.args
|
||||
assert access.openai_api_key is not None
|
||||
self._work_dir = Path(tempfile.mkdtemp(prefix="wfgateway-"))
|
||||
self._gateway = OpenAIGateway(
|
||||
openai_api_key=access.openai_api_key,
|
||||
model_names=models,
|
||||
work_dir=self._work_dir,
|
||||
)
|
||||
try:
|
||||
started = self._gateway.__enter__()
|
||||
except BaseException:
|
||||
self.__exit__(None, None, None)
|
||||
raise
|
||||
self.args.base_url = started.base_url
|
||||
self.args.auth_token = started.auth_token
|
||||
return self.args
|
||||
|
||||
def __exit__(self, *exc: object) -> None:
|
||||
if self._gateway is not None:
|
||||
self._gateway.__exit__(*exc)
|
||||
self._gateway = None
|
||||
if self._work_dir is not None:
|
||||
try:
|
||||
for child in self._work_dir.glob("*"):
|
||||
child.unlink(missing_ok=True)
|
||||
self._work_dir.rmdir()
|
||||
except OSError:
|
||||
# Best-effort: a leftover empty work dir must not hide the
|
||||
# original gateway error or block process teardown.
|
||||
pass
|
||||
self._work_dir = None
|
||||
|
|
@ -24,6 +24,10 @@ MAX_ORACLE_PATH_BYTES = 240
|
|||
MAX_ORACLE_COMMAND_BYTES = 8 * 1024
|
||||
ORACLE_ENV_VAR = "GITNEXUS_BENCH_ORACLE_ROOT"
|
||||
HIDDEN_HARNESS_PATH = PurePosixPath("eval/workflow_bench")
|
||||
# Historical PR diffs can still edit this tree. Sanitization deletes it before
|
||||
# setup, so `git apply` must skip those hunks or it fails with
|
||||
# `error: eval/workflow_bench/<file>: No such file or directory`.
|
||||
HIDDEN_HARNESS_APPLY_EXCLUDE = f"{HIDDEN_HARNESS_PATH.as_posix()}/*"
|
||||
MAX_CLONE_REFS = 1024
|
||||
MAX_CLONE_REF_BYTES = 2 * 1024 * 1024
|
||||
|
||||
|
|
@ -240,6 +244,76 @@ def _git_checked(
|
|||
return result.stdout_tail.strip()
|
||||
|
||||
|
||||
def hidden_harness_dir(clone: Path) -> Path:
|
||||
"""Checkout-relative path of the hidden benchmark harness tree."""
|
||||
|
||||
return clone / HIDDEN_HARNESS_PATH
|
||||
|
||||
|
||||
def hidden_harness_is_visible(clone: Path) -> bool:
|
||||
"""True when the hidden harness still exists as a file, directory, or symlink."""
|
||||
|
||||
leftover = hidden_harness_dir(clone)
|
||||
return leftover.exists() or leftover.is_symlink()
|
||||
|
||||
|
||||
def require_hidden_harness_absent(clone: Path) -> None:
|
||||
"""Fail closed if setup left ``eval/workflow_bench`` visible to the model.
|
||||
|
||||
Review cells stage a historical patch under this tree so ``git apply`` can
|
||||
read it. Overlaying an empty mask on the same path hides that file and
|
||||
aborts every cell with ``can't open patch``. The runner therefore leaves
|
||||
the staged copy visible during sandboxed setup, then requires this tree to
|
||||
be gone before the model session starts.
|
||||
"""
|
||||
|
||||
if hidden_harness_is_visible(clone):
|
||||
raise ValueError(
|
||||
"task setup left the hidden harness visible to the model: "
|
||||
f"{HIDDEN_HARNESS_PATH.as_posix()}"
|
||||
)
|
||||
|
||||
|
||||
def with_hidden_harness_apply_exclude(setup: str) -> str:
|
||||
"""Skip hunks for the harness tree sanitization already deleted.
|
||||
|
||||
Review cells copy a historical PR patch back under ``eval/workflow_bench``
|
||||
and then ``git apply`` it. That patch may still mention harness files
|
||||
(``learnings.jsonl`` on review-pr-2718-defect). Those files are gone from
|
||||
the parentless snapshot, and setup deletes the directory again after apply,
|
||||
so the hunks are never model-visible.
|
||||
|
||||
The runner must not overlay a mask on ``eval/workflow_bench`` before that
|
||||
``git apply``: the staged patch lives at the same path, and a mask makes
|
||||
``git apply`` fail with ``can't open patch``.
|
||||
"""
|
||||
|
||||
if "git apply" not in setup:
|
||||
return setup
|
||||
flag = f"--exclude='{HIDDEN_HARNESS_APPLY_EXCLUDE}'"
|
||||
if flag in setup or f'--exclude="{HIDDEN_HARNESS_APPLY_EXCLUDE}"' in setup:
|
||||
return setup
|
||||
return setup.replace("git apply ", f"git apply {flag} ", 1)
|
||||
|
||||
|
||||
def review_case_setup_command(patch_name: str) -> str:
|
||||
"""Apply one review-case patch, then delete the staged harness copy.
|
||||
|
||||
The patch is staged back under ``eval/workflow_bench`` after sanitization
|
||||
so this command can read it inside the sandbox. The trailing ``rm -rf``
|
||||
is what hides the tree from the model — not an empty overlay on the
|
||||
same path, which would hide the patch from ``git apply``.
|
||||
"""
|
||||
|
||||
if not patch_name or "/" in patch_name or "\\" in patch_name or patch_name in {".", ".."}:
|
||||
raise ValueError(f"review patch name must be a single path segment: {patch_name!r}")
|
||||
patch = f"{HIDDEN_HARNESS_PATH.as_posix()}/review_cases/{patch_name}"
|
||||
return (
|
||||
f"git apply --exclude='{HIDDEN_HARNESS_APPLY_EXCLUDE}' {patch} "
|
||||
f"&& rm -rf {HIDDEN_HARNESS_PATH.as_posix()}"
|
||||
)
|
||||
|
||||
|
||||
def sanitize_clone_for_hidden_oracles(clone: Path) -> str:
|
||||
"""Remove the harness and its recoverable Git history from a disposable clone.
|
||||
|
||||
|
|
|
|||
|
|
@ -0,0 +1,25 @@
|
|||
import { describe, expect, it } from 'vitest';
|
||||
import { SupportedLanguages } from 'gitnexus-shared';
|
||||
|
||||
import { SCOPE_RESOLVERS } from '../gitnexus/src/core/ingestion/scope-resolution/pipeline/registry.js';
|
||||
|
||||
describe('hidden oracle: C system headers do not bind to repository decoys', () => {
|
||||
it('rejects a stdio.h decoy while preserving local-header resolution', () => {
|
||||
const resolver = SCOPE_RESOLVERS.get(SupportedLanguages.C);
|
||||
expect(resolver).toBeDefined();
|
||||
const files = new Set(['src/stdio.h', 'include/util.h', 'src/main.c']);
|
||||
const context = { parsedFiles: [] };
|
||||
|
||||
const external = resolver!.resolveImportTarget(
|
||||
'stdio.h',
|
||||
'src/main.c',
|
||||
files,
|
||||
undefined,
|
||||
context,
|
||||
);
|
||||
const local = resolver!.resolveImportTarget('util.h', 'src/main.c', files, undefined, context);
|
||||
|
||||
expect(external).toBeNull();
|
||||
expect(local).toBe('include/util.h');
|
||||
});
|
||||
});
|
||||
|
|
@ -1,30 +0,0 @@
|
|||
import { describe, expect, it, vi } from 'vitest';
|
||||
|
||||
vi.mock('../gitnexus/src/mcp/local/pdg-impact.js', async (importOriginal) => {
|
||||
const actual = await importOriginal<Record<string, unknown>>();
|
||||
return { ...actual, pdgStampForMode: vi.fn().mockResolvedValue(false) };
|
||||
});
|
||||
|
||||
import { LocalBackend } from '../gitnexus/src/mcp/local/local-backend.js';
|
||||
|
||||
describe('hidden oracle: pdg_query missing sub-layer note', () => {
|
||||
it.each([
|
||||
['controls', 'CDG'],
|
||||
['flows', 'REACHING_DEF'],
|
||||
] as const)("names the %s mode's %s sub-layer", async (mode, expectedLayer) => {
|
||||
const backend = Object.create(LocalBackend.prototype) as LocalBackend & {
|
||||
ensureInitialized: () => Promise<void>;
|
||||
_pdgQueryImpl: (repo: unknown, params: unknown) => Promise<Record<string, unknown>>;
|
||||
};
|
||||
backend.ensureInitialized = vi.fn().mockResolvedValue(undefined);
|
||||
|
||||
const result = await backend._pdgQueryImpl(
|
||||
{ lbugPath: '/unreachable-hidden-oracle-db' },
|
||||
{ mode, target: 'src/example.ts' },
|
||||
);
|
||||
|
||||
expect(result).toMatchObject({ mode, results: [], total: 0 });
|
||||
expect(String(result.note)).toContain(expectedLayer);
|
||||
expect(String(result.note)).toContain('gitnexus analyze --pdg');
|
||||
});
|
||||
});
|
||||
|
|
@ -0,0 +1,53 @@
|
|||
{
|
||||
"schema_version": 1,
|
||||
"findings": [
|
||||
{
|
||||
"id": "pr2108-leakage-classified-as-plateau",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/scripts/bench/fts-evict-reload-rss.mjs",
|
||||
"line_start": 364,
|
||||
"line_end": 364,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2108-batch-failure-silently-demotes-symbols",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/src/mcp/local/local-backend.ts",
|
||||
"line_start": 1134,
|
||||
"line_end": 1178,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2108-build-only-env-leaks-to-runtime",
|
||||
"severity": "low",
|
||||
"path": "Dockerfile.cli",
|
||||
"line_start": 88,
|
||||
"line_end": 88,
|
||||
"category": "other"
|
||||
},
|
||||
{
|
||||
"id": "pr2108-native-benchmark-bypasses-production-open",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/scripts/bench/fts-evict-reload-rss.mjs",
|
||||
"line_start": 122,
|
||||
"line_end": 186,
|
||||
"category": "tests"
|
||||
},
|
||||
{
|
||||
"id": "pr2108-unpinned-build-extension",
|
||||
"severity": "medium",
|
||||
"path": "Dockerfile.cli",
|
||||
"line_start": 86,
|
||||
"line_end": 89,
|
||||
"category": "security"
|
||||
},
|
||||
{
|
||||
"id": "pr2108-verify-flag-parsed-as-extension",
|
||||
"severity": "low",
|
||||
"path": "gitnexus/scripts/install-duckdb-extension.mjs",
|
||||
"line_start": 59,
|
||||
"line_end": 59,
|
||||
"category": "correctness"
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
@ -0,0 +1 @@
|
|||
{ "schema_version": 1, "findings": [] }
|
||||
|
|
@ -0,0 +1,21 @@
|
|||
{
|
||||
"schema_version": 1,
|
||||
"findings": [
|
||||
{
|
||||
"id": "pr2258-vacuous-gated-recall-pass",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/bench/impact-pdg/gate-mutation-recall.mjs",
|
||||
"line_start": 20,
|
||||
"line_end": 29,
|
||||
"category": "tests"
|
||||
},
|
||||
{
|
||||
"id": "pr2258-stale-child-invocation-doc",
|
||||
"severity": "low",
|
||||
"path": "gitnexus/bench/impact-pdg/README.md",
|
||||
"line_start": 233,
|
||||
"line_end": 233,
|
||||
"category": "other"
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
@ -0,0 +1,61 @@
|
|||
{
|
||||
"schema_version": 1,
|
||||
"findings": [
|
||||
{
|
||||
"id": "pr2718-multiline-closure-join",
|
||||
"severity": "high",
|
||||
"path": "gitnexus/src/core/ingestion/scope-resolution/graph-bridge/ids.ts",
|
||||
"line_start": 212,
|
||||
"line_end": 212,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2718-constructor-parameter-property-identity",
|
||||
"severity": "high",
|
||||
"path": "gitnexus/src/core/ingestion/workers/parse-worker.ts",
|
||||
"line_start": 2310,
|
||||
"line_end": 2310,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2718-dart-top-level-closures",
|
||||
"severity": "high",
|
||||
"path": "gitnexus/src/core/ingestion/languages/dart/query.ts",
|
||||
"line_start": 109,
|
||||
"line_end": 109,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2718-ruby-closure-forms",
|
||||
"severity": "high",
|
||||
"path": "gitnexus/src/core/ingestion/languages/ruby/query.ts",
|
||||
"line_start": 106,
|
||||
"line_end": 106,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2718-dart-signature-assumption",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/src/core/ingestion/utils/ast-helpers.ts",
|
||||
"line_start": 1283,
|
||||
"line_end": 1283,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2718-receiver-qualified-lambda",
|
||||
"severity": "low",
|
||||
"path": "gitnexus/src/core/ingestion/languages/ruby/query.ts",
|
||||
"line_start": 104,
|
||||
"line_end": 104,
|
||||
"category": "correctness"
|
||||
},
|
||||
{
|
||||
"id": "pr2718-rust-closure-source-coverage",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/src/core/ingestion/tree-sitter-queries.ts",
|
||||
"line_start": 1346,
|
||||
"line_end": 1346,
|
||||
"category": "tests"
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
@ -0,0 +1 @@
|
|||
{ "schema_version": 1, "findings": [] }
|
||||
|
|
@ -0,0 +1,53 @@
|
|||
{
|
||||
"schema_version": 1,
|
||||
"findings": [
|
||||
{
|
||||
"id": "pr2794-module-global-retains-parsed-files",
|
||||
"severity": "high",
|
||||
"path": "gitnexus/src/core/ingestion/languages/cpp/inline-namespaces.ts",
|
||||
"line_start": 147,
|
||||
"line_end": 147,
|
||||
"category": "performance"
|
||||
},
|
||||
{
|
||||
"id": "pr2794-benchmark-misses-receiver-regression",
|
||||
"severity": "high",
|
||||
"path": "gitnexus/bench/cpp-qualified-ns/measure.mjs",
|
||||
"line_start": 142,
|
||||
"line_end": 142,
|
||||
"category": "tests"
|
||||
},
|
||||
{
|
||||
"id": "pr2794-recursive-quadratic-namespace-walk",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/src/core/ingestion/languages/cpp/inline-namespaces.ts",
|
||||
"line_start": 193,
|
||||
"line_end": 223,
|
||||
"category": "performance"
|
||||
},
|
||||
{
|
||||
"id": "pr2794-scaling-gate-not-wired",
|
||||
"severity": "medium",
|
||||
"path": ".github/workflows/ci-tests.yml",
|
||||
"line_start": 509,
|
||||
"line_end": 509,
|
||||
"category": "tests"
|
||||
},
|
||||
{
|
||||
"id": "pr2794-cross-file-dedup-coverage",
|
||||
"severity": "medium",
|
||||
"path": "gitnexus/test/unit/cpp-qualified-ns-index.test.ts",
|
||||
"line_start": 78,
|
||||
"line_end": 78,
|
||||
"category": "tests"
|
||||
},
|
||||
{
|
||||
"id": "pr2794-false-load-bearing-order-contract",
|
||||
"severity": "low",
|
||||
"path": "gitnexus/src/core/ingestion/languages/cpp/inline-namespaces.ts",
|
||||
"line_start": 138,
|
||||
"line_end": 138,
|
||||
"category": "other"
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
@ -0,0 +1,28 @@
|
|||
import { spawnSync } from 'node:child_process';
|
||||
import path from 'node:path';
|
||||
import { describe, expect, it } from 'vitest';
|
||||
|
||||
describe('hidden oracle: status -j JSON alias', () => {
|
||||
it('returns the same machine-readable status shape as --json', () => {
|
||||
const executable = path.resolve('node_modules/.bin/tsx');
|
||||
const options = {
|
||||
cwd: process.cwd(),
|
||||
encoding: 'utf8' as const,
|
||||
env: { ...process.env, NO_COLOR: '1' },
|
||||
};
|
||||
const short = spawnSync(executable, ['src/cli/index.ts', 'status', '-j'], options);
|
||||
const long = spawnSync(executable, ['src/cli/index.ts', 'status', '--json'], options);
|
||||
|
||||
expect(short.error).toBeUndefined();
|
||||
expect(short.status).toBe(0);
|
||||
expect(short.stderr.trim()).toBe('');
|
||||
expect(long.status).toBe(0);
|
||||
|
||||
const shortPayload = JSON.parse(short.stdout) as Record<string, unknown>;
|
||||
const longPayload = JSON.parse(long.stdout) as Record<string, unknown>;
|
||||
expect(shortPayload.schemaVersion).toBe(1);
|
||||
expect(shortPayload).toHaveProperty('status');
|
||||
expect(shortPayload).toHaveProperty('repository');
|
||||
expect(shortPayload).toEqual(longPayload);
|
||||
});
|
||||
});
|
||||
|
|
@ -1,20 +0,0 @@
|
|||
import { spawnSync } from 'node:child_process';
|
||||
import path from 'node:path';
|
||||
import { describe, expect, it } from 'vitest';
|
||||
|
||||
import pkg from '../gitnexus/package.json';
|
||||
|
||||
describe('hidden oracle: -V version alias', () => {
|
||||
it('prints exactly the installed GitNexus version and exits successfully', () => {
|
||||
const result = spawnSync(path.resolve('node_modules/.bin/tsx'), ['src/cli/index.ts', '-V'], {
|
||||
cwd: process.cwd(),
|
||||
encoding: 'utf8',
|
||||
env: { ...process.env, NO_COLOR: '1' },
|
||||
});
|
||||
|
||||
expect(result.error).toBeUndefined();
|
||||
expect(result.status).toBe(0);
|
||||
expect(result.stdout.trim()).toBe(pkg.version);
|
||||
expect(result.stderr.trim()).toBe('');
|
||||
});
|
||||
});
|
||||
|
|
@ -12,9 +12,12 @@ import os
|
|||
import shutil
|
||||
import signal
|
||||
import subprocess
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
from collections.abc import Mapping, Sequence
|
||||
from collections.abc import Callable, Iterator, Mapping, Sequence
|
||||
from contextlib import contextmanager
|
||||
from contextvars import ContextVar
|
||||
from dataclasses import dataclass, replace
|
||||
from pathlib import Path
|
||||
from typing import BinaryIO, Literal
|
||||
|
|
@ -22,9 +25,42 @@ from typing import BinaryIO, Literal
|
|||
|
||||
MAX_TAIL_BYTES = 64 * 1024
|
||||
DEFAULT_TERMINATE_GRACE = 5.0
|
||||
_CANCELLATION: ContextVar[threading.Event | None] = ContextVar("managed_process_cancellation", default=None)
|
||||
|
||||
|
||||
@contextmanager
|
||||
def cancellation_scope(
|
||||
event: threading.Event | None = None, *, handle_signals: bool = False
|
||||
) -> Iterator[threading.Event]:
|
||||
"""Share one cancellation signal through every managed command in a run.
|
||||
|
||||
Worker contexts must be copied explicitly at executor submission. This also
|
||||
covers clone/setup and evidence helpers which call run_managed indirectly.
|
||||
"""
|
||||
event = event or _CANCELLATION.get() or threading.Event()
|
||||
token = _CANCELLATION.set(event)
|
||||
previous = {}
|
||||
try:
|
||||
if handle_signals and threading.current_thread() is threading.main_thread():
|
||||
for signum in (signal.SIGINT, signal.SIGTERM):
|
||||
previous[signum] = signal.signal(signum, lambda *_: event.set())
|
||||
yield event
|
||||
except BaseException:
|
||||
event.set()
|
||||
raise
|
||||
finally:
|
||||
for signum, handler in previous.items():
|
||||
signal.signal(signum, handler)
|
||||
_CANCELLATION.reset(token)
|
||||
|
||||
|
||||
class _CancellationRequested(Exception):
|
||||
pass
|
||||
|
||||
|
||||
ProcessState = Literal[
|
||||
"exited",
|
||||
"cancelled",
|
||||
"input-failure",
|
||||
"timeout",
|
||||
"forced-kill",
|
||||
|
|
@ -124,12 +160,36 @@ def _drain(
|
|||
pipe: BinaryIO,
|
||||
tail: _TailBuffer,
|
||||
capture: _BoundedCapture | None = None,
|
||||
echo: BinaryIO | None = None,
|
||||
observer: Callable[[bytes], None] | None = None,
|
||||
) -> None:
|
||||
try:
|
||||
while chunk := pipe.read(8192):
|
||||
# read1, not read: on a BufferedReader, read(n) blocks until it has all
|
||||
# n bytes or the pipe closes. A child that emits a line every 45 minutes
|
||||
# never fills 8 KB, so its output would surface only when it exits —
|
||||
# which is precisely what echo_stdout exists to avoid.
|
||||
while chunk := pipe.read1(8192):
|
||||
tail.append(chunk)
|
||||
if capture is not None:
|
||||
capture.append(chunk)
|
||||
if echo is not None:
|
||||
# Progress passthrough for long child runs whose output is
|
||||
# ordinary log text (see run_managed's echo_stdout). Never
|
||||
# enabled for a Claude session, whose stdout is the evidence
|
||||
# stream and is only written out after redaction.
|
||||
try:
|
||||
echo.write(chunk)
|
||||
echo.flush()
|
||||
except (OSError, ValueError):
|
||||
echo = None
|
||||
if observer is not None:
|
||||
# Progress reporting must never be able to break the drain, and
|
||||
# the drain must keep running even if the observer is broken:
|
||||
# a stalled reader is what deadlocks the child.
|
||||
try:
|
||||
observer(chunk)
|
||||
except Exception:
|
||||
observer = None
|
||||
except (OSError, ValueError):
|
||||
# A forced close is part of the reap path. The terminal result records
|
||||
# an actual reap failure; a reader seeing the close is not one itself.
|
||||
|
|
@ -419,11 +479,17 @@ def _run_managed_inner(
|
|||
require_pid_namespace: bool = False,
|
||||
stdin_data: bytes | None = None,
|
||||
capture_stdout_bytes: int | None = None,
|
||||
echo_stdout: bool = False,
|
||||
stdout_observer: Callable[[bytes], None] | None = None,
|
||||
cancel_event: threading.Event | None = None,
|
||||
_ownership_slot: list[tuple[subprocess.Popen[bytes], _WindowsJob | None, int | None]],
|
||||
) -> ManagedProcessResult:
|
||||
"""Implementation registered with an outer post-spawn ownership guard."""
|
||||
|
||||
started = time.monotonic()
|
||||
cancel_event = cancel_event or _CANCELLATION.get()
|
||||
if cancel_event is not None and cancel_event.is_set():
|
||||
return _empty_result("cancelled", started, "cancelled before spawn")
|
||||
if timeout <= 0 or terminate_grace < 0 or tail_bytes <= 0:
|
||||
raise ValueError("timeout and tail_bytes must be positive; terminate_grace must be non-negative")
|
||||
if capture_stdout_bytes is not None and capture_stdout_bytes <= 0:
|
||||
|
|
@ -465,8 +531,13 @@ def _run_managed_inner(
|
|||
stdout = _TailBuffer(tail_bytes)
|
||||
stderr = _TailBuffer(tail_bytes)
|
||||
stdout_capture = _BoundedCapture(capture_stdout_bytes) if capture_stdout_bytes is not None else None
|
||||
echo = getattr(sys.stderr, "buffer", None) if echo_stdout else None
|
||||
readers = [
|
||||
threading.Thread(target=_drain, args=(process.stdout, stdout, stdout_capture), daemon=True),
|
||||
threading.Thread(
|
||||
target=_drain,
|
||||
args=(process.stdout, stdout, stdout_capture, echo, stdout_observer),
|
||||
daemon=True,
|
||||
),
|
||||
threading.Thread(target=_drain, args=(process.stderr, stderr), daemon=True),
|
||||
]
|
||||
for reader in readers:
|
||||
|
|
@ -486,14 +557,28 @@ def _run_managed_inner(
|
|||
detail = None
|
||||
timed_out = False
|
||||
forced_kill = False
|
||||
cancelled = False
|
||||
try:
|
||||
process.wait(timeout=timeout)
|
||||
deadline = time.monotonic() + timeout
|
||||
while True:
|
||||
if cancel_event is not None and cancel_event.is_set():
|
||||
raise _CancellationRequested()
|
||||
remaining = deadline - time.monotonic()
|
||||
if remaining <= 0:
|
||||
raise subprocess.TimeoutExpired(command, timeout)
|
||||
try:
|
||||
process.wait(timeout=min(0.1, remaining) if cancel_event is not None else remaining)
|
||||
break
|
||||
except subprocess.TimeoutExpired:
|
||||
if time.monotonic() >= deadline:
|
||||
raise
|
||||
except (KeyboardInterrupt, SystemExit):
|
||||
_abort_owned_process(process, job, owned_pgid)
|
||||
raise
|
||||
except subprocess.TimeoutExpired:
|
||||
timed_out = True
|
||||
state = "timeout"
|
||||
except (subprocess.TimeoutExpired, _CancellationRequested) as stopped:
|
||||
cancelled = isinstance(stopped, _CancellationRequested)
|
||||
timed_out = not cancelled
|
||||
state = "cancelled" if cancelled else "timeout"
|
||||
try:
|
||||
if job is not None:
|
||||
# Job Object termination is the Windows tree-wide primitive;
|
||||
|
|
@ -656,6 +741,8 @@ def _run_managed_inner(
|
|||
raise
|
||||
|
||||
captured_stdout, capture_overflow = stdout_capture.result() if stdout_capture is not None else (None, False)
|
||||
if cancelled and state in ("timeout", "forced-kill", "cancelled"):
|
||||
state = "cancelled"
|
||||
return ManagedProcessResult(
|
||||
state=state,
|
||||
returncode=process.returncode,
|
||||
|
|
@ -683,8 +770,22 @@ def run_managed(
|
|||
require_pid_namespace: bool = False,
|
||||
stdin_data: bytes | None = None,
|
||||
capture_stdout_bytes: int | None = None,
|
||||
echo_stdout: bool = False,
|
||||
stdout_observer: Callable[[bytes], None] | None = None,
|
||||
cancel_event: threading.Event | None = None,
|
||||
) -> ManagedProcessResult:
|
||||
"""Run one command with bounded output and owned-tree termination."""
|
||||
"""Run one command with bounded output and owned-tree termination.
|
||||
|
||||
`echo_stdout` streams the child's stdout to this process's stderr as it
|
||||
arrives, so a long child (the benchmark sweep) reports progress in the CI
|
||||
log instead of surfacing only its bounded tail after it finishes. Use it
|
||||
only for children whose stdout is log text.
|
||||
|
||||
`stdout_observer` sees the same chunks without copying them anywhere, so a
|
||||
child whose stdout is *not* printable (a Claude session's evidence stream)
|
||||
can still report derived progress. The observer runs on the reader thread:
|
||||
it must not block, and raising only disables further calls.
|
||||
"""
|
||||
|
||||
ownership_slot: list[tuple[subprocess.Popen[bytes], _WindowsJob | None, int | None]] = []
|
||||
try:
|
||||
|
|
@ -699,6 +800,9 @@ def run_managed(
|
|||
require_pid_namespace=require_pid_namespace,
|
||||
stdin_data=stdin_data,
|
||||
capture_stdout_bytes=capture_stdout_bytes,
|
||||
echo_stdout=echo_stdout,
|
||||
stdout_observer=stdout_observer,
|
||||
cancel_event=cancel_event,
|
||||
_ownership_slot=ownership_slot,
|
||||
)
|
||||
except BaseException:
|
||||
|
|
|
|||
|
|
@ -42,6 +42,8 @@ def mirror_targets(relative: PurePosixPath) -> list[PurePosixPath]:
|
|||
rest = PurePosixPath(*relative.parts[3:])
|
||||
targets = [relative]
|
||||
targets += [PurePosixPath(root, skill, rest) for root in MIRROR_SKILL_ROOTS]
|
||||
if skill == "gitnexus-review":
|
||||
targets.append(PurePosixPath("gitnexus-cursor-integration/skills", skill, rest))
|
||||
return targets
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -27,6 +27,8 @@ SANDBOX_TMP = "/tmp"
|
|||
SANDBOX_CLAUDE = "/opt/claude/claude"
|
||||
SANDBOX_SHELL_PREFIX = "/opt/claude/shell-prefix"
|
||||
SANDBOX_PYTHON3 = "/opt/claude/python3"
|
||||
SANDBOX_GITNEXUS_CLI = "/opt/claude/gitnexus"
|
||||
SANDBOX_GIT_EXCLUDES = "/opt/claude/git-excludes"
|
||||
SANDBOX_NODE = "/opt/claude/node"
|
||||
SANDBOX_NODE_PREFIX = "/opt/claude/nodejs"
|
||||
# Vite transpiles a TypeScript config into <node_modules>/.vite-temp before it
|
||||
|
|
@ -42,12 +44,139 @@ SANDBOX_GITNEXUS = "/opt/gitnexus"
|
|||
SANDBOX_GITNEXUS_SHARED = "/opt/gitnexus-shared"
|
||||
SANDBOX_GITNEXUS_REGISTRY = "/opt/gitnexus-registry"
|
||||
SANDBOX_USER_SKILLS = f"{SANDBOX_HOME}/.claude/skills"
|
||||
SANDBOX_EVIDENCE = "/evidence"
|
||||
|
||||
# Claude Code's weaker nested sandbox overlays these absent root paths with
|
||||
# /dev/null devices. They are tool-created mount noise, not model-authored
|
||||
# files. Git must ignore them so provenance never mistakes those synthetic
|
||||
# devices for untracked repository content. Leading "/" anchors every pattern
|
||||
# at the repository root; tracked paths are never hidden by excludes.
|
||||
CLAUDE_SANDBOX_GIT_EXCLUDES = (
|
||||
"/.bash_profile",
|
||||
"/.bashrc",
|
||||
"/.gitconfig",
|
||||
"/.idea",
|
||||
"/.profile",
|
||||
"/.ripgreprc",
|
||||
"/.vscode",
|
||||
"/.zprofile",
|
||||
"/.zshrc",
|
||||
"/scripts",
|
||||
)
|
||||
|
||||
|
||||
class SandboxError(RuntimeError):
|
||||
"""Containment could not be established without weakening the contract."""
|
||||
|
||||
|
||||
# Claude Code 2.1.214's subprocess scrubber and nested sandbox bind these
|
||||
# paths even when absent. Mount targets must exist before sealing the clone.
|
||||
REVIEW_RUNTIME_FILES = (
|
||||
"bunfig.toml",
|
||||
".mcp.json",
|
||||
"package.json",
|
||||
".npmrc",
|
||||
".yarnrc",
|
||||
".yarnrc.yml",
|
||||
".gitmodules",
|
||||
"package-lock.json",
|
||||
"yarn.lock",
|
||||
"pnpm-lock.yaml",
|
||||
".env",
|
||||
".env.local",
|
||||
".env.development",
|
||||
".env.development.local",
|
||||
".env.test",
|
||||
".env.test.local",
|
||||
".env.production",
|
||||
".env.production.local",
|
||||
".git/config",
|
||||
".git/config.lock",
|
||||
".git/config.worktree",
|
||||
".git/config.worktree.lock",
|
||||
".git/info/exclude",
|
||||
".gitconfig",
|
||||
".bash_profile",
|
||||
".bashrc",
|
||||
".profile",
|
||||
".ripgreprc",
|
||||
".zprofile",
|
||||
".zshrc",
|
||||
)
|
||||
REVIEW_RUNTIME_DIRECTORIES = (
|
||||
".git/hooks",
|
||||
".git/info",
|
||||
".git/modules",
|
||||
".git/worktrees",
|
||||
".claude/commands",
|
||||
".claude/agents",
|
||||
"node_modules/.bin",
|
||||
".github",
|
||||
"scripts",
|
||||
".vscode",
|
||||
".idea",
|
||||
)
|
||||
|
||||
|
||||
def prepare_review_workspace(sandbox: SandboxSession, artifact_name: str) -> Path:
|
||||
"""Prepare disposable mount targets; never truncate a pre-existing entry."""
|
||||
|
||||
clone = _real_directory(sandbox.clone, label="review clone")
|
||||
if PurePosixPath(artifact_name).name != artifact_name or "\\" in artifact_name or artifact_name in ("", ".", ".."):
|
||||
raise SandboxError("review artifact must be a root filename")
|
||||
output = clone / artifact_name
|
||||
# No agent runs while this private clone is being prepared. On POSIX the
|
||||
# directory descriptor additionally binds the exclusive create to its owner.
|
||||
directory_fd = None
|
||||
try:
|
||||
if os.name != "nt":
|
||||
directory_fd = os.open(clone, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW)
|
||||
fd = os.open(
|
||||
artifact_name if directory_fd is not None else output,
|
||||
os.O_WRONLY | os.O_CREAT | os.O_EXCL | getattr(os, "O_NOFOLLOW", 0),
|
||||
0o600,
|
||||
dir_fd=directory_fd,
|
||||
)
|
||||
try:
|
||||
if not stat.S_ISREG(os.fstat(fd).st_mode):
|
||||
raise SandboxError("review artifact must be a regular file")
|
||||
finally:
|
||||
os.close(fd)
|
||||
except FileExistsError as exc:
|
||||
raise SandboxError("review artifact already exists") from exc
|
||||
finally:
|
||||
if directory_fd is not None:
|
||||
os.close(directory_fd)
|
||||
|
||||
if sandbox.backend != "bwrap":
|
||||
return output
|
||||
created: list[str] = []
|
||||
for names, directory in ((REVIEW_RUNTIME_DIRECTORIES, True), (REVIEW_RUNTIME_FILES, False)):
|
||||
for name in names:
|
||||
# Record only absent entries. The no-follow traversal validates
|
||||
# every parent before creating anything in the host filesystem.
|
||||
missing = not os.path.lexists(clone / name)
|
||||
_prepare_clone_target(clone, PurePosixPath(name), directory=directory, label="review runtime")
|
||||
if missing:
|
||||
if name == ".mcp.json":
|
||||
(clone / name).write_text("{}\n")
|
||||
created.append(name)
|
||||
common_dir = clone / ".git/commondir"
|
||||
missing = not os.path.lexists(common_dir)
|
||||
_prepare_clone_target(clone, PurePosixPath(".git/commondir"), directory=False, label="review runtime")
|
||||
if missing:
|
||||
common_dir.write_text(".\n")
|
||||
created.append(".git/commondir")
|
||||
# Private status presentation only; never alter the repository's ignores.
|
||||
excludes = sandbox.private_root / "git-excludes"
|
||||
excludes.chmod(0o600)
|
||||
with excludes.open("a") as stream:
|
||||
stream.write("".join(f"/{name}\n" for name in created))
|
||||
excludes.chmod(0o400)
|
||||
(sandbox.private_root / "review-created-paths.json").write_text(json.dumps(created))
|
||||
return output
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ReadOnlyMount:
|
||||
source: Path
|
||||
|
|
@ -61,13 +190,58 @@ class SandboxSession:
|
|||
home: Path
|
||||
temp: Path
|
||||
bwrap_bin: Path
|
||||
backend: str
|
||||
claude_host_bin: Path
|
||||
command_prefix: list[str]
|
||||
read_only_mounts: tuple[ReadOnlyMount, ...]
|
||||
|
||||
@property
|
||||
def require_pid_namespace(self) -> bool:
|
||||
return self.backend == "bwrap"
|
||||
|
||||
def host_path(self, raw: str | Path) -> str:
|
||||
"""Translate a sandbox path to its real host path in unsafe mode."""
|
||||
|
||||
value = str(raw)
|
||||
if self.backend == "bwrap":
|
||||
return value
|
||||
node = Path(shutil.which("node") or "/usr/bin/node").resolve()
|
||||
mappings = [
|
||||
*(self.read_only_mounts),
|
||||
ReadOnlyMount(node.parent.parent, SANDBOX_NODE_PREFIX),
|
||||
ReadOnlyMount(node, SANDBOX_NODE),
|
||||
ReadOnlyMount(self.claude_host_bin, SANDBOX_CLAUDE),
|
||||
ReadOnlyMount(self.clone, SANDBOX_WORKSPACE),
|
||||
ReadOnlyMount(self.home, SANDBOX_HOME),
|
||||
ReadOnlyMount(self.temp, SANDBOX_TMP),
|
||||
]
|
||||
for mount in sorted(mappings, key=lambda item: len(item.target), reverse=True):
|
||||
target = mount.target.rstrip("/")
|
||||
if value == target:
|
||||
return str(mount.source)
|
||||
if value.startswith(f"{target}/"):
|
||||
return str(mount.source / value[len(target) + 1 :])
|
||||
return value
|
||||
|
||||
def host_text(self, value: str) -> str:
|
||||
if self.backend == "bwrap":
|
||||
return value
|
||||
targets = [
|
||||
*(mount.target for mount in self.read_only_mounts),
|
||||
SANDBOX_NODE_PREFIX,
|
||||
SANDBOX_NODE,
|
||||
SANDBOX_CLAUDE,
|
||||
SANDBOX_WORKSPACE,
|
||||
SANDBOX_HOME,
|
||||
SANDBOX_TMP,
|
||||
]
|
||||
ordered = sorted(set(targets), key=len, reverse=True)
|
||||
pattern = re.compile("|".join(re.escape(target) for target in ordered))
|
||||
return pattern.sub(lambda match: self.host_path(match.group(0)), value)
|
||||
|
||||
@property
|
||||
def claude_bin(self) -> str:
|
||||
return SANDBOX_CLAUDE
|
||||
return self.host_path(SANDBOX_CLAUDE)
|
||||
|
||||
@property
|
||||
def transcript_projects(self) -> Path:
|
||||
|
|
@ -75,7 +249,7 @@ class SandboxSession:
|
|||
|
||||
@property
|
||||
def settings_json(self) -> str:
|
||||
return build_claude_settings()
|
||||
return build_claude_settings(sandbox_enabled=self.backend == "bwrap")
|
||||
|
||||
def environment(
|
||||
self,
|
||||
|
|
@ -83,7 +257,17 @@ class SandboxSession:
|
|||
auth_token: str | None = None,
|
||||
base_url: str | None = None,
|
||||
) -> dict[str, str]:
|
||||
return build_sandbox_environment(auth_token=auth_token, base_url=base_url)
|
||||
env = build_sandbox_environment(auth_token=auth_token, base_url=base_url)
|
||||
if self.backend != "bwrap":
|
||||
env = {key: self.host_text(value) for key, value in env.items()}
|
||||
env.pop("CLAUDE_CODE_SUBPROCESS_ENV_SCRUB", None)
|
||||
env.pop("CLAUDE_CODE_SHELL_PREFIX", None)
|
||||
# Sandbox-only bin dirs have no single host counterpart; drop the
|
||||
# entries that survive translation as non-existent paths.
|
||||
env["PATH"] = ":".join(
|
||||
entry for entry in env["PATH"].split(":") if Path(entry).is_dir()
|
||||
)
|
||||
return env
|
||||
|
||||
def run(
|
||||
self,
|
||||
|
|
@ -93,11 +277,17 @@ class SandboxSession:
|
|||
env: Mapping[str, str] | None = None,
|
||||
stdin_data: bytes | None = None,
|
||||
) -> ManagedProcessResult:
|
||||
translated = [self.host_text(str(part)) for part in command]
|
||||
return run_managed(
|
||||
[*self.command_prefix, *command],
|
||||
[*self.command_prefix, *translated],
|
||||
cwd=None if self.command_prefix else self.clone,
|
||||
timeout=timeout,
|
||||
env=dict(env) if env is not None else build_sandbox_environment(),
|
||||
require_pid_namespace=True,
|
||||
env=(
|
||||
{key: self.host_text(value) for key, value in env.items()}
|
||||
if env is not None
|
||||
else self.environment()
|
||||
),
|
||||
require_pid_namespace=self.require_pid_namespace,
|
||||
stdin_data=stdin_data,
|
||||
)
|
||||
|
||||
|
|
@ -108,6 +298,7 @@ class SandboxSession:
|
|||
unshare_network: bool = False,
|
||||
read_only_paths: Sequence[Path] = (),
|
||||
extra_read_only_mounts: Sequence[ReadOnlyMount] = (),
|
||||
extra_writable_mounts: Sequence[ReadOnlyMount] = (),
|
||||
) -> list[str]:
|
||||
"""Build a stricter command boundary from this session's fixed roots.
|
||||
|
||||
|
|
@ -117,6 +308,9 @@ class SandboxSession:
|
|||
for harness-owned, post-session evidence such as hidden oracles.
|
||||
"""
|
||||
|
||||
if self.backend != "bwrap":
|
||||
return []
|
||||
|
||||
additional: list[ReadOnlyMount] = []
|
||||
clone = _real_directory(self.clone, label="sandbox clone")
|
||||
for raw_path in read_only_paths:
|
||||
|
|
@ -158,6 +352,25 @@ class SandboxSession:
|
|||
raise SandboxError(f"extra read-only mount target must be absolute: {mount.target}")
|
||||
additional.append(ReadOnlyMount(source=source, target=target.as_posix()))
|
||||
|
||||
writable: list[ReadOnlyMount] = []
|
||||
for mount in extra_writable_mounts:
|
||||
source = mount.source.expanduser().absolute()
|
||||
try:
|
||||
metadata = source.lstat()
|
||||
resolved = source.resolve(strict=True)
|
||||
except OSError as exc:
|
||||
raise SandboxError(f"writable artifact mount is unavailable: {source}") from exc
|
||||
if (
|
||||
resolved != source
|
||||
or stat.S_ISLNK(metadata.st_mode)
|
||||
or not (stat.S_ISDIR(metadata.st_mode) or stat.S_ISREG(metadata.st_mode))
|
||||
):
|
||||
raise SandboxError(f"writable artifact mount must be real and non-symlink: {source}")
|
||||
target = PurePosixPath(mount.target)
|
||||
if not target.is_absolute() or ".." in target.parts:
|
||||
raise SandboxError(f"writable artifact mount target must be absolute: {mount.target}")
|
||||
writable.append(ReadOnlyMount(source=source, target=target.as_posix()))
|
||||
|
||||
return _sandbox_command_prefix(
|
||||
bwrap=self.bwrap_bin,
|
||||
clone=clone,
|
||||
|
|
@ -165,6 +378,7 @@ class SandboxSession:
|
|||
temp=self.temp,
|
||||
claude_bin=self.claude_host_bin,
|
||||
mounts=(*self.read_only_mounts, *additional),
|
||||
writable_mounts=writable,
|
||||
read_only_workspace=read_only_workspace,
|
||||
unshare_network=unshare_network,
|
||||
)
|
||||
|
|
@ -305,14 +519,27 @@ def build_sandbox_environment(
|
|||
return env
|
||||
|
||||
|
||||
def build_claude_settings() -> str:
|
||||
"""Inline settings: hooks/plugins are absent and every Bash stays sandboxed."""
|
||||
def build_claude_settings(*, sandbox_enabled: bool = True) -> str:
|
||||
"""Inline settings that keep every Bash sandboxed and pre-approve the tools.
|
||||
|
||||
Deliberately hook-free: headless ``claude -p`` (2.1.247) never dispatches
|
||||
``PreToolUse``, whatever source the hook is declared in — inline
|
||||
``--settings``, a settings file, project/user/local ``--setting-sources``,
|
||||
or a trusted project entry in ``~/.claude.json``. Confinement therefore
|
||||
rests only on mechanisms the CLI honors in this mode: the sandbox policy
|
||||
below, ``--tools``/``--allowedTools``, and the bwrap mounts.
|
||||
"""
|
||||
|
||||
permissions = {
|
||||
"allow": ["Read", "Grep", "Glob", "Bash"],
|
||||
}
|
||||
if sandbox_enabled:
|
||||
permissions["disableBypassPermissionsMode"] = "disable"
|
||||
settings = {
|
||||
"sandbox": {
|
||||
"enabled": True,
|
||||
"failIfUnavailable": True,
|
||||
"autoAllowBashIfSandboxed": True,
|
||||
"enabled": sandbox_enabled,
|
||||
"failIfUnavailable": sandbox_enabled,
|
||||
"autoAllowBashIfSandboxed": sandbox_enabled,
|
||||
"allowUnsandboxedCommands": False,
|
||||
"enableWeakerNestedSandbox": True,
|
||||
"network": {
|
||||
|
|
@ -336,6 +563,7 @@ def build_claude_settings() -> str:
|
|||
SANDBOX_GITNEXUS,
|
||||
SANDBOX_GITNEXUS_SHARED,
|
||||
SANDBOX_GITNEXUS_REGISTRY,
|
||||
SANDBOX_EVIDENCE,
|
||||
],
|
||||
},
|
||||
},
|
||||
|
|
@ -347,13 +575,16 @@ def build_claude_settings() -> str:
|
|||
# allow rule, so pre-approve the proposer's exact tool surface. Bash
|
||||
# is the only writable tool under --bare (it writes the candidate
|
||||
# overlay) and stays sandbox-confined by the sandbox.* policy above.
|
||||
"allow": ["Read", "Grep", "Glob", "Bash"],
|
||||
"disableBypassPermissionsMode": "disable",
|
||||
},
|
||||
"env": {
|
||||
"CLAUDE_CODE_SUBPROCESS_ENV_SCRUB": "1",
|
||||
"CLAUDE_CODE_DONT_INHERIT_ENV": "1",
|
||||
**permissions,
|
||||
},
|
||||
"env": (
|
||||
{
|
||||
"CLAUDE_CODE_SUBPROCESS_ENV_SCRUB": "1",
|
||||
"CLAUDE_CODE_DONT_INHERIT_ENV": "1",
|
||||
}
|
||||
if sandbox_enabled
|
||||
else {"CLAUDE_CODE_DONT_INHERIT_ENV": "1"}
|
||||
),
|
||||
}
|
||||
return json.dumps(settings, sort_keys=True, separators=(",", ":"))
|
||||
|
||||
|
|
@ -437,6 +668,8 @@ def _create_shell_prefix_wrapper(private_root: Path) -> Path:
|
|||
"exec /usr/bin/env -i "
|
||||
f"HOME={SANDBOX_HOME} USER=agent LOGNAME=agent TMPDIR={SANDBOX_TMP} "
|
||||
f"PATH={SANDBOX_PATH} LANG=C.UTF-8 LC_ALL=C.UTF-8 TERM=dumb "
|
||||
f"GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=core.excludesFile GIT_CONFIG_VALUE_0={SANDBOX_GIT_EXCLUDES} "
|
||||
"GITNEXUS_INVOCATION=gitnexus "
|
||||
'/bin/bash -c "$1"\n'
|
||||
)
|
||||
wrapper.chmod(0o500)
|
||||
|
|
@ -461,6 +694,24 @@ def _create_python3_wrapper(private_root: Path) -> Path:
|
|||
return wrapper
|
||||
|
||||
|
||||
def _create_gitnexus_wrapper(private_root: Path) -> Path:
|
||||
"""Expose the already-mounted pinned CLI without npm or network access."""
|
||||
|
||||
wrapper = private_root / "gitnexus"
|
||||
wrapper.write_text(f'#!/bin/bash\nset -eu\nexec {SANDBOX_NODE} {SANDBOX_GITNEXUS}/dist/cli/index.js "$@"\n')
|
||||
wrapper.chmod(0o500)
|
||||
return wrapper
|
||||
|
||||
|
||||
def _create_git_excludes(private_root: Path) -> Path:
|
||||
"""Create the immutable excludes for nested-sandbox mount artifacts."""
|
||||
|
||||
excludes = private_root / "git-excludes"
|
||||
excludes.write_text("\n".join(CLAUDE_SANDBOX_GIT_EXCLUDES) + "\n")
|
||||
excludes.chmod(0o400)
|
||||
return excludes
|
||||
|
||||
|
||||
def _resolve_executable(executable: Path | str | None, default: str) -> Path:
|
||||
raw = os.fspath(executable) if executable is not None else shutil.which(default)
|
||||
if not raw:
|
||||
|
|
@ -499,6 +750,12 @@ def preflight_bubblewrap(bwrap_bin: Path | str | None = None) -> Path:
|
|||
return bwrap
|
||||
|
||||
|
||||
def preflight_unsafe_host() -> Path:
|
||||
"""Return a harmless sentinel for the explicit non-containment backend."""
|
||||
|
||||
return _resolve_executable(None, "env")
|
||||
|
||||
|
||||
def pid_namespace_command(
|
||||
command: Sequence[str],
|
||||
*,
|
||||
|
|
@ -652,6 +909,7 @@ def _sandbox_command_prefix(
|
|||
temp: Path,
|
||||
claude_bin: Path,
|
||||
mounts: Sequence[ReadOnlyMount],
|
||||
writable_mounts: Sequence[ReadOnlyMount] = (),
|
||||
read_only_workspace: bool = False,
|
||||
unshare_network: bool = False,
|
||||
) -> list[str]:
|
||||
|
|
@ -700,10 +958,159 @@ def _sandbox_command_prefix(
|
|||
# not carry it, and overlaying them would fail with EROFS.
|
||||
if PurePosixPath(mount.target).name == DEPENDENCY_MOUNT_BASENAME and (mount.source / VITE_TEMP_DIR).is_dir():
|
||||
args += ["--tmpfs", f"{mount.target}/{VITE_TEMP_DIR}"]
|
||||
for mount in writable_mounts:
|
||||
args += ["--bind", str(mount.source), mount.target]
|
||||
args += ["--chdir", SANDBOX_WORKSPACE, "--"]
|
||||
return args
|
||||
|
||||
|
||||
def _drop_host_workspace_write_bits(
|
||||
root: Path,
|
||||
*,
|
||||
writable: Sequence[Path],
|
||||
) -> list[tuple[Path, int]]:
|
||||
"""Clear write bits on a host-unsafe clone except explicit artifact paths."""
|
||||
|
||||
root = root.expanduser().absolute()
|
||||
try:
|
||||
metadata = root.lstat()
|
||||
resolved = root.resolve(strict=True)
|
||||
except OSError as exc:
|
||||
raise SandboxError(f"host workspace lock root is unavailable: {root}: {exc}") from exc
|
||||
if stat.S_ISLNK(metadata.st_mode) or not stat.S_ISDIR(metadata.st_mode) or resolved != root:
|
||||
raise SandboxError(f"host workspace lock root must be a real directory: {root}")
|
||||
|
||||
allowed: set[Path] = set()
|
||||
for raw in writable:
|
||||
path = raw.expanduser().absolute()
|
||||
try:
|
||||
metadata = path.lstat()
|
||||
resolved = path.resolve(strict=True)
|
||||
path.relative_to(root)
|
||||
except (OSError, ValueError) as exc:
|
||||
raise SandboxError(f"writable host path escapes the workspace: {raw}") from exc
|
||||
if stat.S_ISLNK(metadata.st_mode) or resolved != path:
|
||||
raise SandboxError(f"writable host path must be real and non-symlink: {raw}")
|
||||
allowed.add(path)
|
||||
|
||||
records: list[tuple[Path, int]] = []
|
||||
pending = [root]
|
||||
while pending:
|
||||
current = pending.pop()
|
||||
try:
|
||||
metadata = current.lstat()
|
||||
except OSError as exc:
|
||||
raise SandboxError(f"host workspace lock path is unreadable: {current}: {exc}") from exc
|
||||
records.append((current, stat.S_IMODE(metadata.st_mode)))
|
||||
if stat.S_ISLNK(metadata.st_mode):
|
||||
continue
|
||||
if stat.S_ISDIR(metadata.st_mode):
|
||||
if current in allowed:
|
||||
# Match bwrap: a writable directory bind keeps its children writable.
|
||||
continue
|
||||
try:
|
||||
children = [Path(entry.path) for entry in os.scandir(current)]
|
||||
except OSError as exc:
|
||||
raise SandboxError(f"host workspace lock directory is unreadable: {current}: {exc}") from exc
|
||||
pending.extend(children)
|
||||
os.chmod(current, 0o500)
|
||||
continue
|
||||
if current in allowed:
|
||||
os.chmod(current, stat.S_IMODE(metadata.st_mode) | 0o222)
|
||||
continue
|
||||
if stat.S_ISREG(metadata.st_mode):
|
||||
os.chmod(current, stat.S_IMODE(metadata.st_mode) & ~0o222)
|
||||
return records
|
||||
|
||||
|
||||
def _force_rmtree(path: Path) -> None:
|
||||
"""Delete a tree even when leftover copies inherited 0555 directory modes.
|
||||
|
||||
Host-unsafe review sessions drop write bits on the clone. An agent that
|
||||
``copytree``s those directories into the private TMPDIR leaves a tree
|
||||
``shutil.rmtree`` cannot remove: a non-empty 0555 directory raises
|
||||
PermissionError. Restore owner write bits, then delete.
|
||||
"""
|
||||
|
||||
root = Path(path)
|
||||
if not root.exists():
|
||||
return
|
||||
for dirpath, _dirnames, filenames in os.walk(root, topdown=False, followlinks=False):
|
||||
try:
|
||||
os.chmod(
|
||||
dirpath,
|
||||
stat.S_IMODE(os.lstat(dirpath).st_mode) | 0o700,
|
||||
follow_symlinks=False,
|
||||
)
|
||||
except OSError:
|
||||
# Directory may already be gone or refuse chmod; rmtree still tries.
|
||||
pass
|
||||
for name in filenames:
|
||||
child = os.path.join(dirpath, name)
|
||||
try:
|
||||
metadata = os.lstat(child)
|
||||
except OSError:
|
||||
continue
|
||||
if stat.S_ISLNK(metadata.st_mode):
|
||||
continue
|
||||
try:
|
||||
os.chmod(child, stat.S_IMODE(metadata.st_mode) | 0o200, follow_symlinks=False)
|
||||
except OSError:
|
||||
# File vanished or is immutable; skip and let rmtree report.
|
||||
pass
|
||||
shutil.rmtree(root)
|
||||
|
||||
|
||||
def _restore_host_workspace_modes(records: Sequence[tuple[Path, int]]) -> None:
|
||||
errors: list[str] = []
|
||||
for path, mode in reversed(records):
|
||||
try:
|
||||
os.chmod(path, mode)
|
||||
except FileNotFoundError:
|
||||
continue
|
||||
except OSError as exc:
|
||||
errors.append(f"{path}: {exc}")
|
||||
if errors:
|
||||
raise SandboxError("failed to restore host workspace modes: " + "; ".join(errors[:8]))
|
||||
|
||||
|
||||
@contextmanager
|
||||
def host_workspace_write_boundary(
|
||||
root: Path,
|
||||
*,
|
||||
writable: Sequence[Path] = (),
|
||||
) -> Iterator[None]:
|
||||
"""Best-effort host analog of bwrap ``--ro-bind`` plus one writable artifact.
|
||||
|
||||
This is not a security boundary: a session that can ``chmod`` can undo it.
|
||||
It is enough to stop accidental ``npm install`` / analyze writes from
|
||||
invalidating review evidence on a host-unsafe diagnostic run.
|
||||
"""
|
||||
|
||||
records = _drop_host_workspace_write_bits(root, writable=writable)
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
_restore_host_workspace_modes(records)
|
||||
|
||||
|
||||
@contextmanager
|
||||
def sandbox_workspace_write_boundary(
|
||||
sandbox: Any,
|
||||
*,
|
||||
read_only_workspace: bool,
|
||||
writable: Sequence[Path] = (),
|
||||
) -> Iterator[None]:
|
||||
"""Apply the host-unsafe write lock only when bwrap is not the backend."""
|
||||
|
||||
if getattr(sandbox, "backend", "bwrap") != "host-unsafe" or not read_only_workspace:
|
||||
yield
|
||||
return
|
||||
clone = Path(sandbox.clone)
|
||||
with host_workspace_write_boundary(clone, writable=writable):
|
||||
yield
|
||||
|
||||
|
||||
@contextmanager
|
||||
def prepare_sandbox(
|
||||
*,
|
||||
|
|
@ -712,17 +1119,21 @@ def prepare_sandbox(
|
|||
bwrap_bin: Path | str | None = None,
|
||||
read_only_mounts: Sequence[ReadOnlyMount] = (),
|
||||
preflight: bool = True,
|
||||
backend: str = "bwrap",
|
||||
) -> Iterator[SandboxSession]:
|
||||
"""Create private host backing dirs and one immutable Bubblewrap command."""
|
||||
"""Create private host backing dirs and one virtualized command."""
|
||||
|
||||
# Validate the lexical path before resolving it. Resolving first would
|
||||
# erase the evidence that the caller supplied a symlinked clone root.
|
||||
clone = _real_directory(clone, label="sandbox clone")
|
||||
if backend not in ("bwrap", "host-unsafe"):
|
||||
raise SandboxError(f"unknown sandbox backend: {backend}")
|
||||
if preflight:
|
||||
bwrap = preflight_bubblewrap(bwrap_bin)
|
||||
require_claude_sandbox_helpers()
|
||||
bwrap = preflight_bubblewrap(bwrap_bin) if backend == "bwrap" else preflight_unsafe_host()
|
||||
if backend == "bwrap":
|
||||
require_claude_sandbox_helpers()
|
||||
else:
|
||||
bwrap = _resolve_executable(bwrap_bin, "bwrap")
|
||||
bwrap = _resolve_executable(bwrap_bin, "bwrap") if backend == "bwrap" else preflight_unsafe_host()
|
||||
claude = _resolve_executable(claude_bin, "claude")
|
||||
private_root = Path(tempfile.mkdtemp(prefix="wfbench-sandbox-"))
|
||||
private_root.chmod(0o700)
|
||||
|
|
@ -733,6 +1144,8 @@ def prepare_sandbox(
|
|||
directory.chmod(0o700)
|
||||
shell_prefix = _create_shell_prefix_wrapper(private_root)
|
||||
python3_wrapper = _create_python3_wrapper(private_root)
|
||||
gitnexus_wrapper = _create_gitnexus_wrapper(private_root)
|
||||
git_excludes = _create_git_excludes(private_root)
|
||||
# Claude may discover user-level skills below HOME. Keep the rest of HOME
|
||||
# writable for normal CLI state, but overlay an immutable empty skills root
|
||||
# so a model cannot shadow the evaluated repository/plugin skill by name.
|
||||
|
|
@ -744,16 +1157,22 @@ def prepare_sandbox(
|
|||
ReadOnlyMount(source=user_skills, target=SANDBOX_USER_SKILLS),
|
||||
ReadOnlyMount(source=shell_prefix, target=SANDBOX_SHELL_PREFIX),
|
||||
ReadOnlyMount(source=python3_wrapper, target=SANDBOX_PYTHON3),
|
||||
ReadOnlyMount(source=gitnexus_wrapper, target=SANDBOX_GITNEXUS_CLI),
|
||||
ReadOnlyMount(source=git_excludes, target=SANDBOX_GIT_EXCLUDES),
|
||||
)
|
||||
primary: BaseException | None = None
|
||||
try:
|
||||
command_prefix = _sandbox_command_prefix(
|
||||
bwrap=bwrap,
|
||||
clone=clone,
|
||||
home=home,
|
||||
temp=temp,
|
||||
claude_bin=claude,
|
||||
mounts=protected_mounts,
|
||||
command_prefix = (
|
||||
_sandbox_command_prefix(
|
||||
bwrap=bwrap,
|
||||
clone=clone,
|
||||
home=home,
|
||||
temp=temp,
|
||||
claude_bin=claude,
|
||||
mounts=protected_mounts,
|
||||
)
|
||||
if backend == "bwrap"
|
||||
else []
|
||||
)
|
||||
yield SandboxSession(
|
||||
private_root=private_root,
|
||||
|
|
@ -761,6 +1180,7 @@ def prepare_sandbox(
|
|||
home=home,
|
||||
temp=temp,
|
||||
bwrap_bin=bwrap,
|
||||
backend=backend,
|
||||
claude_host_bin=claude,
|
||||
command_prefix=command_prefix,
|
||||
read_only_mounts=protected_mounts,
|
||||
|
|
@ -770,7 +1190,7 @@ def prepare_sandbox(
|
|||
raise
|
||||
finally:
|
||||
try:
|
||||
shutil.rmtree(private_root)
|
||||
_force_rmtree(private_root)
|
||||
except OSError as cleanup:
|
||||
if primary is None:
|
||||
raise
|
||||
|
|
|
|||
66
eval/workflow_bench/review_cases/manifest.json
Normal file
66
eval/workflow_bench/review_cases/manifest.json
Normal file
|
|
@ -0,0 +1,66 @@
|
|||
{
|
||||
"schema_version": 1,
|
||||
"corpus_version": "2026-09-04.3",
|
||||
"cases": [
|
||||
{
|
||||
"id": "review-pr-2718-defect",
|
||||
"patch": "pr-2718.patch",
|
||||
"pr": "https://github.com/abhigyanpatwari/GitNexus/pull/2718",
|
||||
"base_sha": "ff86ccf1e79cd7e4175da437ae8aeaf67b64aaa1",
|
||||
"head_sha": "cfd2434c6ca0e303ac40a896db798f30505390d5",
|
||||
"patch_sha256": "14ca0d5659fb7aa529c32543f194deabc8e25d592037dad1d70ef8f6ab906391",
|
||||
"human_verification_commit": "cebe66b33509021de8d09083c461501cb49a460b",
|
||||
"label_source": "exact-head tri-review 4799215165"
|
||||
},
|
||||
{
|
||||
"id": "review-pr-2794-defect",
|
||||
"patch": "pr-2794.patch",
|
||||
"pr": "https://github.com/abhigyanpatwari/GitNexus/pull/2794",
|
||||
"base_sha": "911151e2304f298a995fcc69c738ad2c6db9393a",
|
||||
"head_sha": "48afb7480778ef2f5e0be455498ace889d03edf9",
|
||||
"patch_sha256": "e2271ede4b4d65993abde16012061191a094847aadf0431a2dec908a59b6af93",
|
||||
"human_verification_commit": "0015b0d64537c9ac97fda5ed094c99059d596cfc",
|
||||
"label_source": "exact-head tri-review 4837878304"
|
||||
},
|
||||
{
|
||||
"id": "review-pr-2108-defect",
|
||||
"patch": "pr-2108.patch",
|
||||
"pr": "https://github.com/abhigyanpatwari/GitNexus/pull/2108",
|
||||
"base_sha": "3a4247ec36b5ad86b1123d3bbce8183a643f7434",
|
||||
"head_sha": "fdabb8a1fa441b8fc7f3471121a9fa5a9d885376",
|
||||
"patch_sha256": "12928510253fe079b179fc4658334b732de8907471ac92a6ed8061fb78aedbf2",
|
||||
"human_verification_commit": "40dff64992ac1e72c7c042b2f508dbdc433f74dd",
|
||||
"label_source": "exact-head tri-review 4456060714"
|
||||
},
|
||||
{
|
||||
"id": "review-pr-2258-defect",
|
||||
"patch": "pr-2258.patch",
|
||||
"pr": "https://github.com/abhigyanpatwari/GitNexus/pull/2258",
|
||||
"base_sha": "78b4077d8acc86f1b0c32e41012174d484e81f12",
|
||||
"head_sha": "c93ca8ce73db7e4f0b083226a166dd9841288140",
|
||||
"patch_sha256": "8e0eb214bd6151d1060cf2c47df7735fe5c57398e08bee9306b7b702abeca843",
|
||||
"human_verification_commit": "00e52fa7fbae21668ae818a41f786d529c4ca390",
|
||||
"label_source": "exact-head tri-review 4538570459"
|
||||
},
|
||||
{
|
||||
"id": "review-pr-2258-clean",
|
||||
"patch": "pr-2258b.patch",
|
||||
"pr": "https://github.com/abhigyanpatwari/GitNexus/pull/2258",
|
||||
"base_sha": "78b4077d8acc86f1b0c32e41012174d484e81f12",
|
||||
"head_sha": "00e52fa7fbae21668ae818a41f786d529c4ca390",
|
||||
"patch_sha256": "4323ec2ff01f612f18a493e055c70eb4307e9c1859b97f5b60da07b177d3a238",
|
||||
"human_verification_commit": "00e52fa7fbae21668ae818a41f786d529c4ca390",
|
||||
"label_source": "production-ready tri-review confirmation"
|
||||
},
|
||||
{
|
||||
"id": "review-pr-2773-clean",
|
||||
"patch": "pr-2773.patch",
|
||||
"pr": "https://github.com/abhigyanpatwari/GitNexus/pull/2773",
|
||||
"base_sha": "84f584449de02376a8ffc096dceac2e8f732cab5",
|
||||
"head_sha": "f584f83bcb93c91752d91144b251f98d39027180",
|
||||
"patch_sha256": "1c0327d4d9725428c1b49ab3586abefdbc084147735b120afd43312e8e48dd91",
|
||||
"human_verification_commit": "f584f83bcb93c91752d91144b251f98d39027180",
|
||||
"label_source": "review fix confirmation 3694979051"
|
||||
}
|
||||
]
|
||||
}
|
||||
845
eval/workflow_bench/review_cases/pr-2108.patch
Normal file
845
eval/workflow_bench/review_cases/pr-2108.patch
Normal file
|
|
@ -0,0 +1,845 @@
|
|||
diff --git a/Dockerfile.cli b/Dockerfile.cli
|
||||
index 3275c8f7e..cfba9bac4 100644
|
||||
--- a/Dockerfile.cli
|
||||
+++ b/Dockerfile.cli
|
||||
@@ -67,6 +67,28 @@ COPY --from=builder --chown=node:node /app/gitnexus/vendor ./gitnexus/vendor
|
||||
# unreachable from $PATH.
|
||||
RUN ln -s /app/gitnexus/dist/cli/index.js /usr/local/bin/gitnexus
|
||||
|
||||
+# Bake the LadybugDB FTS extension into the image so BM25 keyword search works
|
||||
+# at runtime. The server runs the default `load-only` extension policy (the read
|
||||
+# pool pins `{ policy: 'load-only' }`), so a runtime `LOAD EXTENSION fts` never
|
||||
+# INSTALLs — the extension must already exist in the runtime user's HOME
|
||||
+# extension dir, or every keyword search silently degrades (no FTS indexes are
|
||||
+# written and ranking falls back to vector-only with only a `warning` field).
|
||||
+# Run the installer as the `node` user with the SAME HOME the server runs under,
|
||||
+# so `INSTALL fts` materializes the extension under `$HOME/.lbdb/extension` where
|
||||
+# the runtime `LOAD` resolves it offline. `ENV HOME` is pinned because Docker
|
||||
+# does not derive HOME from `USER`, so without it build-install and runtime-load
|
||||
+# would resolve different paths. Requires network egress for the one-time
|
||||
+# INSTALL; the build fails loudly if it cannot fetch the extension. The DB-size
|
||||
+# default comes from GITNEXUS_LBUG_MAX_DB_SIZE (single source of truth, matches
|
||||
+# the runtime) — it only sizes the throwaway scratch DB used to run INSTALL.
|
||||
+# The second `--verify-only` step re-LOADs the extension in a FRESH process
|
||||
+# under the same HOME, so a HOME/extension-dir mismatch fails the build here
|
||||
+# rather than silently degrading keyword search to vector-only at runtime.
|
||||
+ENV HOME=/home/node \
|
||||
+ GITNEXUS_LBUG_MAX_DB_SIZE=17179869184
|
||||
+RUN su node -s /bin/sh -c "HOME=/home/node node /app/gitnexus/scripts/install-duckdb-extension.mjs fts" \
|
||||
+ && su node -s /bin/sh -c "HOME=/home/node node /app/gitnexus/scripts/install-duckdb-extension.mjs fts --verify-only"
|
||||
+
|
||||
USER node
|
||||
|
||||
# The web UI defaults to http://localhost:4747 - keep that contract.
|
||||
diff --git a/gitnexus/scripts/bench/fts-evict-reload-rss.mjs b/gitnexus/scripts/bench/fts-evict-reload-rss.mjs
|
||||
new file mode 100644
|
||||
index 000000000..6240e16b7
|
||||
--- /dev/null
|
||||
+++ b/gitnexus/scripts/bench/fts-evict-reload-rss.mjs
|
||||
@@ -0,0 +1,421 @@
|
||||
+#!/usr/bin/env node
|
||||
+// FTS evict→reload RSS repro (gitnexus-enterprise PR #222 / local U3).
|
||||
+//
|
||||
+// Settles ONE empirical question that no static read can answer: when a
|
||||
+// LadybugDB database that has `LOAD EXTENSION fts` applied is closed and a
|
||||
+// fresh one is opened + re-LOADed (the pool's evict→reload cycle), does the
|
||||
+// native FTS arena get reclaimed by `db.close()` — or is it stranded, so RSS
|
||||
+// climbs without bound over a long-lived MCP `serve` session?
|
||||
+//
|
||||
+// • PLATEAU across cycles → db.close() reclaims the FTS arena; the OSS pool's
|
||||
+// footprint is bounded by MAX_POOL_SIZE (~5 live arenas). No unbounded leak;
|
||||
+// the #222 worker-isolation rewrite (plan U4) is NOT justified for OSS.
|
||||
+// • MONOTONIC CLIMB → the FTS arena is stranded per reopen; the user's
|
||||
+// hypothesis holds and U4 (route FTS reads through a reclaimable worker) is
|
||||
+// justified.
|
||||
+//
|
||||
+// SCOPE OF THE VERDICT (read before citing it). A per-reload FTS-arena leak
|
||||
+// would be PROPORTIONAL to the index size. A small fixture therefore produces a
|
||||
+// small per-cycle increment that an absolute threshold can read as PLATEAU even
|
||||
+// when a production-scale graph would leak visibly. So:
|
||||
+// - `--rows` controls fixture size; run it LARGE (tens of thousands) before
|
||||
+// concluding "no leak". The default is deliberately not tiny.
|
||||
+// - The CLIMB gate combines a per-cycle slope with BOTH an absolute and a
|
||||
+// per-row-relative delta floor, so the sensitivity scales with fixture size.
|
||||
+// - The PLATEAU verdict is only valid for the corpus size it was run at; the
|
||||
+// output states that size. The production-faithful confirmation is a
|
||||
+// `--via-pool` run against a real large analyzed repo over a long session.
|
||||
+//
|
||||
+// Two modes:
|
||||
+// (default) NATIVE — reproduces the native sequence doInitLbug()+closeOne()
|
||||
+// perform (open Database → new Connection → LOAD EXTENSION fts →
|
||||
+// QUERY_FTS_INDEX → close), against K self-built FTS fixtures, with no
|
||||
+// gitnexus build required. `--no-await-close` mirrors the pool's
|
||||
+// fire-and-forget close instead of awaiting (the production close shape).
|
||||
+// --via-pool <lbugPath> — drives the REAL gitnexus pool from compiled dist
|
||||
+// (initLbug → executeParameterized → closeLbug) against an existing analyzed
|
||||
+// repo, exercising the production path + the GITNEXUS_POOL_RSS_TRACE
|
||||
+// instrumentation. Probes ALL FTS indexes the repo has. Forces an explicit
|
||||
+// close+reinit each cycle. Run `node scripts/build.js` first so the dist
|
||||
+// reflects the current pool-adapter (incl. the RSS trace).
|
||||
+//
|
||||
+// Run with --expose-gc so RSS excludes V8-heap noise:
|
||||
+// node --expose-gc gitnexus/scripts/bench/fts-evict-reload-rss.mjs
|
||||
+// node --expose-gc gitnexus/scripts/bench/fts-evict-reload-rss.mjs --rows 40000 --cycles 30
|
||||
+// GITNEXUS_POOL_RSS_TRACE=1 node --expose-gc \
|
||||
+// gitnexus/scripts/bench/fts-evict-reload-rss.mjs --via-pool /path/to/repo/.gitnexus/lbug
|
||||
+//
|
||||
+// Flags by mode: --rows/--repos/--read-write/--no-await-close apply to NATIVE
|
||||
+// only; --cycles applies to both. VIA-POOL warns when a NATIVE-only flag is set.
|
||||
+//
|
||||
+// Memory benches are noisy. Default is 24 cycles; trust the TREND (slope /
|
||||
+// first-third vs last-third), never a single delta. A flat trend at a LARGE
|
||||
+// fixture is a real NEGATIVE result (no unbounded leak), not a failed run.
|
||||
+
|
||||
+import { createRequire } from 'node:module';
|
||||
+import os from 'node:os';
|
||||
+import path from 'node:path';
|
||||
+import fs from 'node:fs';
|
||||
+
|
||||
+const require = createRequire(import.meta.url);
|
||||
+const lbugModule = require('@ladybugdb/core');
|
||||
+const lbug = lbugModule.default ?? lbugModule;
|
||||
+
|
||||
+const LBUG_MAX_DB_SIZE = 16 * 1024 * 1024 * 1024;
|
||||
+
|
||||
+// ── args ──────────────────────────────────────────────────────────────────
|
||||
+function argVal(flag, dflt) {
|
||||
+ const i = process.argv.indexOf(flag);
|
||||
+ return i >= 0 && process.argv[i + 1] ? process.argv[i + 1] : dflt;
|
||||
+}
|
||||
+const CYCLES = Math.max(6, parseInt(argVal('--cycles', '24'), 10) || 24);
|
||||
+const REPOS = Math.max(1, parseInt(argVal('--repos', '6'), 10) || 6); // >5 mirrors LRU thrash
|
||||
+// Fixture size. Default is large enough that a size-proportional leak would be
|
||||
+// visible across cycles; raise it further before trusting a PLATEAU verdict.
|
||||
+const ROWS = Math.max(100, parseInt(argVal('--rows', '8000'), 10) || 8000);
|
||||
+const VIA_POOL = argVal('--via-pool', null);
|
||||
+const READONLY = !process.argv.includes('--read-write');
|
||||
+const AWAIT_CLOSE = !process.argv.includes('--no-await-close');
|
||||
+
|
||||
+if (VIA_POOL) {
|
||||
+ // These flags are consumed only by NATIVE mode; warn rather than ignore
|
||||
+ // silently so a VIA-POOL run is not misread as honoring them.
|
||||
+ const ignored = ['--rows', '--repos', '--read-write', '--no-await-close'].filter((f) =>
|
||||
+ process.argv.includes(f),
|
||||
+ );
|
||||
+ if (ignored.length) {
|
||||
+ console.error(
|
||||
+ `[fts-rss] NOTE: ${ignored.join(', ')} apply to NATIVE mode only; ignored in --via-pool.`,
|
||||
+ );
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+if (typeof global.gc !== 'function') {
|
||||
+ console.error(
|
||||
+ '[fts-rss] WARNING: run with --expose-gc for clean RSS samples ' +
|
||||
+ '(`node --expose-gc <thisfile>`). Continuing without forced GC — results are noisier.',
|
||||
+ );
|
||||
+}
|
||||
+
|
||||
+const gc = () => {
|
||||
+ if (typeof global.gc === 'function') {
|
||||
+ global.gc();
|
||||
+ global.gc();
|
||||
+ }
|
||||
+};
|
||||
+const rssMb = () => Math.round(process.memoryUsage().rss / (1024 * 1024));
|
||||
+const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
|
||||
+
|
||||
+// ── fixture: a minimal FTS-bearing .lbug ────────────────────────────────────
|
||||
+const WORDS = [
|
||||
+ 'login auth session token user password validate verify credential',
|
||||
+ 'parse tree syntax node grammar lexer token ast traversal visitor',
|
||||
+ 'graph query cypher match relation node edge pattern aggregate index',
|
||||
+ 'memory pool buffer arena allocate reclaim evict cache resident heap',
|
||||
+ 'search rank score bm25 fts index stem porter keyword document corpus',
|
||||
+ 'worker fork process spawn kill reclaim isolate native binding addon',
|
||||
+];
|
||||
+
|
||||
+function buildFixture(dir) {
|
||||
+ fs.mkdirSync(dir, { recursive: true });
|
||||
+ const dbPath = path.join(dir, 'fixture.lbug');
|
||||
+ const db = new lbug.Database(dbPath, 0, false, false, LBUG_MAX_DB_SIZE);
|
||||
+ const conn = new lbug.Connection(db);
|
||||
+ return (async () => {
|
||||
+ await conn.query('LOAD EXTENSION fts');
|
||||
+ await conn.query(
|
||||
+ 'CREATE NODE TABLE Doc(id STRING, name STRING, content STRING, PRIMARY KEY(id))',
|
||||
+ );
|
||||
+ // Batch-insert via UNWIND so large fixtures (`--rows`) build in seconds
|
||||
+ // instead of one round-trip per row. The fixture size drives the per-arena
|
||||
+ // FTS allocation, which is what makes a size-proportional leak observable.
|
||||
+ const rows = [];
|
||||
+ for (let i = 0; i < ROWS; i++) {
|
||||
+ const w = WORDS[i % WORDS.length];
|
||||
+ const name = `sym_${i}`;
|
||||
+ const content = `${w} ${name} block number ${i} ${WORDS[(i + 3) % WORDS.length]}`;
|
||||
+ rows.push({ id: `doc:${i}`, name, content });
|
||||
+ }
|
||||
+ const INSERT_CHUNK = 2000;
|
||||
+ for (let i = 0; i < rows.length; i += INSERT_CHUNK) {
|
||||
+ const chunk = rows.slice(i, i + INSERT_CHUNK);
|
||||
+ const stmt = await conn.prepare(
|
||||
+ 'UNWIND $rows AS r CREATE (:Doc {id: r.id, name: r.name, content: r.content})',
|
||||
+ );
|
||||
+ await conn.execute(stmt, { rows: chunk });
|
||||
+ }
|
||||
+ await conn.query(
|
||||
+ "CALL CREATE_FTS_INDEX('Doc', 'doc_fts', ['name', 'content'], stemmer := 'porter')",
|
||||
+ );
|
||||
+ await conn.close();
|
||||
+ await db.close();
|
||||
+ return dbPath;
|
||||
+ })();
|
||||
+}
|
||||
+
|
||||
+const QUERIES = ['login token', 'parse node', 'memory arena', 'search index', 'worker reclaim'];
|
||||
+
|
||||
+// ── NATIVE mode ─────────────────────────────────────────────────────────────
|
||||
+async function runNative() {
|
||||
+ const root = fs.mkdtempSync(path.join(os.tmpdir(), 'fts-rss-'));
|
||||
+ console.error(
|
||||
+ `[fts-rss] NATIVE: ${REPOS} fixtures × ${ROWS} rows × ${CYCLES} cycles ` +
|
||||
+ `(readOnly=${READONLY}, awaitClose=${AWAIT_CLOSE})`,
|
||||
+ );
|
||||
+ console.error(`[fts-rss] building ${REPOS} FTS fixture(s) under ${root} …`);
|
||||
+
|
||||
+ const srcDb = await buildFixture(path.join(root, 'src'));
|
||||
+ const repoPaths = [];
|
||||
+ for (let k = 0; k < REPOS; k++) {
|
||||
+ const dst = path.join(root, `repo-${k}`);
|
||||
+ fs.cpSync(path.dirname(srcDb), dst, { recursive: true });
|
||||
+ repoPaths.push(path.join(dst, 'fixture.lbug'));
|
||||
+ }
|
||||
+
|
||||
+ // Mirror the pool's evict→reload: each visit opens a FRESH Database, makes a
|
||||
+ // Connection, LOADs fts, runs an FTS query, then closes — no caching, so every
|
||||
+ // visit is a reload. K>5 amplifies the LRU-thrash signal the pool would see.
|
||||
+ const series = [];
|
||||
+ gc();
|
||||
+ await sleep(50);
|
||||
+ const baseline = rssMb();
|
||||
+ console.error(`[fts-rss] baseline RSS=${baseline}MB`);
|
||||
+
|
||||
+ for (let cycle = 0; cycle < CYCLES; cycle++) {
|
||||
+ for (let k = 0; k < REPOS; k++) {
|
||||
+ const db = new lbug.Database(repoPaths[k], 0, false, READONLY, LBUG_MAX_DB_SIZE);
|
||||
+ const conn = new lbug.Connection(db);
|
||||
+ try {
|
||||
+ await conn.query('LOAD EXTENSION fts'); // the per-reload re-LOAD under test
|
||||
+ const q = QUERIES[(cycle + k) % QUERIES.length];
|
||||
+ const res = await conn.query(
|
||||
+ `CALL QUERY_FTS_INDEX('Doc', 'doc_fts', '${q}') RETURN node.id AS id, score ORDER BY score DESC LIMIT 20`,
|
||||
+ );
|
||||
+ // Drain so the query actually materializes results.
|
||||
+ if (res && typeof res.getAll === 'function') await res.getAll();
|
||||
+ } catch (e) {
|
||||
+ console.error(`[fts-rss] query error (cycle ${cycle}, repo ${k}): ${e?.message || e}`);
|
||||
+ } finally {
|
||||
+ // AWAIT_CLOSE (default) is the best case for reclamation. --no-await-close
|
||||
+ // mirrors the pool's fire-and-forget close (closeOne: db.close().catch())
|
||||
+ // so a leak that only manifests without awaiting is not hidden.
|
||||
+ if (AWAIT_CLOSE) {
|
||||
+ try {
|
||||
+ await conn.close();
|
||||
+ await db.close();
|
||||
+ } catch {
|
||||
+ /* ignore */
|
||||
+ }
|
||||
+ } else {
|
||||
+ conn.close().catch(() => {});
|
||||
+ db.close().catch(() => {});
|
||||
+ }
|
||||
+ }
|
||||
+ }
|
||||
+ gc();
|
||||
+ // Longer settle when not awaiting close, so fire-and-forget native teardown
|
||||
+ // has a chance to complete before the RSS sample (avoids a false PLATEAU).
|
||||
+ await sleep(AWAIT_CLOSE ? 20 : 200);
|
||||
+ const rss = rssMb();
|
||||
+ series.push(rss);
|
||||
+ console.error(`[fts-rss] cycle ${String(cycle + 1).padStart(3)}/${CYCLES} rssMB=${rss}`);
|
||||
+ }
|
||||
+
|
||||
+ fs.rmSync(root, { recursive: true, force: true });
|
||||
+ return { baseline, series, corpus: `${REPOS}×${ROWS} rows, native, awaitClose=${AWAIT_CLOSE}` };
|
||||
+}
|
||||
+
|
||||
+// ── VIA-POOL mode (real gitnexus pool from compiled dist) ───────────────────
|
||||
+async function runViaPool(lbugPath) {
|
||||
+ if (!fs.existsSync(lbugPath)) {
|
||||
+ console.error(`[fts-rss] --via-pool path not found: ${lbugPath}`);
|
||||
+ process.exit(2);
|
||||
+ }
|
||||
+ // Compiled dist is required (the pool pulls the native addon + many modules).
|
||||
+ const distUrl = new URL('../../dist/core/lbug/pool-adapter.js', import.meta.url);
|
||||
+ let pool;
|
||||
+ try {
|
||||
+ pool = await import(distUrl.href);
|
||||
+ } catch (e) {
|
||||
+ console.error(
|
||||
+ `[fts-rss] could not import compiled pool-adapter (${e?.message}). ` +
|
||||
+ `Run \`node scripts/build.js\` first, or use NATIVE mode.`,
|
||||
+ );
|
||||
+ process.exit(2);
|
||||
+ }
|
||||
+ const { initLbug, executeParameterized, closeLbug } = pool;
|
||||
+ console.error(
|
||||
+ `[fts-rss] VIA-POOL on ${lbugPath} × ${CYCLES} cycles ` +
|
||||
+ `(explicit closeLbug+initLbug per cycle = forced evict→reload)`,
|
||||
+ );
|
||||
+
|
||||
+ // Probe ALL FTS indexes the analyzed graph carries (mirrors fts-schema.ts
|
||||
+ // FTS_INDEXES) so the per-cycle FTS arena load matches production, not a
|
||||
+ // 2-of-5 subset that would understate it.
|
||||
+ const FTS_INDEXES = [
|
||||
+ { table: 'File', indexName: 'file_fts' },
|
||||
+ { table: 'Function', indexName: 'function_fts' },
|
||||
+ { table: 'Class', indexName: 'class_fts' },
|
||||
+ { table: 'Method', indexName: 'method_fts' },
|
||||
+ { table: 'Interface', indexName: 'interface_fts' },
|
||||
+ ];
|
||||
+
|
||||
+ const series = [];
|
||||
+ gc();
|
||||
+ const baseline = rssMb();
|
||||
+ console.error(`[fts-rss] baseline RSS=${baseline}MB`);
|
||||
+
|
||||
+ for (let cycle = 0; cycle < CYCLES; cycle++) {
|
||||
+ try {
|
||||
+ await initLbug(lbugPath, lbugPath);
|
||||
+ const q = QUERIES[cycle % QUERIES.length];
|
||||
+ for (const { table, indexName } of FTS_INDEXES) {
|
||||
+ await executeParameterized(
|
||||
+ lbugPath,
|
||||
+ `CALL QUERY_FTS_INDEX('${table}', '${indexName}', $q) RETURN node.id AS id, score ORDER BY score DESC LIMIT 20`,
|
||||
+ { q },
|
||||
+ ).catch(() => []); // index may not exist for this graph — that's fine
|
||||
+ }
|
||||
+ await closeLbug(lbugPath); // force eviction → next cycle reopens + re-LOADs fts
|
||||
+ } catch (e) {
|
||||
+ console.error(`[fts-rss] pool cycle ${cycle} error: ${e?.message || e}`);
|
||||
+ }
|
||||
+ gc();
|
||||
+ // closeLbug fires a fire-and-forget native close (pool closeOne:
|
||||
+ // db.close().catch()), so settle longer than NATIVE's awaited close to let
|
||||
+ // native teardown finish before sampling — else a real leak reads PLATEAU.
|
||||
+ await sleep(200);
|
||||
+ const rss = rssMb();
|
||||
+ series.push(rss);
|
||||
+ console.error(`[fts-rss] cycle ${String(cycle + 1).padStart(3)}/${CYCLES} rssMB=${rss}`);
|
||||
+ }
|
||||
+ await closeLbug().catch(() => {});
|
||||
+ return { baseline, series, corpus: `via-pool ${path.basename(path.dirname(lbugPath))}` };
|
||||
+}
|
||||
+
|
||||
+// ── verdict ─────────────────────────────────────────────────────────────────
|
||||
+function median(xs) {
|
||||
+ const s = [...xs].sort((a, b) => a - b);
|
||||
+ const m = Math.floor(s.length / 2);
|
||||
+ return s.length % 2 ? s[m] : Math.round((s[m - 1] + s[m]) / 2);
|
||||
+}
|
||||
+function slopeMbPerCycle(series) {
|
||||
+ // Least-squares slope of rss vs cycle index.
|
||||
+ const n = series.length;
|
||||
+ const xs = series.map((_, i) => i);
|
||||
+ const xMean = xs.reduce((a, b) => a + b, 0) / n;
|
||||
+ const yMean = series.reduce((a, b) => a + b, 0) / n;
|
||||
+ let num = 0,
|
||||
+ den = 0;
|
||||
+ for (let i = 0; i < n; i++) {
|
||||
+ num += (xs[i] - xMean) * (series[i] - yMean);
|
||||
+ den += (xs[i] - xMean) ** 2;
|
||||
+ }
|
||||
+ return den === 0 ? 0 : num / den;
|
||||
+}
|
||||
+
|
||||
+function verdict({ baseline, series, corpus }) {
|
||||
+ const third = Math.max(1, Math.floor(series.length / 3));
|
||||
+ const firstMed = median(series.slice(0, third));
|
||||
+ const lastMed = median(series.slice(-third));
|
||||
+ const delta = lastMed - firstMed;
|
||||
+ const slope = slopeMbPerCycle(series);
|
||||
+ const peak = Math.max(...series);
|
||||
+
|
||||
+ // The discriminant between a real leak and allocator warmup is SLOPE
|
||||
+ // DECELERATION, not total delta. Both a leak and a warmup-to-plateau climb;
|
||||
+ // they differ in whether the per-cycle increment is SUSTAINED or DECAYS:
|
||||
+ // - true per-reload leak (stranded FTS arena): RSS rises ~linearly, so the
|
||||
+ // second-half slope ≈ the first-half slope (the increment does not decay).
|
||||
+ // - allocator working-set warmup (native free-pool growing to the working
|
||||
+ // set, freed pages retained-then-reused): RSS rises then flattens, so the
|
||||
+ // second-half slope is a small FRACTION of the first-half slope. Larger
|
||||
+ // fixtures warm up over MORE cycles, which a fixed absolute-delta gate
|
||||
+ // misreads as a leak — the slope ratio is scale-invariant and does not.
|
||||
+ const half = Math.max(1, Math.floor(series.length / 2));
|
||||
+ const firstHalfSlope = slopeMbPerCycle(series.slice(0, half));
|
||||
+ const secondHalfSlope = slopeMbPerCycle(series.slice(-half));
|
||||
+
|
||||
+ // Detect a STEP DISCONTINUITY — a single cycle-to-cycle jump far larger than
|
||||
+ // the typical per-cycle delta. A one-time allocator/arena reservation jump
|
||||
+ // (then flat) is NOT a per-reload leak, but it inflates the second-half slope
|
||||
+ // and would fool a pure slope test; it also signals a noisy run.
|
||||
+ const deltas = series.slice(1).map((v, i) => v - series[i]);
|
||||
+ const absDeltas = deltas.map(Math.abs).sort((a, b) => a - b);
|
||||
+ const medAbsDelta = absDeltas.length ? absDeltas[Math.floor(absDeltas.length / 2)] : 0;
|
||||
+ const maxJump = deltas.length ? Math.max(...deltas) : 0;
|
||||
+ const stepDiscontinuity = maxJump > Math.max(30, 5 * Math.max(medAbsDelta, 1));
|
||||
+
|
||||
+ // The discriminant between a real leak and allocator warmup is SLOPE
|
||||
+ // DECELERATION, not total delta. A true per-reload leak (stranded FTS arena)
|
||||
+ // rises ~linearly: the second-half slope stays ≈ the first-half slope. An
|
||||
+ // allocator working-set warmup rises then flattens: the second-half slope is
|
||||
+ // a small FRACTION of the first-half. Larger fixtures warm up over MORE
|
||||
+ // cycles, which a fixed absolute-delta gate misreads as a leak — the slope
|
||||
+ // ratio is scale-invariant. Below SUSTAIN_FLOOR (~0.5 MB/cycle) the tail is
|
||||
+ // effectively flat (noise).
|
||||
+ const SUSTAIN_FLOOR = 0.5;
|
||||
+ const decelRatio = secondHalfSlope / Math.max(firstHalfSlope, 1e-9);
|
||||
+ let label;
|
||||
+ if (stepDiscontinuity) {
|
||||
+ // A discrete jump (then flat) is not linear accumulation, but the run is
|
||||
+ // noisy — don't claim a clean result either way.
|
||||
+ label = 'INCONCLUSIVE';
|
||||
+ } else if (secondHalfSlope < SUSTAIN_FLOOR) {
|
||||
+ label = 'PLATEAU';
|
||||
+ } else if (decelRatio >= 0.6) {
|
||||
+ label = 'CLIMB';
|
||||
+ } else {
|
||||
+ // Tail slope above the flat floor but clearly decelerating — converging,
|
||||
+ // but not yet flat. Honest answer at this corpus is "not resolved".
|
||||
+ label = 'INCONCLUSIVE';
|
||||
+ }
|
||||
+
|
||||
+ console.log('\n==================== FTS evict→reload RSS verdict ====================');
|
||||
+ console.log(`corpus: ${corpus}`);
|
||||
+ console.log(`samples (MB): ${series.join(' ')}`);
|
||||
+ console.log(
|
||||
+ `baseline=${baseline} firstThirdMed=${firstMed} lastThirdMed=${lastMed} delta=${delta}MB ` +
|
||||
+ `peak=${peak} overallSlope=${slope.toFixed(2)} firstHalfSlope=${firstHalfSlope.toFixed(2)} ` +
|
||||
+ `secondHalfSlope=${secondHalfSlope.toFixed(2)}MB/cycle maxJump=${maxJump}MB step=${stepDiscontinuity} cycles=${series.length}`,
|
||||
+ );
|
||||
+ if (label === 'CLIMB') {
|
||||
+ console.log(
|
||||
+ 'VERDICT: CLIMB — the per-cycle increment is SUSTAINED (second-half slope ≈ first-half),\n' +
|
||||
+ ' i.e. RSS rises ~linearly with no decay. The native FTS arena is NOT reclaimed\n' +
|
||||
+ ' by db.close(); the leak is real over a long-lived session.\n' +
|
||||
+ ' → plan U4 (worker/process isolation of the FTS read path) is JUSTIFIED.',
|
||||
+ );
|
||||
+ } else if (label === 'PLATEAU') {
|
||||
+ console.log(
|
||||
+ `VERDICT: PLATEAU at this corpus (${corpus}) — the per-cycle increment DECAYS to flat\n` +
|
||||
+ ' (second-half slope below the noise floor). db.close() reclaims the FTS arena;\n' +
|
||||
+ ' footprint is bounded (and the pool further caps it at MAX_POOL_SIZE). No\n' +
|
||||
+ ' unbounded leak. Caveat: synthetic fixture — confirm with a --via-pool run\n' +
|
||||
+ ' against a real large analyzed repo before fully closing plan U4.',
|
||||
+ );
|
||||
+ } else {
|
||||
+ console.log(
|
||||
+ `VERDICT: INCONCLUSIVE at this corpus (${corpus}) — the run is noisy (step discontinuity)\n` +
|
||||
+ ' or still decelerating without reaching flat, so neither a clean PLATEAU nor a\n' +
|
||||
+ ' sustained linear CLIMB can be asserted. NATIVE synthetic runs do not resolve\n' +
|
||||
+ ' this reliably at scale. The definitive test is a --via-pool run against a real\n' +
|
||||
+ ' large analyzed repo over many cycles (with GITNEXUS_POOL_RSS_TRACE=1). Plan U4\n' +
|
||||
+ ' stays GATED — neither closed nor built on this evidence.',
|
||||
+ );
|
||||
+ }
|
||||
+ console.log(
|
||||
+ `MACHINE: ${JSON.stringify({ mode: VIA_POOL ? 'via-pool' : 'native', corpus, baseline, firstMed, lastMed, delta, overallSlope: Number(slope.toFixed(3)), firstHalfSlope: Number(firstHalfSlope.toFixed(3)), secondHalfSlope: Number(secondHalfSlope.toFixed(3)), maxJump, stepDiscontinuity, peak, cycles: series.length, verdict: label })}`,
|
||||
+ );
|
||||
+ console.log('=====================================================================\n');
|
||||
+}
|
||||
+
|
||||
+// ── main ────────────────────────────────────────────────────────────────────
|
||||
+(async () => {
|
||||
+ const result = VIA_POOL ? await runViaPool(VIA_POOL) : await runNative();
|
||||
+ verdict(result);
|
||||
+ process.exit(0);
|
||||
+})().catch((e) => {
|
||||
+ console.error('[fts-rss] fatal:', e?.stack || e);
|
||||
+ process.exit(1);
|
||||
+});
|
||||
diff --git a/gitnexus/scripts/install-duckdb-extension.mjs b/gitnexus/scripts/install-duckdb-extension.mjs
|
||||
index 2bc65a05e..7492e084f 100644
|
||||
--- a/gitnexus/scripts/install-duckdb-extension.mjs
|
||||
+++ b/gitnexus/scripts/install-duckdb-extension.mjs
|
||||
@@ -14,7 +14,7 @@ function parseLbugMaxDbSize(raw) {
|
||||
return Math.floor(parsed);
|
||||
}
|
||||
|
||||
-async function installDuckDbExtension(extensionName) {
|
||||
+async function installDuckDbExtension(extensionName, verifyOnly = false) {
|
||||
if (!extensionName || !EXTENSION_NAME_PATTERN.test(extensionName)) {
|
||||
throw new Error(`Invalid DuckDB extension name: ${extensionName ?? '<missing>'}`);
|
||||
}
|
||||
@@ -22,9 +22,11 @@ async function installDuckDbExtension(extensionName) {
|
||||
const require = createRequire(import.meta.url);
|
||||
const lbugModule = require('@ladybugdb/core');
|
||||
const lbug = lbugModule.default ?? lbugModule;
|
||||
- const lbugMaxDbSize = parseLbugMaxDbSize(
|
||||
- process.argv[3] ?? process.env.GITNEXUS_LBUG_MAX_DB_SIZE,
|
||||
- );
|
||||
+ // argv[3] is the optional positional size; ignore it when it is actually a
|
||||
+ // flag token (e.g. `--verify-only`) and fall back to the env default.
|
||||
+ const sizeArg =
|
||||
+ process.argv[3] && !process.argv[3].startsWith('--') ? process.argv[3] : undefined;
|
||||
+ const lbugMaxDbSize = parseLbugMaxDbSize(sizeArg ?? process.env.GITNEXUS_LBUG_MAX_DB_SIZE);
|
||||
|
||||
const tmpDir = await fs.mkdtemp(path.join(os.tmpdir(), 'gitnexus-ext-install-'));
|
||||
const dbPath = path.join(tmpDir, 'install.lbug');
|
||||
@@ -34,7 +36,18 @@ async function installDuckDbExtension(extensionName) {
|
||||
try {
|
||||
db = new lbug.Database(dbPath, 0, false, false, lbugMaxDbSize);
|
||||
conn = new lbug.Connection(db);
|
||||
- await conn.query(`INSTALL ${extensionName}`);
|
||||
+ if (verifyOnly) {
|
||||
+ // Prove a previously-baked extension is resolvable by a FRESH process
|
||||
+ // under the current HOME (the runtime `LOAD EXTENSION` path) — no INSTALL,
|
||||
+ // no network. Used as a Docker build-time gate so a HOME/extension-dir
|
||||
+ // mismatch fails the build instead of silently degrading search at runtime.
|
||||
+ await conn.query(`LOAD EXTENSION ${extensionName}`);
|
||||
+ console.log(
|
||||
+ `[install-ext] LOAD-only verify OK for '${extensionName}' (HOME=${process.env.HOME})`,
|
||||
+ );
|
||||
+ } else {
|
||||
+ await conn.query(`INSTALL ${extensionName}`);
|
||||
+ }
|
||||
} finally {
|
||||
if (conn) await conn.close().catch(() => {});
|
||||
if (db) await db.close().catch(() => {});
|
||||
@@ -42,7 +55,10 @@ async function installDuckDbExtension(extensionName) {
|
||||
}
|
||||
}
|
||||
|
||||
-installDuckDbExtension(process.argv[2] ?? process.env.GITNEXUS_LBUG_EXTENSION_NAME).catch((err) => {
|
||||
+installDuckDbExtension(
|
||||
+ process.argv[2] ?? process.env.GITNEXUS_LBUG_EXTENSION_NAME,
|
||||
+ process.argv.includes('--verify-only'),
|
||||
+).catch((err) => {
|
||||
console.error(err instanceof Error ? (err.stack ?? err.message) : String(err));
|
||||
process.exitCode = 1;
|
||||
});
|
||||
diff --git a/gitnexus/src/core/lbug/pool-adapter.ts b/gitnexus/src/core/lbug/pool-adapter.ts
|
||||
index 030688d14..e5338e3d0 100644
|
||||
--- a/gitnexus/src/core/lbug/pool-adapter.ts
|
||||
+++ b/gitnexus/src/core/lbug/pool-adapter.ts
|
||||
@@ -103,6 +103,19 @@ const IDLE_TIMEOUT_MS = 5 * 60 * 1000; // 5 minutes
|
||||
/** Max connections per repo (caps concurrent queries per repo) */
|
||||
const MAX_CONNS_PER_REPO = 8;
|
||||
|
||||
+// Behavior-neutral RSS tracing for the FTS evict→reload memory repro
|
||||
+// (gitnexus/scripts/bench/fts-evict-reload-rss.mjs). Two invariants keep it safe
|
||||
+// in the pool init/close hot path: it writes ONLY to stderr (stdout is the MCP
|
||||
+// JSON-RPC channel), and the GITNEXUS_POOL_RSS_TRACE gate makes it a no-op — one
|
||||
+// env-var compare per call, nothing else — unless a harness explicitly enables it.
|
||||
+function traceRss(event: 'init' | 'close', repoId: string): void {
|
||||
+ if (process.env.GITNEXUS_POOL_RSS_TRACE !== '1') return;
|
||||
+ const rssMb = Math.round(process.memoryUsage().rss / (1024 * 1024));
|
||||
+ process.stderr.write(
|
||||
+ `[pool-rss] ${event} repo=${repoId} pool=${pool.size} dbCache=${dbCache.size} rssMB=${rssMb}\n`,
|
||||
+ );
|
||||
+}
|
||||
+
|
||||
let idleTimer: ReturnType<typeof setInterval> | null = null;
|
||||
|
||||
// Stdout-capture state lives in `gitnexus/src/mcp/stdio-capture.ts` — a leaf
|
||||
@@ -240,6 +253,8 @@ function closeOne(repoId: string): void {
|
||||
// Isolate listener failures — teardown must complete.
|
||||
}
|
||||
}
|
||||
+
|
||||
+ traceRss('close', repoId);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -611,6 +626,7 @@ async function doInitLbug(repoId: string, dbPath: string): Promise<void> {
|
||||
closed: false,
|
||||
});
|
||||
ensureIdleTimer();
|
||||
+ traceRss('init', repoId);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -673,6 +689,7 @@ export async function initLbugWithDb(
|
||||
closed: false,
|
||||
});
|
||||
ensureIdleTimer();
|
||||
+ traceRss('init', repoId);
|
||||
}
|
||||
|
||||
/**
|
||||
diff --git a/gitnexus/src/mcp/local/local-backend.ts b/gitnexus/src/mcp/local/local-backend.ts
|
||||
index 77411b107..c0a30eb58 100644
|
||||
--- a/gitnexus/src/mcp/local/local-backend.ts
|
||||
+++ b/gitnexus/src/mcp/local/local-backend.ts
|
||||
@@ -1112,73 +1112,120 @@ export class LocalBackend {
|
||||
>();
|
||||
const definitions: any[] = []; // standalone symbols not in any process
|
||||
|
||||
- for (const [_, item] of merged) {
|
||||
- const sym = item.data;
|
||||
- if (!sym.nodeId) {
|
||||
- // File-level results go to definitions
|
||||
- definitions.push({
|
||||
- name: sym.name,
|
||||
- type: sym.type || 'File',
|
||||
- filePath: sym.filePath,
|
||||
- });
|
||||
- continue;
|
||||
- }
|
||||
-
|
||||
- // Find processes this symbol participates in
|
||||
- let processRows: any[] = [];
|
||||
+ // Batch-fetch process participation, cohesion, and (optionally) content for
|
||||
+ // ALL matched symbols in 2-3 graph queries instead of 2-3 *per symbol*. The
|
||||
+ // previous per-symbol loop issued up to 3N sequential pool round-trips
|
||||
+ // (searchLimit symbols × {STEP_IN_PROCESS, MEMBER_OF, content}); on a warm
|
||||
+ // repo the IPC + query-setup overhead of those round-trips dominated query
|
||||
+ // latency. Collapsing to `WHERE n.id IN $nodeIds` preserves identical output
|
||||
+ // (the aggregation loop below is unchanged) while cutting the round-trips.
|
||||
+ // Array params bind through the pool exactly as bm25Search's
|
||||
+ // `WHERE n.id IN $nodeIds` already does. (Ported from gitnexus-enterprise
|
||||
+ // PR #222 — N+1 → 2-3 batched queries.)
|
||||
+ const nodeIds = merged.map(([, m]) => m.data?.nodeId).filter((id): id is string => !!id);
|
||||
+
|
||||
+ const processRowsByNode = new Map<string, any[]>();
|
||||
+ const cohesionByNode = new Map<string, { cohesion: number; module?: string }>();
|
||||
+ const contentByNode = new Map<string, string>();
|
||||
+
|
||||
+ // Chunk the IN-list like the impact path (CHUNK_SIZE=100) so a large result
|
||||
+ // set never builds an unbounded `IN` parameter. Default batch is
|
||||
+ // processLimit*maxSymbolsPerProcess (≤ one chunk), but chunk for robustness.
|
||||
+ const QUERY_CHUNK_SIZE = 100;
|
||||
+ for (let i = 0; i < nodeIds.length; i += QUERY_CHUNK_SIZE) {
|
||||
+ const ids = nodeIds.slice(i, i + QUERY_CHUNK_SIZE);
|
||||
+
|
||||
+ // Processes each symbol participates in. `n.id AS nodeId` is prepended as
|
||||
+ // column 0 so rows from many symbols can be re-associated to their symbol.
|
||||
try {
|
||||
- processRows = await executeParameterized(
|
||||
+ const rows = await executeParameterized(
|
||||
repo.lbugPath,
|
||||
`
|
||||
- MATCH (n {id: $nodeId})-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
|
||||
- RETURN p.id AS pid, p.label AS label, p.heuristicLabel AS heuristicLabel, p.processType AS processType, p.stepCount AS stepCount, r.step AS step
|
||||
+ MATCH (n)-[r:CodeRelation {type: 'STEP_IN_PROCESS'}]->(p:Process)
|
||||
+ WHERE n.id IN $nodeIds
|
||||
+ RETURN n.id AS nodeId, p.id AS pid, p.label AS label, p.heuristicLabel AS heuristicLabel, p.processType AS processType, p.stepCount AS stepCount, r.step AS step
|
||||
`,
|
||||
- { nodeId: sym.nodeId },
|
||||
+ { nodeIds: ids },
|
||||
);
|
||||
+ for (const row of rows) {
|
||||
+ const nid = row.nodeId ?? row[0];
|
||||
+ let list = processRowsByNode.get(nid);
|
||||
+ if (!list) processRowsByNode.set(nid, (list = []));
|
||||
+ list.push(row);
|
||||
+ }
|
||||
} catch (e) {
|
||||
logQueryError('query:process-lookup', e);
|
||||
}
|
||||
|
||||
- // Get cluster membership + cohesion (cohesion used as internal ranking signal)
|
||||
- let cohesion = 0;
|
||||
- let module: string | undefined;
|
||||
+ // Cluster membership + cohesion. Keep the FIRST community row per node to
|
||||
+ // mirror the prior per-symbol `LIMIT 1` (each symbol keeps ITS community,
|
||||
+ // not one community for the whole batch).
|
||||
try {
|
||||
- const cohesionRows = await executeParameterized(
|
||||
+ const rows = await executeParameterized(
|
||||
repo.lbugPath,
|
||||
`
|
||||
- MATCH (n {id: $nodeId})-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
|
||||
- RETURN c.cohesion AS cohesion, c.heuristicLabel AS module
|
||||
- LIMIT 1
|
||||
+ MATCH (n)-[:CodeRelation {type: 'MEMBER_OF'}]->(c:Community)
|
||||
+ WHERE n.id IN $nodeIds
|
||||
+ RETURN n.id AS nodeId, c.cohesion AS cohesion, c.heuristicLabel AS module
|
||||
`,
|
||||
- { nodeId: sym.nodeId },
|
||||
+ { nodeIds: ids },
|
||||
);
|
||||
- if (cohesionRows.length > 0) {
|
||||
- cohesion = (cohesionRows[0].cohesion ?? cohesionRows[0][0]) || 0;
|
||||
- module = cohesionRows[0].module ?? cohesionRows[0][1];
|
||||
+ for (const row of rows) {
|
||||
+ const nid = row.nodeId ?? row[0];
|
||||
+ if (!cohesionByNode.has(nid)) {
|
||||
+ cohesionByNode.set(nid, {
|
||||
+ cohesion: (row.cohesion ?? row[1]) || 0,
|
||||
+ module: row.module ?? row[2],
|
||||
+ });
|
||||
+ }
|
||||
}
|
||||
} catch (e) {
|
||||
logQueryError('query:cluster-info', e);
|
||||
}
|
||||
|
||||
- // Optionally fetch content
|
||||
- let content: string | undefined;
|
||||
+ // Optionally fetch content for every matched symbol.
|
||||
if (includeContent) {
|
||||
try {
|
||||
- const contentRows = await executeParameterized(
|
||||
+ const rows = await executeParameterized(
|
||||
repo.lbugPath,
|
||||
`
|
||||
- MATCH (n {id: $nodeId})
|
||||
- RETURN n.content AS content
|
||||
+ MATCH (n)
|
||||
+ WHERE n.id IN $nodeIds
|
||||
+ RETURN n.id AS nodeId, n.content AS content
|
||||
`,
|
||||
- { nodeId: sym.nodeId },
|
||||
+ { nodeIds: ids },
|
||||
);
|
||||
- if (contentRows.length > 0) {
|
||||
- content = contentRows[0].content ?? contentRows[0][0];
|
||||
+ for (const row of rows) {
|
||||
+ const nid = row.nodeId ?? row[0];
|
||||
+ contentByNode.set(nid, row.content ?? row[1]);
|
||||
}
|
||||
} catch (e) {
|
||||
logQueryError('query:content-fetch', e);
|
||||
}
|
||||
}
|
||||
+ }
|
||||
+
|
||||
+ // Aggregation is unchanged from the per-symbol version — it now reads the
|
||||
+ // pre-fetched maps instead of issuing a query per symbol. Iterating `merged`
|
||||
+ // in the same (sorted) order preserves processMap insertion order, the
|
||||
+ // definitions order, and the item.score association exactly.
|
||||
+ for (const [_, item] of merged) {
|
||||
+ const sym = item.data;
|
||||
+ if (!sym.nodeId) {
|
||||
+ // File-level results go to definitions
|
||||
+ definitions.push({
|
||||
+ name: sym.name,
|
||||
+ type: sym.type || 'File',
|
||||
+ filePath: sym.filePath,
|
||||
+ });
|
||||
+ continue;
|
||||
+ }
|
||||
+
|
||||
+ const processRows = processRowsByNode.get(sym.nodeId) ?? [];
|
||||
+ const coh = cohesionByNode.get(sym.nodeId);
|
||||
+ const cohesion = coh?.cohesion ?? 0;
|
||||
+ const module = coh?.module;
|
||||
+ const content = includeContent ? contentByNode.get(sym.nodeId) : undefined;
|
||||
|
||||
const symbolEntry = {
|
||||
id: sym.nodeId,
|
||||
@@ -1197,12 +1244,13 @@ export class LocalBackend {
|
||||
} else {
|
||||
// Add to each process it belongs to
|
||||
for (const row of processRows) {
|
||||
- const pid = row.pid ?? row[0];
|
||||
- const label = row.label ?? row[1];
|
||||
- const hLabel = row.heuristicLabel ?? row[2];
|
||||
- const pType = row.processType ?? row[3];
|
||||
- const stepCount = row.stepCount ?? row[4];
|
||||
- const step = row.step ?? row[5];
|
||||
+ // Positional fallbacks shift +1 because `n.id AS nodeId` is column 0.
|
||||
+ const pid = row.pid ?? row[1];
|
||||
+ const label = row.label ?? row[2];
|
||||
+ const hLabel = row.heuristicLabel ?? row[3];
|
||||
+ const pType = row.processType ?? row[4];
|
||||
+ const stepCount = row.stepCount ?? row[5];
|
||||
+ const step = row.step ?? row[6];
|
||||
|
||||
if (!processMap.has(pid)) {
|
||||
processMap.set(pid, {
|
||||
diff --git a/gitnexus/test/fixtures/local-backend-seed.ts b/gitnexus/test/fixtures/local-backend-seed.ts
|
||||
index 3f299046e..4878348b2 100644
|
||||
--- a/gitnexus/test/fixtures/local-backend-seed.ts
|
||||
+++ b/gitnexus/test/fixtures/local-backend-seed.ts
|
||||
@@ -35,6 +35,12 @@ export const LOCAL_BACKEND_SEED_DATA = [
|
||||
CREATE (a)-[:CodeRelation {type: 'STEP_IN_PROCESS', confidence: 1.0, reason: '', step: 1}]->(p)`,
|
||||
`MATCH (a:Function), (p:Process) WHERE a.id = 'func:validate' AND p.id = 'proc:login-flow'
|
||||
CREATE (a)-[:CodeRelation {type: 'STEP_IN_PROCESS', confidence: 1.0, reason: '', step: 2}]->(p)`,
|
||||
+ // func:validate is the terminalId of proc:beta-flow too — wiring its second
|
||||
+ // STEP_IN_PROCESS edge makes it a genuine MULTI-process symbol, which the
|
||||
+ // batched-query test uses to exercise the full row[1..6] positional shift
|
||||
+ // (a single-process symbol can't expose an off-by-one in those fallbacks).
|
||||
+ `MATCH (a:Function), (p:Process) WHERE a.id = 'func:validate' AND p.id = 'proc:beta-flow'
|
||||
+ CREATE (a)-[:CodeRelation {type: 'STEP_IN_PROCESS', confidence: 1.0, reason: '', step: 3}]->(p)`,
|
||||
`MATCH (h:Function), (t:Tool) WHERE h.id = 'func:alpha' AND t.id = 'Tool:alpha'
|
||||
CREATE (h)-[:CodeRelation {type: 'HANDLES_TOOL', confidence: 1.0, reason: 'tool-definition', step: 0}]->(t)`,
|
||||
`MATCH (h:Function), (t:Tool) WHERE h.id = 'func:beta' AND t.id = 'Tool:beta'
|
||||
diff --git a/gitnexus/test/integration/local-backend-calltool.test.ts b/gitnexus/test/integration/local-backend-calltool.test.ts
|
||||
index e2640deab..176e061b2 100644
|
||||
--- a/gitnexus/test/integration/local-backend-calltool.test.ts
|
||||
+++ b/gitnexus/test/integration/local-backend-calltool.test.ts
|
||||
@@ -113,6 +113,72 @@ withTestLbugDB(
|
||||
expect(result.timing.bm25 ?? result.timing.vector).toBeGreaterThanOrEqual(0);
|
||||
});
|
||||
|
||||
+ // PR #222 port: the query tool batches per-symbol process/cohesion/content
|
||||
+ // lookups (N+1 → 2-3 `WHERE n.id IN $nodeIds` queries). These assertions
|
||||
+ // guard the batch-adaptation hazards that a naive cherry-pick would break:
|
||||
+ // (1) each symbol keeps ITS OWN community (the per-node first-row pick that
|
||||
+ // replaced the per-symbol `LIMIT 1`), and (2) content maps to the right
|
||||
+ // node — both depend on the +1 positional-index shift after prepending
|
||||
+ // `n.id AS nodeId`. func:login is MEMBER_OF comm:auth ("Authentication");
|
||||
+ // func:validate has no community, so it must NOT inherit login's.
|
||||
+ it('query batches per-symbol enrichment without cross-assigning community/content', async () => {
|
||||
+ const findSym = (res: any, id: string) =>
|
||||
+ (res.process_symbols ?? []).find((s: any) => s.id === id) ??
|
||||
+ (res.definitions ?? []).find((s: any) => s.id === id);
|
||||
+
|
||||
+ const loginRes = await backend.callTool('query', {
|
||||
+ query: 'login',
|
||||
+ include_content: true,
|
||||
+ });
|
||||
+ expect(loginRes).not.toHaveProperty('error');
|
||||
+ const login = findSym(loginRes, 'func:login');
|
||||
+ expect(login).toBeDefined();
|
||||
+ // Community correctly associated to its own node (not dropped, not leaked).
|
||||
+ expect(login.module).toBe('Authentication');
|
||||
+ // Content correctly mapped to its own node (positional [1] after nodeId).
|
||||
+ expect(login.content).toBe('function login() {}');
|
||||
+
|
||||
+ const validateRes = await backend.callTool('query', {
|
||||
+ query: 'validate',
|
||||
+ include_content: true,
|
||||
+ });
|
||||
+ expect(validateRes).not.toHaveProperty('error');
|
||||
+ const validate = findSym(validateRes, 'func:validate');
|
||||
+ expect(validate).toBeDefined();
|
||||
+ // validate has no MEMBER_OF edge — a flat batched `LIMIT 1` would have
|
||||
+ // leaked some other node's community onto it. It must have none.
|
||||
+ expect(validate.module).toBeUndefined();
|
||||
+ expect(validate.content).toBe('function validate() {}');
|
||||
+ });
|
||||
+
|
||||
+ // PR #222 port: a symbol in MULTIPLE processes is what fully exercises the
|
||||
+ // +1 positional shift in the batched STEP_IN_PROCESS aggregation — with a
|
||||
+ // single process row, `row.pid ?? row[1]` succeeds whether the shift is
|
||||
+ // right or wrong. func:validate is a step in BOTH proc:login-flow (step 2)
|
||||
+ // and proc:beta-flow (step 3), so both rows for the one node must be parsed
|
||||
+ // (pid=row[1], step=row[6]); an off-by-one would drop a process or mis-pair
|
||||
+ // pid↔step. Also pins process ranking (totalScore via the regroup-by-nodeId).
|
||||
+ it('query batches a multi-process symbol and ranks processes (positional shift across rows)', async () => {
|
||||
+ const res = await backend.callTool('query', { query: 'validate' });
|
||||
+ expect(res).not.toHaveProperty('error');
|
||||
+ const processIds = (res.processes ?? []).map((p: any) => p.id);
|
||||
+ // Both of validate's processes must appear — both STEP_IN_PROCESS rows
|
||||
+ // were parsed and grouped by the correct pid (row[1]).
|
||||
+ expect(processIds).toContain('proc:login-flow');
|
||||
+ expect(processIds).toContain('proc:beta-flow');
|
||||
+
|
||||
+ // process_symbols dedups by id, so validate appears once carrying the
|
||||
+ // pid+step of its top-ranked process — they must come from the SAME
|
||||
+ // shifted row: login-flow⇒step 2, beta-flow⇒step 3.
|
||||
+ const v = (res.process_symbols ?? []).find((s: any) => s.id === 'func:validate');
|
||||
+ expect(v).toBeDefined();
|
||||
+ expect(v.step_index).toBe(v.process_id === 'proc:beta-flow' ? 3 : 2);
|
||||
+
|
||||
+ // Ranking: 'login' surfaces proc:login-flow as the top process.
|
||||
+ const loginRes = await backend.callTool('query', { query: 'login' });
|
||||
+ expect((loginRes.processes ?? [])[0]?.id).toBe('proc:login-flow');
|
||||
+ });
|
||||
+
|
||||
it('tool_map returns per-tool flows without cross-attributing same-file tools', async () => {
|
||||
const result = await backend.callTool('tool_map', {});
|
||||
expect(result).not.toHaveProperty('error');
|
||||
414
eval/workflow_bench/review_cases/pr-2258.patch
Normal file
414
eval/workflow_bench/review_cases/pr-2258.patch
Normal file
|
|
@ -0,0 +1,414 @@
|
|||
diff --git a/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs b/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs
|
||||
index 5d3e44a0d..49953f8ab 100644
|
||||
--- a/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs
|
||||
+++ b/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs
|
||||
@@ -17,7 +17,14 @@ const floor = Number(process.env.MUTATION_RECALL_FLOOR ?? '0.5');
|
||||
|
||||
const report = JSON.parse(fs.readFileSync(reportPath, 'utf8'));
|
||||
const checks = Array.isArray(report?.mutation?.checks) ? report.mutation.checks : [];
|
||||
-const scored = checks.filter((c) => typeof c.recall === 'number');
|
||||
+// Gate only the checks the oracle marked recall-gated. measure.mjs sets
|
||||
+// `recallGated: false` for cases a forward value-diff oracle cannot fairly
|
||||
+// score against the PDG slice: UPSTREAM fixtures (the oracle runs in its native
|
||||
+// downstream sense, so its behavioral AIS can never intersect a reverse slice —
|
||||
+// recall is 0 by construction) and id-discrimination corroboration fixtures.
|
||||
+// Those still carry a numeric `recall` for the report, so the legacy
|
||||
+// `typeof c.recall === 'number'` filter wrongly tripped the floor on them.
|
||||
+const scored = checks.filter((c) => c.recallGated === true && typeof c.recall === 'number');
|
||||
const recalls = scored.map((c) => c.recall);
|
||||
const min = recalls.length ? Math.min(...recalls) : null;
|
||||
const mean = recalls.length ? recalls.reduce((a, b) => a + b, 0) / recalls.length : null;
|
||||
diff --git a/gitnexus/bench/impact-pdg/measure.mjs b/gitnexus/bench/impact-pdg/measure.mjs
|
||||
index 65d3a7675..f6158f1da 100644
|
||||
--- a/gitnexus/bench/impact-pdg/measure.mjs
|
||||
+++ b/gitnexus/bench/impact-pdg/measure.mjs
|
||||
@@ -25,11 +25,16 @@
|
||||
* `repo-manager.getGlobalDir()` — it roots the registry; the per-repo DB
|
||||
* lands in `<fixtureCopy>/.gitnexus/`, so fixtures are copied to a temp
|
||||
* working dir to keep the source tree clean);
|
||||
- * 2. SHELL OUT to the real CLI as a child process:
|
||||
- * node --import tsx src/cli/index.ts analyze <copy> --pdg --skip-git --index-only
|
||||
- * (child-process isolation sidesteps `process.exit`; real `saveMeta` +
|
||||
- * `registerRepo` land in the temp home; workers spawn from `dist/`, so the
|
||||
- * harness builds `dist/` first — run `node scripts/build.js`);
|
||||
+ * 2. SHELL OUT to the real CLI as a child process (see `cliChildArgs`):
|
||||
+ * node dist/cli/index.js analyze <copy> --pdg --skip-git --index-only
|
||||
+ * preferring the BUILT `dist/` CLI when present — plain JS, no tsx, and the
|
||||
+ * parse workers it spawns also load from `dist/`. The mutation workflow
|
||||
+ * builds `dist/` first (`node scripts/build.js`); `node --import tsx
|
||||
+ * src/cli/index.ts` is NOT used because Node >=22.18 native type-stripping
|
||||
+ * breaks the `.js`->`.ts` entry resolution (ERR_MODULE_NOT_FOUND on
|
||||
+ * `lazy-action.js`). Build-free runs fall back to tsx's own CLI over src.
|
||||
+ * (Child-process isolation sidesteps `process.exit`; real `saveMeta` +
|
||||
+ * `registerRepo` land in the temp home);
|
||||
* 3. `new LocalBackend(); await init()` resolves the fixture via the REAL
|
||||
* registry (the parent process ALSO sets `GITNEXUS_HOME` so init reads the
|
||||
* temp registry, not the user's ~/.gitnexus);
|
||||
@@ -53,6 +58,7 @@ import os from 'node:os';
|
||||
import path from 'node:path';
|
||||
import crypto from 'node:crypto';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
+import { createRequire } from 'node:module';
|
||||
import { fileURLToPath, pathToFileURL } from 'node:url';
|
||||
|
||||
import {
|
||||
@@ -95,6 +101,33 @@ const REPO_ROOT = path.resolve(__dirname, '..', '..'); // gitnexus/
|
||||
const FIXTURES_DIR = path.join(__dirname, 'fixtures');
|
||||
const BASELINE_PATH = path.join(__dirname, 'baselines.json');
|
||||
const CLI_ENTRY = path.join(REPO_ROOT, 'src', 'cli', 'index.ts');
|
||||
+// Shipped CLI entry (package.json `bin`). PREFERRED for the child analyze: it's
|
||||
+// plain compiled JS, so the analyze process — AND the parse workers it spawns,
|
||||
+// which resolve relative to the running entry — load from `dist/` with no tsx in
|
||||
+// the loop. The build-free path below stays as a fallback.
|
||||
+const DIST_CLI = path.join(REPO_ROOT, 'dist', 'cli', 'index.js');
|
||||
+// Build-free fallback: tsx's OWN cli entry (resolved from this package), NOT
|
||||
+// `node --import tsx <entry>.ts`. On Node >=22.18 native TypeScript type-
|
||||
+// stripping is enabled by default and intercepts the `.ts` entry before tsx's
|
||||
+// `--import` resolve hook applies; native stripping does NOT remap `./foo.js`
|
||||
+// specifiers to `foo.ts` (tsx does), so `node --import tsx src/cli/index.ts`
|
||||
+// crashes resolving `./lazy-action.js` (ERR_MODULE_NOT_FOUND) on newer Node.
|
||||
+// The tsx CLI takes over module loading and is version-agnostic across the
|
||||
+// declared engines range (node >=22.0, where `--no-experimental-strip-types`
|
||||
+// is not a universally-recognized flag). Workers still spawn from src via tsx on
|
||||
+// this path, so it is only robust on the older Node devs run locally.
|
||||
+const TSX_CLI = createRequire(import.meta.url).resolve('tsx/cli');
|
||||
+
|
||||
+/**
|
||||
+ * Build the argv that runs the real CLI as a child of `process.execPath`.
|
||||
+ * Prefers the built `dist/` CLI (production-faithful, no tsx, dist workers) when
|
||||
+ * present — this is what the mutation workflow uses (it builds dist first). Falls
|
||||
+ * back to the tsx CLI over src for build-free local runs. Returns the args AFTER
|
||||
+ * the node binary, i.e. ready for `spawnSync(process.execPath, [...args])`.
|
||||
+ */
|
||||
+function cliChildArgs(rest) {
|
||||
+ return fs.existsSync(DIST_CLI) ? [DIST_CLI, ...rest] : [TSX_CLI, CLI_ENTRY, ...rest];
|
||||
+}
|
||||
|
||||
const SCOPES = ['intra', 'inter', 'mixed'];
|
||||
const MODES = ['callgraph', 'pdg'];
|
||||
@@ -138,7 +171,7 @@ async function analyzeAndImpact(fx, home, { pdgOn = true } = {}) {
|
||||
fs.cpSync(path.join(fx.dir, 'src'), path.join(work, 'src'), { recursive: true });
|
||||
|
||||
const env = { ...process.env, GITNEXUS_HOME: home };
|
||||
- const args = ['--import', 'tsx', CLI_ENTRY, 'analyze', work, '--skip-git', '--index-only'];
|
||||
+ const args = cliChildArgs(['analyze', work, '--skip-git', '--index-only']);
|
||||
if (pdgOn) args.push('--pdg');
|
||||
const an = spawnSync(process.execPath, args, {
|
||||
env,
|
||||
diff --git a/gitnexus/package-lock.json b/gitnexus/package-lock.json
|
||||
index b28f005d7..927fe93c3 100644
|
||||
--- a/gitnexus/package-lock.json
|
||||
+++ b/gitnexus/package-lock.json
|
||||
@@ -52,6 +52,10 @@
|
||||
"gitnexus": "dist/cli/index.js"
|
||||
},
|
||||
"devDependencies": {
|
||||
+ "@babel/generator": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/traverse": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
"@types/busboy": "^1.5.4",
|
||||
"@types/cli-progress": "^3.11.6",
|
||||
"@types/cors": "^2.8.17",
|
||||
@@ -76,10 +80,59 @@
|
||||
"typescript": "^6.0.3"
|
||||
}
|
||||
},
|
||||
+ "node_modules/@babel/code-frame": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/code-frame/-/code-frame-7.29.7.tgz",
|
||||
+ "integrity": "sha512-Aup7aUOfpbAUg2ROOJN6Iw5f9DMBlzu0mIkm/malLQFN/YQgO48wCj0Kxa3sEHJvPVFg7siR+qRInwXd2qhQKw==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/helper-validator-identifier": "^7.29.7",
|
||||
+ "js-tokens": "^4.0.0",
|
||||
+ "picocolors": "^1.1.1"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
+ "node_modules/@babel/code-frame/node_modules/js-tokens": {
|
||||
+ "version": "4.0.0",
|
||||
+ "resolved": "https://registry.npmjs.org/js-tokens/-/js-tokens-4.0.0.tgz",
|
||||
+ "integrity": "sha512-RdJUflcE3cUzKiMqQgsCu06FPu9UdIJO0beYbPhHN4k6apgJtifcoCtT9bcxOpYBtpD2kCM6Sbzg4CausW/PKQ==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT"
|
||||
+ },
|
||||
+ "node_modules/@babel/generator": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/generator/-/generator-7.29.7.tgz",
|
||||
+ "integrity": "sha512-DkXD5OJQaAQIdZ1bt3UZdEnHAn9Imd3IVBdX03UFe+ony9Ojw5pzr9YVKGDY1jt+Gcn/FnGkNf8r+Vj5NOJWtQ==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
+ "@jridgewell/gen-mapping": "^0.3.12",
|
||||
+ "@jridgewell/trace-mapping": "^0.3.28",
|
||||
+ "jsesc": "^3.0.2"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
+ "node_modules/@babel/helper-globals": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/helper-globals/-/helper-globals-7.29.7.tgz",
|
||||
+ "integrity": "sha512-3nQVUAtvkKH9zahfWgw96Jc/uFOmjACE1kQz82E2lqWmHBgjzbNlsC22nuQTfahmWeQtTq5nQ/4Nnd2A1wj4zA==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/@babel/helper-string-parser": {
|
||||
- "version": "7.27.1",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/helper-string-parser/-/helper-string-parser-7.27.1.tgz",
|
||||
- "integrity": "sha512-qMlSxKbpRlAridDExk92nSobyDdpPijUq2DW6oDnUqd0iOGxmQjyqhMIihI9+zv4LPyZdRje2cavWPbCbWm3eA==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/helper-string-parser/-/helper-string-parser-7.29.7.tgz",
|
||||
+ "integrity": "sha512-Pb5ijPrZ89GDH8223L4UP8i6QApWxs04RbPQJTeWDV0/keR2E36MeKnyr6LYmUUvqRRI+Iv87SuF1W6ErINzYw==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
@@ -87,9 +140,9 @@
|
||||
}
|
||||
},
|
||||
"node_modules/@babel/helper-validator-identifier": {
|
||||
- "version": "7.28.5",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/helper-validator-identifier/-/helper-validator-identifier-7.28.5.tgz",
|
||||
- "integrity": "sha512-qSs4ifwzKJSV39ucNjsvc6WVHs6b7S03sOh2OcHF9UHfVPqWWALUsNUVzhSBiItjRZoLHx7nIarVjqKVusUZ1Q==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/helper-validator-identifier/-/helper-validator-identifier-7.29.7.tgz",
|
||||
+ "integrity": "sha512-qehxGkRj55h/ff8EMaJ+cYhyaKlHIxqYDn682wQD7RNp9UujOQsHog2uS0r2vzr4pW+sXf90NeeayjcNaX3fFg==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
@@ -97,13 +150,13 @@
|
||||
}
|
||||
},
|
||||
"node_modules/@babel/parser": {
|
||||
- "version": "7.29.2",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/parser/-/parser-7.29.2.tgz",
|
||||
- "integrity": "sha512-4GgRzy/+fsBa72/RZVJmGKPmZu9Byn8o4MoLpmNe1m8ZfYnz5emHLQz3U4gLud6Zwl0RZIcgiLD7Uq7ySFuDLA==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/parser/-/parser-7.29.7.tgz",
|
||||
+ "integrity": "sha512-hnORnjP/1P/zFEndoeX+n+t1RwWRJiJpM/jO7FW32Kn9r5+sJB2JWOdYo4L6k78j15eCwY3Gm/7364B1EMwtNg==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
- "@babel/types": "^7.29.0"
|
||||
+ "@babel/types": "^7.29.7"
|
||||
},
|
||||
"bin": {
|
||||
"parser": "bin/babel-parser.js"
|
||||
@@ -112,15 +165,49 @@
|
||||
"node": ">=6.0.0"
|
||||
}
|
||||
},
|
||||
+ "node_modules/@babel/template": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/template/-/template-7.29.7.tgz",
|
||||
+ "integrity": "sha512-puq+Gf35oI24FeN11LkoUQFqv9uwNeWpxXZi/Ji3rRIoKAzKnxRaZ+Gkj0vKS9ZCiTESfng1N9LyOyXvo+m+Gg==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/code-frame": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
+ "node_modules/@babel/traverse": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/traverse/-/traverse-7.29.7.tgz",
|
||||
+ "integrity": "sha512-EhlfNQtZ+NK22w5BM61ciuiq1m58ed33Wr1Xan//ZRTy6hgjnwyCffRYwzsGXdASJSUJ1guZILsErh1eQcl+zw==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/code-frame": "^7.29.7",
|
||||
+ "@babel/generator": "^7.29.7",
|
||||
+ "@babel/helper-globals": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/template": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
+ "debug": "^4.3.1"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/@babel/types": {
|
||||
- "version": "7.29.0",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/types/-/types-7.29.0.tgz",
|
||||
- "integrity": "sha512-LwdZHpScM4Qz8Xw2iKSzS+cfglZzJGvofQICy7W7v4caru4EaAmyUuO6BGrbyQ2mYV11W0U8j5mBhd14dd3B0A==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/types/-/types-7.29.7.tgz",
|
||||
+ "integrity": "sha512-4zBIxpPzowiZpusoFkyGVwakdRJUyuH5PxQ/PrqghfdFWWasvnCdPfQXHrenDai+gyLARulZjZowCOj6fjT4pA==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
- "@babel/helper-string-parser": "^7.27.1",
|
||||
- "@babel/helper-validator-identifier": "^7.28.5"
|
||||
+ "@babel/helper-string-parser": "^7.29.7",
|
||||
+ "@babel/helper-validator-identifier": "^7.29.7"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=6.9.0"
|
||||
@@ -1128,6 +1215,17 @@
|
||||
"node": ">=18.0.0"
|
||||
}
|
||||
},
|
||||
+ "node_modules/@jridgewell/gen-mapping": {
|
||||
+ "version": "0.3.13",
|
||||
+ "resolved": "https://registry.npmjs.org/@jridgewell/gen-mapping/-/gen-mapping-0.3.13.tgz",
|
||||
+ "integrity": "sha512-2kkt/7niJ6MgEPxF0bYdQ6etZaA+fQvDcLKckhy1yIQOzaoKjBBjSj63/aLVjYE3qhRt5dvM+uUyfCg6UKCBbA==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@jridgewell/sourcemap-codec": "^1.5.0",
|
||||
+ "@jridgewell/trace-mapping": "^0.3.24"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/@jridgewell/resolve-uri": {
|
||||
"version": "3.1.2",
|
||||
"resolved": "https://registry.npmjs.org/@jridgewell/resolve-uri/-/resolve-uri-3.1.2.tgz",
|
||||
@@ -1471,9 +1569,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1491,9 +1586,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1511,9 +1603,6 @@
|
||||
"ppc64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1531,9 +1620,6 @@
|
||||
"s390x"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1551,9 +1637,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1571,9 +1654,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3487,6 +3567,19 @@
|
||||
"js-yaml": "bin/js-yaml.js"
|
||||
}
|
||||
},
|
||||
+ "node_modules/jsesc": {
|
||||
+ "version": "3.1.0",
|
||||
+ "resolved": "https://registry.npmjs.org/jsesc/-/jsesc-3.1.0.tgz",
|
||||
+ "integrity": "sha512-/sM3dO2FOzXjKQhJuo0Q173wf2KOo8t4I8vHy6lF9poUp7bKT0/NHE8fPX23PwfhnykfqnC2xRxOnVw5XuGIaA==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "bin": {
|
||||
+ "jsesc": "bin/jsesc"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/json-bignum": {
|
||||
"version": "0.0.3",
|
||||
"resolved": "https://registry.npmjs.org/json-bignum/-/json-bignum-0.0.3.tgz",
|
||||
@@ -3668,9 +3761,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3692,9 +3782,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3716,9 +3803,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3740,9 +3824,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
diff --git a/gitnexus/package.json b/gitnexus/package.json
|
||||
index 6fa0c9007..82f2cabfd 100644
|
||||
--- a/gitnexus/package.json
|
||||
+++ b/gitnexus/package.json
|
||||
@@ -94,6 +94,10 @@
|
||||
"uuid": "^14.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
+ "@babel/generator": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/traverse": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
"@types/busboy": "^1.5.4",
|
||||
"@types/cli-progress": "^3.11.6",
|
||||
"@types/cors": "^2.8.17",
|
||||
453
eval/workflow_bench/review_cases/pr-2258b.patch
Normal file
453
eval/workflow_bench/review_cases/pr-2258b.patch
Normal file
|
|
@ -0,0 +1,453 @@
|
|||
diff --git a/gitnexus/bench/impact-pdg/README.md b/gitnexus/bench/impact-pdg/README.md
|
||||
index 845debfa2..a7bbac3ce 100644
|
||||
--- a/gitnexus/bench/impact-pdg/README.md
|
||||
+++ b/gitnexus/bench/impact-pdg/README.md
|
||||
@@ -227,10 +227,14 @@ analyze via a temp `GITNEXUS_HOME`, mock-free**. Per fixture:
|
||||
first, keeping the source tree clean).
|
||||
2. **Shell out** to the real CLI as a child process — child-process isolation
|
||||
sidesteps `process.exit`; real `saveMeta` + `registerRepo` land in the temp
|
||||
- home; parse workers spawn from `dist/` (so the harness needs a built `dist/`):
|
||||
+ home. The harness prefers the built `dist/` CLI (plain JS, no tsx; the parse
|
||||
+ workers it spawns also load from `dist/`), so it needs a built `dist/`; it
|
||||
+ falls back to tsx's own CLI over `src/` for build-free local runs. (`node
|
||||
+ --import tsx src/cli/index.ts` is avoided: Node ≥22.18 native type-stripping
|
||||
+ breaks the `.ts` entry's `./lazy-action.js`→`.ts` import resolution.)
|
||||
|
||||
```
|
||||
- node --import tsx src/cli/index.ts analyze <fixtureCopy> --pdg --skip-git --index-only
|
||||
+ node dist/cli/index.js analyze <fixtureCopy> --pdg --skip-git --index-only
|
||||
```
|
||||
3. `new LocalBackend(); await init()` resolves the fixture via the **real**
|
||||
registry (the parent process sets `GITNEXUS_HOME` too, so `init()` reads the
|
||||
diff --git a/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs b/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs
|
||||
index 5d3e44a0d..2ab46072a 100644
|
||||
--- a/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs
|
||||
+++ b/gitnexus/bench/impact-pdg/gate-mutation-recall.mjs
|
||||
@@ -17,7 +17,14 @@ const floor = Number(process.env.MUTATION_RECALL_FLOOR ?? '0.5');
|
||||
|
||||
const report = JSON.parse(fs.readFileSync(reportPath, 'utf8'));
|
||||
const checks = Array.isArray(report?.mutation?.checks) ? report.mutation.checks : [];
|
||||
-const scored = checks.filter((c) => typeof c.recall === 'number');
|
||||
+// Gate only the checks the oracle marked recall-gated. measure.mjs sets
|
||||
+// `recallGated: false` for cases a forward value-diff oracle cannot fairly
|
||||
+// score against the PDG slice: UPSTREAM fixtures (the oracle runs in its native
|
||||
+// downstream sense, so its behavioral AIS can never intersect a reverse slice —
|
||||
+// recall is 0 by construction) and id-discrimination corroboration fixtures.
|
||||
+// Those still carry a numeric `recall` for the report, so the legacy
|
||||
+// `typeof c.recall === 'number'` filter wrongly tripped the floor on them.
|
||||
+const scored = checks.filter((c) => c.recallGated === true && typeof c.recall === 'number');
|
||||
const recalls = scored.map((c) => c.recall);
|
||||
const min = recalls.length ? Math.min(...recalls) : null;
|
||||
const mean = recalls.length ? recalls.reduce((a, b) => a + b, 0) / recalls.length : null;
|
||||
@@ -39,6 +46,17 @@ if (process.env.GITHUB_STEP_SUMMARY) {
|
||||
}
|
||||
process.stdout.write(summary + '\n');
|
||||
|
||||
+// A report that produced checks but gated NONE of them has no recall signal:
|
||||
+// the floor check below would pass vacuously (`min === null`). Fail loudly so a
|
||||
+// degenerate corpus, or a harvest that silently emptied every behavioral AIS,
|
||||
+// surfaces as a red run instead of a green "scored cases: 0 of N".
|
||||
+if (checks.length > 0 && scored.length === 0) {
|
||||
+ console.error(
|
||||
+ `Mutation gate has no signal: 0 of ${checks.length} checks were recall-gated — refusing to pass.`,
|
||||
+ );
|
||||
+ process.exit(1);
|
||||
+}
|
||||
+
|
||||
if (min !== null && min < floor) {
|
||||
console.error(`Mutation recall regression: min realized recall ${fmt(min)} < floor ${floor}`);
|
||||
process.exit(1);
|
||||
diff --git a/gitnexus/bench/impact-pdg/measure.mjs b/gitnexus/bench/impact-pdg/measure.mjs
|
||||
index 65d3a7675..f6158f1da 100644
|
||||
--- a/gitnexus/bench/impact-pdg/measure.mjs
|
||||
+++ b/gitnexus/bench/impact-pdg/measure.mjs
|
||||
@@ -25,11 +25,16 @@
|
||||
* `repo-manager.getGlobalDir()` — it roots the registry; the per-repo DB
|
||||
* lands in `<fixtureCopy>/.gitnexus/`, so fixtures are copied to a temp
|
||||
* working dir to keep the source tree clean);
|
||||
- * 2. SHELL OUT to the real CLI as a child process:
|
||||
- * node --import tsx src/cli/index.ts analyze <copy> --pdg --skip-git --index-only
|
||||
- * (child-process isolation sidesteps `process.exit`; real `saveMeta` +
|
||||
- * `registerRepo` land in the temp home; workers spawn from `dist/`, so the
|
||||
- * harness builds `dist/` first — run `node scripts/build.js`);
|
||||
+ * 2. SHELL OUT to the real CLI as a child process (see `cliChildArgs`):
|
||||
+ * node dist/cli/index.js analyze <copy> --pdg --skip-git --index-only
|
||||
+ * preferring the BUILT `dist/` CLI when present — plain JS, no tsx, and the
|
||||
+ * parse workers it spawns also load from `dist/`. The mutation workflow
|
||||
+ * builds `dist/` first (`node scripts/build.js`); `node --import tsx
|
||||
+ * src/cli/index.ts` is NOT used because Node >=22.18 native type-stripping
|
||||
+ * breaks the `.js`->`.ts` entry resolution (ERR_MODULE_NOT_FOUND on
|
||||
+ * `lazy-action.js`). Build-free runs fall back to tsx's own CLI over src.
|
||||
+ * (Child-process isolation sidesteps `process.exit`; real `saveMeta` +
|
||||
+ * `registerRepo` land in the temp home);
|
||||
* 3. `new LocalBackend(); await init()` resolves the fixture via the REAL
|
||||
* registry (the parent process ALSO sets `GITNEXUS_HOME` so init reads the
|
||||
* temp registry, not the user's ~/.gitnexus);
|
||||
@@ -53,6 +58,7 @@ import os from 'node:os';
|
||||
import path from 'node:path';
|
||||
import crypto from 'node:crypto';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
+import { createRequire } from 'node:module';
|
||||
import { fileURLToPath, pathToFileURL } from 'node:url';
|
||||
|
||||
import {
|
||||
@@ -95,6 +101,33 @@ const REPO_ROOT = path.resolve(__dirname, '..', '..'); // gitnexus/
|
||||
const FIXTURES_DIR = path.join(__dirname, 'fixtures');
|
||||
const BASELINE_PATH = path.join(__dirname, 'baselines.json');
|
||||
const CLI_ENTRY = path.join(REPO_ROOT, 'src', 'cli', 'index.ts');
|
||||
+// Shipped CLI entry (package.json `bin`). PREFERRED for the child analyze: it's
|
||||
+// plain compiled JS, so the analyze process — AND the parse workers it spawns,
|
||||
+// which resolve relative to the running entry — load from `dist/` with no tsx in
|
||||
+// the loop. The build-free path below stays as a fallback.
|
||||
+const DIST_CLI = path.join(REPO_ROOT, 'dist', 'cli', 'index.js');
|
||||
+// Build-free fallback: tsx's OWN cli entry (resolved from this package), NOT
|
||||
+// `node --import tsx <entry>.ts`. On Node >=22.18 native TypeScript type-
|
||||
+// stripping is enabled by default and intercepts the `.ts` entry before tsx's
|
||||
+// `--import` resolve hook applies; native stripping does NOT remap `./foo.js`
|
||||
+// specifiers to `foo.ts` (tsx does), so `node --import tsx src/cli/index.ts`
|
||||
+// crashes resolving `./lazy-action.js` (ERR_MODULE_NOT_FOUND) on newer Node.
|
||||
+// The tsx CLI takes over module loading and is version-agnostic across the
|
||||
+// declared engines range (node >=22.0, where `--no-experimental-strip-types`
|
||||
+// is not a universally-recognized flag). Workers still spawn from src via tsx on
|
||||
+// this path, so it is only robust on the older Node devs run locally.
|
||||
+const TSX_CLI = createRequire(import.meta.url).resolve('tsx/cli');
|
||||
+
|
||||
+/**
|
||||
+ * Build the argv that runs the real CLI as a child of `process.execPath`.
|
||||
+ * Prefers the built `dist/` CLI (production-faithful, no tsx, dist workers) when
|
||||
+ * present — this is what the mutation workflow uses (it builds dist first). Falls
|
||||
+ * back to the tsx CLI over src for build-free local runs. Returns the args AFTER
|
||||
+ * the node binary, i.e. ready for `spawnSync(process.execPath, [...args])`.
|
||||
+ */
|
||||
+function cliChildArgs(rest) {
|
||||
+ return fs.existsSync(DIST_CLI) ? [DIST_CLI, ...rest] : [TSX_CLI, CLI_ENTRY, ...rest];
|
||||
+}
|
||||
|
||||
const SCOPES = ['intra', 'inter', 'mixed'];
|
||||
const MODES = ['callgraph', 'pdg'];
|
||||
@@ -138,7 +171,7 @@ async function analyzeAndImpact(fx, home, { pdgOn = true } = {}) {
|
||||
fs.cpSync(path.join(fx.dir, 'src'), path.join(work, 'src'), { recursive: true });
|
||||
|
||||
const env = { ...process.env, GITNEXUS_HOME: home };
|
||||
- const args = ['--import', 'tsx', CLI_ENTRY, 'analyze', work, '--skip-git', '--index-only'];
|
||||
+ const args = cliChildArgs(['analyze', work, '--skip-git', '--index-only']);
|
||||
if (pdgOn) args.push('--pdg');
|
||||
const an = spawnSync(process.execPath, args, {
|
||||
env,
|
||||
diff --git a/gitnexus/package-lock.json b/gitnexus/package-lock.json
|
||||
index b28f005d7..927fe93c3 100644
|
||||
--- a/gitnexus/package-lock.json
|
||||
+++ b/gitnexus/package-lock.json
|
||||
@@ -52,6 +52,10 @@
|
||||
"gitnexus": "dist/cli/index.js"
|
||||
},
|
||||
"devDependencies": {
|
||||
+ "@babel/generator": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/traverse": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
"@types/busboy": "^1.5.4",
|
||||
"@types/cli-progress": "^3.11.6",
|
||||
"@types/cors": "^2.8.17",
|
||||
@@ -76,10 +80,59 @@
|
||||
"typescript": "^6.0.3"
|
||||
}
|
||||
},
|
||||
+ "node_modules/@babel/code-frame": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/code-frame/-/code-frame-7.29.7.tgz",
|
||||
+ "integrity": "sha512-Aup7aUOfpbAUg2ROOJN6Iw5f9DMBlzu0mIkm/malLQFN/YQgO48wCj0Kxa3sEHJvPVFg7siR+qRInwXd2qhQKw==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/helper-validator-identifier": "^7.29.7",
|
||||
+ "js-tokens": "^4.0.0",
|
||||
+ "picocolors": "^1.1.1"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
+ "node_modules/@babel/code-frame/node_modules/js-tokens": {
|
||||
+ "version": "4.0.0",
|
||||
+ "resolved": "https://registry.npmjs.org/js-tokens/-/js-tokens-4.0.0.tgz",
|
||||
+ "integrity": "sha512-RdJUflcE3cUzKiMqQgsCu06FPu9UdIJO0beYbPhHN4k6apgJtifcoCtT9bcxOpYBtpD2kCM6Sbzg4CausW/PKQ==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT"
|
||||
+ },
|
||||
+ "node_modules/@babel/generator": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/generator/-/generator-7.29.7.tgz",
|
||||
+ "integrity": "sha512-DkXD5OJQaAQIdZ1bt3UZdEnHAn9Imd3IVBdX03UFe+ony9Ojw5pzr9YVKGDY1jt+Gcn/FnGkNf8r+Vj5NOJWtQ==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
+ "@jridgewell/gen-mapping": "^0.3.12",
|
||||
+ "@jridgewell/trace-mapping": "^0.3.28",
|
||||
+ "jsesc": "^3.0.2"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
+ "node_modules/@babel/helper-globals": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/helper-globals/-/helper-globals-7.29.7.tgz",
|
||||
+ "integrity": "sha512-3nQVUAtvkKH9zahfWgw96Jc/uFOmjACE1kQz82E2lqWmHBgjzbNlsC22nuQTfahmWeQtTq5nQ/4Nnd2A1wj4zA==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/@babel/helper-string-parser": {
|
||||
- "version": "7.27.1",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/helper-string-parser/-/helper-string-parser-7.27.1.tgz",
|
||||
- "integrity": "sha512-qMlSxKbpRlAridDExk92nSobyDdpPijUq2DW6oDnUqd0iOGxmQjyqhMIihI9+zv4LPyZdRje2cavWPbCbWm3eA==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/helper-string-parser/-/helper-string-parser-7.29.7.tgz",
|
||||
+ "integrity": "sha512-Pb5ijPrZ89GDH8223L4UP8i6QApWxs04RbPQJTeWDV0/keR2E36MeKnyr6LYmUUvqRRI+Iv87SuF1W6ErINzYw==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
@@ -87,9 +140,9 @@
|
||||
}
|
||||
},
|
||||
"node_modules/@babel/helper-validator-identifier": {
|
||||
- "version": "7.28.5",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/helper-validator-identifier/-/helper-validator-identifier-7.28.5.tgz",
|
||||
- "integrity": "sha512-qSs4ifwzKJSV39ucNjsvc6WVHs6b7S03sOh2OcHF9UHfVPqWWALUsNUVzhSBiItjRZoLHx7nIarVjqKVusUZ1Q==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/helper-validator-identifier/-/helper-validator-identifier-7.29.7.tgz",
|
||||
+ "integrity": "sha512-qehxGkRj55h/ff8EMaJ+cYhyaKlHIxqYDn682wQD7RNp9UujOQsHog2uS0r2vzr4pW+sXf90NeeayjcNaX3fFg==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
@@ -97,13 +150,13 @@
|
||||
}
|
||||
},
|
||||
"node_modules/@babel/parser": {
|
||||
- "version": "7.29.2",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/parser/-/parser-7.29.2.tgz",
|
||||
- "integrity": "sha512-4GgRzy/+fsBa72/RZVJmGKPmZu9Byn8o4MoLpmNe1m8ZfYnz5emHLQz3U4gLud6Zwl0RZIcgiLD7Uq7ySFuDLA==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/parser/-/parser-7.29.7.tgz",
|
||||
+ "integrity": "sha512-hnORnjP/1P/zFEndoeX+n+t1RwWRJiJpM/jO7FW32Kn9r5+sJB2JWOdYo4L6k78j15eCwY3Gm/7364B1EMwtNg==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
- "@babel/types": "^7.29.0"
|
||||
+ "@babel/types": "^7.29.7"
|
||||
},
|
||||
"bin": {
|
||||
"parser": "bin/babel-parser.js"
|
||||
@@ -112,15 +165,49 @@
|
||||
"node": ">=6.0.0"
|
||||
}
|
||||
},
|
||||
+ "node_modules/@babel/template": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/template/-/template-7.29.7.tgz",
|
||||
+ "integrity": "sha512-puq+Gf35oI24FeN11LkoUQFqv9uwNeWpxXZi/Ji3rRIoKAzKnxRaZ+Gkj0vKS9ZCiTESfng1N9LyOyXvo+m+Gg==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/code-frame": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
+ "node_modules/@babel/traverse": {
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/traverse/-/traverse-7.29.7.tgz",
|
||||
+ "integrity": "sha512-EhlfNQtZ+NK22w5BM61ciuiq1m58ed33Wr1Xan//ZRTy6hgjnwyCffRYwzsGXdASJSUJ1guZILsErh1eQcl+zw==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@babel/code-frame": "^7.29.7",
|
||||
+ "@babel/generator": "^7.29.7",
|
||||
+ "@babel/helper-globals": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/template": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
+ "debug": "^4.3.1"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6.9.0"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/@babel/types": {
|
||||
- "version": "7.29.0",
|
||||
- "resolved": "https://registry.npmjs.org/@babel/types/-/types-7.29.0.tgz",
|
||||
- "integrity": "sha512-LwdZHpScM4Qz8Xw2iKSzS+cfglZzJGvofQICy7W7v4caru4EaAmyUuO6BGrbyQ2mYV11W0U8j5mBhd14dd3B0A==",
|
||||
+ "version": "7.29.7",
|
||||
+ "resolved": "https://registry.npmjs.org/@babel/types/-/types-7.29.7.tgz",
|
||||
+ "integrity": "sha512-4zBIxpPzowiZpusoFkyGVwakdRJUyuH5PxQ/PrqghfdFWWasvnCdPfQXHrenDai+gyLARulZjZowCOj6fjT4pA==",
|
||||
"dev": true,
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
- "@babel/helper-string-parser": "^7.27.1",
|
||||
- "@babel/helper-validator-identifier": "^7.28.5"
|
||||
+ "@babel/helper-string-parser": "^7.29.7",
|
||||
+ "@babel/helper-validator-identifier": "^7.29.7"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=6.9.0"
|
||||
@@ -1128,6 +1215,17 @@
|
||||
"node": ">=18.0.0"
|
||||
}
|
||||
},
|
||||
+ "node_modules/@jridgewell/gen-mapping": {
|
||||
+ "version": "0.3.13",
|
||||
+ "resolved": "https://registry.npmjs.org/@jridgewell/gen-mapping/-/gen-mapping-0.3.13.tgz",
|
||||
+ "integrity": "sha512-2kkt/7niJ6MgEPxF0bYdQ6etZaA+fQvDcLKckhy1yIQOzaoKjBBjSj63/aLVjYE3qhRt5dvM+uUyfCg6UKCBbA==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "dependencies": {
|
||||
+ "@jridgewell/sourcemap-codec": "^1.5.0",
|
||||
+ "@jridgewell/trace-mapping": "^0.3.24"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/@jridgewell/resolve-uri": {
|
||||
"version": "3.1.2",
|
||||
"resolved": "https://registry.npmjs.org/@jridgewell/resolve-uri/-/resolve-uri-3.1.2.tgz",
|
||||
@@ -1471,9 +1569,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1491,9 +1586,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1511,9 +1603,6 @@
|
||||
"ppc64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1531,9 +1620,6 @@
|
||||
"s390x"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1551,9 +1637,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -1571,9 +1654,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MIT",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3487,6 +3567,19 @@
|
||||
"js-yaml": "bin/js-yaml.js"
|
||||
}
|
||||
},
|
||||
+ "node_modules/jsesc": {
|
||||
+ "version": "3.1.0",
|
||||
+ "resolved": "https://registry.npmjs.org/jsesc/-/jsesc-3.1.0.tgz",
|
||||
+ "integrity": "sha512-/sM3dO2FOzXjKQhJuo0Q173wf2KOo8t4I8vHy6lF9poUp7bKT0/NHE8fPX23PwfhnykfqnC2xRxOnVw5XuGIaA==",
|
||||
+ "dev": true,
|
||||
+ "license": "MIT",
|
||||
+ "bin": {
|
||||
+ "jsesc": "bin/jsesc"
|
||||
+ },
|
||||
+ "engines": {
|
||||
+ "node": ">=6"
|
||||
+ }
|
||||
+ },
|
||||
"node_modules/json-bignum": {
|
||||
"version": "0.0.3",
|
||||
"resolved": "https://registry.npmjs.org/json-bignum/-/json-bignum-0.0.3.tgz",
|
||||
@@ -3668,9 +3761,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3692,9 +3782,6 @@
|
||||
"arm64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3716,9 +3803,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "glibc"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
@@ -3740,9 +3824,6 @@
|
||||
"x64"
|
||||
],
|
||||
"dev": true,
|
||||
- "libc": [
|
||||
- "musl"
|
||||
- ],
|
||||
"license": "MPL-2.0",
|
||||
"optional": true,
|
||||
"os": [
|
||||
diff --git a/gitnexus/package.json b/gitnexus/package.json
|
||||
index 6fa0c9007..82f2cabfd 100644
|
||||
--- a/gitnexus/package.json
|
||||
+++ b/gitnexus/package.json
|
||||
@@ -94,6 +94,10 @@
|
||||
"uuid": "^14.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
+ "@babel/generator": "^7.29.7",
|
||||
+ "@babel/parser": "^7.29.7",
|
||||
+ "@babel/traverse": "^7.29.7",
|
||||
+ "@babel/types": "^7.29.7",
|
||||
"@types/busboy": "^1.5.4",
|
||||
"@types/cli-progress": "^3.11.6",
|
||||
"@types/cors": "^2.8.17",
|
||||
1096
eval/workflow_bench/review_cases/pr-2718.patch
Normal file
1096
eval/workflow_bench/review_cases/pr-2718.patch
Normal file
File diff suppressed because it is too large
Load diff
1725
eval/workflow_bench/review_cases/pr-2773.patch
Normal file
1725
eval/workflow_bench/review_cases/pr-2773.patch
Normal file
File diff suppressed because it is too large
Load diff
817
eval/workflow_bench/review_cases/pr-2794.patch
Normal file
817
eval/workflow_bench/review_cases/pr-2794.patch
Normal file
|
|
@ -0,0 +1,817 @@
|
|||
diff --git a/.github/workflows/ci-tests.yml b/.github/workflows/ci-tests.yml
|
||||
index b4fbb3d56..6d03b4c73 100644
|
||||
--- a/.github/workflows/ci-tests.yml
|
||||
+++ b/.github/workflows/ci-tests.yml
|
||||
@@ -500,6 +500,17 @@ jobs:
|
||||
run: node --import tsx bench/callable-value-flow/measure.mjs --check
|
||||
working-directory: gitnexus
|
||||
|
||||
+ - name: C++ qualified-namespace resolution guards (#2788)
|
||||
+ # Build-free: asserts resolveCppQualifiedNamespaceMember resolves an
|
||||
+ # unchanged symbol set (fingerprint) and that per-call-site cost stays
|
||||
+ # independent of corpus size. It rescanned every parsed file per
|
||||
+ # qualified `ns::member()` call site until #2788 — 25 of 33 analyze
|
||||
+ # minutes on a 1.5k-file C++ repo. #1990 had already fixed the exact
|
||||
+ # same bug in the sibling ADL path and shipped without a scaling gate,
|
||||
+ # which is how the bug class came back; this step is that gate.
|
||||
+ run: node --import tsx bench/cpp-qualified-ns/measure.mjs --check
|
||||
+ working-directory: gitnexus
|
||||
+
|
||||
- name: Receiver-resolution drop guards
|
||||
# NOT build-free: this one runs the real pipeline, so it needs dist/
|
||||
# (the setup action above builds). ~2m15s.
|
||||
diff --git a/gitnexus/bench/cpp-qualified-ns/baselines.json b/gitnexus/bench/cpp-qualified-ns/baselines.json
|
||||
new file mode 100644
|
||||
index 000000000..01d1cad17
|
||||
--- /dev/null
|
||||
+++ b/gitnexus/bench/cpp-qualified-ns/baselines.json
|
||||
@@ -0,0 +1,6 @@
|
||||
+{
|
||||
+ "_comment": "Baselines for bench/cpp-qualified-ns/measure.mjs --check (#2788). `fingerprint` is a sha256 over every `receiver::member -> outcome` the synthetic corpus resolves at the LARGE scale (hit nodeId, `<ambiguous>` per #1564, or `<none>`); it is a CORRECTNESS gate, so drift means C++ qualified `ns::member()` lookup started resolving a different symbol set and must be explained, never re-baselined to make CI green. `scaling_budget` is a timing gate and carries deliberate headroom for shared CI runners.",
|
||||
+ "fingerprint": "30122b086b90fe1d56edc5acfd39ab7f26c00350f0e672123aaf607c7702e1e2",
|
||||
+ "scaling_budget": 1.8,
|
||||
+ "_scaling_note": "(t_large/t_small)/(1600/400). ~1.0 is linear; measured 0.93-1.21 on the indexed implementation. The pre-#2788 per-call-site workspace rescan measured 3.45 on the same corpus shape (at a reduced 100/400 scale, since 1600 files x 32k call sites of quadratic work does not finish in a CI step) — so the budget sits between the two bands and cannot be met by reintroducing the scan. Resolution is timed alone; the fingerprint's outcome strings are built in a separate untimed pass because their allocation cost grows with the corpus and would otherwise show up as scaling."
|
||||
+}
|
||||
diff --git a/gitnexus/bench/cpp-qualified-ns/measure.mjs b/gitnexus/bench/cpp-qualified-ns/measure.mjs
|
||||
new file mode 100644
|
||||
index 000000000..36da04b7a
|
||||
--- /dev/null
|
||||
+++ b/gitnexus/bench/cpp-qualified-ns/measure.mjs
|
||||
@@ -0,0 +1,258 @@
|
||||
+/**
|
||||
+ * Build-free scaling + identity bench for `resolveCppQualifiedNamespaceMember`,
|
||||
+ * the C++ qualified `ns::member()` receiver resolver (issue #2788).
|
||||
+ *
|
||||
+ * Before #2788 this function re-scanned EVERY parsed file — rebuilding a
|
||||
+ * per-file `scopesById` map each time — once per qualified call site, so the
|
||||
+ * scope-resolution emit phase cost O(callsites × scopes). On a 1,473-file C++
|
||||
+ * repo that was 25.3 min of a 33-min analyze, with 75% of total self-time in
|
||||
+ * this one function. It is the same bug #1990 had already fixed in the sibling
|
||||
+ * ADL path (`pickCppAdlCandidates` → `AdlCandidateIndex`) — the sibling shipped
|
||||
+ * without a scaling gate, and the bug class came straight back here. Hence this
|
||||
+ * bench: a per-call-site workspace scan must not be reintroduced silently.
|
||||
+ *
|
||||
+ * For a synthetic corpus at two scales it reports:
|
||||
+ * - `elapsed_ms` per scale (fastest of REPS, see `fastest`) for resolving
|
||||
+ * every call site once, INCLUDING the one-time index build — that build is
|
||||
+ * the work the per-site scan was traded for, so hiding it would let an
|
||||
+ * index that is itself quadratic pass;
|
||||
+ * - a scaling ratio `(t_large/t_small)/(LARGE/SMALL)`: ~1.0 linear,
|
||||
+ * ~4.x quadratic at this scale gap;
|
||||
+ * - a sha256 fingerprint over every `receiver::member → outcome` the corpus
|
||||
+ * resolves, as the correctness gate. A fingerprint change means qualified
|
||||
+ * lookup started resolving different symbols — a behaviour change, never a
|
||||
+ * performance one.
|
||||
+ *
|
||||
+ * Build-free: imports the `.ts` hotpath through tsx
|
||||
+ * (`node --import tsx bench/cpp-qualified-ns/measure.mjs`).
|
||||
+ *
|
||||
+ * Without args: prints the JSON report.
|
||||
+ * With `--check`: asserts the fingerprint == the committed baseline AND the
|
||||
+ * scaling ratio is within budget; exits non-zero on drift/regression.
|
||||
+ */
|
||||
+import fs from 'node:fs';
|
||||
+import path from 'node:path';
|
||||
+import crypto from 'node:crypto';
|
||||
+import { fileURLToPath } from 'node:url';
|
||||
+
|
||||
+import {
|
||||
+ clearCppInlineNamespaces,
|
||||
+ markCppInlineNamespaceRange,
|
||||
+ populateCppInlineNamespaceScopes,
|
||||
+ resolveCppQualifiedNamespaceMember,
|
||||
+} from '../../src/core/ingestion/languages/cpp/inline-namespaces.ts';
|
||||
+
|
||||
+const __dirname = path.dirname(fileURLToPath(import.meta.url));
|
||||
+const BASELINE_PATH = path.resolve(__dirname, 'baselines.json');
|
||||
+
|
||||
+const SMALL = 400;
|
||||
+const LARGE = 1600;
|
||||
+const CALLS_PER_FILE = 20;
|
||||
+const REPS = 7;
|
||||
+const WARMUP = 3;
|
||||
+
|
||||
+const NO_SCOPES = {};
|
||||
+
|
||||
+/**
|
||||
+ * Deterministic synthetic corpus — no randomness, so the fingerprint is stable.
|
||||
+ *
|
||||
+ * Per file, one `ns_f` namespace shaped like the ABI-versioning idiom the
|
||||
+ * reporter's repo uses (`namespace x { inline namespace v { … } }`):
|
||||
+ *
|
||||
+ * namespace ns_f {
|
||||
+ * void own0(); void own1(); // direct members
|
||||
+ * inline namespace v1 { void inl0(); void dup(); } // transitively visible
|
||||
+ * inline namespace v2 { void dup(); } // → dup is ambiguous
|
||||
+ * namespace detail { void hidden0(); } // NOT inline → invisible
|
||||
+ * }
|
||||
+ *
|
||||
+ * The three outcome classes all appear, because each takes a different exit
|
||||
+ * from the resolver and only exercising the hit path would let a regression in
|
||||
+ * the miss path (the most common outcome in real source) go unmeasured:
|
||||
+ * resolved hit, `'ambiguous'` (#1564), and `undefined` (miss — both a wrong
|
||||
+ * member name and a non-inline nested member).
|
||||
+ */
|
||||
+function buildCorpus(fileCount) {
|
||||
+ const parsedFiles = [];
|
||||
+ for (let f = 0; f < fileCount; f++) {
|
||||
+ const filePath = `src/file${f}.cpp`;
|
||||
+ const scopes = [];
|
||||
+ let line = 1;
|
||||
+ const scope = (id, kind, parent, defs) => {
|
||||
+ const entry = {
|
||||
+ id,
|
||||
+ kind,
|
||||
+ parent,
|
||||
+ ownedDefs: defs,
|
||||
+ range: { startLine: line, startCol: 0, endLine: line + 1, endCol: 0 },
|
||||
+ };
|
||||
+ line += 2;
|
||||
+ scopes.push(entry);
|
||||
+ return entry;
|
||||
+ };
|
||||
+ const def = (type, qualifiedName) => ({
|
||||
+ nodeId: `def:${filePath}#${qualifiedName}`,
|
||||
+ type,
|
||||
+ qualifiedName,
|
||||
+ });
|
||||
+
|
||||
+ const nsId = `sc:${f}:ns`;
|
||||
+ scope(nsId, 'Namespace', null, [
|
||||
+ def('Namespace', `ns_${f}`),
|
||||
+ def('Function', `ns_${f}.own0`),
|
||||
+ def('Function', `ns_${f}.own1`),
|
||||
+ ]);
|
||||
+ const v1 = scope(`sc:${f}:v1`, 'Namespace', nsId, [
|
||||
+ def('Namespace', `ns_${f}.v1`),
|
||||
+ def('Function', `ns_${f}.v1.inl0`),
|
||||
+ def('Function', `ns_${f}.v1.dup`),
|
||||
+ ]);
|
||||
+ const v2 = scope(`sc:${f}:v2`, 'Namespace', nsId, [
|
||||
+ def('Namespace', `ns_${f}.v2`),
|
||||
+ def('Function', `ns_${f}.v2.dup`),
|
||||
+ ]);
|
||||
+ scope(`sc:${f}:detail`, 'Namespace', nsId, [
|
||||
+ def('Namespace', `ns_${f}.detail`),
|
||||
+ def('Function', `ns_${f}.detail.hidden0`),
|
||||
+ ]);
|
||||
+
|
||||
+ parsedFiles.push({ filePath, scopes, inlineRanges: [v1.range, v2.range] });
|
||||
+ }
|
||||
+ return parsedFiles;
|
||||
+}
|
||||
+
|
||||
+/** Capture-time inline marking + `populateOwners`-time scope-id resolution, in
|
||||
+ * the same order the pipeline runs them. Must re-run after every
|
||||
+ * `clearCppInlineNamespaces`, which drops both the marks and the index. */
|
||||
+function populateInlineState(parsedFiles) {
|
||||
+ clearCppInlineNamespaces();
|
||||
+ for (const parsed of parsedFiles) {
|
||||
+ for (const range of parsed.inlineRanges) markCppInlineNamespaceRange(parsed.filePath, range);
|
||||
+ populateCppInlineNamespaceScopes(parsed);
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/** The call sites: a deterministic spread over receivers and member names so
|
||||
+ * each rep resolves the same set, with hits, misses and ambiguities mixed. */
|
||||
+const MEMBERS = ['own0', 'inl0', 'dup', 'hidden0', 'nosuch'];
|
||||
+function callSites(fileCount) {
|
||||
+ const sites = [];
|
||||
+ for (let f = 0; f < fileCount; f++) {
|
||||
+ for (let c = 0; c < CALLS_PER_FILE; c++) {
|
||||
+ sites.push([`ns_${(f * 7 + c * 13) % fileCount}`, MEMBERS[c % MEMBERS.length]]);
|
||||
+ }
|
||||
+ }
|
||||
+ return sites;
|
||||
+}
|
||||
+
|
||||
+/** The timed loop: resolution only. The outcome strings the fingerprint needs
|
||||
+ * are built in a separate untimed pass (`outcomesOf`), so their allocation
|
||||
+ * cost — which grows with the corpus and would inflate the scaling ratio on
|
||||
+ * its own — never lands in the measurement. `sink` keeps the calls live. */
|
||||
+function resolveAll(parsedFiles, sites) {
|
||||
+ let sink = 0;
|
||||
+ for (const [receiver, member] of sites) {
|
||||
+ const hit = resolveCppQualifiedNamespaceMember(receiver, member, parsedFiles, NO_SCOPES);
|
||||
+ if (hit !== undefined) sink++;
|
||||
+ }
|
||||
+ return sink;
|
||||
+}
|
||||
+
|
||||
+function outcomesOf(parsedFiles, sites) {
|
||||
+ const outcomes = [];
|
||||
+ for (const [receiver, member] of sites) {
|
||||
+ const hit = resolveCppQualifiedNamespaceMember(receiver, member, parsedFiles, NO_SCOPES);
|
||||
+ outcomes.push(
|
||||
+ `${receiver}::${member}\u0000${hit === undefined ? '<none>' : hit === 'ambiguous' ? '<ambiguous>' : hit.nodeId}`,
|
||||
+ );
|
||||
+ }
|
||||
+ return outcomes;
|
||||
+}
|
||||
+
|
||||
+/**
|
||||
+ * MIN, not median — same rationale as bench/callable-value-flow: both scales
|
||||
+ * are timed in one process and every error source (scheduler preemption, GC, a
|
||||
+ * noisy neighbour on a shared CI runner) is additive, so the fastest observed
|
||||
+ * run is the closest estimate of the uncontended cost and keeps the derived
|
||||
+ * ratio comparable across machines.
|
||||
+ */
|
||||
+function fastest(values) {
|
||||
+ return Math.min(...values);
|
||||
+}
|
||||
+
|
||||
+/** Time one full pass: index build (lazy, on the first call) + every call
|
||||
+ * site. The corpus state is reset OUTSIDE the timer so the reset's own
|
||||
+ * O(files) cost never lands in the measurement. */
|
||||
+function timeResolution(parsedFiles, sites) {
|
||||
+ for (let w = 0; w < WARMUP; w++) {
|
||||
+ populateInlineState(parsedFiles);
|
||||
+ resolveAll(parsedFiles, sites);
|
||||
+ }
|
||||
+ const samples = [];
|
||||
+ for (let r = 0; r < REPS; r++) {
|
||||
+ populateInlineState(parsedFiles);
|
||||
+ const t0 = performance.now();
|
||||
+ resolveAll(parsedFiles, sites);
|
||||
+ samples.push(performance.now() - t0);
|
||||
+ }
|
||||
+ return { ms: fastest(samples), outcomes: outcomesOf(parsedFiles, sites) };
|
||||
+}
|
||||
+
|
||||
+function fingerprint(outcomes) {
|
||||
+ return crypto
|
||||
+ .createHash('sha256')
|
||||
+ .update([...outcomes].sort().join('\n'))
|
||||
+ .digest('hex');
|
||||
+}
|
||||
+
|
||||
+const scales = {};
|
||||
+for (const [name, fileCount] of [
|
||||
+ ['small', SMALL],
|
||||
+ ['large', LARGE],
|
||||
+]) {
|
||||
+ const parsedFiles = buildCorpus(fileCount);
|
||||
+ const sites = callSites(fileCount);
|
||||
+ const { ms, outcomes } = timeResolution(parsedFiles, sites);
|
||||
+ scales[name] = {
|
||||
+ files: fileCount,
|
||||
+ call_sites: sites.length,
|
||||
+ ms: Number(ms.toFixed(3)),
|
||||
+ fingerprint: fingerprint(outcomes),
|
||||
+ };
|
||||
+}
|
||||
+
|
||||
+const scalingRatio = scales.large.ms / scales.small.ms / (LARGE / SMALL);
|
||||
+
|
||||
+const report = {
|
||||
+ small: scales.small,
|
||||
+ large: scales.large,
|
||||
+ scaling_ratio: Number(scalingRatio.toFixed(3)),
|
||||
+ fingerprint: scales.large.fingerprint,
|
||||
+};
|
||||
+
|
||||
+if (!process.argv.includes('--check')) {
|
||||
+ console.log(JSON.stringify(report, null, 2));
|
||||
+ process.exit(0);
|
||||
+}
|
||||
+
|
||||
+const baseline = JSON.parse(fs.readFileSync(BASELINE_PATH, 'utf-8'));
|
||||
+const failures = [];
|
||||
+if (report.fingerprint !== baseline.fingerprint) {
|
||||
+ failures.push(
|
||||
+ `fingerprint drift: ${report.fingerprint} != ${baseline.fingerprint} — qualified ` +
|
||||
+ `namespace lookup resolved a DIFFERENT symbol set. This is a behaviour change, not a perf one.`,
|
||||
+ );
|
||||
+}
|
||||
+if (report.scaling_ratio > baseline.scaling_budget) {
|
||||
+ failures.push(
|
||||
+ `scaling ${report.scaling_ratio} > budget ${baseline.scaling_budget} — per-call-site cost ` +
|
||||
+ `now grows with corpus size again (#2788).`,
|
||||
+ );
|
||||
+}
|
||||
+
|
||||
+console.log(JSON.stringify(report, null, 2));
|
||||
+if (failures.length > 0) {
|
||||
+ console.error(`[cpp-qualified-ns --check] FAIL\n - ${failures.join('\n - ')}`);
|
||||
+ process.exit(1);
|
||||
+}
|
||||
+console.log('[cpp-qualified-ns --check] PASS');
|
||||
diff --git a/gitnexus/src/core/ingestion/languages/cpp/inline-namespaces.ts b/gitnexus/src/core/ingestion/languages/cpp/inline-namespaces.ts
|
||||
index 134738dfd..ad2749928 100644
|
||||
--- a/gitnexus/src/core/ingestion/languages/cpp/inline-namespaces.ts
|
||||
+++ b/gitnexus/src/core/ingestion/languages/cpp/inline-namespaces.ts
|
||||
@@ -14,12 +14,14 @@
|
||||
* 2. **Transitive qualified visibility.** `outer::foo()` resolves to
|
||||
* `outer::v1::foo()` when `v1` is inline. The qualified-namespace
|
||||
* receiver resolver (`resolveCppQualifiedNamespaceMember`) walks
|
||||
- * inline-namespace children transitively when collecting candidates.
|
||||
+ * inline-namespace children transitively when collecting candidates —
|
||||
+ * once per pipeline run, into {@link QualifiedNsMemberIndex} (#2788).
|
||||
*
|
||||
* State lifecycle: capture-time `markCppInlineNamespaceRange` records each
|
||||
* inline namespace's source range; `populateCppInlineNamespaceScopes`
|
||||
* resolves ranges to `ScopeId`s during `populateOwners`. Cleared via
|
||||
- * `clearCppInlineNamespaces`, called from `clearFileLocalNames`.
|
||||
+ * `clearCppInlineNamespaces`, called from
|
||||
+ * `cppScopeResolver.loadResolutionConfig` at the start of every pass.
|
||||
*
|
||||
* STL idiom this enables: `std::__1::vector` (libc++) and `std::__cxx11`
|
||||
* (libstdc++) are inline namespaces of `std`. With this support,
|
||||
@@ -85,10 +87,13 @@ export function applyCppInlineNamespaceSideChannel(
|
||||
for (const r of ranges) set.add(r);
|
||||
}
|
||||
|
||||
-/** Clear all inline-namespace state. Called from `clearFileLocalNames`. */
|
||||
+/** Clear all inline-namespace state. Called from
|
||||
+ * `cppScopeResolver.loadResolutionConfig` at the start of every pass. */
|
||||
export function clearCppInlineNamespaces(): void {
|
||||
inlineNamespaceRangesByFile.clear();
|
||||
inlineNamespaceScopeIds.clear();
|
||||
+ qualifiedNsIndex = undefined;
|
||||
+ qualifiedNsIndexSource = undefined;
|
||||
}
|
||||
|
||||
/** Resolve captured ranges to actual ScopeIds by matching scope ranges
|
||||
@@ -115,11 +120,136 @@ export function isCppInlineNamespaceScope(scopeId: ScopeId): boolean {
|
||||
}
|
||||
|
||||
/**
|
||||
- * Walk every parsed file looking for a Namespace scope whose qualified
|
||||
- * name matches `receiverName`, collect its callable ownedDefs matching
|
||||
- * `memberName`, transitively descending into any inline-namespace
|
||||
- * children (since they're members of the enclosing namespace under ISO
|
||||
- * C++).
|
||||
+ * Qualified-namespace member index — built **once** per pipeline run from
|
||||
+ * `parsedFiles` and reused by every qualified call site.
|
||||
+ *
|
||||
+ * The legacy lookup re-scanned every parsed file (rebuilding a per-file
|
||||
+ * `scopesById` map each time) once **per qualified call site**, making the
|
||||
+ * scope-resolution emit phase O(callsites × scopes): 25.3 min of a 33-min
|
||||
+ * analyze on a 1,473-file C++ repo, 75% of total self-time in this one
|
||||
+ * function (#2788). Mirrors the same fix #1990 applied to ADL
|
||||
+ * (`pickCppAdlCandidates` → {@link AdlCandidateIndex}); per-site cost drops
|
||||
+ * to two Map lookups.
|
||||
+ *
|
||||
+ * `byReceiver`: namespace simple name → member simple name → callable defs,
|
||||
+ * in the exact order the legacy linear scan produced them (file-major; within
|
||||
+ * a file, `parsed.scopes` declaration order; within a namespace, own
|
||||
+ * `ownedDefs` before inline-namespace children, depth-first). Ordering is
|
||||
+ * load-bearing: the caller returns `allHits[0]` for the single-hit case and
|
||||
+ * `narrowOverloadCandidates` is first-wins.
|
||||
+ */
|
||||
+interface QualifiedNsMemberIndex {
|
||||
+ readonly byReceiver: ReadonlyMap<string, ReadonlyMap<string, readonly SymbolDefinition[]>>;
|
||||
+}
|
||||
+
|
||||
+type NsScope = ParsedFile['scopes'][number];
|
||||
+
|
||||
+let qualifiedNsIndex: QualifiedNsMemberIndex | undefined;
|
||||
+let qualifiedNsIndexSource: readonly ParsedFile[] | undefined;
|
||||
+
|
||||
+/** Build the index in a single pass over the workspace. Visitation order
|
||||
+ * mirrors the legacy scan exactly (see {@link QualifiedNsMemberIndex}). */
|
||||
+function buildQualifiedNsMemberIndex(parsedFiles: readonly ParsedFile[]): QualifiedNsMemberIndex {
|
||||
+ const byReceiver = new Map<string, Map<string, SymbolDefinition[]>>();
|
||||
+ // Legacy dedup was a per-call `seenNodeId` set spanning all files; since a
|
||||
+ // def only ever lands in one `(receiver, member)` bucket, a per-receiver set
|
||||
+ // keyed `member \0 nodeId` reproduces it. Only reachable at all via
|
||||
+ // same-name inline nesting (`namespace ns { inline namespace ns { … } }`),
|
||||
+ // but kept so a def is never double-counted into `'ambiguous'`.
|
||||
+ const seenByReceiver = new Map<string, Set<string>>();
|
||||
+
|
||||
+ for (const parsed of parsedFiles) {
|
||||
+ // parent → inline-namespace children. The legacy transitive walk filtered
|
||||
+ // `scopesById.values()` by `parent` per recursion step — O(scopes) each,
|
||||
+ // and O(scopes²) per file overall; this is the same order, built once.
|
||||
+ const inlineChildrenByParent = new Map<ScopeId, (typeof parsed.scopes)[number][]>();
|
||||
+ for (const sc of parsed.scopes) {
|
||||
+ if (sc.parent === null) continue;
|
||||
+ if (sc.kind !== 'Namespace') continue;
|
||||
+ if (!inlineNamespaceScopeIds.has(sc.id)) continue;
|
||||
+ let kids = inlineChildrenByParent.get(sc.parent);
|
||||
+ if (kids === undefined) {
|
||||
+ kids = [];
|
||||
+ inlineChildrenByParent.set(sc.parent, kids);
|
||||
+ }
|
||||
+ kids.push(sc);
|
||||
+ }
|
||||
+
|
||||
+ for (const scope of parsed.scopes) {
|
||||
+ if (scope.kind !== 'Namespace') continue;
|
||||
+ const nsDef = findNamespaceDefInScope(scope);
|
||||
+ if (nsDef === undefined) continue;
|
||||
+ const nsName = nsDef.qualifiedName?.split('.').pop() ?? nsDef.qualifiedName ?? '';
|
||||
+ let byMember = byReceiver.get(nsName);
|
||||
+ if (byMember === undefined) {
|
||||
+ byMember = new Map();
|
||||
+ byReceiver.set(nsName, byMember);
|
||||
+ }
|
||||
+ let seen = seenByReceiver.get(nsName);
|
||||
+ if (seen === undefined) {
|
||||
+ seen = new Set();
|
||||
+ seenByReceiver.set(nsName, seen);
|
||||
+ }
|
||||
+ collectNamespaceMembers(scope, inlineChildrenByParent, byMember, seen);
|
||||
+ }
|
||||
+ }
|
||||
+ return { byReceiver };
|
||||
+}
|
||||
+
|
||||
+/** Bucket a namespace scope's callable `ownedDefs` by member simple name,
|
||||
+ * then descend into inline-namespace children — the index-build twin of the
|
||||
+ * legacy `findMemberInNamespaceTransitive`, collecting every member name in
|
||||
+ * one walk instead of one walk per `(call site, member name)`. */
|
||||
+function collectNamespaceMembers(
|
||||
+ scope: NsScope,
|
||||
+ inlineChildrenByParent: ReadonlyMap<ScopeId, readonly NsScope[]>,
|
||||
+ byMember: Map<string, SymbolDefinition[]>,
|
||||
+ seen: Set<string>,
|
||||
+): void {
|
||||
+ for (const def of scope.ownedDefs) {
|
||||
+ if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
|
||||
+ const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName ?? '';
|
||||
+ const dedupKey = `${simple}\u0000${def.nodeId}`;
|
||||
+ if (seen.has(dedupKey)) continue;
|
||||
+ seen.add(dedupKey);
|
||||
+ let arr = byMember.get(simple);
|
||||
+ if (arr === undefined) {
|
||||
+ arr = [];
|
||||
+ byMember.set(simple, arr);
|
||||
+ }
|
||||
+ arr.push(def);
|
||||
+ }
|
||||
+ for (const child of inlineChildrenByParent.get(scope.id) ?? []) {
|
||||
+ collectNamespaceMembers(child, inlineChildrenByParent, byMember, seen);
|
||||
+ }
|
||||
+}
|
||||
+
|
||||
+/** Build the index on first use of a given `parsedFiles` set; reuse it for
|
||||
+ * every subsequent call site in the same pipeline run.
|
||||
+ *
|
||||
+ * The index is a function of TWO inputs: `parsedFiles` and the module-level
|
||||
+ * `inlineNamespaceScopeIds` (which inline children get descended into).
|
||||
+ * Reference identity on `parsedFiles` alone is sound here because
|
||||
+ * `populateCppInlineNamespaceScopes` fills `inlineNamespaceScopeIds` during
|
||||
+ * `populateOwners` — strictly before any resolution pass calls in — and
|
||||
+ * {@link clearCppInlineNamespaces} drops the index at the start of every
|
||||
+ * pass. Any future caller that mutates `inlineNamespaceScopeIds` mid-pass
|
||||
+ * while reusing the same `parsedFiles` reference MUST call
|
||||
+ * `clearCppInlineNamespaces` in between. Same contract as `ensureAdlIndex`. */
|
||||
+function qualifiedNsMemberIndex(parsedFiles: readonly ParsedFile[]): QualifiedNsMemberIndex {
|
||||
+ if (qualifiedNsIndex === undefined || qualifiedNsIndexSource !== parsedFiles) {
|
||||
+ qualifiedNsIndex = buildQualifiedNsMemberIndex(parsedFiles);
|
||||
+ qualifiedNsIndexSource = parsedFiles;
|
||||
+ }
|
||||
+ return qualifiedNsIndex;
|
||||
+}
|
||||
+
|
||||
+/**
|
||||
+ * Find the Namespace scopes whose simple name matches `receiverName` and
|
||||
+ * return their callable members matching `memberName`, transitively
|
||||
+ * including inline-namespace children (since they're members of the
|
||||
+ * enclosing namespace under ISO C++). Served from a per-pipeline index
|
||||
+ * ({@link QualifiedNsMemberIndex}), not a per-call-site workspace scan.
|
||||
*
|
||||
* Returns the most specific (innermost) match — for `outer::foo()`
|
||||
* where `inline namespace v1` declares `foo`, returns `v1::foo`. When
|
||||
@@ -134,27 +264,8 @@ export function resolveCppQualifiedNamespaceMember(
|
||||
_scopes: ScopeResolutionIndexes,
|
||||
callsite?: Callsite,
|
||||
): SymbolDefinition | 'ambiguous' | undefined {
|
||||
- const allHits: SymbolDefinition[] = [];
|
||||
- const seenNodeId = new Set<string>();
|
||||
- for (const parsed of parsedFiles) {
|
||||
- const scopesById = new Map<ScopeId, (typeof parsed.scopes)[number]>();
|
||||
- for (const sc of parsed.scopes) scopesById.set(sc.id, sc);
|
||||
- for (const scope of parsed.scopes) {
|
||||
- if (scope.kind !== 'Namespace') continue;
|
||||
- const nsDef = findNamespaceDefInScope(scope);
|
||||
- if (nsDef === undefined) continue;
|
||||
- const nsName = nsDef.qualifiedName?.split('.').pop() ?? nsDef.qualifiedName ?? '';
|
||||
- if (nsName !== receiverName) continue;
|
||||
- // Found a matching namespace scope in this file. Collect ALL
|
||||
- // members transitively through any inline-namespace children.
|
||||
- const hits = findMemberInNamespaceTransitive(scope, scopesById, memberName);
|
||||
- for (const hit of hits) {
|
||||
- if (seenNodeId.has(hit.nodeId)) continue;
|
||||
- seenNodeId.add(hit.nodeId);
|
||||
- allHits.push(hit);
|
||||
- }
|
||||
- }
|
||||
- }
|
||||
+ const allHits =
|
||||
+ qualifiedNsMemberIndex(parsedFiles).byReceiver.get(receiverName)?.get(memberName) ?? [];
|
||||
if (allHits.length === 0) return undefined;
|
||||
if (allHits.length === 1) return allHits[0];
|
||||
|
||||
@@ -182,46 +293,6 @@ export function resolveCppQualifiedNamespaceMember(
|
||||
return 'ambiguous';
|
||||
}
|
||||
|
||||
-/** Recursively search a namespace scope and any inline-namespace
|
||||
- * descendants for callable defs with the given simple name. Non-inline
|
||||
- * nested namespaces are NOT traversed — they require explicit
|
||||
- * qualification (`outer::nested::foo`). Returns ALL matches so the
|
||||
- * caller can detect same-name ambiguity across inline children (#1564). */
|
||||
-function findMemberInNamespaceTransitive(
|
||||
- scope: {
|
||||
- readonly id: ScopeId;
|
||||
- readonly ownedDefs: readonly SymbolDefinition[];
|
||||
- readonly parent: ScopeId | null;
|
||||
- },
|
||||
- scopesById: ReadonlyMap<
|
||||
- ScopeId,
|
||||
- {
|
||||
- readonly id: ScopeId;
|
||||
- readonly kind: string;
|
||||
- readonly parent: ScopeId | null;
|
||||
- readonly ownedDefs: readonly SymbolDefinition[];
|
||||
- }
|
||||
- >,
|
||||
- memberName: string,
|
||||
-): SymbolDefinition[] {
|
||||
- const results: SymbolDefinition[] = [];
|
||||
- // Check this scope's own ownedDefs first.
|
||||
- for (const def of scope.ownedDefs) {
|
||||
- if (def.type !== 'Function' && def.type !== 'Method' && def.type !== 'Constructor') continue;
|
||||
- const simple = def.qualifiedName?.split('.').pop() ?? def.qualifiedName ?? '';
|
||||
- if (simple === memberName) results.push(def);
|
||||
- }
|
||||
- // Descend into inline-namespace children.
|
||||
- for (const childScope of scopesById.values()) {
|
||||
- if (childScope.parent !== scope.id) continue;
|
||||
- if (childScope.kind !== 'Namespace') continue;
|
||||
- if (!inlineNamespaceScopeIds.has(childScope.id)) continue;
|
||||
- const childHits = findMemberInNamespaceTransitive(childScope, scopesById, memberName);
|
||||
- for (const hit of childHits) results.push(hit);
|
||||
- }
|
||||
- return results;
|
||||
-}
|
||||
-
|
||||
function findNamespaceDefInScope(scope: {
|
||||
readonly ownedDefs: readonly SymbolDefinition[];
|
||||
}): SymbolDefinition | undefined {
|
||||
diff --git a/gitnexus/test/unit/cpp-qualified-ns-index.test.ts b/gitnexus/test/unit/cpp-qualified-ns-index.test.ts
|
||||
new file mode 100644
|
||||
index 000000000..a8af02384
|
||||
--- /dev/null
|
||||
+++ b/gitnexus/test/unit/cpp-qualified-ns-index.test.ts
|
||||
@@ -0,0 +1,258 @@
|
||||
+/**
|
||||
+ * #2788 — `resolveCppQualifiedNamespaceMember` serves qualified `ns::member()`
|
||||
+ * lookups from a per-pipeline index instead of rescanning every parsed file per
|
||||
+ * call site. These tests pin the two properties the index must not lose:
|
||||
+ *
|
||||
+ * 1. Transitive inline-namespace collection, ordering and same-name
|
||||
+ * ambiguity (#1564) — the semantics the old linear scan provided.
|
||||
+ * 2. Cache invalidation — a new `parsedFiles` array, or a
|
||||
+ * `clearCppInlineNamespaces()` between passes, must not serve stale hits.
|
||||
+ * This is the failure mode the index introduces; nothing else covers it.
|
||||
+ */
|
||||
+
|
||||
+import type {
|
||||
+ ParsedFile,
|
||||
+ ScopeId,
|
||||
+ ScopeResolutionIndexes,
|
||||
+ SymbolDefinition,
|
||||
+} from 'gitnexus-shared';
|
||||
+import { beforeEach, describe, expect, it } from 'vitest';
|
||||
+import {
|
||||
+ clearCppInlineNamespaces,
|
||||
+ markCppInlineNamespaceRange,
|
||||
+ populateCppInlineNamespaceScopes,
|
||||
+ resolveCppQualifiedNamespaceMember,
|
||||
+} from '../../src/core/ingestion/languages/cpp/inline-namespaces.js';
|
||||
+
|
||||
+const NO_SCOPES = {} as unknown as ScopeResolutionIndexes;
|
||||
+
|
||||
+interface ScopeSpec {
|
||||
+ readonly id: string;
|
||||
+ readonly kind: 'Namespace' | 'Module';
|
||||
+ readonly parent: string | null;
|
||||
+ readonly defs: readonly SymbolDefinition[];
|
||||
+ /** Distinguishes each scope's range so inline marking targets exactly one. */
|
||||
+ readonly line: number;
|
||||
+}
|
||||
+
|
||||
+function def(nodeId: string, type: string, qualifiedName: string): SymbolDefinition {
|
||||
+ return { nodeId, type, qualifiedName } as unknown as SymbolDefinition;
|
||||
+}
|
||||
+
|
||||
+function nsDef(nodeId: string, qualifiedName: string): SymbolDefinition {
|
||||
+ return def(nodeId, 'Namespace', qualifiedName);
|
||||
+}
|
||||
+
|
||||
+function fnDef(nodeId: string, qualifiedName: string): SymbolDefinition {
|
||||
+ return def(nodeId, 'Function', qualifiedName);
|
||||
+}
|
||||
+
|
||||
+function range(line: number): {
|
||||
+ startLine: number;
|
||||
+ startCol: number;
|
||||
+ endLine: number;
|
||||
+ endCol: number;
|
||||
+} {
|
||||
+ return { startLine: line, startCol: 0, endLine: line + 1, endCol: 0 };
|
||||
+}
|
||||
+
|
||||
+/** Build a single-file `parsedFiles` array from scope specs, marking the
|
||||
+ * scopes named in `inlineIds` as inline namespaces (capture-time range mark +
|
||||
+ * `populateOwners`-time scope-id resolution, same order as the pipeline). */
|
||||
+function makeParsedFiles(
|
||||
+ filePath: string,
|
||||
+ specs: readonly ScopeSpec[],
|
||||
+ inlineIds: readonly string[],
|
||||
+): readonly ParsedFile[] {
|
||||
+ const parsed = {
|
||||
+ filePath,
|
||||
+ scopes: specs.map((s) => ({
|
||||
+ id: s.id as unknown as ScopeId,
|
||||
+ kind: s.kind,
|
||||
+ parent: s.parent as unknown as ScopeId | null,
|
||||
+ ownedDefs: s.defs,
|
||||
+ range: range(s.line),
|
||||
+ })),
|
||||
+ } as unknown as ParsedFile;
|
||||
+ markInline(parsed, specs, inlineIds);
|
||||
+ return [parsed];
|
||||
+}
|
||||
+
|
||||
+/** Capture-time inline marking + `populateOwners`-time scope-id resolution,
|
||||
+ * in the same order the pipeline runs them. */
|
||||
+function markInline(
|
||||
+ parsed: ParsedFile,
|
||||
+ specs: readonly ScopeSpec[],
|
||||
+ inlineIds: readonly string[],
|
||||
+): void {
|
||||
+ for (const id of inlineIds) {
|
||||
+ const spec = specs.find((s) => s.id === id);
|
||||
+ if (spec === undefined) throw new Error(`inline scope ${id} must exist`);
|
||||
+ markCppInlineNamespaceRange(parsed.filePath, range(spec.line));
|
||||
+ }
|
||||
+ populateCppInlineNamespaceScopes(parsed);
|
||||
+}
|
||||
+
|
||||
+/** `namespace outer { <ownDefs> inline namespace v1 { <inlineDefs> } }` */
|
||||
+function outerWithInlineChild(
|
||||
+ filePath: string,
|
||||
+ ownDefs: readonly SymbolDefinition[],
|
||||
+ inlineDefs: readonly SymbolDefinition[],
|
||||
+): readonly ParsedFile[] {
|
||||
+ return makeParsedFiles(
|
||||
+ filePath,
|
||||
+ [
|
||||
+ {
|
||||
+ id: 'sc:outer',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: null,
|
||||
+ defs: [nsDef('n:outer', 'outer'), ...ownDefs],
|
||||
+ line: 1,
|
||||
+ },
|
||||
+ {
|
||||
+ id: 'sc:v1',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: 'sc:outer',
|
||||
+ defs: [nsDef('n:v1', 'outer.v1'), ...inlineDefs],
|
||||
+ line: 10,
|
||||
+ },
|
||||
+ ],
|
||||
+ ['sc:v1'],
|
||||
+ );
|
||||
+}
|
||||
+
|
||||
+describe('C++ qualified-namespace member index (#2788)', () => {
|
||||
+ beforeEach(() => {
|
||||
+ clearCppInlineNamespaces();
|
||||
+ });
|
||||
+
|
||||
+ it('resolves outer::foo through an inline-namespace child', () => {
|
||||
+ const files = outerWithInlineChild('a.cpp', [], [fnDef('n:foo@v1', 'outer.v1.foo')]);
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', files, NO_SCOPES)).toMatchObject({
|
||||
+ nodeId: 'n:foo@v1',
|
||||
+ });
|
||||
+ });
|
||||
+
|
||||
+ it('returns undefined for an unknown namespace or member', () => {
|
||||
+ const files = outerWithInlineChild('a.cpp', [], [fnDef('n:foo@v1', 'outer.v1.foo')]);
|
||||
+ expect(resolveCppQualifiedNamespaceMember('nope', 'foo', files, NO_SCOPES)).toBeUndefined();
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'nope', files, NO_SCOPES)).toBeUndefined();
|
||||
+ });
|
||||
+
|
||||
+ it('does not descend into a non-inline nested namespace', () => {
|
||||
+ const files = makeParsedFiles(
|
||||
+ 'a.cpp',
|
||||
+ [
|
||||
+ {
|
||||
+ id: 'sc:outer',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: null,
|
||||
+ defs: [nsDef('n:outer', 'outer')],
|
||||
+ line: 1,
|
||||
+ },
|
||||
+ {
|
||||
+ id: 'sc:nested',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: 'sc:outer',
|
||||
+ defs: [nsDef('n:nested', 'outer.nested'), fnDef('n:foo@nested', 'outer.nested.foo')],
|
||||
+ line: 10,
|
||||
+ },
|
||||
+ ],
|
||||
+ [],
|
||||
+ );
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', files, NO_SCOPES)).toBeUndefined();
|
||||
+ expect(resolveCppQualifiedNamespaceMember('nested', 'foo', files, NO_SCOPES)).toMatchObject({
|
||||
+ nodeId: 'n:foo@nested',
|
||||
+ });
|
||||
+ });
|
||||
+
|
||||
+ it('reports same-name hits across two inline children as ambiguous (#1564)', () => {
|
||||
+ const files = makeParsedFiles(
|
||||
+ 'a.cpp',
|
||||
+ [
|
||||
+ {
|
||||
+ id: 'sc:outer',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: null,
|
||||
+ defs: [nsDef('n:outer', 'outer')],
|
||||
+ line: 1,
|
||||
+ },
|
||||
+ {
|
||||
+ id: 'sc:v1',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: 'sc:outer',
|
||||
+ defs: [nsDef('n:v1', 'outer.v1'), fnDef('n:foo@v1', 'outer.v1.foo')],
|
||||
+ line: 10,
|
||||
+ },
|
||||
+ {
|
||||
+ id: 'sc:v2',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: 'sc:outer',
|
||||
+ defs: [nsDef('n:v2', 'outer.v2'), fnDef('n:foo@v2', 'outer.v2.foo')],
|
||||
+ line: 20,
|
||||
+ },
|
||||
+ ],
|
||||
+ ['sc:v1', 'sc:v2'],
|
||||
+ );
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', files, NO_SCOPES)).toBe('ambiguous');
|
||||
+ });
|
||||
+
|
||||
+ it('keeps scan order: a namespace-owned def precedes its inline child hits', () => {
|
||||
+ // Two candidates with no call-site info narrow to 'ambiguous', so order is
|
||||
+ // asserted through the single-hit path: only the own def is present here,
|
||||
+ // and the inline child contributes a different member name.
|
||||
+ const files = outerWithInlineChild(
|
||||
+ 'a.cpp',
|
||||
+ [fnDef('n:foo@outer', 'outer.foo')],
|
||||
+ [fnDef('n:bar@v1', 'outer.v1.bar')],
|
||||
+ );
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', files, NO_SCOPES)).toMatchObject({
|
||||
+ nodeId: 'n:foo@outer',
|
||||
+ });
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'bar', files, NO_SCOPES)).toMatchObject({
|
||||
+ nodeId: 'n:bar@v1',
|
||||
+ });
|
||||
+ });
|
||||
+
|
||||
+ it('does not serve one parsedFiles array’s index to another', () => {
|
||||
+ const first = outerWithInlineChild('a.cpp', [], [fnDef('n:foo@a', 'outer.v1.foo')]);
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', first, NO_SCOPES)).toMatchObject({
|
||||
+ nodeId: 'n:foo@a',
|
||||
+ });
|
||||
+ const second = outerWithInlineChild('b.cpp', [], [fnDef('n:foo@b', 'outer.v1.foo')]);
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', second, NO_SCOPES)).toMatchObject({
|
||||
+ nodeId: 'n:foo@b',
|
||||
+ });
|
||||
+ });
|
||||
+
|
||||
+ it('rebuilds after clearCppInlineNamespaces even when parsedFiles is reused', () => {
|
||||
+ // Pass 1: `v1` is inline, so `outer::foo` reaches through it.
|
||||
+ const specs: readonly ScopeSpec[] = [
|
||||
+ {
|
||||
+ id: 'sc:outer',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: null,
|
||||
+ defs: [nsDef('n:outer', 'outer')],
|
||||
+ line: 1,
|
||||
+ },
|
||||
+ {
|
||||
+ id: 'sc:v1',
|
||||
+ kind: 'Namespace',
|
||||
+ parent: 'sc:outer',
|
||||
+ defs: [nsDef('n:v1', 'outer.v1'), fnDef('n:foo@v1', 'outer.v1.foo')],
|
||||
+ line: 10,
|
||||
+ },
|
||||
+ ];
|
||||
+ const files = makeParsedFiles('a.cpp', specs, ['sc:v1']);
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', files, NO_SCOPES)).toMatchObject({
|
||||
+ nodeId: 'n:foo@v1',
|
||||
+ });
|
||||
+
|
||||
+ // Pass 2: SAME `parsedFiles` reference (so identity alone would serve the
|
||||
+ // cached index), but `v1` is no longer inline. Without the index reset in
|
||||
+ // `clearCppInlineNamespaces` the stale pass-1 hit survives.
|
||||
+ clearCppInlineNamespaces();
|
||||
+ markInline(files[0], specs, []);
|
||||
+ expect(resolveCppQualifiedNamespaceMember('outer', 'foo', files, NO_SCOPES)).toBeUndefined();
|
||||
+ });
|
||||
+});
|
||||
313
eval/workflow_bench/review_scoring.py
Normal file
313
eval/workflow_bench/review_scoring.py
Normal file
|
|
@ -0,0 +1,313 @@
|
|||
"""Strict review artifacts and deterministic hidden-label scoring."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import math
|
||||
import posixpath
|
||||
from dataclasses import astuple, dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Mapping, Sequence
|
||||
|
||||
from .oracle_assets import TaskOracleSnapshot
|
||||
|
||||
REVIEW_OUTPUT = "review-output.json"
|
||||
REVIEW_SCHEMA_VERSION = 1
|
||||
MAX_REVIEW_BYTES = 256 * 1024
|
||||
MAX_FINDINGS = 100
|
||||
SEVERITIES = ("critical", "high", "medium", "low")
|
||||
BLOCKING_SEVERITIES = frozenset({"critical", "high"})
|
||||
SEVERITY_WEIGHT = {"critical": 8.0, "high": 5.0, "medium": 2.0, "low": 1.0}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ReviewFinding:
|
||||
finding_id: str
|
||||
severity: str
|
||||
path: str
|
||||
line: int
|
||||
end_line: int
|
||||
category: str
|
||||
scenario: str
|
||||
evidence: str
|
||||
recommendation: str
|
||||
blocking: bool
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ExpectedFinding:
|
||||
finding_id: str
|
||||
severity: str
|
||||
path: str
|
||||
line_start: int
|
||||
line_end: int
|
||||
category: str
|
||||
|
||||
|
||||
def _nonblank(value: Any, label: str, *, limit: int = 4096) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ValueError(f"{label} must be a nonblank string")
|
||||
text = value.strip()
|
||||
if len(text.encode()) > limit:
|
||||
raise ValueError(f"{label} exceeds {limit} bytes")
|
||||
return text
|
||||
|
||||
|
||||
def _relative_path(value: Any, label: str) -> str:
|
||||
path = _nonblank(value, label, limit=512).replace("\\", "/")
|
||||
normalized = posixpath.normpath(path)
|
||||
if normalized.startswith("/") or normalized in {".", ".."} or normalized.startswith("../"):
|
||||
raise ValueError(f"{label} must be a repository-relative path")
|
||||
return normalized
|
||||
|
||||
|
||||
def _positive_int(value: Any, label: str) -> int:
|
||||
if isinstance(value, bool) or not isinstance(value, int) or value < 1:
|
||||
raise ValueError(f"{label} must be a positive integer")
|
||||
return value
|
||||
|
||||
|
||||
def _severity(value: Any, label: str) -> str:
|
||||
severity = _nonblank(value, label, limit=16).casefold()
|
||||
if severity not in SEVERITIES:
|
||||
raise ValueError(f"{label} must be one of {', '.join(SEVERITIES)}")
|
||||
return severity
|
||||
|
||||
|
||||
def _parse_review_finding(raw: Any, index: int) -> ReviewFinding:
|
||||
if not isinstance(raw, Mapping):
|
||||
raise ValueError(f"findings[{index}] must be an object")
|
||||
required = {
|
||||
"id",
|
||||
"severity",
|
||||
"path",
|
||||
"line",
|
||||
"end_line",
|
||||
"category",
|
||||
"scenario",
|
||||
"evidence",
|
||||
"recommendation",
|
||||
"blocking",
|
||||
}
|
||||
if set(raw) != required:
|
||||
raise ValueError(f"findings[{index}] requires exactly {sorted(required)}")
|
||||
line = _positive_int(raw["line"], f"findings[{index}].line")
|
||||
end_line = _positive_int(raw["end_line"], f"findings[{index}].end_line")
|
||||
if end_line < line:
|
||||
raise ValueError(f"findings[{index}].end_line precedes line")
|
||||
if not isinstance(raw["blocking"], bool):
|
||||
raise ValueError(f"findings[{index}].blocking must be boolean")
|
||||
return ReviewFinding(
|
||||
finding_id=_nonblank(raw["id"], f"findings[{index}].id", limit=128),
|
||||
severity=_severity(raw["severity"], f"findings[{index}].severity"),
|
||||
path=_relative_path(raw["path"], f"findings[{index}].path"),
|
||||
line=line,
|
||||
end_line=end_line,
|
||||
category=_nonblank(raw["category"], f"findings[{index}].category", limit=128).casefold(),
|
||||
scenario=_nonblank(raw["scenario"], f"findings[{index}].scenario"),
|
||||
evidence=_nonblank(raw["evidence"], f"findings[{index}].evidence"),
|
||||
recommendation=_nonblank(raw["recommendation"], f"findings[{index}].recommendation"),
|
||||
blocking=raw["blocking"],
|
||||
)
|
||||
|
||||
|
||||
def parse_review_output(path: Path) -> tuple[str, tuple[ReviewFinding, ...]]:
|
||||
metadata = path.lstat()
|
||||
if path.is_symlink() or not path.is_file() or metadata.st_size > MAX_REVIEW_BYTES:
|
||||
raise ValueError("review output must be a bounded regular non-symlink file")
|
||||
try:
|
||||
raw = json.loads(path.read_text())
|
||||
except (OSError, UnicodeError, json.JSONDecodeError) as exc:
|
||||
raise ValueError("review output is not valid UTF-8 JSON") from exc
|
||||
if not isinstance(raw, Mapping) or set(raw) != {"schema_version", "verdict", "findings"}:
|
||||
raise ValueError("review output requires exactly schema_version, verdict, and findings")
|
||||
if raw["schema_version"] != REVIEW_SCHEMA_VERSION:
|
||||
raise ValueError(f"review output schema_version must be {REVIEW_SCHEMA_VERSION}")
|
||||
verdict = _nonblank(raw["verdict"], "verdict", limit=32).casefold()
|
||||
if verdict not in {"approve", "comment", "request_changes"}:
|
||||
raise ValueError("verdict must be approve, comment, or request_changes")
|
||||
findings_raw = raw["findings"]
|
||||
if not isinstance(findings_raw, list) or len(findings_raw) > MAX_FINDINGS:
|
||||
raise ValueError(f"findings must be a list of at most {MAX_FINDINGS} items")
|
||||
findings = tuple(_parse_review_finding(item, index) for index, item in enumerate(findings_raw))
|
||||
ids = [item.finding_id for item in findings]
|
||||
if len(set(ids)) != len(ids):
|
||||
raise ValueError("review finding ids must be unique")
|
||||
if verdict == "approve" and findings:
|
||||
raise ValueError("approve verdict cannot contain findings")
|
||||
if verdict == "request_changes" and not any(item.blocking for item in findings):
|
||||
raise ValueError("request_changes requires at least one blocking finding")
|
||||
return verdict, findings
|
||||
|
||||
|
||||
def expected_findings(snapshot: TaskOracleSnapshot) -> tuple[ExpectedFinding, ...]:
|
||||
matches = [item for item in snapshot.files if item.target == "review-labels.json"]
|
||||
if len(matches) != 1:
|
||||
raise ValueError("review task oracle requires exactly one review-labels.json")
|
||||
try:
|
||||
raw = json.loads(matches[0].payload)
|
||||
except (UnicodeError, json.JSONDecodeError) as exc:
|
||||
raise ValueError("hidden review labels are malformed") from exc
|
||||
if not isinstance(raw, Mapping) or set(raw) != {"schema_version", "findings"}:
|
||||
raise ValueError("hidden review labels require schema_version and findings")
|
||||
if raw["schema_version"] != REVIEW_SCHEMA_VERSION or not isinstance(raw["findings"], list):
|
||||
raise ValueError("hidden review labels have an unsupported schema")
|
||||
labels: list[ExpectedFinding] = []
|
||||
for index, item in enumerate(raw["findings"]):
|
||||
if not isinstance(item, Mapping):
|
||||
raise ValueError(f"hidden findings[{index}] must be an object")
|
||||
required = {"id", "severity", "path", "line_start", "line_end", "category"}
|
||||
if set(item) != required:
|
||||
raise ValueError(f"hidden findings[{index}] requires exactly {sorted(required)}")
|
||||
start = _positive_int(item["line_start"], f"hidden findings[{index}].line_start")
|
||||
end = _positive_int(item["line_end"], f"hidden findings[{index}].line_end")
|
||||
if end < start:
|
||||
raise ValueError(f"hidden findings[{index}] has an inverted range")
|
||||
labels.append(
|
||||
ExpectedFinding(
|
||||
finding_id=_nonblank(item["id"], f"hidden findings[{index}].id", limit=128),
|
||||
severity=_severity(item["severity"], f"hidden findings[{index}].severity"),
|
||||
path=_relative_path(item["path"], f"hidden findings[{index}].path"),
|
||||
line_start=start,
|
||||
line_end=end,
|
||||
category=_nonblank(item["category"], f"hidden findings[{index}].category", limit=128).casefold(),
|
||||
)
|
||||
)
|
||||
if len({item.finding_id for item in labels}) != len(labels):
|
||||
raise ValueError("hidden review finding ids must be unique")
|
||||
return tuple(labels)
|
||||
|
||||
|
||||
def _overlaps(actual: ReviewFinding, expected: ExpectedFinding) -> bool:
|
||||
return actual.line <= expected.line_end and actual.end_line >= expected.line_start
|
||||
|
||||
|
||||
def _match_score(actual: ReviewFinding, expected: ExpectedFinding) -> tuple[int, int, int] | None:
|
||||
if actual.path != expected.path or not _overlaps(actual, expected):
|
||||
return None
|
||||
category = int(actual.category == expected.category)
|
||||
severity = int(actual.severity == expected.severity)
|
||||
distance = abs(actual.line - expected.line_start)
|
||||
return category, severity, -distance
|
||||
|
||||
|
||||
def _assign_pairs(
|
||||
candidates: list[tuple[tuple[int, int, int, int, str, str, int], int, int]],
|
||||
) -> list[tuple[int, int]]:
|
||||
"""Maximum cardinality at every size, with deterministic ranked traversal."""
|
||||
|
||||
edges: dict[int, list[int]] = {}
|
||||
for _score, actual_index, expected_index in candidates:
|
||||
edges.setdefault(actual_index, []).append(expected_index)
|
||||
owners: dict[int, int] = {}
|
||||
|
||||
def augment(actual_index: int, seen: set[int]) -> bool:
|
||||
for expected_index in edges[actual_index]:
|
||||
if expected_index in seen:
|
||||
continue
|
||||
seen.add(expected_index)
|
||||
previous = owners.get(expected_index)
|
||||
if previous is None or augment(previous, seen):
|
||||
owners[expected_index] = actual_index
|
||||
return True
|
||||
return False
|
||||
|
||||
for actual_index in sorted(edges):
|
||||
augment(actual_index, set())
|
||||
return sorted((actual_index, expected_index) for expected_index, actual_index in owners.items())
|
||||
|
||||
|
||||
def score_review(
|
||||
verdict: str,
|
||||
actual: Sequence[ReviewFinding],
|
||||
expected: Sequence[ExpectedFinding],
|
||||
) -> dict[str, Any]:
|
||||
actual = sorted(actual, key=astuple)
|
||||
expected = sorted(expected, key=astuple)
|
||||
candidates: list[tuple[tuple[int, int, int, int, str, str, int], int, int]] = []
|
||||
for actual_index, finding in enumerate(actual):
|
||||
for expected_index, label in enumerate(expected):
|
||||
score = _match_score(finding, label)
|
||||
if score is not None:
|
||||
# Content keys, not list indices: JSON finding order must not
|
||||
# change which pairs the matching selects.
|
||||
expected_span = label.line_end - label.line_start
|
||||
candidates.append(
|
||||
(
|
||||
(*score, -expected_span, finding.path, finding.category, finding.line),
|
||||
actual_index,
|
||||
expected_index,
|
||||
)
|
||||
)
|
||||
candidates.sort(reverse=True)
|
||||
pairs = _assign_pairs(candidates)
|
||||
matched_actual = {actual_index for actual_index, _expected_index in pairs}
|
||||
|
||||
tp = len(pairs)
|
||||
fp = len(actual) - tp
|
||||
fn = len(expected) - tp
|
||||
precision = tp / (tp + fp) if tp + fp else None
|
||||
recall = tp / (tp + fn) if tp + fn else None
|
||||
f1 = (
|
||||
2 * precision * recall / (precision + recall)
|
||||
if precision is not None and recall is not None and precision + recall
|
||||
else (0.0 if expected else None)
|
||||
)
|
||||
expected_weight = sum(SEVERITY_WEIGHT[item.severity] for item in expected)
|
||||
matched_weight = sum(
|
||||
min(SEVERITY_WEIGHT[actual[a].severity], SEVERITY_WEIGHT[expected[e].severity]) for a, e in pairs
|
||||
)
|
||||
fp_weight = sum(SEVERITY_WEIGHT[actual[a].severity] for a in range(len(actual)) if a not in matched_actual)
|
||||
weighted_precision = matched_weight / (matched_weight + fp_weight) if matched_weight + fp_weight else None
|
||||
weighted_recall = matched_weight / expected_weight if expected_weight else None
|
||||
weighted_f1 = (
|
||||
2 * weighted_precision * weighted_recall / (weighted_precision + weighted_recall)
|
||||
if weighted_precision is not None and weighted_recall is not None and weighted_precision + weighted_recall
|
||||
else (0.0 if expected else None)
|
||||
)
|
||||
blockers = [index for index, item in enumerate(expected) if item.severity in BLOCKING_SEVERITIES]
|
||||
matched_blockers = {e for a, e in pairs if actual[a].severity in BLOCKING_SEVERITIES}
|
||||
blocker_recall = sum(index in matched_blockers for index in blockers) / len(blockers) if blockers else None
|
||||
severity_accuracy = (
|
||||
sum(actual[a].severity == expected[e].severity for a, e in pairs) / tp if tp else None
|
||||
)
|
||||
category_accuracy = (
|
||||
sum(actual[a].category == expected[e].category for a, e in pairs) / tp if tp else None
|
||||
)
|
||||
grounded = (
|
||||
sum(bool(item.path and item.line > 0 and item.evidence.strip()) for item in actual) / len(actual)
|
||||
if actual
|
||||
else None
|
||||
)
|
||||
correct_verdict = (not expected and verdict == "approve") or (
|
||||
bool(expected)
|
||||
and verdict == ("request_changes" if any(item.severity in BLOCKING_SEVERITIES for item in expected) else "comment")
|
||||
)
|
||||
def rounded(value: float | None) -> float | None:
|
||||
return None if value is None else round(value, 6)
|
||||
return {
|
||||
"true_positives": tp,
|
||||
"false_positives": fp,
|
||||
"false_negatives": fn,
|
||||
"precision": rounded(precision),
|
||||
"recall": rounded(recall),
|
||||
"f1": rounded(f1),
|
||||
"weighted_precision": rounded(weighted_precision),
|
||||
"weighted_recall": rounded(weighted_recall),
|
||||
"weighted_f1": rounded(weighted_f1),
|
||||
"blocker_recall": rounded(blocker_recall),
|
||||
"severity_accuracy": rounded(severity_accuracy),
|
||||
"category_accuracy": rounded(category_accuracy),
|
||||
"grounded_evidence": rounded(grounded),
|
||||
"verdict_correct": correct_verdict,
|
||||
"clean_control": not expected,
|
||||
"clean_pass": not expected and not actual and verdict == "approve",
|
||||
"score_finite": all(
|
||||
value is None or math.isfinite(float(value))
|
||||
for key, value in {
|
||||
"precision": precision,
|
||||
"recall": recall,
|
||||
"weighted_f1": weighted_f1,
|
||||
}.items()
|
||||
),
|
||||
}
|
||||
224
eval/workflow_bench/run-evolution.sh
Executable file
224
eval/workflow_bench/run-evolution.sh
Executable file
|
|
@ -0,0 +1,224 @@
|
|||
#!/usr/bin/env bash
|
||||
# Local and CI entrypoint for the propose → benchmark → gate loop.
|
||||
#
|
||||
# The GitHub skill-evolution job must call this script. Do not inline
|
||||
# `python -m workflow_bench.evolve` in the workflow; that argv lives here so a
|
||||
# laptop run and a self-hosted run cannot drift.
|
||||
#
|
||||
# Usage:
|
||||
# ./workflow_bench/run-evolution.sh # local, no --apply
|
||||
# ./workflow_bench/run-evolution.sh --apply # CI / mutate working tree
|
||||
# ./workflow_bench/run-evolution.sh --dry-run # print argv, no model calls
|
||||
# PROVIDER=openai ./workflow_bench/run-evolution.sh
|
||||
#
|
||||
# Environment (same names the workflow already sets):
|
||||
# MODEL PROPOSER_MODEL EFFORT GENERATIONS RUNS WORKERS PROVIDER
|
||||
# EVOLUTION_PROFILE CE_PLUGIN_DIR CE_PLUGIN_VERSION
|
||||
# INCLUDE_EXPENSIVE SEED_RESULTS CLAUDE_BIN OUT_ROOT
|
||||
# UNSAFE_NO_BWRAP=1 (local review diagnostics only)
|
||||
# GITNEXUS_BENCH_ANTHROPIC_API_KEY (legacy GITNEXUS_BENCH_AUTH_TOKEN)
|
||||
# GITNEXUS_BENCH_OPENAI_API_KEY
|
||||
set -euo pipefail
|
||||
|
||||
usage() {
|
||||
sed -n '2,18p' "$0" | sed 's/^# \{0,1\}//'
|
||||
}
|
||||
|
||||
script_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
eval_dir="$(cd "${script_dir}/.." && pwd)"
|
||||
apply=0
|
||||
dry_run=0
|
||||
passthrough=()
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--apply)
|
||||
apply=1
|
||||
shift
|
||||
;;
|
||||
--dry-run)
|
||||
dry_run=1
|
||||
shift
|
||||
;;
|
||||
-h | --help)
|
||||
usage
|
||||
exit 0
|
||||
;;
|
||||
--)
|
||||
shift
|
||||
passthrough+=("$@")
|
||||
break
|
||||
;;
|
||||
*)
|
||||
passthrough+=("$1")
|
||||
shift
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
MODEL="${MODEL:-gpt-5.6-sol}"
|
||||
PROPOSER_MODEL="${PROPOSER_MODEL:-gpt-5.6-sol}"
|
||||
EFFORT="${EFFORT:-xhigh}"
|
||||
GENERATIONS="${GENERATIONS:-1}"
|
||||
RUNS="${RUNS:-3}"
|
||||
WORKERS="${WORKERS:-1}"
|
||||
PROVIDER="${PROVIDER:-openai}"
|
||||
INCLUDE_EXPENSIVE="${INCLUDE_EXPENSIVE:-}"
|
||||
SEED_RESULTS="${SEED_RESULTS:-}"
|
||||
EVOLUTION_PROFILE="${EVOLUTION_PROFILE:-review}"
|
||||
CE_PLUGIN_DIR="${CE_PLUGIN_DIR:-}"
|
||||
CE_PLUGIN_VERSION="${CE_PLUGIN_VERSION:-}"
|
||||
anthropic_key="${GITNEXUS_BENCH_ANTHROPIC_API_KEY:-${GITNEXUS_BENCH_AUTH_TOKEN:-}}"
|
||||
openai_key="${GITNEXUS_BENCH_OPENAI_API_KEY:-}"
|
||||
|
||||
route_openai=0
|
||||
case "${PROVIDER}" in
|
||||
openai)
|
||||
route_openai=1
|
||||
;;
|
||||
anthropic)
|
||||
route_openai=0
|
||||
;;
|
||||
auto)
|
||||
if [[ -z "${anthropic_key}" && -n "${openai_key}" ]]; then
|
||||
route_openai=1
|
||||
fi
|
||||
;;
|
||||
*)
|
||||
echo "Unknown PROVIDER '${PROVIDER}' (expected auto, openai, or anthropic)." >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
|
||||
if ((route_openai)); then
|
||||
if [[ "${MODEL}" == claude-* ]]; then
|
||||
echo "Routing Claude Code through a loopback OpenAI gateway with model gpt-5.6-sol. Set MODEL to override." >&2
|
||||
MODEL=gpt-5.6-sol
|
||||
fi
|
||||
if [[ "${PROPOSER_MODEL}" == claude-* ]]; then
|
||||
PROPOSER_MODEL=gpt-5.6-sol
|
||||
fi
|
||||
fi
|
||||
|
||||
if ((dry_run == 0)); then
|
||||
if [[ "${PROVIDER}" == openai && -z "${openai_key}" ]]; then
|
||||
echo "PROVIDER=openai requires GITNEXUS_BENCH_OPENAI_API_KEY." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ "${PROVIDER}" == anthropic && -z "${anthropic_key}" ]]; then
|
||||
echo "PROVIDER=anthropic requires GITNEXUS_BENCH_ANTHROPIC_API_KEY (legacy GITNEXUS_BENCH_AUTH_TOKEN is accepted)." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ -z "${anthropic_key}" && -z "${openai_key}" ]]; then
|
||||
echo "Set GITNEXUS_BENCH_ANTHROPIC_API_KEY and/or GITNEXUS_BENCH_OPENAI_API_KEY before a real run." >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
claude_bin="${CLAUDE_BIN:-}"
|
||||
if [[ -z "${claude_bin}" && -n "${RUNNER_TEMP:-}" ]]; then
|
||||
canary="${RUNNER_TEMP}/claude-canary/node_modules/@anthropic-ai/claude-code-linux-x64/claude"
|
||||
if [[ -x "${canary}" ]]; then
|
||||
claude_bin="${canary}"
|
||||
fi
|
||||
fi
|
||||
if [[ -z "${claude_bin}" ]] && command -v claude >/dev/null 2>&1; then
|
||||
claude_bin="$(command -v claude)"
|
||||
fi
|
||||
if [[ -z "${claude_bin}" ]]; then
|
||||
claude_bin=claude
|
||||
fi
|
||||
if ((dry_run == 0)) && ! command -v "${claude_bin}" >/dev/null 2>&1 && [[ ! -x "${claude_bin}" ]]; then
|
||||
echo "Claude Code binary not found (${claude_bin}). Set CLAUDE_BIN or install claude on PATH." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ -n "${OUT_ROOT:-}" ]]; then
|
||||
out_root="${OUT_ROOT}"
|
||||
elif [[ -n "${RUNNER_TEMP:-}" ]]; then
|
||||
out_root="${RUNNER_TEMP}/wfevolve"
|
||||
else
|
||||
out_root="${eval_dir}/results/wfevolve-$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
fi
|
||||
|
||||
cmd=(
|
||||
uv run --locked --extra dev python -m workflow_bench.evolve
|
||||
--model "${MODEL}"
|
||||
--proposer-model "${PROPOSER_MODEL}"
|
||||
--effort "${EFFORT}"
|
||||
--generations "${GENERATIONS}"
|
||||
--runs "${RUNS}"
|
||||
--workers "${WORKERS}"
|
||||
--claude-bin "${claude_bin}"
|
||||
--out-root "${out_root}"
|
||||
)
|
||||
case "${EVOLUTION_PROFILE}" in
|
||||
review)
|
||||
[[ -n "${CE_PLUGIN_DIR}" && -n "${CE_PLUGIN_VERSION}" ]] || {
|
||||
echo "review profile requires CE_PLUGIN_DIR and CE_PLUGIN_VERSION" >&2
|
||||
exit 1
|
||||
}
|
||||
cmd+=(
|
||||
--tasks workflow_bench/tasks.review.scenarios.yaml
|
||||
--arms review
|
||||
--ce-plugin-dir "${CE_PLUGIN_DIR}"
|
||||
--ce-plugin-version "${CE_PLUGIN_VERSION}"
|
||||
)
|
||||
;;
|
||||
implementation)
|
||||
cmd+=(--tasks workflow_bench/tasks.scenarios.yaml)
|
||||
;;
|
||||
*)
|
||||
echo "Unknown EVOLUTION_PROFILE '${EVOLUTION_PROFILE}' (expected review or implementation)." >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
if ((apply)); then
|
||||
cmd+=(--apply)
|
||||
fi
|
||||
if [[ -n "${INCLUDE_EXPENSIVE}" && "${INCLUDE_EXPENSIVE}" != "0" && "${INCLUDE_EXPENSIVE}" != "false" ]]; then
|
||||
cmd+=(--include-expensive)
|
||||
fi
|
||||
if [[ -n "${SEED_RESULTS}" ]]; then
|
||||
cmd+=(--seed-results "${SEED_RESULTS}")
|
||||
fi
|
||||
if [[ -n "${UNSAFE_NO_BWRAP:-}" && "${UNSAFE_NO_BWRAP}" != "0" && "${UNSAFE_NO_BWRAP}" != "false" ]]; then
|
||||
if [[ -n "${CI:-}" ]]; then
|
||||
echo "UNSAFE_NO_BWRAP is forbidden in CI." >&2
|
||||
exit 1
|
||||
fi
|
||||
cmd+=(--unsafe-no-bwrap)
|
||||
fi
|
||||
if ((${#passthrough[@]})); then
|
||||
cmd+=("${passthrough[@]}")
|
||||
fi
|
||||
|
||||
if ((dry_run)); then
|
||||
printf '%q ' "${cmd[@]}"
|
||||
printf '\n'
|
||||
exit 0
|
||||
fi
|
||||
|
||||
mkdir -p "${out_root}"
|
||||
source_sha="$(git -C "${eval_dir}/.." rev-parse HEAD)"
|
||||
runtime_digest="$(
|
||||
{
|
||||
sha256sum "${eval_dir}/../gitnexus/dist/cli/index.js"
|
||||
sha256sum "${eval_dir}/../gitnexus/package-lock.json"
|
||||
sha256sum "${eval_dir}/../gitnexus-shared/package-lock.json"
|
||||
} | sha256sum | cut -d' ' -f1
|
||||
)"
|
||||
unsafe_backend="$([[ -n "${UNSAFE_NO_BWRAP:-}" && "${UNSAFE_NO_BWRAP}" != "0" && "${UNSAFE_NO_BWRAP}" != "false" ]] && echo host-unsafe || echo bwrap)"
|
||||
SOURCE_SHA="${source_sha}" RUNTIME_DIGEST="${runtime_digest}" SANDBOX_BACKEND="${unsafe_backend}" \
|
||||
node -e 'require("fs").writeFileSync(process.argv[1], JSON.stringify({
|
||||
schema_version: 1,
|
||||
source_sha: process.env.SOURCE_SHA,
|
||||
runtime_digest: process.env.RUNTIME_DIGEST,
|
||||
profile: process.env.EVOLUTION_PROFILE,
|
||||
ce_plugin_version: process.env.CE_PLUGIN_VERSION,
|
||||
sandbox_backend: process.env.SANDBOX_BACKEND
|
||||
}, null, 2) + "\n")' "${out_root}/runtime-provenance.json"
|
||||
|
||||
export PYTHONUNBUFFERED=1
|
||||
cd "${eval_dir}"
|
||||
exec "${cmd[@]}"
|
||||
File diff suppressed because it is too large
Load diff
|
|
@ -34,6 +34,8 @@ MAX_WORKSPACE_SNAPSHOT_FILE_BYTES = 1024 * 1024 * 1024
|
|||
WORKSPACE_SNAPSHOT_BOOTSTRAP_NOISE = frozenset(
|
||||
{
|
||||
".claude",
|
||||
".bash_profile",
|
||||
".bashrc",
|
||||
".env",
|
||||
".env.development",
|
||||
".env.development.local",
|
||||
|
|
@ -43,9 +45,16 @@ WORKSPACE_SNAPSHOT_BOOTSTRAP_NOISE = frozenset(
|
|||
".env.test",
|
||||
".env.test.local",
|
||||
".gitmodules",
|
||||
".gitconfig",
|
||||
".idea",
|
||||
".npmrc",
|
||||
".profile",
|
||||
".ripgreprc",
|
||||
".vscode",
|
||||
".yarnrc",
|
||||
".yarnrc.yml",
|
||||
".zprofile",
|
||||
".zshrc",
|
||||
"bunfig.toml",
|
||||
"node_modules",
|
||||
"package-lock.json",
|
||||
|
|
@ -54,6 +63,9 @@ WORKSPACE_SNAPSHOT_BOOTSTRAP_NOISE = frozenset(
|
|||
"yarn.lock",
|
||||
}
|
||||
)
|
||||
# Claude Code may drop a workspace-root `scripts` *file* during bootstrap.
|
||||
# Only that exact entry is noise — a `scripts/` directory is real workspace.
|
||||
WORKSPACE_SNAPSHOT_ROOT_FILE_NOISE = frozenset({"scripts"})
|
||||
|
||||
# The set above is matched at the workspace ROOT only, because most of its
|
||||
# entries (package.json, node_modules, the .env family) are also legitimate
|
||||
|
|
@ -110,12 +122,24 @@ class VerificationResult:
|
|||
yield self.output
|
||||
|
||||
|
||||
def _is_bootstrap_noise(relative: PurePosixPath) -> bool:
|
||||
def _is_bootstrap_noise(
|
||||
relative: PurePosixPath,
|
||||
*,
|
||||
is_dir: bool = False,
|
||||
is_file: bool = False,
|
||||
) -> bool:
|
||||
"""Report whether a walked entry is harness noise rather than workspace change."""
|
||||
|
||||
parts = relative.parts
|
||||
if parts[0] == ".git" or parts[0] in WORKSPACE_SNAPSHOT_BOOTSTRAP_NOISE:
|
||||
return True
|
||||
if (
|
||||
is_file
|
||||
and not is_dir
|
||||
and len(parts) == 1
|
||||
and parts[0] in WORKSPACE_SNAPSHOT_ROOT_FILE_NOISE
|
||||
):
|
||||
return True
|
||||
return len(parts) >= 2 and parts[-2] == CLAUDE_BOOTSTRAP_DIR and parts[-1] in CLAUDE_BOOTSTRAP_ENTRIES
|
||||
|
||||
|
||||
|
|
@ -143,7 +167,11 @@ def workspace_snapshot(worktree: Path) -> dict[str, str]:
|
|||
raise ValueError(f"workspace snapshot directory is unreadable: {directory}: {exc}") from exc
|
||||
for entry in children:
|
||||
relative = relative_dir / entry.name
|
||||
if _is_bootstrap_noise(relative):
|
||||
if _is_bootstrap_noise(
|
||||
relative,
|
||||
is_dir=entry.is_dir(follow_symlinks=False),
|
||||
is_file=entry.is_file(follow_symlinks=False),
|
||||
):
|
||||
continue
|
||||
entry_count += 1
|
||||
path_bytes += len(relative.as_posix().encode())
|
||||
|
|
|
|||
|
|
@ -2,13 +2,17 @@
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import contextlib
|
||||
import hashlib
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
import stat
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
from collections import deque
|
||||
from collections.abc import Sequence
|
||||
from pathlib import Path, PurePosixPath
|
||||
from typing import Any
|
||||
|
|
@ -32,10 +36,248 @@ USAGE_FIELDS = (
|
|||
"output_tokens",
|
||||
)
|
||||
MAX_TRANSCRIPT_BYTES = 8 * 1024 * 1024
|
||||
# Wall-clock ceiling for one headless session, shared by the runner and the
|
||||
# evolution loop so both CLIs kill a session at the same point. 3600s was too
|
||||
# tight for the `workflow` arm: run 29907431284 lost two investigation-task
|
||||
# incumbent runs to SIGTERM at the ceiling while a Bash verification step was
|
||||
# still going, and the promotion gate demands zero excluded runs in both paired
|
||||
# arms — so a single ceiling hit costs the whole generation. Successful
|
||||
# `workflow` rows in that run finished in ~1600-2600s across both sessions, so
|
||||
# this leaves real headroom over observed work rather than over the timeout.
|
||||
SESSION_TIMEOUT_SECONDS = 5400
|
||||
# Provenance tag stamped on every parent-captured transcript artifact. The
|
||||
# evidence preflight (evolve._transcript_artifact_metadata) validates against
|
||||
# this exact value, so producer and consumer stay pinned to one schema.
|
||||
PARENT_EVENT_STREAM_SOURCE = "parent-captured-stream-json"
|
||||
# Progress reporting only. A session can work quietly for many minutes, so the
|
||||
# reporter also speaks up on its own to distinguish "thinking" from "wedged".
|
||||
PROGRESS_HEARTBEAT_SECONDS = 60.0
|
||||
MAX_PROGRESS_LINE_BYTES = 1024 * 1024
|
||||
MAX_PROGRESS_PENDING = 256
|
||||
MAX_PROGRESS_TOOL_ID_CHARS = 256
|
||||
MAX_TOOL_PREVIEW_CHARS = 800
|
||||
_SAFE_TOOL_NAME = re.compile(r"[A-Za-z0-9._:-]{1,64}")
|
||||
|
||||
|
||||
def _safe_tool_name(value: Any) -> str:
|
||||
"""A tool name is an identifier; anything else is treated as content."""
|
||||
|
||||
match = _SAFE_TOOL_NAME.fullmatch(value.strip()) if isinstance(value, str) else None
|
||||
return match.group(0) if match else "tool"
|
||||
|
||||
|
||||
def _debuggable_tool(name: str) -> bool:
|
||||
return name in {"Read", "Bash", "Grep", "Glob"} or name.startswith("mcp__")
|
||||
|
||||
|
||||
def _tool_preview(value: Any, secrets: Sequence[str]) -> str:
|
||||
"""Render one bounded, redacted, single-line tool payload preview."""
|
||||
|
||||
try:
|
||||
raw = json.dumps(value, ensure_ascii=False, separators=(",", ":"), default=str)
|
||||
except (TypeError, ValueError):
|
||||
raw = json.dumps(str(value), ensure_ascii=False)
|
||||
redacted = redact_text(raw, secrets)
|
||||
if len(redacted) <= MAX_TOOL_PREVIEW_CHARS:
|
||||
return redacted
|
||||
omitted = len(redacted) - MAX_TOOL_PREVIEW_CHARS
|
||||
return f"{redacted[:MAX_TOOL_PREVIEW_CHARS]}…[truncated {omitted} chars]"
|
||||
|
||||
|
||||
def _tool_result_text(content: Any) -> str:
|
||||
if isinstance(content, str):
|
||||
return content
|
||||
if isinstance(content, list):
|
||||
return "\n".join(
|
||||
block["text"]
|
||||
for block in content
|
||||
if isinstance(block, dict) and block.get("type") == "text" and isinstance(block.get("text"), str)
|
||||
)
|
||||
return ""
|
||||
|
||||
|
||||
def _mcp_result_has_semantic_error(content: Any) -> bool:
|
||||
"""Recognize GitNexus error envelopes that MCP transported successfully."""
|
||||
|
||||
text = _tool_result_text(content).strip()
|
||||
if not text:
|
||||
return False
|
||||
payload = text.split("\n\n---", 1)[0].strip()
|
||||
try:
|
||||
decoded = json.loads(payload)
|
||||
except (json.JSONDecodeError, ValueError):
|
||||
return bool(re.match(r"^error\s*:", payload, re.IGNORECASE))
|
||||
return isinstance(decoded, dict) and isinstance(decoded.get("error"), str)
|
||||
|
||||
|
||||
def _tool_result_log_status(name: str, block: dict[str, Any]) -> str:
|
||||
if block.get("is_error") is True:
|
||||
return "error"
|
||||
if name.startswith("mcp__") and _mcp_result_has_semantic_error(block.get("content")):
|
||||
return "semantic-error"
|
||||
return "ok"
|
||||
|
||||
|
||||
class SessionProgress:
|
||||
"""Report session metadata and bounded, redacted tool I/O previews.
|
||||
|
||||
A session's stdout is evidence: it is redacted before anything is written
|
||||
out, so raw events and model prose are never streamed to the log. Turn
|
||||
counts, tool activity, API retries, and quiet time distinguish work from
|
||||
a wedged session. Tool arguments and results are content, not metadata.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
label: str,
|
||||
*,
|
||||
stream: Any = None,
|
||||
heartbeat_s: float = PROGRESS_HEARTBEAT_SECONDS,
|
||||
secrets: Sequence[str] = (),
|
||||
) -> None:
|
||||
self.label = label
|
||||
self.heartbeat_s = heartbeat_s
|
||||
self._secrets = tuple(secret for secret in secrets if secret)
|
||||
# stdout, not stderr: the benchmark sweep runs as a child of the
|
||||
# evolution loop, which echoes only the child's stdout as it arrives
|
||||
# (run_managed(echo_stdout=True)). Its stderr surfaces as a bounded
|
||||
# tail after the fact, which is exactly the blind spot this closes.
|
||||
self._stream = stream if stream is not None else sys.stdout
|
||||
self._lock = threading.Lock()
|
||||
self._buffer = bytearray()
|
||||
self._started = time.monotonic()
|
||||
self._last_spoke = self._started
|
||||
self._events = 0
|
||||
self._turns = 0
|
||||
self._tools = 0
|
||||
self._pending_tools: dict[str, str] = {}
|
||||
self._last_activity = "starting"
|
||||
self._pending_messages: deque[str] = deque(maxlen=MAX_PROGRESS_PENDING)
|
||||
self._timer: threading.Thread | None = None
|
||||
self._done = threading.Event()
|
||||
|
||||
def __enter__(self) -> SessionProgress:
|
||||
self._say(f"started (heartbeat every {self.heartbeat_s:g}s)")
|
||||
self._emit_pending()
|
||||
self._timer = threading.Thread(target=self._heartbeat, daemon=True)
|
||||
self._timer.start()
|
||||
return self
|
||||
|
||||
def __exit__(self, *exc: object) -> None:
|
||||
self._done.set()
|
||||
timer, self._timer = self._timer, None
|
||||
if timer is not None:
|
||||
timer.join(timeout=2)
|
||||
self._emit_pending()
|
||||
|
||||
def _elapsed(self) -> str:
|
||||
seconds = int(time.monotonic() - self._started)
|
||||
return f"{seconds // 60}m{seconds % 60:02d}s"
|
||||
|
||||
def _say(self, message: str) -> None:
|
||||
# Queue only: the stdout drain thread calls observe() and must not
|
||||
# block on a full log pipe (process_control.stdout_observer contract).
|
||||
self._pending_messages.append(f"[{self.label} {self._elapsed()}] {message}")
|
||||
self._last_spoke = time.monotonic()
|
||||
|
||||
def _emit_pending(self) -> None:
|
||||
with self._lock:
|
||||
messages = list(self._pending_messages)
|
||||
self._pending_messages.clear()
|
||||
for message in messages:
|
||||
try:
|
||||
print(message, file=self._stream, flush=True)
|
||||
except (OSError, ValueError):
|
||||
return
|
||||
|
||||
def _heartbeat(self) -> None:
|
||||
tick = min(1.0, max(self.heartbeat_s / 2, 0.01))
|
||||
while not self._done.wait(tick):
|
||||
with self._lock:
|
||||
quiet = time.monotonic() - self._last_spoke
|
||||
if quiet >= self.heartbeat_s:
|
||||
self._say(
|
||||
f"still running · {self._events} events · {self._turns} turns · "
|
||||
f"{self._tools} tool calls · last: {self._last_activity}"
|
||||
)
|
||||
self._emit_pending()
|
||||
|
||||
def observe(self, chunk: bytes) -> None:
|
||||
"""Consume one stdout chunk. Never raises; never blocks on I/O."""
|
||||
|
||||
with self._lock:
|
||||
self._buffer.extend(chunk)
|
||||
# Bound the partial line: a single enormous event must not grow the
|
||||
# buffer without limit just because it has no newline yet.
|
||||
if len(self._buffer) > MAX_PROGRESS_LINE_BYTES:
|
||||
del self._buffer[:-MAX_PROGRESS_LINE_BYTES]
|
||||
while (newline := self._buffer.find(b"\n")) >= 0:
|
||||
line = bytes(self._buffer[:newline])
|
||||
del self._buffer[: newline + 1]
|
||||
self._observe_line(line)
|
||||
|
||||
def _observe_line(self, line: bytes) -> None:
|
||||
if not line.strip():
|
||||
return
|
||||
try:
|
||||
event = json.loads(line.decode("utf-8", errors="replace"))
|
||||
except (json.JSONDecodeError, ValueError):
|
||||
return
|
||||
if not isinstance(event, dict):
|
||||
return
|
||||
self._events += 1
|
||||
kind = event.get("type")
|
||||
if kind == "assistant":
|
||||
self._turns += 1
|
||||
uses = [
|
||||
block for block in _event_content(event) if isinstance(block, dict) and block.get("type") == "tool_use"
|
||||
]
|
||||
names = [_safe_tool_name(block.get("name")) for block in uses]
|
||||
if names:
|
||||
self._tools += len(names)
|
||||
self._last_activity = ", ".join(names[:4])
|
||||
self._say(f"turn {self._turns} · {self._last_activity}")
|
||||
for block, name in zip(uses, names, strict=True):
|
||||
tool_id = block.get("id")
|
||||
if isinstance(tool_id, str) and len(tool_id) <= MAX_PROGRESS_TOOL_ID_CHARS:
|
||||
self._pending_tools.pop(tool_id, None)
|
||||
self._pending_tools[tool_id] = name
|
||||
if len(self._pending_tools) > MAX_PROGRESS_PENDING:
|
||||
del self._pending_tools[next(iter(self._pending_tools))]
|
||||
if _debuggable_tool(name):
|
||||
self._say(f"tool {name} input={_tool_preview(block.get('input'), self._secrets)}")
|
||||
else:
|
||||
self._last_activity = "model reply"
|
||||
elif kind == "user":
|
||||
for block in _event_content(event):
|
||||
if not isinstance(block, dict) or block.get("type") != "tool_result":
|
||||
continue
|
||||
tool_id = block.get("tool_use_id")
|
||||
name = self._pending_tools.pop(tool_id, "tool") if isinstance(tool_id, str) else "tool"
|
||||
if not _debuggable_tool(name):
|
||||
continue
|
||||
status = _tool_result_log_status(name, block)
|
||||
self._last_activity = f"{name} result {status}"
|
||||
self._say(f"tool {name} result={status} output={_tool_preview(block.get('content'), self._secrets)}")
|
||||
elif kind == "system" and event.get("subtype") == "api_retry":
|
||||
# The signature of the gateway wedging: say it loudly and at once.
|
||||
attempt = event.get("attempt")
|
||||
limit = event.get("max_retries")
|
||||
delay = event.get("retry_delay_ms")
|
||||
wait = f" in {float(delay) / 1000:.0f}s" if isinstance(delay, (int, float)) else ""
|
||||
self._last_activity = f"API retry {attempt}/{limit}"
|
||||
self._say(f"API retry {attempt}/{limit}{wait} — no response from the model endpoint")
|
||||
elif kind == "system" and event.get("subtype") == "init":
|
||||
self._last_activity = "session init"
|
||||
self._say("session initialized")
|
||||
elif kind == "result":
|
||||
cost = event.get("total_cost_usd")
|
||||
self._last_activity = "result"
|
||||
self._say(
|
||||
f"finished · {event.get('num_turns', 0)} turns · "
|
||||
f"{'error' if event.get('is_error') else 'ok'}"
|
||||
+ (f" · ${float(cost):.2f}" if isinstance(cost, (int, float)) else "")
|
||||
)
|
||||
|
||||
|
||||
def measured_cost(raw: Any) -> float | None:
|
||||
|
|
@ -53,6 +295,16 @@ def measured_cost(raw: Any) -> float | None:
|
|||
return float(raw)
|
||||
|
||||
|
||||
def _na(value: Any) -> Any:
|
||||
"""Render an unmeasured metric as ``n/a`` instead of a misleading number.
|
||||
|
||||
Lives next to ``measured_cost`` because it renders exactly what that
|
||||
returns: every caller reporting a cost has to distinguish "never measured"
|
||||
from a real 0.0.
|
||||
"""
|
||||
return "n/a" if value is None else value
|
||||
|
||||
|
||||
SANDBOX_GITNEXUS_ENTRYPOINT = f"{SANDBOX_GITNEXUS}/dist/cli/index.js"
|
||||
SENSITIVE_EVENT_KEYS = frozenset(
|
||||
{
|
||||
|
|
@ -126,8 +378,13 @@ def sandbox_mcp_config() -> str:
|
|||
return json.dumps(config, sort_keys=True, separators=(",", ":"))
|
||||
|
||||
|
||||
def allowed_agent_tools(*, implementation: bool, include_mcp: bool = True) -> list[str]:
|
||||
tools = [*BUILTIN_AGENT_TOOLS]
|
||||
def allowed_agent_tools(
|
||||
*,
|
||||
implementation: bool,
|
||||
include_mcp: bool = True,
|
||||
allow_edit: bool = True,
|
||||
) -> list[str]:
|
||||
tools = [tool for tool in BUILTIN_AGENT_TOOLS if allow_edit or tool != "Edit"]
|
||||
if include_mcp:
|
||||
tools.extend(GITNEXUS_READ_ONLY_TOOLS)
|
||||
if include_mcp and implementation:
|
||||
|
|
@ -226,7 +483,13 @@ def _persist_parent_event_stream(
|
|||
|
||||
|
||||
def _normalized_skill_identifier(value: Any) -> str | None:
|
||||
"""Return the exact identifier token accepted by the Skill tool."""
|
||||
"""Return the exact identifier token accepted by the Skill tool.
|
||||
|
||||
Plugin skills are requested as ``plugin:skill`` (observed in review
|
||||
evolution: ``compound-engineering:ce-code-review``). Compare on the
|
||||
skill token after the last colon so a successful plugin-qualified
|
||||
invocation still counts as the expected skill.
|
||||
"""
|
||||
|
||||
if not isinstance(value, str):
|
||||
return None
|
||||
|
|
@ -236,6 +499,8 @@ def _normalized_skill_identifier(value: Any) -> str | None:
|
|||
token = stripped.split(maxsplit=1)[0]
|
||||
if token.startswith("/"):
|
||||
token = token[1:]
|
||||
if ":" in token:
|
||||
token = token.rsplit(":", 1)[-1]
|
||||
return token or None
|
||||
|
||||
|
||||
|
|
@ -336,6 +601,7 @@ def run_claude(
|
|||
timeout: int,
|
||||
disallowed_tools: list[str] | None = None,
|
||||
model: str | None = None,
|
||||
effort: str | None = None,
|
||||
env: dict[str, str] | None = None,
|
||||
permission_mode: str | None = None,
|
||||
expected_skill: str | None = None,
|
||||
|
|
@ -354,6 +620,7 @@ def run_claude(
|
|||
transcript_output_prefix: str | None = None,
|
||||
transcript_secrets: tuple[str, ...] = (),
|
||||
plugin_dirs: Sequence[str] = (),
|
||||
progress_label: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Run one headless session and return its usage record."""
|
||||
|
||||
|
|
@ -394,19 +661,33 @@ def run_claude(
|
|||
cmd += ["--permission-mode", permission_mode]
|
||||
if model:
|
||||
cmd += ["--model", model]
|
||||
if effort:
|
||||
cmd += ["--effort", effort]
|
||||
for tool in disallowed_tools or []:
|
||||
cmd += ["--disallowedTools", tool]
|
||||
managed_cmd = [*(command_prefix or []), *cmd]
|
||||
started = time.monotonic()
|
||||
proc = run_managed(
|
||||
managed_cmd,
|
||||
cwd=None if command_prefix else cwd,
|
||||
timeout=timeout,
|
||||
env=env,
|
||||
require_pid_namespace=require_pid_namespace,
|
||||
stdin_data=prompt.encode(),
|
||||
capture_stdout_bytes=MAX_TRANSCRIPT_BYTES,
|
||||
)
|
||||
with contextlib.ExitStack() as progress_stack:
|
||||
progress = (
|
||||
progress_stack.enter_context(
|
||||
SessionProgress(
|
||||
progress_label,
|
||||
secrets=transcript_secrets,
|
||||
)
|
||||
)
|
||||
if progress_label
|
||||
else None
|
||||
)
|
||||
proc = run_managed(
|
||||
managed_cmd,
|
||||
cwd=None if command_prefix else cwd,
|
||||
timeout=timeout,
|
||||
env=env,
|
||||
require_pid_namespace=require_pid_namespace,
|
||||
stdin_data=prompt.encode(),
|
||||
capture_stdout_bytes=MAX_TRANSCRIPT_BYTES,
|
||||
stdout_observer=progress.observe if progress is not None else None,
|
||||
)
|
||||
wall_s = time.monotonic() - started
|
||||
event_stream_error: str | None = None
|
||||
events: list[dict[str, Any]] = []
|
||||
|
|
@ -416,12 +697,23 @@ def run_claude(
|
|||
if proc.stdout_capture_overflow:
|
||||
raise ValueError(f"parent-captured event stream exceeds {MAX_TRANSCRIPT_BYTES} bytes")
|
||||
events = _parse_parent_event_stream(proc.stdout_capture)
|
||||
result_events = [event for event in events if event.get("type") == "result"]
|
||||
if len(result_events) != 1:
|
||||
raise ValueError(f"expected exactly one final result event, observed {len(result_events)}")
|
||||
data = result_events[0]
|
||||
if events[-1] is not data:
|
||||
result_indexes = [index for index, event in enumerate(events) if event.get("type") == "result"]
|
||||
if len(result_indexes) != 1:
|
||||
raise ValueError(f"expected exactly one final result event, observed {len(result_indexes)}")
|
||||
result_index = result_indexes[0]
|
||||
data = events[result_index]
|
||||
# A session that used a background task drains its bookkeeping after
|
||||
# the final result event (`background_tasks_changed`, `task_updated`,
|
||||
# `task_notification` — all `type: "system"`), so the result is last
|
||||
# only among the events that carry evidence. `system` events hold no
|
||||
# tool_use/tool_result/usage payload and so cannot forge skill or cost
|
||||
# evidence; any other event after the result still fails closed.
|
||||
if any(event.get("type") != "system" for event in events[result_index + 1 :]):
|
||||
raise ValueError("final result event is not the last event in the captured stream")
|
||||
# Nothing after the result is evidence. Cut the window here so that is
|
||||
# a property of what the evidence readers below can see, rather than an
|
||||
# assumption that a `system` event never carries a tool_use block.
|
||||
events = events[: result_index + 1]
|
||||
except (UnicodeError, ValueError) as exc:
|
||||
event_stream_error = str(exc)
|
||||
data = {}
|
||||
|
|
@ -435,9 +727,15 @@ def run_claude(
|
|||
or str(subtype).startswith("error")
|
||||
or not well_formed
|
||||
)
|
||||
tool_use_counts: dict[str, int] = {}
|
||||
for event in events:
|
||||
for block in _event_content(event):
|
||||
if isinstance(block, dict) and block.get("type") == "tool_use":
|
||||
name = _safe_tool_name(block.get("name"))
|
||||
tool_use_counts[name] = tool_use_counts.get(name, 0) + 1
|
||||
record = {
|
||||
"ok": not session_error,
|
||||
"error_kind": "session-error" if session_error else None,
|
||||
"error_kind": "cancelled" if proc.state == "cancelled" else ("session-error" if session_error else None),
|
||||
"error_detail": (
|
||||
{
|
||||
"subtype": subtype,
|
||||
|
|
@ -463,6 +761,8 @@ def run_claude(
|
|||
"cost_usd": measured_cost(data.get("total_cost_usd")),
|
||||
"duration_s": round(data.get("duration_ms", wall_s * 1000) / 1000, 1),
|
||||
"transcript_missing": False,
|
||||
"tool_use_counts": tool_use_counts,
|
||||
"gitnexus_tool_uses": sum(count for name, count in tool_use_counts.items() if name.startswith("mcp__gitnexus")),
|
||||
**{field: usage.get(field, 0) for field in USAGE_FIELDS},
|
||||
}
|
||||
needs_evidence = expected_skill is not None or transcript_output_dir is not None
|
||||
|
|
@ -552,4 +852,10 @@ def sum_sessions(sessions: list[dict[str, Any]]) -> dict[str, Any]:
|
|||
total["evidence_diagnostics"] = [
|
||||
diagnostic for session in sessions for diagnostic in session.get("evidence_diagnostics", [])
|
||||
]
|
||||
total["gitnexus_tool_uses"] = sum(session.get("gitnexus_tool_uses", 0) for session in sessions)
|
||||
tool_use_counts: dict[str, int] = {}
|
||||
for session in sessions:
|
||||
for name, count in session.get("tool_use_counts", {}).items():
|
||||
tool_use_counts[name] = tool_use_counts.get(name, 0) + int(count)
|
||||
total["tool_use_counts"] = tool_use_counts
|
||||
return total
|
||||
|
|
|
|||
|
|
@ -28,7 +28,6 @@ from .proposer_sandbox import (
|
|||
SandboxError,
|
||||
)
|
||||
|
||||
PINNED_GITNEXUS_VERSION = "1.6.9"
|
||||
HARNESS_ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
CE_ARMS = frozenset({"ce_workflow", "ce_workflow_direct", "ce_review"})
|
||||
|
|
@ -147,24 +146,91 @@ def _validated_runtime_root(path: Path, *, label: str) -> Path:
|
|||
return root
|
||||
|
||||
|
||||
def _primary_checkout_root(harness_root: Path) -> Path | None:
|
||||
"""Return the main worktree when *harness_root* is a linked git worktree.
|
||||
|
||||
Linked worktrees store a regular ``.git`` file pointing at
|
||||
``<primary>/.git/worktrees/<name>``. A symlink or oversized file is
|
||||
ignored so this helper cannot be used to follow an unexpected path.
|
||||
"""
|
||||
|
||||
git_file = harness_root / ".git"
|
||||
try:
|
||||
metadata = git_file.lstat()
|
||||
except OSError:
|
||||
return None
|
||||
if stat.S_ISLNK(metadata.st_mode) or not stat.S_ISREG(metadata.st_mode):
|
||||
return None
|
||||
if metadata.st_size > 4096:
|
||||
return None
|
||||
try:
|
||||
text = git_file.read_text(encoding="utf-8").strip()
|
||||
except (OSError, UnicodeError):
|
||||
return None
|
||||
if not text.startswith("gitdir:"):
|
||||
return None
|
||||
raw = text.split(":", 1)[1].strip()
|
||||
if not raw:
|
||||
return None
|
||||
gitdir = Path(raw)
|
||||
if not gitdir.is_absolute():
|
||||
gitdir = harness_root / gitdir
|
||||
if gitdir.parent.name != "worktrees" or gitdir.parent.parent.name != ".git":
|
||||
return None
|
||||
primary = gitdir.parent.parent.parent
|
||||
try:
|
||||
primary_git = (primary / ".git").lstat()
|
||||
except OSError:
|
||||
return None
|
||||
if stat.S_ISLNK(primary_git.st_mode) or not stat.S_ISDIR(primary_git.st_mode):
|
||||
return None
|
||||
resolved_primary = primary.resolve()
|
||||
if resolved_primary == harness_root.resolve():
|
||||
return None
|
||||
return resolved_primary
|
||||
|
||||
|
||||
def _validated_runtime_component(
|
||||
root: Path,
|
||||
relative: str,
|
||||
target: str,
|
||||
*,
|
||||
directory: bool,
|
||||
allow_primary_worktree_symlink: bool = False,
|
||||
) -> ReadOnlyMount:
|
||||
"""Validate one direct runtime component before exposing only that path."""
|
||||
|
||||
source = root / relative
|
||||
kind = "directory" if directory else "file"
|
||||
try:
|
||||
mode = source.lstat().st_mode
|
||||
resolved = source.resolve(strict=True)
|
||||
except OSError as exc:
|
||||
raise SandboxError(f"pinned GitNexus runtime component is unavailable: {source}: {exc}") from exc
|
||||
if stat.S_ISLNK(mode) and allow_primary_worktree_symlink and directory:
|
||||
primary = _primary_checkout_root(HARNESS_ROOT)
|
||||
if primary is None:
|
||||
raise SandboxError(f"pinned GitNexus runtime component must be a real {kind}: {source}")
|
||||
expected = primary / "gitnexus" / relative
|
||||
try:
|
||||
expected_mode = expected.lstat().st_mode
|
||||
expected_resolved = expected.resolve(strict=True)
|
||||
except OSError as exc:
|
||||
raise SandboxError(
|
||||
f"pinned GitNexus runtime component symlink must point at the primary checkout: {source}"
|
||||
) from exc
|
||||
if (
|
||||
stat.S_ISLNK(expected_mode)
|
||||
or not stat.S_ISDIR(expected_mode)
|
||||
or expected_resolved != expected
|
||||
or resolved != expected_resolved
|
||||
):
|
||||
raise SandboxError(
|
||||
f"pinned GitNexus runtime component symlink must point at the primary checkout: {source}"
|
||||
)
|
||||
return ReadOnlyMount(source=expected_resolved, target=target)
|
||||
expected_type = stat.S_ISDIR(mode) if directory else stat.S_ISREG(mode)
|
||||
if stat.S_ISLNK(mode) or not expected_type or resolved != source:
|
||||
kind = "directory" if directory else "file"
|
||||
raise SandboxError(f"pinned GitNexus runtime component must be a real {kind}: {source}")
|
||||
return ReadOnlyMount(source=source, target=target)
|
||||
|
||||
|
|
@ -180,6 +246,39 @@ def trusted_gitnexus_runtime_mounts() -> tuple[ReadOnlyMount, ...]:
|
|||
HARNESS_ROOT / "gitnexus-shared",
|
||||
label="pinned GitNexus shared runtime",
|
||||
)
|
||||
node_modules = _validated_runtime_component(
|
||||
runtime,
|
||||
"node_modules",
|
||||
f"{SANDBOX_GITNEXUS}/node_modules",
|
||||
directory=True,
|
||||
allow_primary_worktree_symlink=True,
|
||||
)
|
||||
primary = _primary_checkout_root(HARNESS_ROOT)
|
||||
if primary is not None:
|
||||
try:
|
||||
primary_shared = _validated_runtime_root(
|
||||
primary / "gitnexus-shared",
|
||||
label="primary GitNexus shared runtime",
|
||||
)
|
||||
except SandboxError:
|
||||
# Keep the already validated local shared runtime when primary is unavailable.
|
||||
pass
|
||||
else:
|
||||
reuse_primary_shared = node_modules.source == (primary / "gitnexus" / "node_modules")
|
||||
if not reuse_primary_shared:
|
||||
linked_shared = node_modules.source / "gitnexus-shared"
|
||||
try:
|
||||
reuse_primary_shared = (
|
||||
linked_shared.is_symlink()
|
||||
and linked_shared.resolve(strict=True) == primary_shared
|
||||
)
|
||||
except OSError:
|
||||
reuse_primary_shared = False
|
||||
if reuse_primary_shared:
|
||||
# Reused primary node_modules (or its inner gitnexus-shared
|
||||
# link) still points at that checkout. Mount the same tree or
|
||||
# the sandbox inner link is a host path.
|
||||
shared = primary_shared
|
||||
mounts = (
|
||||
_validated_runtime_component(
|
||||
runtime,
|
||||
|
|
@ -193,12 +292,7 @@ def trusted_gitnexus_runtime_mounts() -> tuple[ReadOnlyMount, ...]:
|
|||
f"{SANDBOX_GITNEXUS}/package.json",
|
||||
directory=False,
|
||||
),
|
||||
_validated_runtime_component(
|
||||
runtime,
|
||||
"node_modules",
|
||||
f"{SANDBOX_GITNEXUS}/node_modules",
|
||||
directory=True,
|
||||
),
|
||||
node_modules,
|
||||
_validated_runtime_component(
|
||||
runtime,
|
||||
"vendor",
|
||||
|
|
@ -233,14 +327,23 @@ def trusted_gitnexus_runtime_mounts() -> tuple[ReadOnlyMount, ...]:
|
|||
raise SandboxError(f"pinned GitNexus runtime metadata is invalid: {exc}") from exc
|
||||
if stat.S_ISLNK(entrypoint_mode) or not stat.S_ISREG(entrypoint_mode):
|
||||
raise SandboxError(f"pinned GitNexus runtime entrypoint must be regular and non-symlink: {entrypoint}")
|
||||
if package.get("version") != PINNED_GITNEXUS_VERSION:
|
||||
raise SandboxError(
|
||||
"pinned GitNexus runtime version drifted: "
|
||||
f"expected {PINNED_GITNEXUS_VERSION}, got {package.get('version')!r}"
|
||||
)
|
||||
if not isinstance(package.get("version"), str) or not package["version"]:
|
||||
raise SandboxError("pinned GitNexus runtime package.json has no version")
|
||||
|
||||
linked_shared = mounts[2].source / "gitnexus-shared"
|
||||
if not linked_shared.is_symlink() or linked_shared.resolve(strict=True) != shared:
|
||||
allowed_shared = {shared}
|
||||
if primary is not None:
|
||||
try:
|
||||
allowed_shared.add(
|
||||
_validated_runtime_root(
|
||||
primary / "gitnexus-shared",
|
||||
label="primary GitNexus shared runtime",
|
||||
)
|
||||
)
|
||||
except SandboxError:
|
||||
# Primary checkout may be absent or unreadable on a standalone eval tree.
|
||||
pass
|
||||
if not linked_shared.is_symlink() or linked_shared.resolve(strict=True) not in allowed_shared:
|
||||
raise SandboxError("pinned GitNexus runtime has an unexpected gitnexus-shared dependency")
|
||||
try:
|
||||
shared_package = json.loads(mounts[5].source.read_text())
|
||||
|
|
|
|||
|
|
@ -20,11 +20,12 @@ from .proposer_sandbox import (
|
|||
SANDBOX_WORKSPACE,
|
||||
ReadOnlyMount,
|
||||
SandboxError,
|
||||
SandboxSession,
|
||||
build_sandbox_environment,
|
||||
prepare_sandbox,
|
||||
)
|
||||
from .runner_artifacts import make_worktree, remove_clone
|
||||
from .task_assets import TaskAssetCache, TaskAssetSnapshot
|
||||
from .task_assets import TaskAssetCache, TaskAssetSnapshot, _is_harness_sandbox_copy
|
||||
|
||||
GRAPH_ASSET_PATHS = (
|
||||
".gitnexus/gitnexus.json",
|
||||
|
|
@ -41,6 +42,12 @@ GRAPH_MARKERS = (
|
|||
)
|
||||
GRAPH_BUILD_TIMEOUT_SECONDS = 3600
|
||||
GRAPH_QUERY_TIMEOUT_SECONDS = 300
|
||||
# CLI default is 5s. Parse-worker top-of-script init loads every required
|
||||
# tree-sitter binding before it can post `{type:'ready'}`; on a loaded WSL
|
||||
# host that handshake is >15s, and a GitNexus-sized --pdg analyze can slow it
|
||||
# further. Integration tests already stub 60s; graph prep uses 120s so a
|
||||
# replacement worker is not classified as deterministic-startup (#2649).
|
||||
GRAPH_WORKER_READY_TIMEOUT_MS = 120_000
|
||||
MAX_GRAPH_SCRUB_ENTRIES = 250_000
|
||||
MAX_GRAPH_SCRUB_FILE_BYTES = 512 * 1024
|
||||
MAX_GRAPH_SCRUB_TOTAL_BYTES = 2 * 1024 * 1024 * 1024
|
||||
|
|
@ -88,7 +95,13 @@ def validate_no_prebuilt_graph_assets(task: Mapping[str, Any]) -> None:
|
|||
if not isinstance(sandbox_copy, list):
|
||||
raise SandboxError("sandbox_copy must be a list")
|
||||
for value in sandbox_copy:
|
||||
if isinstance(value, str) and _is_restricted_path(value):
|
||||
if not isinstance(value, str):
|
||||
continue
|
||||
# Review corpus patches are harness-owned and applied in setup, then
|
||||
# deleted with eval/workflow_bench. They are not a prebuilt graph.
|
||||
if _is_harness_sandbox_copy(PurePosixPath(value)):
|
||||
continue
|
||||
if _is_restricted_path(value):
|
||||
raise SandboxError(f"sandbox_copy cannot import prebuilt graph or harness data: {value}")
|
||||
|
||||
dependencies = task.get("sandbox_dependencies", [])
|
||||
|
|
@ -235,6 +248,7 @@ def _graph_environment() -> dict[str, str]:
|
|||
"GITNEXUS_NO_GITIGNORE": "1",
|
||||
"GITNEXUS_WORKER_POOL_SIZE": "1",
|
||||
"GITNEXUS_PARSE_CHUNK_CONCURRENCY": "1",
|
||||
"GITNEXUS_WORKER_READY_TIMEOUT_MS": str(GRAPH_WORKER_READY_TIMEOUT_MS),
|
||||
}
|
||||
)
|
||||
return env
|
||||
|
|
@ -244,20 +258,23 @@ def _run_graph_cli(
|
|||
prefix: Sequence[str],
|
||||
arguments: Sequence[str],
|
||||
*,
|
||||
sandbox: SandboxSession | None = None,
|
||||
timeout: int,
|
||||
capture_stdout: bool = False,
|
||||
) -> bytes | None:
|
||||
host_path = sandbox.host_path if sandbox is not None else str
|
||||
host_text = sandbox.host_text if sandbox is not None else str
|
||||
command = [
|
||||
*prefix,
|
||||
SANDBOX_NODE,
|
||||
SANDBOX_GITNEXUS_ENTRYPOINT,
|
||||
*arguments,
|
||||
host_path(SANDBOX_NODE),
|
||||
host_path(SANDBOX_GITNEXUS_ENTRYPOINT),
|
||||
*(host_text(argument) for argument in arguments),
|
||||
]
|
||||
result = run_managed(
|
||||
command,
|
||||
timeout=timeout,
|
||||
env=_graph_environment(),
|
||||
require_pid_namespace=True,
|
||||
env={key: host_text(value) for key, value in _graph_environment().items()},
|
||||
require_pid_namespace=(sandbox.require_pid_namespace if sandbox is not None else True),
|
||||
capture_stdout_bytes=(2 * 1024 * 1024 if capture_stdout else None),
|
||||
)
|
||||
if not result.ok:
|
||||
|
|
@ -286,12 +303,14 @@ def _parse_empty_query(raw: bytes, *, label: str) -> None:
|
|||
raise SandboxError(f"{label} found recoverable benchmark harness references")
|
||||
|
||||
|
||||
def _scrub_and_verify_graph(prefix: Sequence[str]) -> None:
|
||||
def _scrub_and_verify_graph(prefix: Sequence[str], sandbox: SandboxSession | None = None) -> None:
|
||||
node_predicate = _marker_predicate("n")
|
||||
relation_predicate = _marker_predicate("r")
|
||||
sandbox_kwargs = {"sandbox": sandbox} if sandbox is not None else {}
|
||||
node_result = _run_graph_cli(
|
||||
prefix,
|
||||
("cypher", f"MATCH (n) WHERE {node_predicate} RETURN n LIMIT 1", "-r", "benchmark-target", "--limit", "1"),
|
||||
**sandbox_kwargs,
|
||||
timeout=GRAPH_QUERY_TIMEOUT_SECONDS,
|
||||
capture_stdout=True,
|
||||
)
|
||||
|
|
@ -305,6 +324,7 @@ def _scrub_and_verify_graph(prefix: Sequence[str]) -> None:
|
|||
"--limit",
|
||||
"1",
|
||||
),
|
||||
**sandbox_kwargs,
|
||||
timeout=GRAPH_QUERY_TIMEOUT_SECONDS,
|
||||
capture_stdout=True,
|
||||
)
|
||||
|
|
@ -339,6 +359,7 @@ def prepare_sanitized_graph(
|
|||
claude_bin: Path | str,
|
||||
bwrap_bin: Path | str,
|
||||
runtime_mounts: Sequence[ReadOnlyMount],
|
||||
sandbox_backend: str = "bwrap",
|
||||
) -> SanitizedGraphSnapshot:
|
||||
"""Sanitize, index offline once, scrub, and freeze graph assets for all arms."""
|
||||
|
||||
|
|
@ -355,8 +376,10 @@ def prepare_sanitized_graph(
|
|||
bwrap_bin=bwrap_bin,
|
||||
read_only_mounts=runtime_mounts,
|
||||
preflight=False,
|
||||
backend=sandbox_backend,
|
||||
) as sandbox:
|
||||
prefix = sandbox.command_prefix_for(unshare_network=True)
|
||||
unsafe_sandbox = sandbox if getattr(sandbox, "backend", "bwrap") == "host-unsafe" else None
|
||||
_run_graph_cli(
|
||||
prefix,
|
||||
(
|
||||
|
|
@ -375,9 +398,10 @@ def prepare_sanitized_graph(
|
|||
"--workers",
|
||||
"1",
|
||||
),
|
||||
**({"sandbox": unsafe_sandbox} if unsafe_sandbox is not None else {}),
|
||||
timeout=GRAPH_BUILD_TIMEOUT_SECONDS,
|
||||
)
|
||||
_scrub_and_verify_graph(prefix)
|
||||
_scrub_and_verify_graph(prefix, unsafe_sandbox)
|
||||
_validate_graph_metadata(seed, sanitized_head)
|
||||
assets = cache.prepare(
|
||||
{"sandbox_copy": list(GRAPH_ASSET_PATHS)},
|
||||
|
|
|
|||
|
|
@ -24,6 +24,7 @@ from dataclasses import dataclass
|
|||
from pathlib import Path, PurePosixPath
|
||||
from typing import Any
|
||||
|
||||
from . import runtime_mounts
|
||||
from .proposer_sandbox import (
|
||||
DEPENDENCY_MOUNT_BASENAME,
|
||||
SANDBOX_WORKSPACE,
|
||||
|
|
@ -63,6 +64,10 @@ _REFLINK_UNAVAILABLE = {
|
|||
|
||||
DEPENDENCY_CONTENT_BINDING_FIELD = "sandbox_dependency_content_digest"
|
||||
DEPENDENCY_MANIFEST_BINDING_FIELD = "sandbox_dependency_manifest_digest"
|
||||
# Review corpus patches are harness-owned. The task repo (`~/GitNexus`) is the
|
||||
# subject checkout and often a different worktree or branch, so these paths
|
||||
# must be read from the package that defined the benchmark.
|
||||
HARNESS_SANDBOX_COPY_PREFIXES = (PurePosixPath("eval/workflow_bench/review_cases"),)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
|
|
@ -191,7 +196,7 @@ class TaskAssetCache:
|
|||
self.root = root.expanduser().absolute()
|
||||
self.root.mkdir(mode=0o700, parents=True, exist_ok=False)
|
||||
self._by_definition: dict[
|
||||
tuple[str, str, tuple[str, ...], tuple[tuple[str, str], ...]],
|
||||
tuple[str, str, tuple[str, ...], tuple[tuple[str, str], ...], str],
|
||||
TaskAssetSnapshot,
|
||||
] = {}
|
||||
self._closed = False
|
||||
|
|
@ -218,7 +223,14 @@ class TaskAssetCache:
|
|||
declarations, relative_paths = _sandbox_copy_declarations(task)
|
||||
dependency_declarations = _sandbox_dependency_declarations(task)
|
||||
dependency_identity = tuple((declaration.source, declaration.target) for declaration in dependency_declarations)
|
||||
definition = (str(repo_identity), resolved_sha, declarations, dependency_identity)
|
||||
harness_identity = _harness_sandbox_copy_identity(relative_paths)
|
||||
definition = (
|
||||
str(repo_identity),
|
||||
resolved_sha,
|
||||
declarations,
|
||||
dependency_identity,
|
||||
harness_identity,
|
||||
)
|
||||
existing = self._by_definition.get(definition)
|
||||
if existing is not None:
|
||||
if expected_dependency_binding is not None:
|
||||
|
|
@ -238,9 +250,22 @@ class TaskAssetCache:
|
|||
repo_identity,
|
||||
os.O_RDONLY | os.O_DIRECTORY | getattr(os, "O_CLOEXEC", 0) | getattr(os, "O_NOFOLLOW", 0),
|
||||
)
|
||||
harness_fd: int | None = None
|
||||
try:
|
||||
for relative in relative_paths:
|
||||
descriptor = _open_relative(repo_fd, relative)
|
||||
if _is_harness_sandbox_copy(relative):
|
||||
if harness_fd is None:
|
||||
harness_root = _harness_sandbox_copy_root()
|
||||
harness_fd = os.open(
|
||||
harness_root,
|
||||
os.O_RDONLY
|
||||
| os.O_DIRECTORY
|
||||
| getattr(os, "O_CLOEXEC", 0)
|
||||
| getattr(os, "O_NOFOLLOW", 0),
|
||||
)
|
||||
descriptor = _open_relative(harness_fd, relative)
|
||||
else:
|
||||
descriptor = _open_relative(repo_fd, relative)
|
||||
try:
|
||||
builder.copy_descriptor(descriptor, relative)
|
||||
finally:
|
||||
|
|
@ -299,6 +324,8 @@ class TaskAssetCache:
|
|||
)
|
||||
finally:
|
||||
os.close(repo_fd)
|
||||
if harness_fd is not None:
|
||||
os.close(harness_fd)
|
||||
|
||||
entries = builder.finished_entries()
|
||||
manifest_digest = _manifest_digest(entries)
|
||||
|
|
@ -317,6 +344,7 @@ class TaskAssetCache:
|
|||
manifest_digest=manifest_digest,
|
||||
dependency_content_digest=dependency_content_digest,
|
||||
dependency_manifest_digest=dependency_manifest_digest,
|
||||
harness_identity=harness_identity,
|
||||
)
|
||||
destination = self.root / digest
|
||||
if destination.exists():
|
||||
|
|
@ -537,6 +565,22 @@ class _SnapshotBuilder:
|
|||
return tuple(sorted(self.entries.values(), key=lambda entry: entry.path.as_posix()))
|
||||
|
||||
|
||||
def _is_harness_sandbox_copy(relative: PurePosixPath) -> bool:
|
||||
if relative.is_absolute() or not relative.parts or ".." in relative.parts:
|
||||
return False
|
||||
return any(relative == prefix or prefix in relative.parents for prefix in HARNESS_SANDBOX_COPY_PREFIXES)
|
||||
|
||||
|
||||
def _harness_sandbox_copy_root() -> Path:
|
||||
return _real_directory(runtime_mounts.HARNESS_ROOT, label="harness sandbox_copy root")
|
||||
|
||||
|
||||
def _harness_sandbox_copy_identity(relative_paths: tuple[PurePosixPath, ...]) -> str:
|
||||
if not any(_is_harness_sandbox_copy(relative) for relative in relative_paths):
|
||||
return ""
|
||||
return str(_harness_sandbox_copy_root())
|
||||
|
||||
|
||||
def _sandbox_copy_declarations(
|
||||
task: Mapping[str, Any],
|
||||
) -> tuple[tuple[str, ...], tuple[PurePosixPath, ...]]:
|
||||
|
|
@ -969,15 +1013,17 @@ def _snapshot_digest(
|
|||
manifest_digest: str,
|
||||
dependency_content_digest: str,
|
||||
dependency_manifest_digest: str,
|
||||
harness_identity: str = "",
|
||||
) -> str:
|
||||
payload = {
|
||||
"declarations": declarations,
|
||||
"dependency_content_digest": dependency_content_digest,
|
||||
"dependency_manifest_digest": dependency_manifest_digest,
|
||||
"harness_identity": harness_identity,
|
||||
"manifest_digest": manifest_digest,
|
||||
"repo_identity": str(repo_identity),
|
||||
"resolved_sha": resolved_sha,
|
||||
"schema_version": 2,
|
||||
"schema_version": 3,
|
||||
}
|
||||
return hashlib.sha256(json.dumps(payload, sort_keys=True, separators=(",", ":")).encode()).hexdigest()
|
||||
|
||||
|
|
|
|||
76
eval/workflow_bench/tasks.review.scenarios.yaml
Normal file
76
eval/workflow_bench/tasks.review.scenarios.yaml
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
# Exact-head historical PR corpus. Hidden labels are captured by the harness
|
||||
# before clone sanitization and never enter a reviewer mount.
|
||||
tasks:
|
||||
- &review_case
|
||||
id: review-pr-2718-defect
|
||||
class: review-defect
|
||||
repo: ~/GitNexus
|
||||
ref: ff86ccf1e79cd7e4175da437ae8aeaf67b64aaa1
|
||||
sandbox_copy: [eval/workflow_bench/review_cases/pr-2718.patch]
|
||||
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2718.patch && rm -rf eval/workflow_bench
|
||||
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2718. Report only actionable defects introduced by the local diff.
|
||||
verify: test -s review-output.json
|
||||
oracle:
|
||||
command: test -s review-output.json
|
||||
files: [{ source: review-pr-2718-defect.labels.json, target: review-labels.json }]
|
||||
sandbox_dependencies: &deps
|
||||
- { source: node_modules, target: node_modules }
|
||||
- { source: gitnexus/node_modules, target: gitnexus/node_modules }
|
||||
- { source: gitnexus-shared/node_modules, target: gitnexus-shared/node_modules }
|
||||
|
||||
- <<: *review_case
|
||||
id: review-pr-2794-defect
|
||||
ref: 911151e2304f298a995fcc69c738ad2c6db9393a
|
||||
sandbox_copy: [eval/workflow_bench/review_cases/pr-2794.patch]
|
||||
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2794.patch && rm -rf eval/workflow_bench
|
||||
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2794. Report only actionable defects introduced by the local diff.
|
||||
oracle:
|
||||
command: test -s review-output.json
|
||||
files: [{ source: review-pr-2794-defect.labels.json, target: review-labels.json }]
|
||||
sandbox_dependencies: *deps
|
||||
|
||||
- <<: *review_case
|
||||
id: review-pr-2108-defect
|
||||
ref: 3a4247ec36b5ad86b1123d3bbce8183a643f7434
|
||||
sandbox_copy: [eval/workflow_bench/review_cases/pr-2108.patch]
|
||||
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2108.patch && rm -rf eval/workflow_bench
|
||||
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2108. Report only actionable defects introduced by the local diff.
|
||||
oracle:
|
||||
command: test -s review-output.json
|
||||
files: [{ source: review-pr-2108-defect.labels.json, target: review-labels.json }]
|
||||
sandbox_dependencies: *deps
|
||||
|
||||
- <<: *review_case
|
||||
id: review-pr-2258-defect
|
||||
ref: 78b4077d8acc86f1b0c32e41012174d484e81f12
|
||||
sandbox_copy: [eval/workflow_bench/review_cases/pr-2258.patch]
|
||||
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2258.patch && rm -rf eval/workflow_bench
|
||||
prompt: Review the historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2258. Report only actionable defects introduced by the local diff.
|
||||
oracle:
|
||||
command: test -s review-output.json
|
||||
files: [{ source: review-pr-2258-defect.labels.json, target: review-labels.json }]
|
||||
sandbox_dependencies: *deps
|
||||
|
||||
- <<: *review_case
|
||||
id: review-pr-2258-clean
|
||||
class: review-clean
|
||||
ref: 78b4077d8acc86f1b0c32e41012174d484e81f12
|
||||
sandbox_copy: [eval/workflow_bench/review_cases/pr-2258b.patch]
|
||||
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2258b.patch && rm -rf eval/workflow_bench
|
||||
prompt: Review this historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2258. Report only actionable defects introduced by the local diff.
|
||||
oracle:
|
||||
command: test -s review-output.json
|
||||
files: [{ source: review-pr-2258-clean.labels.json, target: review-labels.json }]
|
||||
sandbox_dependencies: *deps
|
||||
|
||||
- <<: *review_case
|
||||
id: review-pr-2773-clean
|
||||
class: review-clean
|
||||
ref: 84f584449de02376a8ffc096dceac2e8f732cab5
|
||||
sandbox_copy: [eval/workflow_bench/review_cases/pr-2773.patch]
|
||||
setup: git apply --exclude='eval/workflow_bench/*' eval/workflow_bench/review_cases/pr-2773.patch && rm -rf eval/workflow_bench
|
||||
prompt: Review this historical snapshot of https://github.com/abhigyanpatwari/GitNexus/pull/2773. Report only actionable defects introduced by the local diff.
|
||||
oracle:
|
||||
command: test -s review-output.json
|
||||
files: [{ source: review-pr-2773-clean.labels.json, target: review-labels.json }]
|
||||
sandbox_dependencies: *deps
|
||||
|
|
@ -33,7 +33,7 @@ tasks:
|
|||
# Overhead floor — measured 2026-07-11 (see README calibration): the
|
||||
# workflow is EXPECTED to lose here. Kept in the suite so regressions in
|
||||
# the overhead floor stay visible.
|
||||
- id: trivial-version-alias
|
||||
- id: trivial-status-json-alias
|
||||
class: trivial
|
||||
repo: ~/GitNexus
|
||||
ref: main
|
||||
|
|
@ -45,46 +45,45 @@ tasks:
|
|||
- source: gitnexus-shared/node_modules
|
||||
target: gitnexus-shared/node_modules
|
||||
prompt: >
|
||||
Add -V as a short alias for --version to the gitnexus CLI
|
||||
Add -j as a short alias for --json on the gitnexus status command
|
||||
(gitnexus/src/cli/index.ts), and cover the alias with a unit test in
|
||||
gitnexus/test/unit/cli-commands.test.ts.
|
||||
verify: cd gitnexus && npx tsc --noEmit && npx vitest run test/unit/cli-commands.test.ts
|
||||
gitnexus/test/unit/status-json-alias.test.ts.
|
||||
verify: cd gitnexus && npx tsc --noEmit && npx vitest run test/unit/status-json-alias.test.ts
|
||||
oracle:
|
||||
command: >-
|
||||
cd gitnexus && ./node_modules/.bin/vitest run
|
||||
--config "$GITNEXUS_BENCH_ORACLE_ROOT/vitest.config.mts"
|
||||
"$GITNEXUS_BENCH_ORACLE_ROOT/trivial-version-alias.oracle.test.ts"
|
||||
"$GITNEXUS_BENCH_ORACLE_ROOT/trivial-status-json-alias.oracle.test.ts"
|
||||
files:
|
||||
- source: vitest.config.mts
|
||||
target: vitest.config.mts
|
||||
- source: trivial-version-alias.oracle.test.ts
|
||||
target: trivial-version-alias.oracle.test.ts
|
||||
- source: trivial-status-json-alias.oracle.test.ts
|
||||
target: trivial-status-json-alias.oracle.test.ts
|
||||
|
||||
# Investigation-heavy bug: requires locating the degraded-result path in
|
||||
# local-backend, understanding the layer probe, and changing a contract
|
||||
# message without breaking existing consumers.
|
||||
- id: inv-bug-pdg-note
|
||||
# Investigation-heavy bug (#2965): requires following the C import marker
|
||||
# through the shared scope-resolution pipeline without breaking local
|
||||
# quoted-header resolution.
|
||||
- id: inv-bug-c-system-include
|
||||
class: investigation-bug
|
||||
repo: ~/GitNexus
|
||||
ref: main
|
||||
sandbox_dependencies: *gitnexus_sandbox_dependencies
|
||||
prompt: >
|
||||
When the pdg_query MCP tool returns its "no PDG layer" note, make the
|
||||
note say WHICH sub-layer is missing (CDG vs REACHING_DEF) instead of a
|
||||
generic message, keeping the existing degraded-result contract intact.
|
||||
Cover both modes with unit tests in
|
||||
gitnexus/test/unit/pdg-note-sublayer.test.ts.
|
||||
verify: cd gitnexus && npx tsc --noEmit && npx vitest run test/unit/pdg-note-sublayer.test.ts
|
||||
Fix the C scope resolver so a system include such as <stdio.h> never
|
||||
resolves to an unrelated same-named header inside the repository, while
|
||||
quoted/local includes still resolve. Remove C from the conformance
|
||||
suite's KNOWN_GAPS and cover the negative and positive cases there.
|
||||
verify: cd gitnexus && npx tsc --noEmit && npx vitest run test/unit/scope-resolution/external-import-conformance.test.ts
|
||||
oracle:
|
||||
command: >-
|
||||
cd gitnexus && ./node_modules/.bin/vitest run
|
||||
--config "$GITNEXUS_BENCH_ORACLE_ROOT/vitest.config.mts"
|
||||
"$GITNEXUS_BENCH_ORACLE_ROOT/inv-bug-pdg-note.oracle.test.ts"
|
||||
"$GITNEXUS_BENCH_ORACLE_ROOT/inv-bug-c-system-include.oracle.test.ts"
|
||||
files:
|
||||
- source: vitest.config.mts
|
||||
target: vitest.config.mts
|
||||
- source: inv-bug-pdg-note.oracle.test.ts
|
||||
target: inv-bug-pdg-note.oracle.test.ts
|
||||
- source: inv-bug-c-system-include.oracle.test.ts
|
||||
target: inv-bug-c-system-include.oracle.test.ts
|
||||
|
||||
# Investigation-heavy feature: touches tool schema, backend filtering, and
|
||||
# pagination totals — three seams that must stay consistent.
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
{
|
||||
"name": "gitnexus",
|
||||
"description": "Code intelligence powered by a knowledge graph. Provides execution flow tracing, blast radius analysis, and augmented search across your codebase.",
|
||||
"version": "1.6.9",
|
||||
"version": "1.6.11",
|
||||
"author": {
|
||||
"name": "GitNexus"
|
||||
},
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
{
|
||||
"name": "gitnexus",
|
||||
"description": "Code intelligence powered by a knowledge graph. Provides execution flow tracing, blast radius analysis, and augmented search across your codebase.",
|
||||
"version": "1.6.9",
|
||||
"version": "1.6.11",
|
||||
"skills": "./skills",
|
||||
"mcpServers": "./.mcp.json",
|
||||
"hooks": "./hooks/hooks.json",
|
||||
|
|
|
|||
|
|
@ -19,14 +19,22 @@ node .gitnexus/run.cjs analyze
|
|||
|
||||
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates CLAUDE.md / AGENTS.md context files.
|
||||
|
||||
| Flag | Effect |
|
||||
|------|--------|
|
||||
| `--force` | Force full re-index even if up to date |
|
||||
| Flag | Effect |
|
||||
| -------------- | ---------------------------------------------------------------- |
|
||||
| `--watch` | Keep a Git repository index current with serialized refreshes |
|
||||
| `--debounce <ms>` | Watch quiet period before refresh (default: 300 ms) |
|
||||
| `--force` | Force full re-index even if up to date |
|
||||
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
|
||||
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
|
||||
| `--pdg` | Build the program-dependence layers used by `explain` and `pdg_query` (taint, CDG, and REACHING_DEF). |
|
||||
| `--spring-actuator <path>` | Import opt-in Spring Boot Actuator mappings, beans, conditions, configprops, and env snapshots. Forces a full rebuild; unsupported with `--watch`. |
|
||||
| `--asyncapi-spec <path>` | Read opt-in AsyncAPI 3.x documents (directory or single file) and mint `Destination` nodes from their operations. 2.x is refused, not mapped. Unsupported with `--watch`. |
|
||||
|
||||
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale.
|
||||
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
|
||||
|
||||
For Spring runtime enrichment, pass a JSON bundle, one endpoint JSON file, or a directory containing endpoint files. Route evidence is authoritative only when `runtimeConfirmed === true`; `runtimeSource` records provenance and may also accompany `handler-conflict`. Env/configprops values are never persisted.
|
||||
|
||||
Use `node .gitnexus/run.cjs analyze --watch` for a long-lived local Git repository. It performs an initial analysis, queues scanner-admitted file changes, and retries intact failed batches with bounded backoff. Watch refreshes update only the graph: they skip AGENTS.md / CLAUDE.md injection and standard skill installation, so run a one-shot `analyze` when those generated files need updating. Watch rejects one-shot or context-output flags including `--force`, embedding flags, `--skills`, `--default-branch`, `--skip-agents-md`, `--skip-skills`, `--no-stats`, `--self-commit`, `--index-only`, and `--skip-git`. It never pulls remotes. Scheduled remote clone/pull is a different command: `gitnexus auto-sync`. Bare `gitnexus watch` is reserved and does not start either job. Running MCP and `serve` processes periodically check for a published replacement and reopen it without a restart. MCP checks are throttled to once every five seconds, so a tool call before the next check can briefly use the previous index.
|
||||
|
||||
### status — Check index freshness
|
||||
|
||||
|
|
@ -44,10 +52,10 @@ node .gitnexus/run.cjs clean
|
|||
|
||||
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
|
||||
|
||||
| Flag | Effect |
|
||||
|------|--------|
|
||||
| `--force` | Skip confirmation prompt |
|
||||
| `--all` | Clean all indexed repos, not just the current one |
|
||||
| Flag | Effect |
|
||||
| --------- | ------------------------------------------------- |
|
||||
| `--force` | Skip confirmation prompt |
|
||||
| `--all` | Clean all indexed repos, not just the current one |
|
||||
|
||||
### wiki — Generate documentation from the graph
|
||||
|
||||
|
|
@ -55,19 +63,21 @@ Deletes the `.gitnexus/` directory and unregisters the repo from the global regi
|
|||
node .gitnexus/run.cjs wiki
|
||||
```
|
||||
|
||||
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
|
||||
Generates repository documentation from the knowledge graph using an LLM. HTTP providers require an API key (saved to `~/.gitnexus/config.json` on first use). Local CLI providers (`--provider cursor|claude|codex|opencode|grok`) use your existing CLI login.
|
||||
|
||||
| Flag | Effect |
|
||||
|------|--------|
|
||||
| `--force` | Force full regeneration, also required to re-gerenate an existing wiki in a different language |
|
||||
| `--model <model>` | LLM model (default: MiniMax-M3) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
| `--gist` | Publish wiki as a public GitHub Gist |
|
||||
| Flag | Effect |
|
||||
| ------------------- | ----------------------------------------- |
|
||||
| `--force` | Force full regeneration, also required to re-generate an existing wiki in a different language |
|
||||
| `--provider <name>` | LLM provider: minimax, openai, openrouter, azure, custom, cursor, claude, codex, opencode, or grok (default: minimax). Local CLIs (`cursor`, `claude`, `codex`, `opencode`, `grok`) use your existing CLI login and skip `--api-key`. |
|
||||
| `--model <model>` | LLM model (default: MiniMax-M3) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
| `--timeout <seconds>` | LLM request timeout in seconds (default: disabled) |
|
||||
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
|
||||
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese)|
|
||||
| `--retries <n>` | Max LLM retry attempts per request (default: 3) |
|
||||
| `--lang <lang>` | Output language for generated documentation (e.g. english, chinese, spanish, japanese) |
|
||||
| `--gist` | Publish wiki as a public GitHub Gist |
|
||||
|
||||
### list — Show all indexed repos
|
||||
|
||||
```bash
|
||||
|
|
@ -84,5 +94,5 @@ Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_
|
|||
## Troubleshooting
|
||||
|
||||
- **"Not inside a git repository"**: Run from a directory inside a git repo
|
||||
- **Index is stale after re-analyzing**: Restart Claude Code to reload the MCP server
|
||||
- **Index is stale after re-analyzing**: Wait for the next MCP tool call to reopen the published index; this normally takes no more than five seconds
|
||||
- **Embeddings slow**: Omit `--embeddings` (it's off by default) or set `OPENAI_API_KEY` for faster API-based embedding
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
"mcpServers": {
|
||||
"gitnexus": {
|
||||
"command": "npx",
|
||||
"args": ["-y", "gitnexus@1.6.9", "mcp"]
|
||||
"args": ["-y", "gitnexus@1.6.11", "mcp"]
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Reference in a new issue