docs: invalidate contaminated router benchmark results

This commit is contained in:
Tin Chi Lo 2026-09-15 10:35:35 -07:00
parent 31afc3d5e0
commit 04c4e6284b
7 changed files with 25 additions and 2 deletions

View file

@ -1,3 +1,5 @@
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
# Card and calibration ablations
These comparisons hold the boundary at the original default: capability base 0.5 with step 0.1, or V2 quality gap 0.05. All coefficients were fitted on the training split. This table reports every raw card and the fixed middle regularization strength of 10 for calibrated variants; the JSON contains all strengths. No variant is chosen using these held-out results

View file

@ -1,3 +1,5 @@
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
# Auto Router training experiment
Profiles were fitted on 83 DeepSWE tasks and selected on 30 repository-disjoint validation tasks before inspecting live grades. The live comparison uses 25 native-image-eligible SWE-bench Verified tasks and four current solver models at high effort. Each solver runs once per task; frozen task-pinned policies reuse those attempts and add their judge cost

View file

@ -1,3 +1,5 @@
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
# What the training experiment showed
Both implementations now support runnable, opt-in trained profiles. Training covered three cards, raw probabilities, per-model calibration, task-dependent calibration, and pair-specific boundaries, including combinations. All fitting and policy selection used the DeepSWE training and validation splits. The 25 fresh SWE-bench tasks were used only for evaluation

View file

@ -1,3 +1,5 @@
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
# Experimental Auto Router training snapshots
These opt-in profiles compare the original cards, a research-informed card rewrite, and cards with training-derived capability priors. They also compare raw probabilities, per-model logit calibration, and regularized task-dependent logit calibration. The router makes one judge call; all learned probability adjustments and threshold comparisons run locally

View file

@ -2395,5 +2395,10 @@
"lost_capable_successes": 0,
"gained_efficient_successes": 0
}
]
],
"validation_status": {
"live_benchmark_valid": false,
"reason": "Solver-visible Git history exposed upstream fixes; live quality and savings claims withdrawn",
"replacement": "Fresh isolated training and evaluation in progress"
}
}

View file

@ -4779,5 +4779,10 @@
}
}
}
]
],
"validation_status": {
"live_benchmark_valid": false,
"reason": "Solver-visible Git history exposed upstream fixes; live quality and savings claims withdrawn",
"replacement": "Fresh isolated training and evaluation in progress"
}
}

View file

@ -51,5 +51,10 @@
"ABLATIONS.md",
"ablation_diagnostics.json"
]
},
"validation_status": {
"live_benchmark_valid": false,
"reason": "Solver-visible Git history exposed upstream fixes; live quality and savings claims withdrawn",
"replacement": "Fresh isolated training and evaluation in progress"
}
}