mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-06 02:48:13 +00:00
docs(router): publish held-out primary results and diagnostic profiles
This commit is contained in:
parent
c08f537d89
commit
734a1b18d0
9 changed files with 8687 additions and 3 deletions
31
cookbook/auto_router_selective_training/PRIMARY_FINDINGS.md
Normal file
31
cookbook/auto_router_selective_training/PRIMARY_FINDINGS.md
Normal file
|
|
@ -0,0 +1,31 @@
|
|||
# Primary findings: quality retained versus money saved
|
||||
|
||||
Training improved some comparisons, but it did not approach the ideal Luna-first policy across all 125 held-out tasks. The profiles selected on validation do not combine preservation of every Sol success with lower total cost across all four benchmarks
|
||||
|
||||
Sol solved 87/125 at $20.8964. Luna solved 84/125 at at least $2.6531. There were 11 Sol-only successes, eight Luna-only successes, 76 successes shared by both, and 30 tasks both attempts failed. A perfect Luna-first decision would therefore reach 95/125 at at least $5.2653. That is an unattainable hindsight reference, with at most 74.8% savings using the recorded bills
|
||||
|
||||
| Frozen comparison | Result | Cost | Sol successes lost |
|
||||
|---|---|---:|---:|
|
||||
| Capability retention profile, SWE-bench | 18/25, matching Sol | $6.2493, 41.7% savings | 0 |
|
||||
| V2 quality profile, Terminal-Bench | 19/25 versus Sol's 18/25 | $8.1712, 3.3% savings | 0 |
|
||||
| V2 retention profile, all tasks | 85/125 versus Sol's 87/125 | $14.0036, 33.0% savings | 4 |
|
||||
| V2 review retention profile, all tasks | 88/125 versus Sol's 87/125 | At least $20.9754, higher than Sol | 0 |
|
||||
| V2 review quality profile, all tasks | 88/125 versus Sol's 87/125 | At least $8.6493, at most 58.6% savings | 7 |
|
||||
|
||||
These benchmark-specific successes are separate frozen policies. Combining the winning profile for each benchmark would be a new policy requiring fresh evaluation
|
||||
|
||||
Relative to the original V2 baseline, the trained V2 retention profile improves aggregate quality from 84 to 85 solves and reduces cost from at least $17.0029 to $14.0036. That is at least a 17.6% cost reduction against that baseline. It still misses four Sol successes, so it does not meet the stronger preservation goal
|
||||
|
||||
Among the 40 prespecified family controls, three capability variants retained every Sol success and gained one additional solve. The per-model adjustment control achieved 88/125 at $20.3797, a 2.5% saving. The paired and rescue controls saved 0.2% and 0.6%. These are exploratory comparisons across many variants, not a new winner selection. All twelve upfront family configurations per classifier are provided for reproduction and fresh benchmarking; none was refitted on these outcomes
|
||||
|
||||
The expensive V2 review policy captured all 11 Sol-only successes by escalating 115 of 125 tasks. Its cheaper review policy escalated 15 tasks but captured only four of those 11 rescues. This is the remaining problem: identify the rare tasks where Sol changes failure into success without also sending most other tasks to Sol
|
||||
|
||||
Calibration improved more than routing precision. V2's selected scalar calibration reduced the Luna Brier score from 0.223 to 0.205 and the Sol Brier score from 0.237 to 0.211. However, probability-gap MSE only moved from 0.153 to 0.151. Ranking Sol-only rescues by the estimated gap improved from ROC AUC 0.618 to 0.689, still far from perfect separation. The empirical review variant reached 0.742 gap-ranking AUC, but its frozen economical threshold missed seven rescues. A more accurate average success probability does not by itself isolate the rescue tasks
|
||||
|
||||
The selected upfront profiles retain the original cards. Their improvements come from fitted adjustments and decision boundaries. Card rewrites alone gave mixed results. Training contained only 13 Sol-only successes, including two agentic examples, which limits evidence for agentic rescue decisions. Additional paired agentic training data and stronger verification signals are plausible next directions; this study does not establish their effect
|
||||
|
||||
All policies and family controls were frozen before final outcomes. These are task-level replays of independent attempts, not per-turn routing or Sol continuation from a Luna patch. Small samples and the search across development candidates limit generalization. One attempt per model does not establish a stable solve probability for an individual task
|
||||
|
||||
Two Luna solver requests interrupted by host sleep have unknown charges. An earlier unavailable DNA-assembly review has one additional unmetered connection-timeout request. Thus Luna-only comparisons contain two unmetered requests, and review cascades contain three. The report marks affected costs as lower bounds and savings as upper bounds. The three upfront comparisons quoted with exact costs above are fully metered
|
||||
|
||||
The additional Sonnet/Opus controls are still running. Their completion will extend the comparison; it cannot change these frozen Luna/Sol results
|
||||
116
cookbook/auto_router_selective_training/PRIMARY_REPORT.md
Normal file
116
cookbook/auto_router_selective_training/PRIMARY_REPORT.md
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
# Held-out Luna/Sol routing results
|
||||
|
||||
All 125 held-out tasks and 250 Luna/Sol attempts. Frozen task-level policy replay; additional Anthropic controls remain in the full report. No refitting or reselection.
|
||||
|
||||
Retention and quality denote development-selection objectives, not guarantees. Costs include classifier/reviewer inference and discarded Luna attempts on escalation. These are independent paired attempts; Sol starts from the original task, not from Luna's partial work.
|
||||
|
||||
Costs use recorded gateway bills. A ≥ cost and ≤ savings mark unmetered transport requests: cost is a lower bound and savings versus fully metered Sol is an upper bound. The JSON counts these requests per policy and task. They are not assumed free.
|
||||
|
||||
## all
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol | Sol-only rescues captured |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| Luna only | 84/125 | ≥$2.6531 | ≤87.3% | +8 / -11 | 0/11 |
|
||||
| Sol only | 87/125 | $20.8964 | 0.0% | +0 / -0 | 11/11 |
|
||||
| Capability upfront, retention objective | 86/125 | $13.8156 | 33.9% | +1 / -2 | 9/11 |
|
||||
| Capability upfront, quality objective | 88/125 | $18.5237 | 11.4% | +5 / -4 | 7/11 |
|
||||
| V2 upfront, retention objective | 85/125 | $14.0036 | 33.0% | +2 / -4 | 7/11 |
|
||||
| V2 upfront, quality objective | 86/125 | $18.1999 | 12.9% | +4 / -5 | 6/11 |
|
||||
| Capability baseline | 85/125 | $6.7642 | 67.6% | +8 / -10 | 1/11 |
|
||||
| Capability with learned cards only | 84/125 | $7.2496 | 65.3% | +7 / -10 | 1/11 |
|
||||
| V2 baseline | 84/125 | ≥$17.0029 | ≤18.6% | +3 / -6 | 5/11 |
|
||||
| V2 with learned cards only | 85/125 | ≥$14.6487 | ≤29.9% | +6 / -8 | 3/11 |
|
||||
| Capability review, retention objective | 81/125 | ≥$7.5830 | ≤63.7% | +0 / -6 | 5/11 |
|
||||
| Capability review, quality objective | 87/125 | ≥$10.1853 | ≤51.3% | +7 / -7 | 4/11 |
|
||||
| V2 review, retention objective | 88/125 | ≥$20.9754 | ≤-0.4% | +1 / -0 | 11/11 |
|
||||
| V2 review, quality objective | 88/125 | ≥$8.6493 | ≤58.6% | +8 / -7 | 4/11 |
|
||||
| Hindsight direct routing (not deployable) | 95/125 | ≥$4.5190 | ≤78.4% | +8 / -0 | 11/11 |
|
||||
| Hindsight Luna-first routing (not deployable) | 95/125 | ≥$5.2653 | ≤74.8% | +8 / -0 | 11/11 |
|
||||
|
||||
## mbpp
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol | Sol-only rescues captured |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| Luna only | 38/50 | $0.0174 | 91.6% | +3 / -2 | 0/2 |
|
||||
| Sol only | 37/50 | $0.2064 | 0.0% | +0 / -0 | 2/2 |
|
||||
| Capability upfront, retention objective | 37/50 | $0.2180 | -5.6% | +0 / -0 | 2/2 |
|
||||
| Capability upfront, quality objective | 38/50 | $0.0444 | 78.5% | +3 / -2 | 0/2 |
|
||||
| V2 upfront, retention objective | 37/50 | $0.2298 | -11.3% | +0 / -0 | 2/2 |
|
||||
| V2 upfront, quality objective | 38/50 | $0.0465 | 77.4% | +3 / -2 | 0/2 |
|
||||
| Capability baseline | 38/50 | $0.0290 | 85.9% | +3 / -2 | 0/2 |
|
||||
| Capability with learned cards only | 38/50 | $0.0284 | 86.2% | +3 / -2 | 0/2 |
|
||||
| V2 baseline | 38/50 | $0.0562 | 72.8% | +3 / -2 | 0/2 |
|
||||
| V2 with learned cards only | 38/50 | $0.0396 | 80.8% | +3 / -2 | 0/2 |
|
||||
| Capability review, retention objective | 37/50 | $0.2503 | -21.3% | +0 / -0 | 2/2 |
|
||||
| Capability review, quality objective | 38/50 | $0.0434 | 79.0% | +3 / -2 | 0/2 |
|
||||
| V2 review, retention objective | 37/50 | $0.2441 | -18.3% | +0 / -0 | 2/2 |
|
||||
| V2 review, quality objective | 38/50 | $0.0488 | 76.4% | +3 / -2 | 0/2 |
|
||||
| Luna with public-test escalation | 38/50 | $0.0240 | 88.4% | +2 / -1 | 1/2 |
|
||||
| Hindsight direct routing (not deployable) | 40/50 | $0.0353 | 82.9% | +3 / -0 | 2/2 |
|
||||
| Hindsight Luna-first routing (not deployable) | 40/50 | $0.0361 | 82.5% | +3 / -0 | 2/2 |
|
||||
|
||||
## lcb
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol | Sol-only rescues captured |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| Luna only | 13/25 | $0.1165 | 92.4% | +2 / -3 | 0/3 |
|
||||
| Sol only | 14/25 | $1.5299 | 0.0% | +0 / -0 | 3/3 |
|
||||
| Capability upfront, retention objective | 14/25 | $1.5476 | -1.2% | +0 / -0 | 3/3 |
|
||||
| Capability upfront, quality objective | 14/25 | $0.9200 | 39.9% | +1 / -1 | 2/3 |
|
||||
| V2 upfront, retention objective | 14/25 | $1.3923 | 9.0% | +0 / -0 | 3/3 |
|
||||
| V2 upfront, quality objective | 12/25 | $1.1104 | 27.4% | +0 / -2 | 1/3 |
|
||||
| Capability baseline | 13/25 | $0.2918 | 80.9% | +2 / -3 | 0/3 |
|
||||
| Capability with learned cards only | 12/25 | $0.4860 | 68.2% | +1 / -3 | 0/3 |
|
||||
| V2 baseline | 12/25 | $0.9924 | 35.1% | +0 / -2 | 1/3 |
|
||||
| V2 with learned cards only | 13/25 | $0.4799 | 68.6% | +2 / -3 | 0/3 |
|
||||
| Capability review, retention objective | 14/25 | $1.6846 | -10.1% | +0 / -0 | 3/3 |
|
||||
| Capability review, quality objective | 14/25 | $1.3257 | 13.4% | +2 / -2 | 1/3 |
|
||||
| V2 review, retention objective | 15/25 | $1.6076 | -5.1% | +1 / -0 | 3/3 |
|
||||
| V2 review, quality objective | 14/25 | $1.0584 | 30.8% | +2 / -2 | 1/3 |
|
||||
| Luna with public-test escalation | 15/25 | $1.3645 | 10.8% | +2 / -1 | 2/3 |
|
||||
| Hindsight direct routing (not deployable) | 16/25 | $0.3436 | 77.5% | +2 / -0 | 3/3 |
|
||||
| Hindsight Luna-first routing (not deployable) | 16/25 | $0.3657 | 76.1% | +2 / -0 | 3/3 |
|
||||
|
||||
## swe
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol | Sol-only rescues captured |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| Luna only | 16/25 | $0.6371 | 94.1% | +0 / -2 | 0/2 |
|
||||
| Sol only | 18/25 | $10.7132 | 0.0% | +0 / -0 | 2/2 |
|
||||
| Capability upfront, retention objective | 18/25 | $6.2493 | 41.7% | +0 / -0 | 2/2 |
|
||||
| Capability upfront, quality objective | 18/25 | $9.7546 | 8.9% | +0 / -0 | 2/2 |
|
||||
| V2 upfront, retention objective | 16/25 | $5.3537 | 50.0% | +0 / -2 | 0/2 |
|
||||
| V2 upfront, quality objective | 17/25 | $8.8718 | 17.2% | +0 / -1 | 1/2 |
|
||||
| Capability baseline | 16/25 | $1.3893 | 87.0% | +0 / -2 | 0/2 |
|
||||
| Capability with learned cards only | 16/25 | $1.6826 | 84.3% | +0 / -2 | 0/2 |
|
||||
| V2 baseline | 17/25 | $9.1304 | 14.8% | +0 / -1 | 1/2 |
|
||||
| V2 with learned cards only | 17/25 | $7.4181 | 30.8% | +0 / -1 | 1/2 |
|
||||
| Capability review, retention objective | 16/25 | $2.4752 | 76.9% | +0 / -2 | 0/2 |
|
||||
| Capability review, quality objective | 16/25 | $0.9753 | 90.9% | +0 / -2 | 0/2 |
|
||||
| V2 review, retention objective | 18/25 | $9.1263 | 14.8% | +0 / -0 | 2/2 |
|
||||
| V2 review, quality objective | 16/25 | $0.9775 | 90.9% | +0 / -2 | 0/2 |
|
||||
| Hindsight direct routing (not deployable) | 18/25 | $1.4342 | 86.6% | +0 / -0 | 2/2 |
|
||||
| Hindsight Luna-first routing (not deployable) | 18/25 | $1.5478 | 85.6% | +0 / -0 | 2/2 |
|
||||
|
||||
## terminal
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol | Sol-only rescues captured |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| Luna only | 17/25 | ≥$1.8822 | ≤77.7% | +3 / -4 | 0/4 |
|
||||
| Sol only | 18/25 | $8.4468 | 0.0% | +0 / -0 | 4/4 |
|
||||
| Capability upfront, retention objective | 17/25 | $5.8008 | 31.3% | +1 / -2 | 2/4 |
|
||||
| Capability upfront, quality objective | 18/25 | $7.8047 | 7.6% | +1 / -1 | 3/4 |
|
||||
| V2 upfront, retention objective | 18/25 | $7.0279 | 16.8% | +2 / -2 | 2/4 |
|
||||
| V2 upfront, quality objective | 19/25 | $8.1712 | 3.3% | +1 / -0 | 4/4 |
|
||||
| Capability baseline | 18/25 | $5.0541 | 40.2% | +3 / -3 | 1/4 |
|
||||
| Capability with learned cards only | 18/25 | $5.0526 | 40.2% | +3 / -3 | 1/4 |
|
||||
| V2 baseline | 17/25 | ≥$6.8238 | ≤19.2% | +0 / -1 | 3/4 |
|
||||
| V2 with learned cards only | 17/25 | ≥$6.7112 | ≤20.5% | +1 / -2 | 2/4 |
|
||||
| Capability review, retention objective | 14/25 | ≥$3.1729 | ≤62.4% | +0 / -4 | 0/4 |
|
||||
| Capability review, quality objective | 19/25 | ≥$7.8410 | ≤7.2% | +2 / -1 | 3/4 |
|
||||
| V2 review, retention objective | 18/25 | ≥$9.9975 | ≤-18.4% | +0 / -0 | 4/4 |
|
||||
| V2 review, quality objective | 20/25 | ≥$6.5646 | ≤22.3% | +3 / -1 | 3/4 |
|
||||
| Hindsight direct routing (not deployable) | 21/25 | ≥$2.7059 | ≤68.0% | +3 / -0 | 4/4 |
|
||||
| Hindsight Luna-first routing (not deployable) | 21/25 | ≥$3.3157 | ≤60.7% | +3 / -0 | 4/4 |
|
||||
|
||||
The hindsight rows use hidden outcomes and are unattainable routing references. One attempt per model and small per-benchmark samples do not establish future reliability. JSON records exploratory paired uncertainty, every frozen ablation and per-task decisions. Study training and discarded harness-run spending is separate from deployed-policy cost.
|
||||
|
|
@ -4,6 +4,19 @@ The target is preserving stronger-model task quality while reducing total infere
|
|||
|
||||
The primary model pair is GPT-5.6 Luna and GPT-5.6 Sol at high effort. Sonnet-5 and Opus-5 are additional final-evaluation controls. Model identities and gateway billing must match each response. The original capability classifier and fused V2 are baselines, both using the same Luna judge
|
||||
|
||||
The final allocation below supersedes the initial targets and intermediate preparation counts recorded in the amendment history. Training uses 160 tasks, validation uses 64 eligible tasks, and final evaluation uses 125 tasks. Each final task has one independent attempt from each of the four models, for 500 attempts
|
||||
|
||||
| Benchmark | Training | Validation | Final |
|
||||
|---|---:|---:|---:|
|
||||
| MBPP+ | 80 | 29 | 50 |
|
||||
| LiveCodeBench | 50 | 20 | 25 |
|
||||
| SWE-bench Verified | 25 | 10 | 25 |
|
||||
| Terminal-Bench | 5 | 5 | 25 |
|
||||
|
||||
The initial final runner used three concurrent SWE task workers and two concurrent Terminal trials. The corrected Anthropic-control rerun uses one SWE worker and two Terminal trials, prioritizing remaining Luna/Sol Terminal attempts. These limits share the machine with other work and do not change individual model budgets. The benchmark selection and offline environment controls make this an adapted subset, not an official leaderboard submission
|
||||
|
||||
The following paragraphs preserve the initial plan and its amendments, including why some preparation targets changed
|
||||
|
||||
Development uses 25 SWE-bench Verified training tasks and 10 validation tasks, plus 80 MBPP+ training tasks and 30 validation tasks. Terminal-Bench task counts will be fixed after environment-only eligibility checks, targeting at least 15 training and 10 validation tasks. Final evaluation uses 25 SWE-bench, 25 Terminal-Bench and 50 MBPP+ tasks. SWE-bench repositories are disjoint across splits; previously inspected tasks are excluded. Terminal-Bench variants of the same mechanism must remain in the same split
|
||||
|
||||
Solver outcomes are new independent paired attempts. No gold fixes, future Git objects, hidden tests, prior trajectories or benchmark results are exposed to solvers. Docker hosts, credentials and task sources are not mounted into solver environments. SWE-bench uses a one-commit Git database with no unreachable objects. Terminal-Bench tests and reference solution are uploaded only in isolated control or grading phases. MBPP solutions are generated from prompts only and executed in a sandbox by EvalPlus
|
||||
|
|
@ -49,3 +62,49 @@ Final solver scheduling uses up to four concurrent tasks/trials per agentic benc
|
|||
Development allocation is now fixed at 160 paired training tasks and 64 eligible validation tasks: 80/29 MBPP+, 50/20 LiveCodeBench, 25/10 SWE-bench and 5/5 Terminal-Bench. Additional control-eligible Terminal tasks are reserved for final evaluation or left unused. This keeps card synthesis and all calibration fits on one fixed training corpus while final environment preparation continues
|
||||
|
||||
Before final inference, freeze a diagnostic ablation per classifier, card variant, stage and training family. Each uses the same validation criterion of zero Sol-only losses, no per-benchmark quality loss and cost no greater than Sol. If no family candidate qualifies, record an explicit always-Sol fallback. Report all of these ablations, without selecting another winner from final scores, to distinguish card changes, probability fitting and threshold fitting
|
||||
|
||||
During final execution, a cost audit read a completed Terminal attempt ledger while the periodic exporter was copying that same ledger over its destination. A subsequent read matched the saved run cost exactly. Exported records now use atomic file replacement so concurrent readers never see a partially copied file. No solver attempt, billed response, grade, fitted policy or routing decision was changed
|
||||
|
||||
|
||||
## Terminal runner interruption, September 16 UTC
|
||||
|
||||
At 00:06 UTC the Terminal runner, review watcher and keep-awake processes were absent. No terminating error was recorded; the cause remains unknown. The two active Scheme-interpreter containers were still running and neither had a completed agent record or trial result. Their partial transcripts, 95 billed responses ($4.252549 known), outstanding request IDs and container diffs were archived in `infrastructure_interruptions/20260916-terminal-process-loss`. Two outstanding requests have unknown billing and are not assumed free
|
||||
|
||||
The 58 completed Terminal trial results were hashed and preserved, without consulting their quality labels. Harbor resumed the identical job configuration and reran only the two interrupted attempts and 40 unstarted trials. The custom agent has no supported trajectory-resume path. The restarted processes use independent process sessions; inference budgets, model settings, task allocation and frozen policies are unchanged. Archived interruption costs belong to research overhead, separately from deployment policy cost. No completed attempt was rerun because of its outcome
|
||||
|
||||
|
||||
## Response-limit diagnostics
|
||||
|
||||
Metadata inspection during the still-blinded final run found repeated length-stopped responses that used all 8,192 completion tokens without producing a tool call or visible text. Further inspection discovered the Anthropic message-preservation defect documented below. The affected agentic controls are superseded; their costs remain research overhead, and their responses are excluded from final quality and execution tables. The defect's effect on empty responses has not been established. `execution_diagnostics.json` describes only retained attempts. Empty length stops in corrected attempts are billed behavior under the frozen budget and do not themselves trigger a retry
|
||||
|
||||
## Anthropic tool-continuation correction, September 16 UTC
|
||||
|
||||
The benchmark client reconstructed assistant messages from content, tool calls and OpenAI reasoning items. This discarded Anthropic `thinking_blocks` and `provider_specific_fields`, which must be preserved on tool continuation. The corrected adapter returns the complete Anthropic assistant message unchanged. Both Sonnet and Opus passed live thinking-plus-tool round trips through the gateway. See the [provider's thinking/tool documentation](https://platform.claude.com/docs/en/build-with-claude/thinking-tool-workflows) and `harness_corrections/anthropic_thinking/live_probe.json`
|
||||
|
||||
All 50 SWE and 50 Terminal Anthropic controls are run under the correction, regardless of old outcomes. Fifty completed SWE attempts, 32 completed Terminal attempts and two in-progress Terminal trials were archived with hashes before restarting. The old final quality labels were not inspected. Single-response Anthropic MBPP+/LiveCodeBench controls are unaffected because they have no tool continuation. The OpenAI message merger is unchanged across 3,218 saved responses; 2,691 saved OpenAI tool turns were checked, including 2,678 containing intact reasoning items. All Luna/Sol training, validation, final attempts, frozen coefficients, thresholds and route selections are retained
|
||||
|
||||
The archive records 5,360 unaffected files verified unchanged, the preserved trial-result locations, canceled request IDs and the replacement schedule. Superseded controls, interrupted requests and adapter probes are included separately in research spending. Unknown outstanding-request bills are not assumed zero. SWE reruns use the retained peer's image digest and a fresh model-specific grader run prefix. No task, individual inference budget, prompt, grading criterion or fitted policy changed. The corrected Terminal agent identifies itself as adapter3
|
||||
|
||||
Once all 250 Luna/Sol final attempts and their grades are complete and the post-attempt routing plan is frozen, a primary-pair report may be generated while the additional Anthropic controls finish. This is an early release of the same prespecified comparison, with all 125 final tasks included; it does not select new policies or refit using final outcomes. The complete report still requires all 500 valid final attempts
|
||||
|
||||
|
||||
## Review-route readiness guard
|
||||
|
||||
Before freezing the post-attempt routing plan, the runner now requires each Terminal Luna attempt to have completed grading with no infrastructure exception. Normal agent timeouts remain eligible. This guard inspects only completion and exception metadata; success labels and rewards are not inputs to any routing decision. A synthetic regression check verifies the same readiness result for successful and unsuccessful attempts and rejects transport failures. Learned coefficients, thresholds, features and candidate selections are unchanged
|
||||
|
||||
|
||||
## Host low-power sleep and partial billing, September 16 UTC
|
||||
|
||||
The host entered low-power sleep at 02:42:53 UTC with 1% battery and woke on AC at 03:07:43 UTC, a 1,490-second pause despite the existing idle-sleep assertion. All benchmark processes survived. The in-flight Luna requests on regex-chess and path-tracing returned transport ReadError; the pinned mini-swe-agent model wrapper retried them automatically and both attempts continued. No whole attempt was rerun and no global power settings were changed. Power events and request IDs are retained in infrastructure_interruptions/20260916-low-power-sleep
|
||||
|
||||
Known billed responses stay attached to their attempts. The two interrupted requests have unknown charges. Policy reports count every unmetered request in the selected solver attempt, classifier/reviewer calls and discarded Luna attempt on escalation. Affected costs are lower bounds; savings versus a fully metered Sol baseline are upper bounds. Missing bills are not assumed zero. If a baseline is also unmetered, its cost comparison is marked unresolved. Cost uncertainty intervals use recorded bills only. This accounting change does not alter coefficients, thresholds or routes
|
||||
|
||||
Saved wall durations include host sleep. Elapsed duration is absent from both deterministic route features and reviewer evidence, so the pause itself is not used as a predictive feature. The report makes no latency comparison based on these durations
|
||||
|
||||
|
||||
## Execution budget clarification
|
||||
|
||||
SWE solver containers have a 45-minute configured lifetime in both the original and corrected runner. Each shell command has a 90-second timeout with a five-second kill grace. Terminal agent and verifier timeouts come from each pinned task configuration. These environment limits apply in addition to the 150-call, USD 5 pre-query and 8,192-token response limits; this study does not measure unrestricted model performance.
|
||||
|
||||
|
||||
The final primary billing audit also found one earlier DNA-assembly review ConnectTimeout with no successful response ledger. This is separate from the two host-sleep solver interruptions. Research accounting now discovers transport-only ledgers as well as billed-response ledgers. Review-cascade cost comparisons therefore carry three unmetered requests, while Luna-only comparisons carry two
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
These optional profiles learn when GPT-5.6 Sol adds a successful solve over GPT-5.6 Luna. They use one existing classifier response and a deterministic local prediction, with no second classifier call. They replace the earlier ROI cookbook, whose benchmark results were invalidated and are not the source of these fits
|
||||
|
||||
Training uses 160 fresh paired tasks and selection uses 64 separate validation tasks across SWE-bench, Terminal-Bench, MBPP+ and LiveCodeBench. The fitted policies were frozen before the 125-task held-out evaluation. Held-out results are still in progress
|
||||
Training uses 160 fresh paired tasks and selection uses 64 separate validation tasks across SWE-bench, Terminal-Bench, MBPP+ and LiveCodeBench. The fitted policies were frozen before the 125-task held-out evaluation. The complete 125-task Luna/Sol comparison is available in PRIMARY_FINDINGS.md and PRIMARY_REPORT.md. Additional Sonnet/Opus controls are still running
|
||||
|
||||
`strong-success-retention` minimizes validation cost while preserving every Sol-only success and meeting Sol quality within each benchmark. `quality-first-under-sol-budget` maximizes validation solves under Sol's total inference cost and permits task swaps. These names describe validation objectives, not guarantees on new requests
|
||||
|
||||
|
|
@ -15,3 +15,21 @@ The example config describes the SWE-bench agent harness. Match its execution co
|
|||
`selective_policy` replaces the existing threshold decision and cannot be combined with the older probability calibration. Its `target` identifies the score: `scalar_calibration` and `per_model` estimate a success-probability difference, `paired` estimates Sol-only probability minus Luna-only probability, `rescue` estimates Sol-only probability, and `benefit_per_dollar` divides the paired difference by predicted incremental inference cost. Threshold units depend on this target. Raw classifier probabilities remain available in diagnostics
|
||||
|
||||
Defaults are unchanged when `selective_policy` is omitted. The local implementation validates coefficient dimensions and rejects another classifier's feature schema. Runtime parity checks compare scores and routing choices with the training implementation
|
||||
|
||||
The twelve prespecified upfront family controls are also available in `diagnostic_profiles.json` and `diagnostic_proxy.yaml`. Start the latter config to benchmark aliases such as `selective-v3-diagnostic-original-upfront-per-model`. These controls were frozen before final inference and are included for reproduction and fresh benchmarking. A family with no eligible validation candidate uses a direct Sol deployment that bypasses the classifier. Review cascades still require a separate agent workflow
|
||||
## Held-out results
|
||||
|
||||
These profiles were selected before these results. They are not refitted or selected again on the held-out tasks
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost Sol solves |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Luna only | 84/125 | ≥$2.6531 | ≤87.3% | +8 / -11 |
|
||||
| Sol only | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| Capability baseline | 85/125 | $6.7642 | 67.6% | +8 / -10 |
|
||||
| Capability upfront, retention objective | 86/125 | $13.8156 | 33.9% | +1 / -2 |
|
||||
| Capability upfront, quality objective | 88/125 | $18.5237 | 11.4% | +5 / -4 |
|
||||
| ablation_cap_original_upfront_per_model | 88/125 | $20.3797 | 2.5% | +1 / -0 |
|
||||
|
||||
Costs use recorded gateway bills. A ≥ cost and ≤ savings mark unmetered transport requests: cost is a lower bound and savings versus fully metered Sol is an upper bound. The JSON counts these requests per policy and task. They are not assumed free.
|
||||
|
||||
The per-model family control is an exploratory comparison among forty prespecified variants, not a newly selected winner. No validation-selected profile combines preservation of all Sol successes and lower cost across all 125 tasks. Task-level replay does not measure per-turn routing or Sol continuation from a Luna patch
|
||||
|
|
|
|||
1561
cookbook/auto_router_selective_training/diagnostic_profiles.json
Normal file
1561
cookbook/auto_router_selective_training/diagnostic_profiles.json
Normal file
File diff suppressed because it is too large
Load diff
1039
cookbook/auto_router_selective_training/diagnostic_proxy.yaml
Normal file
1039
cookbook/auto_router_selective_training/diagnostic_proxy.yaml
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -0,0 +1,8 @@
|
|||
{
|
||||
"classifier": "cap",
|
||||
"profiles": 12,
|
||||
"classifier_decisions_verified": 500,
|
||||
"direct_sol_decisions_verified": 1000,
|
||||
"matches_frozen_routes": true,
|
||||
"inputs": "Frozen forecasts and route plans only; no grades or outcome-based selection"
|
||||
}
|
||||
|
|
@ -4,11 +4,14 @@
|
|||
"training_tasks": 160,
|
||||
"validation_tasks": 64,
|
||||
"held_out_tasks": 125,
|
||||
"held_out_status": "in progress",
|
||||
"held_out_status": "primary_pair_complete_anthropic_controls_in_progress",
|
||||
"current_model_pair": [
|
||||
"openai/gpt-5.6-luna",
|
||||
"openai/gpt-5.6-sol"
|
||||
],
|
||||
"judge_effort": "low",
|
||||
"solver_effort": "high"
|
||||
"solver_effort": "high",
|
||||
"diagnostic_upfront_profiles": 12,
|
||||
"diagnostic_profiles_sha256": "468de3b22e309ca3ef800ef90bd9b03a8aca3eebef32bcc2155377b4cf7e1f88",
|
||||
"primary_results_sha256": "e91ec5578d496ad37c1e8cf5760882042f4b7768bd3c2ff5d65d92f19df7a12c"
|
||||
}
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
Loading…
Add table
Reference in a new issue