mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-19 00:01:29 +00:00
docs(router): publish complete selective routing evaluation
This commit is contained in:
parent
734a1b18d0
commit
5523acc0c4
13 changed files with 22264 additions and 7 deletions
|
|
@ -28,4 +28,17 @@ All policies and family controls were frozen before final outcomes. These are ta
|
|||
|
||||
Two Luna solver requests interrupted by host sleep have unknown charges. An earlier unavailable DNA-assembly review has one additional unmetered connection-timeout request. Thus Luna-only comparisons contain two unmetered requests, and review cascades contain three. The report marks affected costs as lower bounds and savings as upper bounds. The three upfront comparisons quoted with exact costs above are fully metered
|
||||
|
||||
The additional Sonnet/Opus controls are still running. Their completion will extend the comparison; it cannot change these frozen Luna/Sol results
|
||||
The additional Sonnet/Opus controls are complete. All 500 final attempts and grades are present, and the complete reporter verified that the previously published Luna/Sol results and calibration diagnostics are unchanged
|
||||
|
||||
| Fixed model | Solved | Recorded inference cost |
|
||||
|---|---:|---:|
|
||||
| Luna | 84/125 | At least $2.6531 |
|
||||
| Sol | 87/125 | $20.8964 |
|
||||
| Sonnet 5 | 81/125 | At least $37.9961 |
|
||||
| Opus 5 | 99/125 | At least $46.0550 |
|
||||
|
||||
Opus solved 22/25 SWE-bench tasks at $10.5296, compared with Sol's 18/25 at $10.7132. On Terminal-Bench it solved 21/25 at at least $32.9157, compared with Sol's 18/25 at $8.4468. These are observed results under the frozen harness and budgets, not unrestricted model rankings. Across all tasks, even perfect selection between the saved Luna and Sol attempts reaches only 95/125, below Opus's 99/125. Changing the candidate model pair is a separate training and evaluation question
|
||||
|
||||
All four models had an 8,192-token response cap and high reasoning effort. Corrected Sonnet and Opus attempts still incurred 116 and 83 length stops with no action or visible text, costing $10.3656 and $17.9127 respectively. Those responses remain in the results. Increasing the response cap could change quality and spending, but this experiment did not test that change. Host sleep and unresolved transport charges also prevent latency conclusions and exact affected cost comparisons
|
||||
|
||||
The practical next experiment is to collect substantially more paired agentic training examples where Sol rescues Luna, then evaluate rescue prediction and verification on fresh tasks. The present fitting set contains only two such agentic cases. The current final set is now diagnostic evidence and must not be reused as an unseen test after training on its failures. Better average probability calibration alone did not deliver the requested selective-rescue precision
|
||||
|
|
|
|||
|
|
@ -108,3 +108,23 @@ SWE solver containers have a 45-minute configured lifetime in both the original
|
|||
|
||||
|
||||
The final primary billing audit also found one earlier DNA-assembly review ConnectTimeout with no successful response ledger. This is separate from the two host-sleep solver interruptions. Research accounting now discovers transport-only ledgers as well as billed-response ledgers. Review-cascade cost comparisons therefore carry three unmetered requests, while Luna-only comparisons carry two
|
||||
|
||||
|
||||
## Lid-closed host sleep, September 16 UTC
|
||||
|
||||
During the remaining Anthropic Terminal controls, the host entered Clamshell Sleep on AC power at full battery and repeatedly returned to Maintenance Sleep. The workers survived; no whole attempt was restarted and no power setting changed. At the 06:12 UTC inspection, fix-ocaml-gc and model-extraction-relu-logits Opus trials remained active. The latter had a pending gateway request whose billing remains unresolved until the response or transport ledger completes. Evidence is retained in infrastructure_interruptions/20260916-clamshell-sleep. Host wall durations include these pauses and are not used as routing features or latency comparisons
|
||||
|
||||
|
||||
## DNS recovery and exact Terminal image reuse, September 16 UTC
|
||||
|
||||
After host sleep, two active Anthropic solver attempts ended with gateway DNS ConnectError and 29 queued controls failed to resolve Docker Hub during environment setup. The runner exited. All 31 failed trial records and exported partial attempts were archived in infrastructure_interruptions/20260916-dns-outage. The two interrupted attempts incurred $1.35120725 in known inference cost and 20 requests with unknown billing. These costs are research overhead, not policy inference cost. Connectivity probes verified host DNS, authenticated gateway access and Docker DNS before resuming
|
||||
|
||||
Recovery validation found that eleven completed Anthropic controls had rebuilt different filesystem layers from the retained Luna/Sol images. Every mismatched control was archived and scheduled again, regardless of score. These controls cost $11.05990725, retained separately as research overhead. The 25 Luna/Sol image pairs match in filesystem layers and runtime image configuration. All retained primary attempts and fitted policies are unchanged
|
||||
|
||||
The Terminal environment now supports an explicit map of the exact saved Sol image for each final task. It verifies local image identity and filesystem layers before startup and verifies the started container image before agent inference. This replaces rebuilds from mutable base tags with the already-cached peer image. The map is terminal_image_pins.json. Task files, solver prompts, budgets, network isolation, and grading rules remain unchanged. The recovery schedules exactly 42 missing controls and preserves 58 valid Terminal results. No live trial was interrupted for this correction
|
||||
|
||||
The model-extraction-relu-logits Opus attempt completed normally after the model wrapper retried the previously pending ReadError. It is retained, with that request's unknown charge included in its cost bounds. The entire attempt was not rerun. Host wall durations remain unsuitable for latency comparisons
|
||||
|
||||
## Final completion
|
||||
|
||||
All 500 held-out model attempts and grades completed on September 17, 2026 UTC. The final audit checked 950 completed development/final attempts and 1,400 forecasts, including paired image identity, Git isolation, archived file hashes and recorded costs. Full reporting verified that primary Luna/Sol results and calibration diagnostics are identical to the earlier primary release. No fitting, policy selection or boundary change followed the final outcomes
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
These optional profiles learn when GPT-5.6 Sol adds a successful solve over GPT-5.6 Luna. They use one existing classifier response and a deterministic local prediction, with no second classifier call. They replace the earlier ROI cookbook, whose benchmark results were invalidated and are not the source of these fits
|
||||
|
||||
Training uses 160 fresh paired tasks and selection uses 64 separate validation tasks across SWE-bench, Terminal-Bench, MBPP+ and LiveCodeBench. The fitted policies were frozen before the 125-task held-out evaluation. The complete 125-task Luna/Sol comparison is available in PRIMARY_FINDINGS.md and PRIMARY_REPORT.md. Additional Sonnet/Opus controls are still running
|
||||
Training uses 160 fresh paired tasks and selection uses 64 separate validation tasks across SWE-bench, Terminal-Bench, MBPP+ and LiveCodeBench. The fitted policies were frozen before the 125-task held-out evaluation. The complete 125-task Luna/Sol comparison is available in PRIMARY_FINDINGS.md and PRIMARY_REPORT.md. The additional Sonnet/Opus controls are complete in REPORT.md
|
||||
|
||||
`strong-success-retention` minimizes validation cost while preserving every Sol-only success and meeting Sol quality within each benchmark. `quality-first-under-sol-budget` maximizes validation solves under Sol's total inference cost and permits task swaps. These names describe validation objectives, not guarantees on new requests
|
||||
|
||||
|
|
@ -21,15 +21,19 @@ The twelve prespecified upfront family controls are also available in `diagnosti
|
|||
|
||||
These profiles were selected before these results. They are not refitted or selected again on the held-out tasks
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost Sol solves |
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Luna only | 84/125 | ≥$2.6531 | ≤87.3% | +8 / -11 |
|
||||
| Sol only | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| Capability baseline | 85/125 | $6.7642 | 67.6% | +8 / -10 |
|
||||
| Capability upfront, retention objective | 86/125 | $13.8156 | 33.9% | +1 / -2 |
|
||||
| Capability upfront, quality objective | 88/125 | $18.5237 | 11.4% | +5 / -4 |
|
||||
| ablation_cap_original_upfront_per_model | 88/125 | $20.3797 | 2.5% | +1 / -0 |
|
||||
| Per-model control (exploratory) | 88/125 | $20.3797 | 2.5% | +1 / -0 |
|
||||
|
||||
Costs use recorded gateway bills. A ≥ cost and ≤ savings mark unmetered transport requests: cost is a lower bound and savings versus fully metered Sol is an upper bound. The JSON counts these requests per policy and task. They are not assumed free.
|
||||
|
||||
The per-model family control is an exploratory comparison among forty prespecified variants, not a newly selected winner. No validation-selected profile combines preservation of all Sol successes and lower cost across all 125 tasks. Task-level replay does not measure per-turn routing or Sol continuation from a Luna patch
|
||||
Costs include the classifier. These are task-pinned replays of fresh independent attempts, with one attempt per model per task. No guarantee of quality retention follows from a 125-task sample. The aggregate combines different benchmarks; inspect each benchmark before using a profile
|
||||
|
||||
The per-model control is an exploratory comparison among forty prespecified variants. No validation-selected profile combines preservation of all Sol successes and lower cost across all 125 tasks. See PRIMARY_FINDINGS.md for the interpretation
|
||||
|
||||

|
||||
|
|
|
|||
204
cookbook/auto_router_selective_training/REPORT.md
Normal file
204
cookbook/auto_router_selective_training/REPORT.md
Normal file
|
|
@ -0,0 +1,204 @@
|
|||
# Clean selective-router evaluation
|
||||
|
||||
Policies were frozen before final solver outcomes. Upfront routing uses one classifier response; the post-attempt variants also use Luna output, optional public-example checks and a Luna review. These are task-level replays over matched independent attempts, not live per-turn routing or continuation from a cheaper model's patch
|
||||
|
||||
The configured original-card baselines receive the same model identity and execution conditions as the trained variants. Capability uses base_threshold=0.5 and threshold_step=0.1; V2 uses max_quality_gap=0.05. The empirical-card baselines change the card supplement while retaining those boundaries
|
||||
|
||||
Retention profiles were selected to preserve Sol successes at lower validation cost. Quality profiles were selected to maximize validation solves within the Sol budget. These names describe selection objectives, not guarantees about the held-out results. Exact policy IDs remain in final_results.json
|
||||
|
||||
Costs use recorded gateway bills. A ≥ cost and ≤ savings mark unmetered transport requests: cost is a lower bound and savings versus fully metered Sol is an upper bound. The JSON counts these requests per policy and task. They are not assumed free.
|
||||
|
||||
## Observed opportunity for routing
|
||||
|
||||
The Sol-only column counts the rescues a perfect decision would capture. The hindsight cascade pays for Luna on every task and adds Sol only for those rescues. Tasks where both attempts failed cannot be recovered by choosing between these two saved outputs. This is a bound on this replay, not a guarantee about future attempts
|
||||
|
||||
| Benchmark | Both solve | Luna only | Sol only | Both fail | Hindsight cascade cost |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| all | 76 | 8 | 11 | 30 | ≥$5.2653 |
|
||||
| lcb | 11 | 2 | 3 | 9 | $0.3657 |
|
||||
| mbpp | 35 | 3 | 2 | 10 | $0.0361 |
|
||||
| swe | 16 | 0 | 2 | 7 | $1.5478 |
|
||||
| terminal | 14 | 3 | 4 | 4 | ≥$3.3157 |
|
||||
|
||||
## all
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Luna only | 84/125 | ≥$2.6531 | ≤87.3% | +8 / -11 |
|
||||
| Sol only | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| Sonnet 5 only | 81/125 | ≥$37.9961 | ≤-81.8% | +8 / -14 |
|
||||
| Opus 5 only | 99/125 | ≥$46.0550 | ≤-120.4% | +19 / -7 |
|
||||
| Capability upfront, retention objective | 86/125 | $13.8156 | 33.9% | +1 / -2 |
|
||||
| Capability upfront, quality objective | 88/125 | $18.5237 | 11.4% | +5 / -4 |
|
||||
| V2 upfront, retention objective | 85/125 | $14.0036 | 33.0% | +2 / -4 |
|
||||
| V2 upfront, quality objective | 86/125 | $18.1999 | 12.9% | +4 / -5 |
|
||||
| Capability baseline | 85/125 | $6.7642 | 67.6% | +8 / -10 |
|
||||
| Capability with learned cards only | 84/125 | $7.2496 | 65.3% | +7 / -10 |
|
||||
| V2 baseline | 84/125 | ≥$17.0029 | ≤18.6% | +3 / -6 |
|
||||
| V2 with learned cards only | 85/125 | ≥$14.6487 | ≤29.9% | +6 / -8 |
|
||||
| Capability review, retention objective | 81/125 | ≥$7.5830 | ≤63.7% | +0 / -6 |
|
||||
| Capability review, quality objective | 87/125 | ≥$10.1853 | ≤51.3% | +7 / -7 |
|
||||
| V2 review, retention objective | 88/125 | ≥$20.9754 | ≤-0.4% | +1 / -0 |
|
||||
| V2 review, quality objective | 88/125 | ≥$8.6493 | ≤58.6% | +8 / -7 |
|
||||
| Hindsight direct routing (not deployable) | 95/125 | ≥$4.5190 | ≤78.4% | +8 / -0 |
|
||||
| Hindsight Luna-first routing (not deployable) | 95/125 | ≥$5.2653 | ≤74.8% | +8 / -0 |
|
||||
|
||||
## lcb
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Luna only | 13/25 | $0.1165 | 92.4% | +2 / -3 |
|
||||
| Sol only | 14/25 | $1.5299 | 0.0% | +0 / -0 |
|
||||
| Sonnet 5 only | 9/25 | $1.3790 | 9.9% | +1 / -6 |
|
||||
| Opus 5 only | 14/25 | $2.3849 | -55.9% | +2 / -2 |
|
||||
| Capability upfront, retention objective | 14/25 | $1.5476 | -1.2% | +0 / -0 |
|
||||
| Capability upfront, quality objective | 14/25 | $0.9200 | 39.9% | +1 / -1 |
|
||||
| V2 upfront, retention objective | 14/25 | $1.3923 | 9.0% | +0 / -0 |
|
||||
| V2 upfront, quality objective | 12/25 | $1.1104 | 27.4% | +0 / -2 |
|
||||
| Capability baseline | 13/25 | $0.2918 | 80.9% | +2 / -3 |
|
||||
| Capability with learned cards only | 12/25 | $0.4860 | 68.2% | +1 / -3 |
|
||||
| V2 baseline | 12/25 | $0.9924 | 35.1% | +0 / -2 |
|
||||
| V2 with learned cards only | 13/25 | $0.4799 | 68.6% | +2 / -3 |
|
||||
| Capability review, retention objective | 14/25 | $1.6846 | -10.1% | +0 / -0 |
|
||||
| Capability review, quality objective | 14/25 | $1.3257 | 13.4% | +2 / -2 |
|
||||
| V2 review, retention objective | 15/25 | $1.6076 | -5.1% | +1 / -0 |
|
||||
| V2 review, quality objective | 14/25 | $1.0584 | 30.8% | +2 / -2 |
|
||||
| Luna with public-test escalation | 15/25 | $1.3645 | 10.8% | +2 / -1 |
|
||||
| Hindsight direct routing (not deployable) | 16/25 | $0.3436 | 77.5% | +2 / -0 |
|
||||
| Hindsight Luna-first routing (not deployable) | 16/25 | $0.3657 | 76.1% | +2 / -0 |
|
||||
|
||||
## mbpp
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Luna only | 38/50 | $0.0174 | 91.6% | +3 / -2 |
|
||||
| Sol only | 37/50 | $0.2064 | 0.0% | +0 / -0 |
|
||||
| Sonnet 5 only | 36/50 | $0.0988 | 52.1% | +2 / -3 |
|
||||
| Opus 5 only | 42/50 | $0.2249 | -9.0% | +7 / -2 |
|
||||
| Capability upfront, retention objective | 37/50 | $0.2180 | -5.6% | +0 / -0 |
|
||||
| Capability upfront, quality objective | 38/50 | $0.0444 | 78.5% | +3 / -2 |
|
||||
| V2 upfront, retention objective | 37/50 | $0.2298 | -11.3% | +0 / -0 |
|
||||
| V2 upfront, quality objective | 38/50 | $0.0465 | 77.4% | +3 / -2 |
|
||||
| Capability baseline | 38/50 | $0.0290 | 85.9% | +3 / -2 |
|
||||
| Capability with learned cards only | 38/50 | $0.0284 | 86.2% | +3 / -2 |
|
||||
| V2 baseline | 38/50 | $0.0562 | 72.8% | +3 / -2 |
|
||||
| V2 with learned cards only | 38/50 | $0.0396 | 80.8% | +3 / -2 |
|
||||
| Capability review, retention objective | 37/50 | $0.2503 | -21.3% | +0 / -0 |
|
||||
| Capability review, quality objective | 38/50 | $0.0434 | 79.0% | +3 / -2 |
|
||||
| V2 review, retention objective | 37/50 | $0.2441 | -18.3% | +0 / -0 |
|
||||
| V2 review, quality objective | 38/50 | $0.0488 | 76.4% | +3 / -2 |
|
||||
| Luna with public-test escalation | 38/50 | $0.0240 | 88.4% | +2 / -1 |
|
||||
| Hindsight direct routing (not deployable) | 40/50 | $0.0353 | 82.9% | +3 / -0 |
|
||||
| Hindsight Luna-first routing (not deployable) | 40/50 | $0.0361 | 82.5% | +3 / -0 |
|
||||
|
||||
## swe
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Luna only | 16/25 | $0.6371 | 94.1% | +0 / -2 |
|
||||
| Sol only | 18/25 | $10.7132 | 0.0% | +0 / -0 |
|
||||
| Sonnet 5 only | 19/25 | $11.0537 | -3.2% | +2 / -1 |
|
||||
| Opus 5 only | 22/25 | $10.5296 | 1.7% | +4 / -0 |
|
||||
| Capability upfront, retention objective | 18/25 | $6.2493 | 41.7% | +0 / -0 |
|
||||
| Capability upfront, quality objective | 18/25 | $9.7546 | 8.9% | +0 / -0 |
|
||||
| V2 upfront, retention objective | 16/25 | $5.3537 | 50.0% | +0 / -2 |
|
||||
| V2 upfront, quality objective | 17/25 | $8.8718 | 17.2% | +0 / -1 |
|
||||
| Capability baseline | 16/25 | $1.3893 | 87.0% | +0 / -2 |
|
||||
| Capability with learned cards only | 16/25 | $1.6826 | 84.3% | +0 / -2 |
|
||||
| V2 baseline | 17/25 | $9.1304 | 14.8% | +0 / -1 |
|
||||
| V2 with learned cards only | 17/25 | $7.4181 | 30.8% | +0 / -1 |
|
||||
| Capability review, retention objective | 16/25 | $2.4752 | 76.9% | +0 / -2 |
|
||||
| Capability review, quality objective | 16/25 | $0.9753 | 90.9% | +0 / -2 |
|
||||
| V2 review, retention objective | 18/25 | $9.1263 | 14.8% | +0 / -0 |
|
||||
| V2 review, quality objective | 16/25 | $0.9775 | 90.9% | +0 / -2 |
|
||||
| Hindsight direct routing (not deployable) | 18/25 | $1.4342 | 86.6% | +0 / -0 |
|
||||
| Hindsight Luna-first routing (not deployable) | 18/25 | $1.5478 | 85.6% | +0 / -0 |
|
||||
|
||||
## terminal
|
||||
|
||||
| Policy | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Luna only | 17/25 | ≥$1.8822 | ≤77.7% | +3 / -4 |
|
||||
| Sol only | 18/25 | $8.4468 | 0.0% | +0 / -0 |
|
||||
| Sonnet 5 only | 17/25 | ≥$25.4647 | ≤-201.5% | +3 / -4 |
|
||||
| Opus 5 only | 21/25 | ≥$32.9157 | ≤-289.7% | +6 / -3 |
|
||||
| Capability upfront, retention objective | 17/25 | $5.8008 | 31.3% | +1 / -2 |
|
||||
| Capability upfront, quality objective | 18/25 | $7.8047 | 7.6% | +1 / -1 |
|
||||
| V2 upfront, retention objective | 18/25 | $7.0279 | 16.8% | +2 / -2 |
|
||||
| V2 upfront, quality objective | 19/25 | $8.1712 | 3.3% | +1 / -0 |
|
||||
| Capability baseline | 18/25 | $5.0541 | 40.2% | +3 / -3 |
|
||||
| Capability with learned cards only | 18/25 | $5.0526 | 40.2% | +3 / -3 |
|
||||
| V2 baseline | 17/25 | ≥$6.8238 | ≤19.2% | +0 / -1 |
|
||||
| V2 with learned cards only | 17/25 | ≥$6.7112 | ≤20.5% | +1 / -2 |
|
||||
| Capability review, retention objective | 14/25 | ≥$3.1729 | ≤62.4% | +0 / -4 |
|
||||
| Capability review, quality objective | 19/25 | ≥$7.8410 | ≤7.2% | +2 / -1 |
|
||||
| V2 review, retention objective | 18/25 | ≥$9.9975 | ≤-18.4% | +0 / -0 |
|
||||
| V2 review, quality objective | 20/25 | ≥$6.5646 | ≤22.3% | +3 / -1 |
|
||||
| Hindsight direct routing (not deployable) | 21/25 | ≥$2.7059 | ≤68.0% | +3 / -0 |
|
||||
| Hindsight Luna-first routing (not deployable) | 21/25 | ≥$3.3157 | ≤60.7% | +3 / -0 |
|
||||
|
||||
## Frozen family comparisons
|
||||
|
||||
Each row was selected on development data before final inference. These are diagnostic comparisons, not a new winner selection on the held-out tasks. An always-Sol fallback means that no candidate in that requested family met the validation criterion; it bypasses the classifier, Luna attempt and reviewer
|
||||
|
||||
| Classifier | Requested stage | Cards | Fitted target | Boundary | Solved | Inference cost | Savings vs Sol | Gained / lost vs Sol |
|
||||
|---|---|---|---|---:|---:|---:|---:|---:|
|
||||
| cap | upfront | empirical | benefit_per_dollar | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | upfront | empirical | paired | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | upfront | empirical | per_model | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | upfront | empirical | raw | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | upfront | empirical | rescue | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | upfront | empirical | scalar_calibration | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | review | empirical | paired | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | review | empirical | per_model | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | review | empirical | rescue | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | upfront | original | benefit_per_dollar | 1 | 86/125 | $13.8156 | 33.9% | +1 / -2 |
|
||||
| cap | upfront | original | paired | -0.02 | 88/125 | $20.8563 | 0.2% | +1 / -0 |
|
||||
| cap | upfront | original | per_model | -0.05 | 88/125 | $20.3797 | 2.5% | +1 / -0 |
|
||||
| cap | upfront | original | raw | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | upfront | original | rescue | 0.03 | 88/125 | $20.7799 | 0.6% | +1 / -0 |
|
||||
| cap | upfront | original | scalar_calibration | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | review | original | paired | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | review | original | per_model | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | review | original | rescue | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | upfront | empirical | benefit_per_dollar | 0.5 | 87/125 | $19.2534 | 7.9% | +2 / -2 |
|
||||
| v2 | upfront | empirical | paired | 0.01 | 85/125 | $19.7030 | 5.7% | +1 / -3 |
|
||||
| v2 | upfront | empirical | per_model | 0.01 | 86/125 | $19.7351 | 5.6% | +1 / -2 |
|
||||
| v2 | upfront | empirical | raw | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | upfront | empirical | rescue | 0.05 | 86/125 | $18.0048 | 13.8% | +2 / -3 |
|
||||
| v2 | upfront | empirical | scalar_calibration | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | review | empirical | paired | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | review | empirical | per_model | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | review | empirical | rescue | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | upfront | original | benefit_per_dollar | 1 | 85/125 | $14.0036 | 33.0% | +2 / -4 |
|
||||
| v2 | upfront | original | paired | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | upfront | original | per_model | 0 | 86/125 | $20.1971 | 3.3% | +0 / -1 |
|
||||
| v2 | upfront | original | raw | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | upfront | original | rescue | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | upfront | original | scalar_calibration | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| v2 | review | original | paired | always Sol fallback | 87/125 | $20.8964 | 0.0% | +0 / -0 |
|
||||
| cap | review | empirical | benefit_per_dollar | 0.1 | 82/125 | ≥$16.8075 | ≤19.6% | +0 / -5 |
|
||||
| cap | review | original | benefit_per_dollar | 0.25 | 81/125 | ≥$7.5830 | ≤63.7% | +0 / -6 |
|
||||
| v2 | review | empirical | benefit_per_dollar | 0.25 | 83/125 | ≥$16.3243 | ≤21.9% | +0 / -4 |
|
||||
| v2 | review | original | benefit_per_dollar | 0.1 | 85/125 | ≥$19.8241 | ≤5.1% | +0 / -2 |
|
||||
| v2 | review | original | per_model | -0.02 | 88/125 | ≥$20.9754 | ≤-0.4% | +1 / -0 |
|
||||
| v2 | review | original | rescue | 0.02 | 86/125 | ≥$23.4088 | ≤-12.0% | +0 / -1 |
|
||||
|
||||
## Response budget effects
|
||||
|
||||
Every solver uses the same 8,192-token response cap and high reasoning effort. Length stops that produce neither a tool call nor visible text still consume inference cost. They are retained as model behavior under this budget, not retried as infrastructure failures. These counts describe completed final attempts across all four benchmarks
|
||||
|
||||
| Model | Responses | Length stops | Length stops without action or text | Cost of those empty length stops |
|
||||
|---|---:|---:|---:|---:|
|
||||
| gpt-5.6-luna | 1341 | 50 | 42 | $0.4217 |
|
||||
| gpt-5.6-sol | 819 | 7 | 4 | $0.6676 |
|
||||
| claude-sonnet-5 | 1703 | 124 | 116 | $10.3656 |
|
||||
| claude-opus-5 | 870 | 87 | 83 | $17.9127 |
|
||||
|
||||
## Experiment spending
|
||||
|
||||
Recorded study inference cost is $225.5557. This includes training, validation, final solver/judge/reviewer calls, superseded development forecasts, archived infrastructure-interrupted attempts, superseded Anthropic agentic controls, controls with mismatched Terminal images, and adapter round-trip probes. It is separate from the deployed-policy costs above. Unknown transport billing, setup, live-proxy QA, local Docker and CPU expenses are not included. See research_costs.json for the breakdown
|
||||
|
||||
## Interpretation limits
|
||||
|
||||
The oracle is a hindsight reference using known costs. Cost intervals use recorded bills and exclude unmetered requests. A small sample with one attempt per model cannot establish the same routing advantage across workloads or repeated runs. Paired bootstrap intervals and exact discordant-pair tests in the JSON are exploratory; policy selection searched many development candidates. Inference costs include classifier/reviewer calls and discarded Luna attempts on escalation. Local checker CPU time and experiment training/transport overhead are separate. Terminal-Bench and SWE-bench use documented offline/native-image eligibility adaptations
|
||||
49
cookbook/auto_router_selective_training/TRAINING.md
Normal file
49
cookbook/auto_router_selective_training/TRAINING.md
Normal file
|
|
@ -0,0 +1,49 @@
|
|||
# Training and selection record
|
||||
|
||||
The learning target is model complementarity: route cheaply when Luna can do the job, and use Sol when its extra cost is likely to produce an additional solve. Better marginal probability accuracy alone is not the selection objective
|
||||
|
||||
## Paired examples
|
||||
|
||||
| Split | Benchmark | Tasks | Both fail | Sol only | Luna only | Both solve |
|
||||
|---|---|---:|---:|---:|---:|---:|
|
||||
| train | lcb | 50 | 14 | 9 | 3 | 24 |
|
||||
| train | mbpp | 80 | 7 | 2 | 3 | 68 |
|
||||
| train | swe | 25 | 5 | 1 | 0 | 19 |
|
||||
| train | terminal | 5 | 0 | 1 | 0 | 4 |
|
||||
| validation | lcb | 20 | 2 | 6 | 2 | 10 |
|
||||
| validation | mbpp | 29 | 4 | 2 | 2 | 21 |
|
||||
| validation | swe | 10 | 1 | 1 | 1 | 7 |
|
||||
| validation | terminal | 5 | 2 | 0 | 1 | 2 |
|
||||
|
||||
There are 13 Sol-only examples in fitting and 9 in validation. Only two fitting examples are agentic rescues, one SWE-bench and one Terminal-Bench. Terminal validation has no Sol-only example. That limits evidence about rescue detection on unseen terminal tasks; zero validation losses there cannot establish rescue recall
|
||||
|
||||
## Variations
|
||||
|
||||
One training-only synthesis produces the empirical card supplement. It uses measured outcomes and fallible classifier/reviewer descriptions, without task identities, solutions or validation-driven rewrites. Original cards are a separate control
|
||||
|
||||
The deterministic head families are raw threshold tuning, scalar probability calibration, task-dependent per-model calibration, four-way paired outcomes, Sol-only rescue probability and paired benefit per predicted incremental dollar. Regularization and fixed threshold grids are selected using validation only. Upfront inputs come from one classifier response. Post-attempt heads additionally use Luna output, a fallible Luna review and task-supplied public checks where available
|
||||
|
||||
The strong-success-retention objective minimizes validation cost while losing no Sol-only successes and preserving Sol aggregate quality in each benchmark. The quality-first-under-sol-budget objective maximizes validation solves under Sol total cost, permitting task swaps. Always-Sol is an explicit fallback. None of these objective names guarantees performance on unseen tasks
|
||||
|
||||
## Frozen selected profiles
|
||||
|
||||
| Classifier | Stage | Objective | Cards | Family | Threshold | Validation solved | Validation cost | Sol-only losses |
|
||||
|---|---|---|---|---|---:|---:|---:|---:|
|
||||
| cap | upfront | strong_success_retention | original | benefit_per_dollar | 1.0 | 50/64 | $5.8544 | 0 |
|
||||
| cap | upfront | quality_first_under_sol_budget | original | scalar_calibration | 0.1 | 52/64 | $5.4692 | 2 |
|
||||
| cap | review | strong_success_retention | original | benefit_per_dollar | 0.25 | 49/64 | $5.6530 | 0 |
|
||||
| cap | review | quality_first_under_sol_budget | empirical | paired | 0.075 | 53/64 | $3.7201 | 2 |
|
||||
| cap | either | strong_success_retention | original | benefit_per_dollar | 0.25 | 49/64 | $5.6530 | 0 |
|
||||
| cap | either | quality_first_under_sol_budget | empirical | paired | 0.075 | 53/64 | $3.7201 | 2 |
|
||||
| v2 | upfront | strong_success_retention | original | benefit_per_dollar | 1.0 | 51/64 | $5.4640 | 0 |
|
||||
| v2 | upfront | quality_first_under_sol_budget | original | scalar_calibration | 0.075 | 53/64 | $5.4524 | 2 |
|
||||
| v2 | review | strong_success_retention | original | per_model | -0.02 | 51/64 | $6.7043 | 0 |
|
||||
| v2 | review | quality_first_under_sol_budget | empirical | per_model | 0.15 | 53/64 | $3.3074 | 2 |
|
||||
| v2 | either | strong_success_retention | original | benefit_per_dollar | 1.0 | 51/64 | $5.4640 | 0 |
|
||||
| v2 | either | quality_first_under_sol_budget | empirical | per_model | 0.15 | 53/64 | $3.3074 | 2 |
|
||||
|
||||
These are development results. Compare held-out results in REPORT.md before interpreting them as an improvement. Both selected upfront classifier profiles keep the original card and use fitted local decisions. Empirical cards remain a measured variation and are selected by some review cascades
|
||||
|
||||
There are 1,892 policy candidates, 118 serialized models, 12 selected profiles and 40 prespecified family ablations. An independent refit reproduced every coefficient, candidate statistic and selection exactly. The selection hash and source/data hashes are recorded in selection_frozen.json; reproduction evidence is in training_reproduction.json
|
||||
|
||||
The threshold unit depends on the target. Marginal and paired gaps are probability differences, rescue is a Sol-only probability, and benefit_per_dollar is expected additional success per predicted dollar. A threshold of 0.05 is not interchangeable across these families
|
||||
11707
cookbook/auto_router_selective_training/benchmark_summary.json
Normal file
11707
cookbook/auto_router_selective_training/benchmark_summary.json
Normal file
File diff suppressed because it is too large
Load diff
5849
cookbook/auto_router_selective_training/calibration_results.json
Normal file
5849
cookbook/auto_router_selective_training/calibration_results.json
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -0,0 +1,62 @@
|
|||
{
|
||||
"train": {
|
||||
"lcb": {
|
||||
"tasks": 50,
|
||||
"both_fail": 14,
|
||||
"sol_only": 9,
|
||||
"luna_only": 3,
|
||||
"both_solve": 24
|
||||
},
|
||||
"mbpp": {
|
||||
"tasks": 80,
|
||||
"both_fail": 7,
|
||||
"sol_only": 2,
|
||||
"luna_only": 3,
|
||||
"both_solve": 68
|
||||
},
|
||||
"swe": {
|
||||
"tasks": 25,
|
||||
"both_fail": 5,
|
||||
"sol_only": 1,
|
||||
"luna_only": 0,
|
||||
"both_solve": 19
|
||||
},
|
||||
"terminal": {
|
||||
"tasks": 5,
|
||||
"both_fail": 0,
|
||||
"sol_only": 1,
|
||||
"luna_only": 0,
|
||||
"both_solve": 4
|
||||
}
|
||||
},
|
||||
"validation": {
|
||||
"lcb": {
|
||||
"tasks": 20,
|
||||
"both_fail": 2,
|
||||
"sol_only": 6,
|
||||
"luna_only": 2,
|
||||
"both_solve": 10
|
||||
},
|
||||
"mbpp": {
|
||||
"tasks": 29,
|
||||
"both_fail": 4,
|
||||
"sol_only": 2,
|
||||
"luna_only": 2,
|
||||
"both_solve": 21
|
||||
},
|
||||
"swe": {
|
||||
"tasks": 10,
|
||||
"both_fail": 1,
|
||||
"sol_only": 1,
|
||||
"luna_only": 1,
|
||||
"both_solve": 7
|
||||
},
|
||||
"terminal": {
|
||||
"tasks": 5,
|
||||
"both_fail": 2,
|
||||
"sol_only": 0,
|
||||
"luna_only": 1,
|
||||
"both_solve": 2
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -0,0 +1,315 @@
|
|||
{
|
||||
"scope": "Execution diagnostics for completed attempts, without reading grades. A length stop is a model output-budget event, not an infrastructure failure. All recorded responses remain billed.",
|
||||
"groups": [
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "lcb",
|
||||
"model": "anthropic/claude-opus-5",
|
||||
"attempts": 25,
|
||||
"responses": 25,
|
||||
"response_cost": 2.384855,
|
||||
"token_limit_responses": 7,
|
||||
"token_limit_response_cost": 1.4849325,
|
||||
"token_limit_without_action_or_text": 7,
|
||||
"token_limit_without_action_or_text_cost": 1.4849325
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "lcb",
|
||||
"model": "anthropic/claude-sonnet-5",
|
||||
"attempts": 25,
|
||||
"responses": 25,
|
||||
"response_cost": 1.3789665000000002,
|
||||
"token_limit_responses": 14,
|
||||
"token_limit_response_cost": 1.1779045000000001,
|
||||
"token_limit_without_action_or_text": 14,
|
||||
"token_limit_without_action_or_text_cost": 1.1779045000000001
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "lcb",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 25,
|
||||
"responses": 25,
|
||||
"response_cost": 0.1164804,
|
||||
"token_limit_responses": 7,
|
||||
"token_limit_response_cost": 0.06988079999999999,
|
||||
"token_limit_without_action_or_text": 7,
|
||||
"token_limit_without_action_or_text_cost": 0.06988079999999999
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "lcb",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 25,
|
||||
"responses": 25,
|
||||
"response_cost": 1.529944,
|
||||
"token_limit_responses": 4,
|
||||
"token_limit_response_cost": 0.667568,
|
||||
"token_limit_without_action_or_text": 4,
|
||||
"token_limit_without_action_or_text_cost": 0.667568
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "mbpp",
|
||||
"model": "anthropic/claude-opus-5",
|
||||
"attempts": 50,
|
||||
"responses": 50,
|
||||
"response_cost": 0.224865
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "mbpp",
|
||||
"model": "anthropic/claude-sonnet-5",
|
||||
"attempts": 50,
|
||||
"responses": 50,
|
||||
"response_cost": 0.09884600000000003
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "mbpp",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 50,
|
||||
"responses": 50,
|
||||
"response_cost": 0.017396
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "mbpp",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 50,
|
||||
"responses": 50,
|
||||
"response_cost": 0.20637600000000006
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "swe",
|
||||
"model": "anthropic/claude-opus-5",
|
||||
"attempts": 25,
|
||||
"responses": 381,
|
||||
"response_cost": 10.529565750000002
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "swe",
|
||||
"model": "anthropic/claude-sonnet-5",
|
||||
"attempts": 25,
|
||||
"responses": 833,
|
||||
"response_cost": 11.053656000000016,
|
||||
"token_limit_responses": 4,
|
||||
"token_limit_response_cost": 0.4144736,
|
||||
"token_limit_without_action_or_text": 4,
|
||||
"token_limit_without_action_or_text_cost": 0.4144736
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "swe",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 25,
|
||||
"responses": 427,
|
||||
"response_cost": 0.6371064399999993
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "swe",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 25,
|
||||
"responses": 403,
|
||||
"response_cost": 10.713241600000002
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "terminal",
|
||||
"model": "anthropic/claude-opus-5",
|
||||
"attempts": 25,
|
||||
"responses": 414,
|
||||
"response_cost": 32.915727249999996,
|
||||
"token_limit_responses": 80,
|
||||
"token_limit_response_cost": 17.381599,
|
||||
"token_limit_without_action_or_text": 76,
|
||||
"token_limit_without_action_or_text_cost": 16.427720750000002
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "terminal",
|
||||
"model": "anthropic/claude-sonnet-5",
|
||||
"attempts": 25,
|
||||
"responses": 795,
|
||||
"response_cost": 25.464663800000025,
|
||||
"token_limit_responses": 106,
|
||||
"token_limit_response_cost": 9.517311799999998,
|
||||
"token_limit_without_action_or_text": 98,
|
||||
"token_limit_without_action_or_text_cost": 8.773203099999998
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "terminal",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 25,
|
||||
"responses": 839,
|
||||
"response_cost": 1.8821569099999986,
|
||||
"token_limit_responses": 43,
|
||||
"token_limit_response_cost": 0.43733629999999996,
|
||||
"token_limit_without_action_or_text": 35,
|
||||
"token_limit_without_action_or_text_cost": 0.35178958
|
||||
},
|
||||
{
|
||||
"split": "test",
|
||||
"benchmark": "terminal",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 25,
|
||||
"responses": 341,
|
||||
"response_cost": 8.446844399999998,
|
||||
"token_limit_responses": 3,
|
||||
"token_limit_response_cost": 0.5060674000000001
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "lcb",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 50,
|
||||
"responses": 50,
|
||||
"response_cost": 0.2573764,
|
||||
"token_limit_responses": 16,
|
||||
"token_limit_response_cost": 0.1596568,
|
||||
"token_limit_without_action_or_text": 16,
|
||||
"token_limit_without_action_or_text_cost": 0.1596568
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "lcb",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 50,
|
||||
"responses": 50,
|
||||
"response_cost": 2.630148,
|
||||
"token_limit_responses": 5,
|
||||
"token_limit_response_cost": 0.836052,
|
||||
"token_limit_without_action_or_text": 5,
|
||||
"token_limit_without_action_or_text_cost": 0.836052
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "mbpp",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 80,
|
||||
"responses": 80,
|
||||
"response_cost": 0.0218284
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "mbpp",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 80,
|
||||
"responses": 80,
|
||||
"response_cost": 0.289156
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "swe",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 25,
|
||||
"responses": 398,
|
||||
"response_cost": 0.5176403399999997
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "swe",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 25,
|
||||
"responses": 346,
|
||||
"response_cost": 7.403714199999999
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "terminal",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 5,
|
||||
"responses": 87,
|
||||
"response_cost": 0.12343140999999999
|
||||
},
|
||||
{
|
||||
"split": "train",
|
||||
"benchmark": "terminal",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 5,
|
||||
"responses": 34,
|
||||
"response_cost": 0.7672414000000001
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "lcb",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 20,
|
||||
"responses": 20,
|
||||
"response_cost": 0.0921004,
|
||||
"token_limit_responses": 6,
|
||||
"token_limit_response_cost": 0.0598586,
|
||||
"token_limit_without_action_or_text": 6,
|
||||
"token_limit_without_action_or_text_cost": 0.0598586
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "lcb",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 20,
|
||||
"responses": 20,
|
||||
"response_cost": 1.0552880000000002,
|
||||
"token_limit_responses": 1,
|
||||
"token_limit_response_cost": 0.166696,
|
||||
"token_limit_without_action_or_text": 1,
|
||||
"token_limit_without_action_or_text_cost": 0.166696
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "mbpp",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 30,
|
||||
"responses": 30,
|
||||
"response_cost": 0.010364400000000001
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "mbpp",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 30,
|
||||
"responses": 30,
|
||||
"response_cost": 0.14699600000000002
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "swe",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 10,
|
||||
"responses": 149,
|
||||
"response_cost": 0.2096464700000001
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "swe",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 10,
|
||||
"responses": 141,
|
||||
"response_cost": 3.702219799999998
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "terminal",
|
||||
"model": "openai/gpt-5.6-luna",
|
||||
"attempts": 5,
|
||||
"responses": 66,
|
||||
"response_cost": 0.10991243999999999,
|
||||
"token_limit_responses": 1,
|
||||
"token_limit_response_cost": 0.009907
|
||||
},
|
||||
{
|
||||
"split": "validation",
|
||||
"benchmark": "terminal",
|
||||
"model": "openai/gpt-5.6-sol",
|
||||
"attempts": 5,
|
||||
"responses": 63,
|
||||
"response_cost": 2.5312083999999997,
|
||||
"token_limit_responses": 6,
|
||||
"token_limit_response_cost": 0.9978314000000001
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
@ -4,7 +4,7 @@
|
|||
"training_tasks": 160,
|
||||
"validation_tasks": 64,
|
||||
"held_out_tasks": 125,
|
||||
"held_out_status": "primary_pair_complete_anthropic_controls_in_progress",
|
||||
"held_out_status": "complete",
|
||||
"current_model_pair": [
|
||||
"openai/gpt-5.6-luna",
|
||||
"openai/gpt-5.6-sol"
|
||||
|
|
@ -13,5 +13,6 @@
|
|||
"solver_effort": "high",
|
||||
"diagnostic_upfront_profiles": 12,
|
||||
"diagnostic_profiles_sha256": "468de3b22e309ca3ef800ef90bd9b03a8aca3eebef32bcc2155377b4cf7e1f88",
|
||||
"primary_results_sha256": "e91ec5578d496ad37c1e8cf5760882042f4b7768bd3c2ff5d65d92f19df7a12c"
|
||||
"primary_results_sha256": "e91ec5578d496ad37c1e8cf5760882042f4b7768bd3c2ff5d65d92f19df7a12c",
|
||||
"final_results_sha256": "b724f5607b532e1d8b354c95129f5f6d95514926ca5000e9a93a9104b13a844d"
|
||||
}
|
||||
|
|
|
|||
BIN
cookbook/auto_router_selective_training/quality_cost.png
Normal file
BIN
cookbook/auto_router_selective_training/quality_cost.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 225 KiB |
4023
cookbook/auto_router_selective_training/quality_cost.svg
Normal file
4023
cookbook/auto_router_selective_training/quality_cost.svg
Normal file
File diff suppressed because it is too large
Load diff
|
After Width: | Height: | Size: 147 KiB |
|
|
@ -0,0 +1,10 @@
|
|||
{
|
||||
"status": "passed",
|
||||
"models": 118,
|
||||
"candidates": 1892,
|
||||
"selected_profiles": 12,
|
||||
"ablations": 40,
|
||||
"exact_coefficient_and_selection_match": true,
|
||||
"data_scope": "Frozen training and validation only; no final tasks loaded",
|
||||
"selection_sha256": "1a7c9241c72c299c8b27fd8171881f94fda069fb7191ac6bc3a2eab9de811853"
|
||||
}
|
||||
Loading…
Add table
Reference in a new issue