mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-05 02:41:56 +00:00
docs: invalidate contaminated router benchmark results
This commit is contained in:
parent
31afc3d5e0
commit
04c4e6284b
7 changed files with 25 additions and 2 deletions
|
|
@ -1,3 +1,5 @@
|
|||
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
|
||||
|
||||
# Card and calibration ablations
|
||||
|
||||
These comparisons hold the boundary at the original default: capability base 0.5 with step 0.1, or V2 quality gap 0.05. All coefficients were fitted on the training split. This table reports every raw card and the fixed middle regularization strength of 10 for calibrated variants; the JSON contains all strengths. No variant is chosen using these held-out results
|
||||
|
|
|
|||
|
|
@ -1,3 +1,5 @@
|
|||
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
|
||||
|
||||
# Auto Router training experiment
|
||||
|
||||
Profiles were fitted on 83 DeepSWE tasks and selected on 30 repository-disjoint validation tasks before inspecting live grades. The live comparison uses 25 native-image-eligible SWE-bench Verified tasks and four current solver models at high effort. Each solver runs once per task; frozen task-pinned policies reuse those attempts and add their judge cost
|
||||
|
|
|
|||
|
|
@ -1,3 +1,5 @@
|
|||
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
|
||||
|
||||
# What the training experiment showed
|
||||
|
||||
Both implementations now support runnable, opt-in trained profiles. Training covered three cards, raw probabilities, per-model calibration, task-dependent calibration, and pair-specific boundaries, including combinations. All fitting and policy selection used the DeepSWE training and validation splits. The 25 fresh SWE-bench tasks were used only for evaluation
|
||||
|
|
|
|||
|
|
@ -1,3 +1,5 @@
|
|||
> **Historical benchmark invalidated:** the 25-task live run exposed upstream fixes through Git objects outside the task checkout. Its quality and savings figures are withdrawn. These files retain the audit record, not evidence of routing improvements. Fresh isolated experiments are in progress
|
||||
|
||||
# Experimental Auto Router training snapshots
|
||||
|
||||
These opt-in profiles compare the original cards, a research-informed card rewrite, and cards with training-derived capability priors. They also compare raw probabilities, per-model logit calibration, and regularized task-dependent logit calibration. The router makes one judge call; all learned probability adjustments and threshold comparisons run locally
|
||||
|
|
|
|||
|
|
@ -2395,5 +2395,10 @@
|
|||
"lost_capable_successes": 0,
|
||||
"gained_efficient_successes": 0
|
||||
}
|
||||
]
|
||||
],
|
||||
"validation_status": {
|
||||
"live_benchmark_valid": false,
|
||||
"reason": "Solver-visible Git history exposed upstream fixes; live quality and savings claims withdrawn",
|
||||
"replacement": "Fresh isolated training and evaluation in progress"
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -4779,5 +4779,10 @@
|
|||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
],
|
||||
"validation_status": {
|
||||
"live_benchmark_valid": false,
|
||||
"reason": "Solver-visible Git history exposed upstream fixes; live quality and savings claims withdrawn",
|
||||
"replacement": "Fresh isolated training and evaluation in progress"
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -51,5 +51,10 @@
|
|||
"ABLATIONS.md",
|
||||
"ablation_diagnostics.json"
|
||||
]
|
||||
},
|
||||
"validation_status": {
|
||||
"live_benchmark_valid": false,
|
||||
"reason": "Solver-visible Git history exposed upstream fixes; live quality and savings claims withdrawn",
|
||||
"replacement": "Fresh isolated training and evaluation in progress"
|
||||
}
|
||||
}
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue