mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-05 02:41:56 +00:00
A client behind an auto-router sends one max_tokens for every tier, so a value sized for the smallest tier starves a bigger tier's thinking budget and a value sized for the biggest is rejected by the smallest. After the complexity router picks a tier, its per-tier litellm_params now carry max_tokens set to the smallest max_output_tokens across that tier model's deployments (model_info, then the cost map), applied the same way a per-tier reasoning_effort already is, on every routing exit including plan mode, the empty-ask default and the classifier fallback. The router seam collapses whichever ceiling alias a tier carries onto the surface's own name, so one tier max_tokens reaches chat, /v1/messages and /v1/responses alike, drops the caller's other carriers of the same setting before the merge, and stamps the caller's original once so a fallback into a group no tier owns gets it back instead of a ceiling sized for the tier that failed. Proxy-level reservations were sized from the caller's cap before routing, so a raised cap left them short. Both owners now re-validate at the deployment hook: the v3 limiter tops up its combined-TPM and project-OTPM reservations to the final cap or writes the admitted cap back, and the budget limiter re-estimates on the chosen deployment and grows the reservation or writes the admitted cap back. An auto-router alias also reserves budget at its priciest tier model now instead of pricing to zero. An explicit per-tier max_tokens, max_completion_tokens or max_output_tokens still wins, and max_tokens_from_tier_model: false forwards the caller's value unchanged. |
||
|---|---|---|
| .. | ||
| public | ||
| scripts | ||
| src | ||
| tests | ||
| .env.development | ||
| .env.production | ||
| .npmrc | ||
| .nvmrc | ||
| .prettierignore | ||
| .prettierrc | ||
| build_release_ui.sh | ||
| build_ui.sh | ||
| build_ui_custom_path.sh | ||
| CLAUDE.md | ||
| components.json | ||
| eslint-budgets.json | ||
| eslint-suppressions.json | ||
| eslint.config.mjs | ||
| knip.json | ||
| next.config.mjs | ||
| package-lock.json | ||
| package.json | ||
| postcss.config.js | ||
| README.md | ||
| tsconfig.json | ||
| tsconfig.tsbuildinfo | ||
| vitest.config.ts | ||
This is a Next.js project bootstrapped with create-next-app.
Getting Started
First, run the development server:
npm run dev
# or
yarn dev
# or
pnpm dev
# or
bun dev
Open http://localhost:3000 with your browser to see the result.
You can start editing the page by modifying app/page.tsx. The page auto-updates as you edit the file.
This project uses next/font to automatically optimize and load Inter, a custom Google Font.
Learn More
To learn more about Next.js, take a look at the following resources:
- Next.js Documentation - learn about Next.js features and API.
- Learn Next.js - an interactive Next.js tutorial.
You can check out the Next.js GitHub repository - your feedback and contributions are welcome!
Deploy on Vercel
The easiest way to deploy your Next.js app is to use the Vercel Platform from the creators of Next.js.
Check out our Next.js deployment documentation for more details.