mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-04 02:31:27 +00:00
Azure PTU deployments carry zeroed per-token pricing because the reservation is billed flat by the hour. When Azure spills a request onto pay-as-you-go capacity it returns x-ms-is-spilled-over: true, and that traffic was still priced at zero. The response cost calculator now detects the spillover header on the result's hidden params or the logged provider response headers and skips the zeroed custom pricing only for genuine PTU deployments while the feature flag is on. Azure sync streaming now also records response headers on the logging object, matching the async paths. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| batches | ||
| chat | ||
| image_edit | ||
| image_generation | ||
| passthrough | ||
| realtime | ||
| response | ||
| search | ||
| text_to_speech | ||
| videos | ||
| test_audio_transcriptions.py | ||
| test_azure.py | ||
| test_azure_common_utils.py | ||
| test_azure_cost_calculation.py | ||
| test_azure_embedding.py | ||
| test_azure_exception_mapping.py | ||
| test_azure_fine_tuning_api.py | ||
| test_azure_speech_audio_transcription.py | ||