Sameer Kankute
bbd8ca3b3d
feat(prometheus): add metrics for managed batch lifecycle
...
- Add Prometheus metrics for managed batch and file operations
- Track batch creation, file size, duration, and deletion events
- Add CheckBatchCost polling metrics (jobs polled/processed, errors)
- Record metrics in managed_files hook and check_batch_cost utility
- Metrics include labels for model, provider, user, and status
Made-with: Cursor
2026-03-27 20:30:09 +05:30
Sameer Kankute
32ded9b2f8
fix double-billing issue
2026-03-18 12:47:42 +05:30
Sameer Kankute
7660f39fdb
fix(file_search): promote DB helper, suppress sub-call billing, add queries-plural test
...
- Promote _fetch_managed_vector_stores_by_uuids from @staticmethod to a module-level
async helper get_managed_vector_store_rows_by_uuids, following the same standalone
helper pattern as get_team_object / get_key_object so the hot-path DB read is a
named importable function rather than an inline prisma_client.db.* call
- Pass no-log=True to both inner _call_aresponses sub-calls so they do not fire
independent billing/monitoring callbacks; cost is accumulated in the synthesized
response's _hidden_params for the outer responses() call
- Add test_H11b covering the primary queries (plural array) function-tool schema,
complementing H11 which exercises only the backward-compat singular query path
Made-with: Cursor
2026-03-18 11:38:49 +05:30
Sameer Kankute
76176f2a64
fix(file_search): restore should_use_emulated helper, fix dedup, extract DB helper, clean docstring
...
- Re-add should_use_emulated_file_search() to emulated_handler.py so H5/H6/H7/H13 tests don't fail with ImportError
- Remove per-file-id deduplication from _build_search_results_for_include so all chunks are returned (matching OpenAI native file_search behaviour); update test_H14 to assert 2 results
- Extract raw prisma DB query in check_vector_store_ids_access into a static _fetch_managed_vector_stores_by_uuids helper so the hot request path uses a named, testable function instead of an inline prisma_client.db.* call
- Remove developer-local path from test module docstring
Made-with: Cursor
2026-03-18 11:26:27 +05:30
Sameer Kankute
c735251570
feat(responses): file_search support — Phase 1 native passthrough + Phase 2 emulated fallback
...
Phase 1 (native passthrough):
- _decode_vector_store_ids_in_tools(): decode LiteLLM-managed unified
vector_store_ids to provider-native IDs in file_search tools
- Split update_responses_tools_with_model_file_ids() into decode pass
(always runs) + code_interpreter mapping pass (guarded)
- BaseResponsesAPIConfig.supports_native_file_search() → False by default;
OpenAIResponsesAPIConfig overrides to True
- ManagedFiles.async_pre_call_hook(): batch team-level access check for
unified vector_store_ids in file_search tools (no N+1)
- Docs: file_search section in response_api.md
Phase 2 (emulated fallback for non-native providers):
- litellm/responses/file_search/emulated_handler.py: converts file_search
tool → function tool, intercepts tool call, runs asearch(), makes
follow-up call, synthesizes OpenAI-format output (file_search_call +
message + file_citation annotations)
- responses/main.py: routes to emulated handler when provider doesn't
support file_search natively
Tests: 41 unit tests across 8 families (A-H) in test_file_search_responses.py
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 11:41:44 +05:30
Sameer Kankute
22b333cae6
Fix downloading vertex ai files
2026-03-16 12:08:06 +05:30
Ephrim Stanley
35f6fd4223
Managed batches fixes for Gemini/Vertex
2026-02-28 20:21:04 -05:00
Ephrim Stanley
27a33565e7
State management fixes for CheckBatchCost - Address greptile comments
2026-02-23 07:52:58 -05:00
Ephrim Stanley
7b5dc3fb9c
State management fixes for CheckBatchCost
2026-02-23 07:16:25 -05:00
Sameer Kankute
03f5717456
Fixes based on greptile reviews
2026-02-18 12:19:11 +05:30
Sameer Kankute
9f5580fddd
Fixes based on greptile reviews
2026-02-18 11:55:06 +05:30
Sameer Kankute
8f80b1085e
Add File deletion criteria with batch references
2026-02-18 11:39:32 +05:30
Sameer Kankute
791cef6d99
fix test_chat_completion
2026-02-17 20:26:28 +05:30
Ephrim Stanley
a3762e7d49
Addressed greptile comments to extract common helpers and return 404
2026-02-16 07:58:04 -05:00
Ephrim Stanley
a5626768a3
Add comments
2026-02-14 19:46:56 -05:00
Ephrim Stanley
4d87cb8fe3
Fix deleted managed files returning 403 instead of 404
2026-02-14 19:12:03 -05:00
Ephrim Stanley
e59c8d22af
fix: afile_retrieve returns unified ID for batch output files
2026-02-14 10:07:12 -05:00
Ephrim Stanley
358180eb2d
Fix: pass deployment credentials to afile_retrieve in managed_files post-call hook
2026-02-14 00:16:34 -05:00
Sameer Kankute
4ad0ecd9eb
Fix mypy issues
2026-02-13 22:02:27 +05:30
Sameer Kankute
fa48166b10
Add _PROXY_LiteLLMManagedVectorStores class
2026-02-13 09:57:17 +05:30
Sameer Kankute
ce4bebbedf
Changed asyncio.create_task() to await for storing batch objects
2026-02-10 12:42:39 +05:30
Sameer Kankute
9bdb163269
Add error file ids as managed files
2026-02-10 12:06:40 +05:30
Sameer Kankute
8b3213ce5c
Add mapping for responses tools in file ids
2026-02-04 13:12:45 +05:30
Sameer Kankute
410e54648c
Fix: Managed Batches: Inconsistent State Management for list and cancel batches
2026-02-03 14:47:28 +05:30
Sameer Kankute
70684ca86f
Fix File access permissions for .retreive and .delete
2026-01-29 11:19:24 +05:30
Sameer Kankute
833cf6a2cf
Fix: Batch cancellation ownership bug
2026-01-29 10:54:42 +05:30
Ephrim Stanley
88280d9cca
Fix batch creation to return the input file's expires_at attribute
2026-01-26 12:02:35 -05:00
Ephrim Stanley
caa2c57619
Fix /batches to return encoded ids (from managed objects table)
2026-01-26 10:36:43 -05:00
Ephrim Stanley
479e40672e
Add end to end integration tests for batches
2025-12-23 12:55:09 -05:00
Ephrim Stanley
73083a1f5b
Add end to end integration tests for batches
2025-12-23 11:54:46 -05:00
Ephrim Stanley
b84aafbab6
Add end to end integration tests for batches
2025-12-23 10:24:32 -05:00
Sameer Kankute
7d0f41f437
Add cost tracking for responses api in background mode
2025-12-19 13:35:48 +05:30
Sameer Kankute
755f024087
remove print statments
2025-12-16 21:37:09 +05:30
Sameer Kankute
3222b4e3a8
Add batch output file in managed table
2025-12-16 15:44:30 +05:30
Sameer Kankute
7b1cef86a7
Add support for target_storage param
2025-12-11 15:08:17 +05:30
Krish Dholakia
70e1e83102
feat(managed_files.py): support /delete for files + feat(managed_batches): support /cancel for batches ( #16387 )
...
* feat(managed_files.py): initial commit fixing managed file delete on litellm
* fix(managed_files.py): fix file delete
* feat(batches_endpoints/endpoints.py): fix cancelling a batch
ensures managed batches works
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
2025-11-18 17:36:26 -08:00
Sameer Kankute
7cebc151b0
Add managed files support for responses API ( #16733 )
...
* Fix responses api with managed files
* fix litellm/llms/vertex_ai/gemini/vertex_and_google_ai_studio_gemini.py mypy
* fix litellm/llms/vertex_ai/gemini/vertex_and_google_ai_studio_gemini.py mypy
* fix mypy errors
2025-11-17 18:41:26 -08:00
Krish Dholakia
06906534b3
feat(audio_transcriptions/): calculate duration of audio file for cost calculation + feat (image_generations): cost tracking accuracy improved with output_format, quality, size values fixed per openai model
...
* feat(audio_transcriptions/): calculate duration of audio file for cost calculation
Fixes https://github.com/BerriAI/litellm/issues/11846
Closes https://github.com/BerriAI/litellm/issues/14605
* fix(cost_calculator.py): correctly use base model, when set
Fixes issue where azure base model was being ignored
* feat(cost_calculator.py): fix default cost tracking quality param for image generation
* feat(image_generations/): return output_format, quality, size
aligns response to openai spec and improves cost tracking accuracy
* fix(cost_calculator.py): refactor cost calculation for image generation to use image response instead of hidden params
* build: update build
* fix: fix cost calculation
* build: update poetry lock
* fix: fix ruff checks
* fix: fix aembedding
* fix: fix ruff errors
* fix: modify to catch errors
* fix: test
* fix: loosen test to handle openai lib out of sync
* fix: fix base models
* fix: fix usage object
2025-11-08 16:24:31 -08:00
Krish Dholakia
202eaeb1a2
Revert "(feat) Audio transcription - cost tracking + (feat) image generation …" ( #16409 )
...
This reverts commit c96da44265 .
2025-11-08 15:38:16 -08:00
Krish Dholakia
c96da44265
(feat) Audio transcription - cost tracking + (feat) image generation - accurate cost tracking based on output_format/quality/size
...
* feat(audio_transcriptions/): calculate duration of audio file for cost calculation
Fixes https://github.com/BerriAI/litellm/issues/11846
Closes https://github.com/BerriAI/litellm/issues/14605
* fix(cost_calculator.py): correctly use base model, when set
Fixes issue where azure base model was being ignored
* feat(cost_calculator.py): fix default cost tracking quality param for image generation
* feat(image_generations/): return output_format, quality, size
aligns response to openai spec and improves cost tracking accuracy
* fix(cost_calculator.py): refactor cost calculation for image generation to use image response instead of hidden params
* build: update build
* fix: fix cost calculation
* build: update poetry lock
* fix: fix ruff checks
* fix: fix aembedding
* fix: fix ruff errors
* fix: modify to catch errors
* fix: test
* fix: loosen test to handle openai lib out of sync
2025-11-08 15:30:46 -08:00
Ishaan Jaffer
0a1fc0eeb2
Revert "fix: fix ruff errors"
...
This reverts commit eef864360e .
2025-11-08 14:33:51 -08:00
Krrish Dholakia
eef864360e
fix: fix ruff errors
2025-11-08 14:15:12 -08:00
Krish Dholakia
f8d6a6edb9
fix(managed_files.py): don't raise error if managed object is not found + (Feat) Azure AI - Search Vector Stores + (Fix) Batches - “User default_user_id does not have access to the object” when object not in db + (fix) Vector Stores - show config.yaml vector stores on UI ( #15873 )
...
* fix(managed_files.py): don't raise error if managed object is not found
* feat(vector_stores): add azure ai search vector store support
Enables direct querying a vector store on azure
* fix(azure/vector_stores): working azure ai search api vector stores
allows azure direct querying on vector stores
* test: update env vars
* docs(docs/): document new azure ai vector store search
* docs(azure_ai_vector_stores.md): add table
* docs: clarify support for 'create' vector stores
* fix(vector_stores/endpoints.py): Fixes https://github.com/BerriAI/litellm/issues/14606
* fix: fix linting errors
2025-10-25 12:06:24 -07:00
Ishaan Jaff
cea318330e
[Feat] Add Guardrails for /v1/messages and /v1/responses API ( #15686 )
...
* add get_guardrails_messages_for_call_type
* fix call type for /messages
* add anthropic endpoints
* fix bedrock guardrails
* fix config.yaml
* fix types
* fix async_pre_call_hook
* ruff fix
* fix guard
* fix test bedrock guardrail
* fix linting
* fix linting
* docs guardrails
* fix mypy linting
2025-10-17 18:09:00 -07:00
Alexsander Hamir
eaa04cd8ce
fix: use fastuuid helper ( #14903 )
...
* fix: use fastuuid helper across the codebase
First batch of changes, simple drop in replacement.
* second batch of changes
* fixed: script mistake on helper file
2025-09-25 15:47:01 -07:00
Jugal D. Bhatt
900c7f45c0
[MCP Gateway] Litellm mcp pre and during guardrails ( #13188 )
...
* add guardrail support
* add guardrail support
* guardrails for MCP
* added changes
* add mcp guardrails
* added test
* add ui
* fix guardrail form
* working with cursor
* remvoe print
* fix mcp servertests
* fix mypy and remove console logs
* fix mypy and remove console logs
* fix mypy tests
2025-08-01 20:02:25 -07:00
yeahyung
782969bb05
( #11794 ) use upsert for managed object table rather than create to avoid UniqueViolationError ( #11795 )
...
* (#11794 ) use upsert for managed object table rather than create to avoid UniqueViolationError
* (#11794 ) use upsert for managed object table rather than create to avoid UniqueViolationError
2025-07-15 20:20:01 -07:00
Krrish Dholakia
c3857e60f2
Store batch output file id in DB + Store batch file status in DB + (experimental) BATCH API COST TRACKING
2025-06-25 22:41:22 -07:00