litellm/litellm/llms/base_llm
mubashir1osmani 45b6ece18c
fix(vertex/files): stream OpenAI->Vertex batch JSONL uploads (#31036)
* fix(vertex/files): stream OpenAI->Vertex batch JSONL uploads to fix OOM on large files

Large (1GB+) batch JSONL uploads to Vertex AI / GCS caused OOM or killed the worker
because the request body was buffered and multiplied 2-3x in size. The create-file
path is now streaming end-to-end: transform_create_file_request returns a
ResumableChunkedUploadConfig carrying a lazy _OpenAIToVertexBatchUploadStream, and the
HTTP handler opens a GCS resumable session and PUTs the body in bounded 8 MiB chunks
(Content-Range, 308 between chunks) so the transformed payload is never held in full.
The proxy /v1/files endpoint streams from Starlette's spooled upload handle instead of
reading the whole body, and batch rate limiting counts tokens and models in a single
streaming pass.

Only gcs_bucket_name is supported for the GCS target; the legacy bucket_name key is
intentionally not read.

Also removes the unreachable VertexAIFilesHandler create path and everything only it
kept alive (VertexAIJsonlFilesTransformation, _stream_openai_jsonl_to_vertex, the legacy
transform helpers), plus the orphaned batch_utils helpers the streaming rewrite replaced.

* fix(batches): return original JSONL on unparseable row to avoid silent batch truncation

The streaming rewrite of replace_model_in_jsonl accumulated physical lines and
skipped a row on JSONDecodeError to support multi-line objects, but a genuinely
malformed or truncated row never completes: it poisons the buffer, swallows every
following row, and the function still returned the partial rewrite (the rows before
the bad one, already model-rewritten) as if the batch were complete. That turned the
pre-rewrite behavior of returning the original file unchanged (so the provider rejects
the bad batch loudly) into a silent partial submission.

Restore the original-content fallback: when an unparseable remainder is left after the
loop, return the original file_content (rewinding a consumed seekable source) instead of
the truncated output. The multi-line happy path is unchanged.

* test(batches): mock resumable GCS upload in vertex batch prediction test

The vertex batch file-create path now streams to a GCS resumable session via
_aresumable_chunked_upload (httpx send) instead of AsyncHTTPHandler.post, so the
existing test's post mock no longer intercepted the upload and a real request hit
GCS (401). Mock _aresumable_chunked_upload to return the GCS object response; the
resumable protocol itself is covered in test_vertex_ai_files_streaming.py.

* fix(batches): resilient per-row token accounting; no hard-block on count failure

The batch input-file pass iterated a generator whose json.loads raised on a
malformed line; the outer except caught it and stopped the loop, so any body.model
on rows after a bad line was never collected and the model allowlist check ran
against a partial set. It also hard-blocked the batch with a 400 whenever token
counting raised, a backwards-incompatible change from the prior swallow-and-proceed
behavior that breaks legitimate rows the token counter cannot measure (e.g. some
multimodal content).

Iterate the JSONL line-by-line and account each row independently. A malformed line
is skipped (its request cannot run upstream anyway) and a row the counter cannot
measure falls back to a conservative size-based estimate. The loop never aborts, so
the allowlist check always sees every parseable model, and the token total is never
zeroed, so a crafted uncountable row still cannot evade the TPM limit, without
hard-rejecting a legitimate batch.

* perf(vertex/files): unblock async upload; drop empty finalize; widen batch MIME types

Three review follow-ups on the resumable batch upload:
- _aresumable_chunked_upload pulled chunks from a synchronous generator that runs
  the per-row transform inline on the event loop thread, blocking other requests
  between PUTs on large uploads. Each chunk is now produced via asyncio.to_thread.
- _iter_resumable_chunks no longer yields a trailing empty chunk, so an exactly
  chunk-aligned upload finalizes on its last data chunk instead of an extra
  zero-byte PUT; a 0-byte stream still finalizes via the caller's empty request.
- valid_content_type now accepts the MIME types clients label .jsonl batch uploads
  with (text/plain, application/json, ndjson, ...), so such a batch file no longer
  silently bypasses the streaming path into the buffered media upload.

* fix(vertex/files): keep legacy bucket_name as GCS bucket fallback

The rename to gcs_bucket_name dropped the legacy bucket_name key entirely, so an SDK caller passing bucket_name to a Vertex AI file create/retrieve/content call with GCS_BUCKET_NAME unset got ValueError("GCS bucket_name is required") where it previously resolved the bucket. _get_configured_bucket_name now reads gcs_bucket_name, then bucket_name, then the env var, and bucket_name is restored to OPTIONAL_KWARGS_KEYS so it survives get_litellm_params on the retrieve and content paths. gcs_bucket_name keeps precedence when both are present

* style: sort imports in llm_http_handler to satisfy I001 budget

---------

Co-authored-by: Yuneng Jiang <yuneng@berri.ai>
(cherry picked from commit 56825926af)
2026-06-29 17:56:01 -07:00
..
agents Gemini managed agents support (#28270) 2026-05-19 16:02:03 -07:00
anthropic_messages style: black format anthropic_messages transformation.py 2026-04-15 18:18:41 -07:00
audio_transcription fix: fix linting errors 2025-09-12 18:03:50 -07:00
batches style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
bridges Logging: prevent double logging logs when bridge is used (anthropic <-> chat completion OR chat completion <-> responses api) (#11687) 2025-06-12 23:07:36 -07:00
chat fix(vertex): propagate Vertex AI metadata in streaming success callbacks (#29899) 2026-06-08 16:14:30 -07:00
completion Add Google AI Studio /v1/files upload API support (#9645) 2025-04-02 08:56:58 -07:00
containers style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
embedding Add Google AI Studio /v1/files upload API support (#9645) 2025-04-02 08:56:58 -07:00
evals Add eval run endpoints methods 2026-02-17 19:32:15 +05:30
files fix(vertex/files): stream OpenAI->Vertex batch JSONL uploads (#31036) 2026-06-29 17:56:01 -07:00
google_genai Revert "QA: improve gpt-5.4 code/bugs" 2026-03-13 10:15:47 -07:00
guardrail_translation feat(guardrails): optional skip tool message in unified guardrail inputs 2026-05-07 18:39:26 -07:00
image_edit refactor(azure): move image gen JSON helper; rename image edit finalize hook 2026-05-05 08:49:46 +05:30
image_generation style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
image_variations VertexAI non-jsonl file storage support (#9781) 2025-04-09 14:01:48 -07:00
interactions style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
managed_resources feat(pass_through): extend passthrough_managed_object_ids to Azure (#29160) 2026-05-30 16:30:10 -07:00
ocr Litellm oss staging 04 21 2026 2 (#26569) 2026-05-20 21:25:19 -07:00
passthrough style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
realtime feat: litellm oss 110626 (#30202) 2026-06-11 22:30:26 -07:00
rerank feat(vertex_ai): propagate metadata labels to embedding, Imagen, rerank 2026-04-02 09:56:56 +05:30
responses fix(responses): presidio PII masking for Azure WebSocket and streaming (#30003) 2026-06-12 07:27:03 -07:00
sandbox feat(sandbox): e2b code execution primitive (#30898) 2026-06-20 16:30:01 -07:00
search style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
skills style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
text_to_speech style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
vector_store Fix extra body error 2026-04-29 08:34:31 +05:30
vector_store_files style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
videos fix(vertex-ai): use DB credentials in video handlers + implement Veo video edit (#29098) 2026-05-28 11:45:41 -07:00
__init__.py [Feat] Add Initial support for Bedrock Batches API (#14190) 2025-09-03 17:19:58 -07:00
base_model_iterator.py chore(oss): litellm oss staging 150626 (#30463) 2026-06-16 12:06:41 -07:00
base_utils.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00