Commit graph

49 commits

Author SHA1 Message Date
devin-ai-integration[bot]
0e26edfdb9
test: move 81 legacy live tests in pass-through, spend, batches, audio, search, guardrails, image and ocr dirs offline (#45288)
* test: make legacy live tests in spend, batches, openai endpoints and audio dirs offline (partial)

* test: migrate wave-1b live tests offline (guardrails, images, ocr, search, openai endpoints)

* test: fix wave-1b review items, add responses/ocr integration tests and firecrawl unit test

* test: anthropic messages router/bedrock/openai-bridge unit tests for wave-1b nodes

* test: anthropic messages logging, prompt-caching and tool-search unit tests; drop migrated base nodes

* test: finish pass_through_unit_tests nodes, logging drain fix and mutations

* test: migrate anthropic passthrough tests to integration wire tests

* test: fix passthrough migration wire spend row lookup and wildcard config

* test: migrate hosted vllm and openai file passthrough tests offline

* test: move assemblyai and vertex passthrough nodes to in-process unit tests

* test: restore unlisted router node and fix logging worker drain in passthrough unit tests

* test: drop spend-row BUG skip and sharpen non-streaming skip reason for anthropic messages

* test: use public presidio alias after merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: restore the batch and file unit tests the migration rewrote away

* test: restore full legacy intent in anthropic messages router unit tests

Drop the false BUG skip on non-streaming aanthropic_messages logging (success
callbacks do fire; the skipped body filtered on the wrong model), assert the
logged model_group, messages, cost and usage, cover streaming logging for both
Anthropic and Bedrock invoke, assert dict content blocks for Anthropic, Bedrock
invoke and the OpenAI bridge, add the Bedrock invoke leg of the router test,
fall back from a real 401, test system-prompt caching and streaming
message_start cache fields on converse and invoke, send the legacy tool-search
tools and beta header, remove the type: ignore and bare dict helper, and add
in-process native /anthropic passthrough spend logging tests

* test: fix passthrough migration wire tests and drop false native spend BUG skip

The native spend row test read rows by the shared master-key digest, so it
matched other tests' rows; read each row by its own message id instead and
assert tokens, total, spend, tags, provider, api_base and end_user on both the
non-streaming and streaming native routes. Merge the streaming test that never
checked spend, assert exact tags from litellm_metadata, stop rebinding a Final
in a loop, require cost > 0, and add the chat-completions bridge cost case the
legacy test covered

* test: harden the spend, OCR, image, search and guardrail migration replacements

The OCR spend tests now build fresh kwargs per case instead of mutating a
shared fixture, and every payload case asserts the exact logged spend. The
OCR wire test reads its spend row by request id and checks exact page
pricing. Image edit, Nova Canvas, DuckDuckGo, Firecrawl, Bedrock guardrail
and Presidio replacements now fake only the provider HTTP boundary (respx or
an in-process aiohttp server) and assert the outbound request, so the
DuckDuckGo limit, Azure base_model pricing and guardrail masking are proven
rather than assumed

* test: cover Exa and Perplexity search structure and max_results offline and retire the two base search methods

* test: drive batch and file replacements through the provider HTTP boundary and real logging callback

Replace monkeypatched litellm.afile_content and AsyncHTTPHandler doubles with respx routes,
read batch logging metadata from a registered success callback instead of get_logging_payload,
use the real managed-files hook for the GEN-2166 regression, assert outbound request bodies,
pin poller ownership explicitly in the migrated DB-sync tests, and require the scripted
upstream to be hit in the responses error-status wire tests

* test: assert file content download headers pass through the proxy

* test: point spend coverage references at the tests that replaced the retired spend job

* test: assert passthrough identity, spend and route dispatch from the code under test

The AssemblyAI non-admin test asserted metadata it wrote itself and leaked a
background poll to the real AssemblyAI host. It now drives assemblyai_proxy_route
with a real Request and waits for the success callback for its own transcript id.
The Vertex spend test matches its log by call id instead of taking the first event.
The OpenAI files wire test hit the native /{provider}/v1/files route; it now calls
/openai/files so the passthrough is what forwards the upload and delete.

* test: mock only the HTTP boundary in the migrated audio tests

Vertex TTS tests no longer replace _ensure_access_token or AsyncHTTPHandler.post;
the token comes from a mocked Google OAuth endpoint and the synthesize call from
respx. Speech tests assert the outbound body, the transcription cache test polls
for the cache write instead of relying on test ordering, and the model pass-through
test checks the multipart model field per model.

* test: wait on a logger event instead of polling the clock in anthropic messages unit tests

Recorders keep payloads in a rebound tuple and set an asyncio.Event; tests
await it with asyncio.wait_for instead of a sleep-and-deadline poll loop

* test: freeze module-level batch and file response fixtures as Final MappingProxyType

* test: fake the presidio analyzer with an in-memory aiohttp connector

The blocked-entity tests started an aiohttp TestServer, which binds a local
socket. They now hand the guardrail a ClientSession whose connector answers
/analyze and /anonymize in process, so no socket is opened and the outbound
analyze text and entities are still asserted

* test: type anthropic messages router test helpers with LiteLLM's Anthropic TypedDicts

Messages, cached system blocks and tool-search tools now use
AnthropicMessagesUserMessageParam, AnthropicMessagesTextParam,
AnthropicToolSearchToolRegex and AnthropicMessagesTool instead of bare
dict shapes; tools are converted to plain dicts only at the acreate call,
whose tools parameter is list[dict]

* test: type batch limiter helpers with TypedDicts and wait on the logging callback event instead of polling

* test: give the migrated OCR, image and presidio helpers precise types

OCR spend helpers take ReadOnly TypedDicts for kwargs and responses and use
LiteLLM's OCRResponse/OCRUsageInfo instead of local pydantic stand-ins; spend
metadata is validated with a TypeAdapter. The presidio fake uses LiteLLM's
PresidioAnalyzeRequest/ResponseItem types, and the image-edit logger validates
the logged payload instead of storing an untyped dict

* test: signal callback and cache events instead of polling

Recorders keep tuples and set an asyncio.Event, thread-safely, when the payload
for this test's transcript id or upstream URL arrives. The transcription cache
test waits on a Cache subclass that signals after async_add_cache. No clock
polling or sleeps remain in these tests.

* test: assert the batch limiter hook updates the caller's request in place

* test: tolerate model-list probes and read native passthrough rows by owned key

The router's OpenAI-compatible model-info refresh (litellm/router.py:10710)
sends GET /v1/models to configured openai api_bases, so the wire answers it
with an empty list and excludes it from the provider-call assertions. Native
/anthropic spend rows are now read by a per-request virtual key digest and
call_type, then the row's request_id is checked against the message id

* test: expect the OCR alias in the proxy response model

The proxy restamps every OpenAI-compatible response model to the name the
client requested (_override_openai_response_model), so /v1/ocr returns the
scenario alias. The upstream model is now checked on the drained request body
instead of inside the peer, where a failed assert never reached the test

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 02:38:12 -07:00
devin-ai-integration[bot]
9e871f2736
test: remove dead imports and helpers left behind by legacy test deletion (#45208)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 18:04:17 -07:00
devin-ai-integration[bot]
0734e35024
test: delete 73 legacy tests covered by e2e, unable to fail, or dead in CI (#45195)
* test: delete 77 legacy tests covered by e2e, unable to fail, or dead in CI

* test: keep router helper tests and coverage ignore list unchanged

* test: keep cohere error handling tests

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 17:11:36 -07:00
devin-ai-integration[bot]
837c6a7481
fix(security): remove the publicly known master key from the repo (#44718)
* fix(security): hash the publicly known master key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: replace weak master key examples and regenerate artifacts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: replace weak key fixtures with generated test keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: generate master keys for proxy startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve lens dev key entropy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore proxy key compatibility in scrub examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: scrub merged SSO fixture and refresh dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ci): stabilize test keys and metadata collection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: drop the rebuilt dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:55:24 -07:00
devin-ai-integration[bot]
6c2ede00ac
test: remove 130 legacy tests owned by stronger unit proofs (#44157)
* test: remove 130 legacy tests owned by stronger unit proofs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: list router _embedding and _aembedding as covered via public embedding calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 10:18:50 -07:00
yuneng-jiang
7d50a31eb5
test(e2e): move live-provider legacy tests into tests/e2e (#44120)
* test(e2e): move live-provider legacy tests into tests/e2e

Port legacy tests that exercise real providers into the tests/e2e suites that own them, using the harness (/model/new plus deferred cleanup) and asserting on what the caller receives. Delete legacy tests already covered at equal or stronger strength by e2e, integration or unit tests, and drop the now empty ocr_testing CircleCI job

* test(e2e): address review on the live-provider test move

Assert the SSE error frame a client actually receives when a post_call guardrail blocks a stream, and require a tool call for every requested city before checking the answer. Restore the OCR matrix and its CircleCI job, the Claude Agent SDK streaming test, and test_async_create_batch, since their SDK-level and callback assertions have no equivalent in tests/e2e

* test(e2e): accept both guardrail block shapes on a blocked stream

A post_call block before the first chunk reaches the client as HTTP 400 with either a JSON error body or a single SSE error frame, depending on whether the block surfaced as an exception or an error chunk. Assert the policy message is present and the blocked output is absent in both

* test(realtime): restore direct SDK realtime tests against OpenAI

The e2e realtime tests go through the proxy and the remaining SDK tests either mock the upstream or assert less, so keep the direct litellm._arealtime tests with and without intent, and TestOpenAIRealtime::test_realtime_connection, in place

* test: make realtime and Nova stream checks deterministic

The direct SDK realtime tests now fail on a refused connection instead of skipping. The with-intent test asserts OpenAI rejects the exact intent value sent, which only happens when the intent is forwarded. The Nova /v1/messages stream test asserts stream structure, stop reason and usage instead of model wording

* test(realtime): own intent forwarding with a unit test instead of a live rejection

Assert litellm._arealtime passes the intent query param into the OpenAI realtime websocket URL, which is the behavior LiteLLM owns, and drop the live test that depended on OpenAI's rejection wording
2026-10-02 00:02:18 -07:00
yuneng-jiang
a93c396a5c
test(integration): move legacy proxy, router and Redis tests into tests/integration (#44128)
* test(integration): move legacy proxy, router and Redis tests into tests/integration

Port 39 legacy tests to the integration tier that owns them, running against
the scripted upstream, local Postgres and Redis, test-owned wire peers and
owned proxies. Delete 5 legacy tests whose contract is already owned by an
existing integration test, and remove the legacy functions, files and helpers
left unused.

* test(integration): cover recovery of a spent key after its budget is raised
2026-10-01 23:00:54 -07:00
Yuneng Jiang
4423876857
test(responses): fix stale Anthropic smoke request 2026-09-11 13:56:27 -07:00
mateo-berri
0904a9223b test(responses): collect the admitted stream events without local mutation 2026-09-03 12:56:56 -07:00
mateo-berri
b38516da88 test(responses): make the background stream cancel deterministic
A five-token response can complete before the cancel lands, which put the
test back on the "Cannot cancel a completed response" path it used to
swallow. Ask for a long generation so the cancel always beats completion,
and assert the cancelled status unconditionally
2026-09-03 12:37:31 -07:00
mateo-berri
36143b53f9 test(responses): bound the background stream cancel e2e so an upstream stall skips fast
test_cancel_streaming_response drained the whole background stream before
cancelling, so an OpenAI keepalive stall held the e2e_openai_endpoints job for
301s and failed it on a generic APIError, and on a healthy day it cancelled an
already completed response and swallowed the 400 without verifying a cancel.
Cancel at the first event carrying a response id, bound admission to 90s, skip
naming the stall when only keepalives arrived, and assert status == cancelled
2026-09-03 12:34:27 -07:00
mateo-berri
d33fe95d19 test(responses): expect the 404 OpenAI now returns for an unknown model 2026-09-02 18:08:45 -07:00
mateo-berri
66947fcf7c test(responses): expect the bad-temperature 400 on a non-reasoning model 2026-08-29 01:58:15 -07:00
ryan-crabbe-berri
e9d40a8f73 test: enforce F811 so a duplicate definition cannot silently replace the first
A name bound twice keeps only the second binding. In `tests/` that is nearly
always a repeated import, harmless but misleading, and the same rule is what
catches the cases that are not harmless: a local that shadows an import the
module still calls, and a second `def test_x` that quietly replaces the first.

311 of the 344 sites were repeated imports and came out with ruff's own fix.
The remaining 33 needed a decision. Four modules imported a name they never
used because a local definition below already shadowed it. Two comprehensions
bound `call` over `unittest.mock.call`, which those modules import and use.
One test rebound the two module handles its nested reload closure had captured.
One class attribute shadowed an unused `status` import.

The load-test fixtures move to a conftest, which is how pytest is meant to share
them, so the test module no longer imports three fixture names it never calls.
The nine `prisma_client` parameters keep a narrow `noqa`: pytest resolves that
fixture by name before the body runs, so the parameter never shadows anything.
2026-08-21 12:06:19 -07:00
ryan-crabbe-berri
680bcfd8aa
test(lint): ban blind pytest.raises(Exception) with ruff B017 (#37731)
* test(lint): ban blind pytest.raises(Exception) with ruff B017

A bare pytest.raises(Exception) accepts whatever the body throws. The TypeError
a refactor introduces satisfies it exactly as well as the rejection the test was
written for, so the crash reads as a pass and the test never goes red.

All 111 existing sites are narrowed here. A runtime probe recorded the concrete
exception each one actually catches, and each site now names that type. Where
the code under test genuinely raises a bare Exception, the site pins a stable
slice of the message with match= instead.

Two sites tell on themselves. The shared responses-API cancel test raises
"custom_llm_provider is required but passed as None" rather than talking to a
provider at all, because cancel_responses takes a provider, not a model. And
test_bedrock_guardrails_with_streaming was the only test in its file still
passing without AWS credentials, because the NoCredentialsError boto3 raised
long before the guardrail ran satisfied the blind raises.

* fix(test): widen the openai batch-dispatch assertion to OpenAIError

The narrowed NotFoundError only holds where OPENAI_API_KEY is set. Without one
the SDK raises OpenAIError while building the client, long before any 404, so CI
went red. OpenAIError covers both and still rejects a TypeError from a refactor.
2026-08-20 18:09:42 -07:00
mateo-berri
eef908d4ad fix(batches): register managed output files on batch cancel
update_batch_in_database now fetches the batch row by unified_object_id
when the caller omits db_batch_object, so the cancel endpoint attributes
newly registered output and error files to the batch owner and returns
unified managed ids instead of raw provider ids. Idempotent cancels that
do not change the stored status also skip the redundant DB write now.

Repair two pre-existing mock tests in test_openai_batches_endpoint.py
that asserted values inside lazy percent-format log strings, and give
the cancel test's prisma mock an awaitable find_first.
2026-08-05 18:28:52 -07:00
Mateo Wang
2c733c00f5
chore(ci): modernize model references in tests and configs (#27856)
* test: modernize models used in CircleCI e2e test suites

Replaces obsolete models (gpt-4o, gpt-4o-mini, gpt-3.5-turbo,
claude-3-5-sonnet-20240620, claude-sonnet-4-20250514) with current
equivalents across the e2e_openai_endpoints and
proxy_e2e_anthropic_messages_tests CircleCI jobs.

- gpt-4o -> gpt-5.5 (responses API e2e tests)
- gpt-4o-mini -> gpt-5-mini (websocket responses, oai_misc_config)
- gpt-4o-mini-2024-07-18 -> gpt-4.1-mini-2025-04-14 (fine-tuning,
  still actively fine-tunable)
- gpt-4 / gpt-3.5-turbo target_model_names example -> gpt-5.5 /
  gpt-5-mini
- bedrock claude-3-5-sonnet-20240620 batch entry -> haiku-4-5-20251001
  (also aligning oai_misc_config model_name with what
  test_bedrock_batches_api.py actually requests)
- bedrock claude-sonnet-4-20250514 (deprecated, retires 2026-06-15)
  -> claude-sonnet-4-5-20250929

* test: point bedrock-claude-sonnet-4 alias at Sonnet 4.6, not 4.5

Greptile/Cursor flagged that after the previous commit, the
bedrock-claude-sonnet-4 alias collided with bedrock-claude-sonnet-4.5
(both pointed to claude-sonnet-4-5-20250929). Rename to
bedrock-claude-sonnet-4.6 and point it at the Sonnet 4.6 Bedrock ID
(us.anthropic.claude-sonnet-4-6, already in the litellm model
registry) so the alias name matches the underlying model version.

* test: modernize models across remaining CI-mounted configs & tests

Expands the modernization sweep to all CircleCI-mounted proxy configs
and to test directories where the model literal is a fixture/route key
(not the test's subject).

Config changes:
- proxy_server_config.yaml: bump gpt-3.5-turbo / gpt-3.5-turbo-1106 /
  gpt-4o / gemini-1.5-flash / dall-e-3 underlying models; rename
  gpt-3.5-turbo-end-user-test alias to gpt-5-mini-end-user-test; bump
  text-embedding-ada-002 underlying to text-embedding-3-small. User-
  facing aliases (gpt-3.5-turbo, gpt-4, text-embedding-ada-002, etc.)
  preserved for backward compatibility with tests.
- simple_config.yaml, otel_test_config.yaml, spend_tracking_config.yaml:
  bump gpt-3.5-turbo underlying to gpt-5-mini.
- pass_through_config.yaml: claude-3-5-sonnet / claude-3-7-sonnet /
  claude-3-haiku entries replaced with claude-sonnet-4-5 / claude-
  haiku-4-5 / claude-opus-4-7.
- oai_misc_config.yaml: align alias name with the gpt-5-mini rename.

Test changes (proactive: claude-sonnet-4-20250514 / claude-opus-4-
20250514 retire 2026-06-15):
- tests/llm_translation/test_anthropic_completion.py: bump 3 references
  + paired Vertex AI ID to claude-sonnet-4-5.
- tests/llm_translation/test_optional_params.py: bump 2 references.
- tests/pass_through_unit_tests/test_anthropic_messages_passthrough.py
  and test_bedrock_anthropic_messages_test.py: bump router fixtures
  using the deprecated model IDs.
- tests/pass_through_unit_tests/base_anthropic_messages_tool_search_test.py:
  modernize docstring examples.
- tests/test_end_users.py: update references to renamed alias.

* test: modernize placeholder model literals in router_unit_tests

Mass replace_all on fixture/placeholder model literals across the
router_unit_tests/ suite (model name is a routing key / label, not the
test subject). Sub-agent sweep so far — additional commits will follow
for logging_callback_tests/, enterprise/, top-level tests/test_*.py,
and other CI-mounted dirs.

Mappings applied:
- gpt-3.5-turbo -> gpt-5-mini
- gpt-4 (bare) -> gpt-5.5
- gpt-4o (bare) -> gpt-5
- text-embedding-ada-002 -> text-embedding-3-small
- claude-3-sonnet-20240229 / claude-3-opus-20240229 /
  claude-3-haiku-20240307 / claude-3-5-sonnet-20240620 ->
  claude-sonnet-4-5-20250929 / claude-opus-4-7 /
  claude-haiku-4-5-20251001 as appropriate

Explicitly preserved:
- gpt-4o-mini-* variants (transcribe, tts, etc.) where they're current
- gpt-4-turbo / gpt-4-vision-preview / gpt-4-0613 (subject literals)
- JSONL batch body literals
- Mock LLM response model fields (must match upstream)
- Fake/mock identifiers

* test: modernize placeholder model literals across remaining CI suites

Sub-agent sweep across logging_callback_tests/, guardrails_tests/,
enterprise/, pass_through_unit_tests/, otel_tests/,
llm_responses_api_testing/, batches_tests/, spend_tracking_tests/,
litellm_utils_tests/, unified_google_tests/, and a few top-level
tests/test_*.py files where the model literal is a fixture or
placeholder (router model_list, mock standard logging payload, mock
callback data) rather than the test's subject.

Mappings applied (see scope notes below):
- gpt-3.5-turbo -> gpt-5-mini
- gpt-4 (bare) -> gpt-5.5
- gpt-4o (bare) -> gpt-5.5 (corrected from initial gpt-5 — bare gpt-5
  is not a valid OpenAI alias; only gpt-5.5 / gpt-5.4 / gpt-5.2-codex
  / gpt-5-mini exist)
- gpt-4o-mini (bare) -> gpt-5-mini
- text-embedding-ada-002 -> text-embedding-3-small
- claude-3-sonnet-20240229 -> claude-sonnet-4-5-20250929
- claude-3-opus-20240229 -> claude-opus-4-7
- claude-3-haiku-20240307 -> claude-haiku-4-5-20251001
- claude-3-5-sonnet-20240620/20241022 -> claude-sonnet-4-5-20250929
- claude-3-7-sonnet-20250219 -> claude-sonnet-4-6
- gemini-1.5-flash -> gemini-2.5-flash
- gemini-1.5-pro -> gemini-2.5-pro

Explicitly preserved (not modernized):
- llm_translation/ tests where model is the SUBJECT (provider-specific
  translation/transformation logic). Only the deprecated 20250514
  references were already bumped in a prior commit.
- Cost-calc / tokenizer subject tests in test_utils.py (skip-ranges
  documented by the sub-agent).
- Bedrock model IDs in test_health_check.py path-stripping tests.
- JSONL batch request bodies and mock LLM response bodies (must match
  upstream literal).
- Langfuse expected-request-body JSON fixtures (cost values are exact-
  match-asserted; changing the model would shift response_cost).
- gpt-3.5-turbo-instruct (text-completion endpoint; no modern OpenAI
  equivalent).
- Top-level tests calling the proxy through user-facing aliases
  (gpt-3.5-turbo, gpt-4, text-embedding-ada-002, dall-e-3) — aliases
  in proxy_server_config.yaml stay; only the underlying model was
  bumped.
- tests/test_gpt5_azure_temperature_support.py (the test's whole point
  is model-name handling).
- Fake / mock / openai/fake identifiers.

Notable side fixes:
- test_spend_accuracy_tests.py: UPSTREAM_MODEL now matches what
  spend_tracking_config.yaml's proxy actually routes to (gpt-5-mini),
  resolving a latent inconsistency.
- proxy_server_config.yaml: bare `gpt-5` alias renamed to `gpt-5.5`
  (bare gpt-5 is not a valid OpenAI alias).
- test_batches_logging_unit_tests.py: explicit_models list entries
  kept distinct (gpt-5-mini + gpt-5.5) after bulk rename.

* test: fix CI failures from model modernization sweep

CI surfaced 4 categories of regression from the bulk modernization:

1. Azure deployment names are customer-specific. Reverted:
   - tests/litellm_utils_tests/test_health_check.py: azure/text-
     embedding-3-small -> azure/text-embedding-ada-002 (the CI Azure
     account does not have a text-embedding-3-small deployment).
   - tests/logging_callback_tests/test_custom_callback_router.py:
     same revert for two router fixtures driving aembedding.

2. gpt-5 family does not accept temperature != 1. Tests that pass a
   custom temperature swapped from gpt-5-mini to gpt-4.1-mini (modern
   non-reasoning OpenAI mini that still accepts temperature/logprobs):
   - tests/logging_callback_tests/test_datadog.py
   - tests/logging_callback_tests/test_langsmith_unit_test.py
   - tests/logging_callback_tests/test_otel_logging.py

3. proxy_server_config.yaml's gpt-3.5-turbo-large alias was routing to
   gpt-5.5 (a reasoning model that rejects logprobs). The proxy test
   tests/test_openai_endpoints.py::test_chat_completion_streaming
   exercises logprobs/top_logprobs through that alias. Bumped the
   underlying model to gpt-4.1 (non-reasoning, still modern).

4. tests/logging_callback_tests/test_gcs_pub_sub.py asserts against a
   pinned JSON fixture (gcs_pub_sub_body/spend_logs_payload.json) with
   hardcoded model="gpt-4o" and a model-specific spend value. Reverted
   the litellm.acompletion calls in the test to model="gpt-4o" so the
   fixture's exact-match assertions still hold.

5. tests/pass_through_unit_tests/test_anthropic_messages_passthrough.py:
   anthropic.messages.create routing to openai/gpt-5-mini returned an
   empty content[0] with max_tokens=100 (reasoning-token consumption).
   Swapped to openai/gpt-4.1-mini.

* test: fix Assistants API model + 2 cursor[bot] review nits

1. pass_through_unit_tests/test_custom_logger_passthrough.py: gpt-5.5
   isn't accepted by the /v1/assistants endpoint
   ("unsupported_model"). Switch to gpt-4.1-mini (modern, Assistants-
   API-supported, non-reasoning).

2. example_config_yaml/pass_through_config.yaml: the previous sweep
   bumped the claude-3-7-sonnet alias to claude-opus-4-7, which is a
   tier change (Sonnet -> Opus). Map to claude-sonnet-4-6 to keep the
   Sonnet tier intact. (Cursor bugbot review.)

3. example_config_yaml/simple_config.yaml: model_name was left as
   gpt-3.5-turbo while the underlying was bumped to gpt-5-mini, which
   muddles the "simple" example. Make both sides gpt-5-mini so the
   most basic example is a straight 1:1 mapping again. (Cursor bugbot
   review.)

* fix: revert gpt-4/gpt-3.5-turbo alias underlying to non-reasoning models

tests/test_openai_endpoints.py::test_completion calls the proxy alias
"gpt-4" with temperature=0, and other tests call gpt-3.5-turbo with
custom temperature / logprobs / the legacy /v1/completions endpoint.
The earlier modernization mapped both aliases to gpt-5.5 / gpt-5-mini,
which are reasoning models that reject temperature != 1 and don't
expose /v1/completions. Map the aliases to gpt-4.1 / gpt-4.1-mini
(modern non-reasoning OpenAI models) instead — keeps user-facing
aliases preserved while picking a current underlying that still
supports the parameters/endpoints the tests exercise.
2026-05-15 15:44:28 -07:00
Ishaan Jaffer
e8461b5b97
style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
harish876
1c74e17bed E2E test to assert response headers from the openai files change 2026-04-10 18:45:00 +00:00
David Chen
d1df4e838b
Litellm fix update bedrock models (#24947)
* update bedrock models in tests

* updated more tests and model_prices_and_context_window

* fix model id and pricing

* replace more sonnet models

* update tests

* git push

* update pricing

* flaky total cost

* monkey patch

* relax the cost change

* fix and revert some changes

* revert the pricing

* chore: move cost/pricing changes to bedrock-cost-fixes branch

* chore: split Bedrock file-api beta stripping to separate branch

Removes strip_unsupported_file_api_betas_for_bedrock_invoke from this branch;
see litellm_bedrock_invoke_strip_file_api_betas for that fix.

Made-with: Cursor
2026-04-01 19:22:54 -07:00
ishaan-berri
e4442a4d98
test fix us.anthropic.claude-haiku-4-5-20251001-v1:0 (#24931)
* test fix us.anthropic.claude-haiku-4-5-20251001-v1:0

* ignore mypy cache files

---------

Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
Co-authored-by: David Chen <clfhhc@gmail.com>
2026-04-01 11:01:03 -07:00
yuneng-jiang
4fc0975d22 Fix flaky e2e batch test: set batch_processed=True on completion in retrieve_batch
The retrieve_batch endpoint sets batch status to "complete" but never set
batch_processed=True, permanently blocking file deletion. CheckBatchCost
(the safety net) also excluded completed batches from its primary query,
so batch_processed was never set by either path.

Three fixes:
1. update_batch_in_database sets batch_processed=True when status reaches
   "complete", with old-schema fallback retry
2. CheckBatchCost primary query no longer excludes complete/completed
   (batch_processed=False filter prevents reprocessing)
3. retrieve_batch early-return now includes "complete" (DB-normalized
   spelling) to avoid unnecessary provider re-polls

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 18:18:32 -07:00
Sameer Kankute
82f5055d89 test(responses): add end-to-end test for responses API WebSocket mode
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-02 17:24:39 +05:30
Sameer Kankute
410e54648c Fix: Managed Batches: Inconsistent State Management for list and cancel batches 2026-02-03 14:47:28 +05:30
Ishaan Jaffer
96a0f960a9 test_e2e_batches_files 2025-11-27 09:14:43 -08:00
Ishaan Jaff
31ecd4ce49
Revert "Respect custom llm provider in header" (#17211) 2025-11-27 09:12:44 -08:00
Sameer Kankute
d4bc1cf23d Respect custom llm provider in header 2025-11-27 16:38:37 +05:30
Sameer Kankute
82dc0354ce
Litellm sameer nov 3 stable branch (#16963)
* Add openai metadata filed in the request

* Add docs related to openai metadata

* Add utils

* test_completion_openai_metadata[True]

* Added support for though signature for gemini 3 in responses api (#16872)

* Added support for though signature for gemini 3

* Update docs with all supported endpoints and cost tracking

* Added config based routing support for batches and files

* fix lint errors

* Litellm anthropic image url support (#16868)

* Add image as url support to anthropic

* fix mypy errors

* fix tests

* Fix: Populate spend_logs_metadata in batch and files endpoints (#16921)

* Add spend-logs-metadata to the metadata

* Add tests for spend logs metadata in batches

* use better names

* Remove support for penalty param for gemini 3 (#16907)

* Remove support for penalty param

* remove halucinated model names

* fix mypy/test errors

* fix tests

* fix too many lines error

* fix too many lines error

* Add config for cicd test case

* Fix final tests

* fix batch tests

* fix batch tests
2025-11-22 09:35:05 -08:00
Ishaan Jaffer
94c2c28f3d claude-sonnet-4-5-20250929 fix 2025-10-31 18:20:52 -07:00
Ishaan Jaff
aea78b8d1a
[Feat] Add support for Batch API Rate limiting - PR1 adds support for input based rate limits (#16075)
* add count_input_file_usage

* add count_input_file_usage

* fix count_input_file_usage

* _get_batch_job_input_file_usage

* fixes imports

* use _get_batch_job_input_file_usage

* test_batch_rate_limits

* add _check_and_increment_batch_counters

* add get_rate_limiter_for_call_type

* test_batch_rate_limit_multiple_requests

* fixes for batch limits

* fix linting

* fix MYPY linting
2025-10-29 18:28:52 -07:00
Timothée Lecomte
3ef9b2015a
feat: read from custom-llm-provider header (#15528) 2025-10-18 22:04:53 -07:00
Ishaan Jaffer
68105ce1a7 fix type 2025-09-16 15:41:52 -07:00
Ishaan Jaff
8e22cf5d65
[Fix] /responses API - add cancel endpoint + allow non-admins to use this as an llm api endpoint (#14594)
* fix: ensure /responses/cancel works for non admins

* test: cancel endpoint

* fix responses API  cancel endpoint

* test fix

* TestGoogleAIStudioResponsesAPITest
2025-09-15 18:49:54 -07:00
Sameer Kankute
110ce543c2
[Feat]Add cancel endpoint support for openai and azure (#14561)
* Add cancel endpoint support for openai
 and azure

* fix lint error

* fix cancel url contruction azure

* readd changes
2025-09-15 07:08:56 -07:00
Ishaan Jaff
93af8fd6ba
[QA] E2E - Testing for bedrock batches api (#14525)
* add bedrock/batch-anthropic.claude-3-5-sonnet-20240620-v1:0

* test_bedrock_batches_api

* fix

* fix import

* test_bedrock_batches_api
2025-09-12 19:31:19 -07:00
Ishaan Jaff
04dc1a5351
[Feat] Add support for returning images with gemini/gemini-2.5-flash-image-preview with /chat/completions (#13983)
* add gemini-2.5-flash-image-preview

* add gemini-2.5-flash-image-preview

* add image in ChatCompletionResponseMessage

* test_gemini_image_generation_async

* Revert "Merge pull request #13394 from Deviad/feature/enhance_logging_for_containers"

This reverts commit 539b94ad4e, reversing
changes made to 71af7bcf9c.

* include `image` in Delta

* fix _process_candidates should show the image response

* fix: _handle_special_delta_attributes

* test_gemini_image_generation_async_stream

* image_generation_chat

* UI - allow looking at generated images from /chat/completions

* _create_streaming_choice

* fix import StreamingChoices

* fix ChatCompletionResponseMessage

* test_gemini_image_generation

* add gemini img migration

* fix _extract_candidate_metadata

* ui fix

* fix batch endpoint test
2025-08-27 16:16:19 -07:00
Krish Dholakia
22d28f5853
Batches - support batch retrieve with target model Query Param + Anthropic - completion bridge, yield content_block_stop chunk (#12228)
* fix(batches_endpoints/endpoints.py): support passing target model names for batch list as a query param

Fixes issue where cloud run fails calls because GET can't contain request body

* test(test_openai_batches_endpoints.py): add unit test

* docs(managed_batches.md): update docs

* feat(spend_tracking_utils.py): support STORE_PROMPTS_IN_SPEND_LOGS env var

ensures prompt is stored in spend logs

* fix(streaming_iterator.py): fix anthropic - completion streaming iterator to yield content block stop

ensures claude code renders messages

* test: skip local test
2025-07-01 22:13:48 -07:00
Ishaan Jaff
4e7115bc34
Bug Fix - responses api fix got multiple values for keyword argument 'litellm_trace_id' (#12225)
* fix - handling trace id arg on responses api

* test_async_response_api_handler_merges_trace_id_without_error

* test_anthropic_with_responses_api
2025-07-01 18:12:22 -07:00
Ishaan Jaff
2e3a0222e6 Revert "test_anthropic_with_responses_api"
This reverts commit 2f0bdf887e80f85669fb0bae9fcede16863d658e.
2025-07-01 17:43:21 -07:00
Ishaan Jaff
648e8d6533 test_anthropic_with_responses_api 2025-07-01 17:43:21 -07:00
Ishaan Jaff
e3094c2249 set flaky tests as flaky 2025-06-14 13:51:52 -07:00
Ishaan Jaff
2e9fb2f6ba # expect an error when getting the response again since 2025-04-25 09:42:35 -07:00
Ishaan Jaff
cb87dbbd51 fix responses test 2025-04-24 21:23:25 -07:00
Ishaan Jaff
5de101ab7b
[Feat] Add GET, DELETE Responses endpoints on LiteLLM Proxy (#10297)
* add GET responses endpoints on router

* add GET responses endpoints on router

* add GET responses endpoints on router

* add DELETE responses endpoints on proxy

* fixes for testing GET, DELETE endpoints

* test_basic_responses api e2e
2025-04-24 17:34:26 -07:00
Ishaan Jaff
d41856ce3d test_bad_request_bad_param_error 2025-03-13 16:02:21 -07:00
Ishaan Jaff
3b632ac825 test_async_bad_request_bad_param_error 2025-03-13 15:57:19 -07:00
Ishaan Jaff
63783383e7 add basic validation tests for e2e responses create endpoint 2025-03-13 15:25:50 -07:00
Ishaan Jaff
e9f3d97eb0 working e2e tests for responses api 2025-03-13 15:17:47 -07:00
Ishaan Jaff
08f4a6844b rename folder to test openai endpoints 2025-03-13 15:13:48 -07:00