Commit graph

111 commits

Author SHA1 Message Date
Devin AI
8978b4562f test(main): drop unrelated reformatting
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:26:54 +00:00
Devin AI
fd2fb4c44e fix(http): address review on outbound HTTP/2
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 18:25:26 +00:00
yassin
e2e1d36804 fix(openai): keep extra_headers out of the chat request body on the httpx handler path
Forwarded client headers on bridged /v1/responses calls were serialized into the
OpenAI JSON body as extra_headers when EXPERIMENTAL_OPENAI_BASE_LLM_HTTP_HANDLER
was set, and OpenAI rejected the request with unknown_parameter. The headers are
already merged into the outgoing HTTP headers, so only set the SDK-style
optional param on the SDK client path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-14 22:27:35 +00:00
Mateo Wang
fe5ff9d3b0
Merge pull request #40771 from BerriAI/litellm_regression_coverage_followup
test: tighten regression tests added in #37974
2026-09-12 15:29:36 -07:00
devin-ai-integration[bot]
880ccc76a5
fix(streaming): keep admitted mock streams alive with empty stream_options and honor zero prompt counts (#40650)
* fix(streaming): keep usage-only chunks from crashing streams with empty stream_options

The usage-only chunk branch in CustomStreamWrapper.chunk_creator indexed stream_options["include_usage"] directly, so a caller passing stream_options={} hit a KeyError that surfaced as MidStreamFallbackError. Streaming mock_response with an admission input_tokens count (#40637) now always emits such a chunk, which made the crash reachable. Reuse the send_stream_usage policy computed at init instead. Also annotate the #40637 test bindings with Final.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): report admitted zero prompt tokens instead of recounting in mock streams

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-10 18:37:39 -07:00
devin-ai-integration[bot]
c7a41c35d5
perf(mock): emit admission-time usage chunk on streaming mock_response (#40637)
* perf(mock): emit admission-time usage chunk on streaming mock_response

Streaming mock_response chunks carried no usage, so the chunk builder re-tokenized the whole prompt in Python after the stream ended even when budget reservation had already counted it at admission. The mock streaming generators now yield a final usage-only chunk carrying the admission prompt count (same completion count as the non-streaming path). Without an admission count the old tokenizer fallback stays.

* fix(mock): type the mock stream generators and keep the usage chunk on the content stream id

Review follow-up: the usage-only chunk was built with a fresh id, so CustomStreamWrapper switched response_id for the finish-reason and usage chunks. It now copies the content stream id. The generators also get full parameter and return annotations.

---------

Co-authored-by: yassin <yassin@berri.ai>
2026-09-11 00:48:25 +00:00
devin-ai-integration[bot]
f84034f500
feat(mock): report admission-time input token count in mock_response usage (#40590)
* feat(mock): report admission-time input token count in mock_response usage

Mock completions always reported prompt_tokens=10, so spend tracking, TPM metrics, budgets and the tokens-per-minute autoscaling signal saw 10 tokens for a 100k-token request. Budget reservation now carries the admission-time input token count in the reservation record, and mock_completion reads it back so mock traffic exercises the same spend and TPM paths as real traffic without any extra tokenizer work.

* fix(mock): keep a zero admission input token count instead of falling back to 10

---------

Co-authored-by: yassin <yassin@berri.ai>
2026-09-10 11:22:26 -07:00
Mateo Wang
aefa1040a0
Merge pull request #40249 from BerriAI/litellm_fix_responses_bridge_reasoning_effort
fix: keep reasoning_effort for mode: responses bridge deployments
2026-09-09 18:46:58 -07:00
mateo-berri
35def27e7b fix(convert_dict_to_response): accept only a real list as choices and keep /v1/messages alive on an empty one
Narrows the no-choices guard so a dict, string, or None still raises the APIError while an empty list passes through,
guards the non-stream Anthropic bridge against indexing an empty choices list, and repairs test_completion_missing_role,
whose raw-response mock was patched in as the create() callable itself so the handler only ever saw a MagicMock
2026-09-09 12:37:29 -07:00
jesus
6f3b3c7957 test: cover responses bridge reasoning at HTTP boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-08 14:27:12 +00:00
jesus
17dd6afbaa fix: keep reasoning_effort for mode: responses bridge deployments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-08 13:51:58 +00:00
mateo-berri
6d01ed803d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mistral_voxtral_tts_speech 2026-09-04 16:45:33 -07:00
Yujong Lee
fae3d224eb Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci
# Conflicts:
#	basedpyright-code-budget.json
#	tests/sdk_function_trace/profiler.py
#	tests/sdk_function_trace/test_profiler.py
2026-09-04 09:01:13 -07:00
Mateo Wang
025a3ca42f
Merge pull request #39631 from BerriAI/litellm_gpt_6_astra_detection
fix: treat gpt-6 names as the gpt-5 request family in OpenAI and Azure configs
2026-09-03 14:00:59 -07:00
mateo-berri
108f558946 test: drop the internal patch from the gpt-6-astra bridge test 2026-09-03 13:46:33 -07:00
mateo-berri
bba75c7ce9 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mistral_voxtral_tts_speech
# Conflicts:
#	tests/test_litellm/test_cost_calculator.py
#	tests/test_litellm/test_main.py
2026-09-03 13:35:30 -07:00
mateo-berri
ab515dbc90 fix: treat gpt-6 names as the gpt-5 request family in OpenAI and Azure configs 2026-09-03 13:20:13 -07:00
Mateo Wang
e4b8caeb36
Merge pull request #38975 from BerriAI/litellm_fix_azure_ai_reclassify
fix(azure_ai): don't reclassify Foundry deployments as azure provider
2026-09-03 13:15:40 -07:00
mateo-berri
7a5e4b9ba4 fix(openai): bridge gpt-5.4+ tool calls to /v1/responses on every api.openai.com host
The auto-bridge that moves gpt-5.4+ requests carrying function tools and no
reasoning_effort onto /v1/responses only fired when the resolved api_base was
the literal https://api.openai.com/v1, so a deployment pointed at an OpenAI
PrivateLink hostname (<region>.privatelink.api.openai.com) or a port-qualified
or trailing-slash default stayed on Chat Completions and got OpenAI's 400 back.
Gate on the resolved URL's hostname instead: api.openai.com or any subdomain of
it bridges, every other custom base still stays on chat
2026-09-03 10:04:18 -07:00
mateo-berri
6edb72f79c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_azure_ai_reclassify
# Conflicts:
#	tests/test_litellm/test_main.py
2026-09-03 00:47:14 -07:00
mateo-berri
0a2581c14c fix(mistral): keep the deployment voice default and drop the unreachable api base fallback
Review turned up two real problems in the TTS path.

Router.aspeech forwarded voice=None whenever the caller omitted it, which overwrote a
voice set in the deployment's litellm_params, so a configured fallback voice was
ignored on voice-less requests. It now leaves the key alone when no voice is passed.

get_complete_url also fell back to MISTRAL_API_BASE, but speech() always receives a
non-null api_base from get_llm_provider, whose mistral branch only reads
MISTRAL_AZURE_API_BASE and otherwise hardcodes the public host. That branch could
never run, and its unit test asserted a behavior the real path does not have. The
working override is api_base on the deployment, now pinned by an end-to-end test
2026-09-03 00:30:07 -07:00
mateo-berri
71f5b89499 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_python_version_ci
Resolves two conflicts:

- tests/test_litellm/vector_stores/test_main.py: staging moved search() to a
  RouterVectorStoreEmbeddingExecutor while this branch parametrized the same
  test over query; keep both the executor assertions and the parametrize.
- tests/logging_callback_tests/test_bedrock_knowledgebase_hook.py: staging
  carries a duplicate embedding_executor kwarg that makes the file a
  SyntaxError; drop the trailing duplicate.
2026-09-03 00:16:19 -07:00
mateo-berri
75e7f4c4a5 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mistral_voxtral_tts_speech
Resolves the tests/test_litellm/test_main.py collision, where both sides appended a
new test at the end of the file, by keeping both.

Also carries the one-line fix from #39502: staging arrived with a duplicate
embedding_executor kwarg in the Bedrock KB fake handler, which ruff rejects as a
syntax error, so every commit here would otherwise fail lint. The change is byte
identical to #39502, so that PR merges cleanly once it lands.
2026-09-03 00:07:31 -07:00
mateo-berri
95da998779 fix(xai): merge litellm_internal_staging and keep the stream builder from short-circuiting xAI's reported cost 2026-09-02 16:46:35 -07:00
mateo-berri
1a1d459701 fix(xai): keep streamed and custom-priced billing inside the cost calculator
Restate xAI's usage.cost_in_usd_ticks as usage.cost on chat and responses
replies, streamed ones included, then let the cost calculator own the
figure: a deployment with its own input_cost_per_token and
output_cost_per_token keeps that price, cost margins apply on chat streams
as they already did on non-streamed calls, and only OpenRouter's usage
cost becomes the llm_provider-x-litellm-response-cost header, so xAI
streams no longer skip the calculator through the header or the
stream_chunk_builder hidden response_cost.
2026-09-02 16:13:31 -07:00
Yujong Lee
77d6aedf0a fix: address cross-version CI failures 2026-09-02 14:17:19 -07:00
mateo-berri
25bb8c92e3 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_mistral_voxtral_tts_speech 2026-09-02 12:06:17 -07:00
mateo-berri
df24dab7c9 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_azure_ai_reclassify
# Conflicts:
#	tests/test_litellm/test_main.py
2026-09-02 11:31:12 -07:00
mateo-berri
7b942fd983 fix(azure_ai): route audio and realtime calls on Foundry hosts through the Azure OpenAI handlers 2026-09-02 11:05:38 -07:00
mateo-berri
9e25dd708f feat(streaming): carry final response cost on streamed usage by default
Streamed responses through the proxy previously exposed no usable cost:
the x-litellm-response-cost header is unreadable mid-stream and the final
usage chunk carried only tokens, priced against an alias model name the
client cannot resolve. The include_cost_in_streaming_usage flag existed
but was off by default and only fixed the wire, not SDK clients.

Stamp usage.cost into the joined streaming response by default wherever a
final usage object is built: the chat-completions stream_chunk_builder,
the native /v1/responses RESPONSE_COMPLETED event, and synthetic response
events. Provider-reported cost always wins over the computed value, and
only positive computed costs are stamped so unpriceable alias responses
keep deferring to the logging object's own calculation. Per-chunk SSE
cost injection (/v1/messages, generateContent, passthrough) stays behind
the flag.

Also normalize non-litellm usage objects in stream_chunk_builder: openai
CompletionUsage lacks Usage.__contains__, so membership probes silently
returned False and client-side rebuilds dropped the wire cost and
recounted token usage locally. Wire token counts and cost now survive.

Resolves LIT-6427
2026-08-31 21:47:13 -07:00
mateo-berri
676f841534 feat(mistral): add text-to-speech support for /v1/audio/speech 2026-08-29 03:47:38 -07:00
mateo-berri
a17bc1f22c fix(streaming): price proxy-aliased models from the model map in stream_chunk_builder 2026-08-28 12:59:42 -07:00
mateo-berri
349a653c29 fix(main): report response_cost and Anthropic citations from stream_chunk_builder 2026-08-28 12:39:37 -07:00
mateo-berri
fdeab570a1 fix(speech): forward api_key to the TTS bridge and isolate response hidden params 2026-08-26 16:23:33 -07:00
mateo-berri
4d1d7b446f test(speech): type the bridge spend regression test helpers 2026-08-26 15:27:38 -07:00
mateo-berri
3418d7baf9 fix(speech): keep proxy metadata and completion cost through the TTS completion bridge 2026-08-26 15:12:41 -07:00
mateo-berri
0bd4d323da fix(router): resolve provider from api_base in deployment validation and acompletion
Router._add_deployment called get_llm_provider without the deployment's api_base, so a config entry with a bare model plus a known OpenAI-compatible endpoint failed startup validation with LLM Provider NOT provided and the proxy returned 400 no healthy deployments for that model group. acompletion had the same gap at request time: it forwarded only base_url into its get_llm_provider call, dropping the api_base kwarg the router passes. Both now forward api_base so endpoint matching resolves the provider the same way sync completion already does
2026-08-25 10:33:40 -07:00
mateo-berri
8dda73c935 test: tighten regression tests added in #37974
Five of the tests could pass without the behavior they guard being
correct. The passthrough spend tests derived their expected spend in
setup_method from whatever cost map was live rather than the pinned
checked-in one. The test_main.py cost fixture cleared only one of the
two price caches, leaving billing to read stale prices while the
assertions read the pinned map. The gpt-5.6 bridge test parametrized
over two suffixes the version check discards, so both cases were
identical. The anthropic flush helper swallowed the loop-binding
RuntimeError it exists to report. The cache-write test pinned a literal
1.25 rate ratio unrelated to the bug it guards.
2026-08-22 16:35:57 -07:00
Mateo Wang
75613bf22f
test: add regression coverage for twelve closed issues (#37974)
* test: add regression coverage for twelve closed issues

Adds targeted regression tests for behavior that was fixed but left ungated,
so the fixes cannot silently regress:

- #33772 openai cache_write_tokens cost
- #34309 Responses API cache cost_breakdown
- #35363 /v1/responses batch spend
- #36619 auto-router api_base/api_key leak on a shared model name
- #35359 batch fallbacks within the owning model group
- #36523 passthrough streamed Responses spend log
- #36646 passthrough embeddings spend log
- #37147 non-object metadata on create_batch is a 400
- #35362 unscoped list files reads the managed-file store
- #33221 gpt-5.6 bridges to Responses on function tools alone
- #34487 LLM complexity classifier runs for every caller metadata shape
- #35124 streamed /v1/messages emits success logging on both bridges

Cost assertions read rates from litellm.model_cost rather than hardcoding
dollar amounts, so they do not drift on repricing.

* fix: stop the new regression tests polluting and tripping over shared global state

Two shard failures, both from global state the new tests share with their
neighbours rather than from the behaviour under test.

test_main.py's local_cost_map pinned litellm.model_cost but left the
get_model_info lru_cache warm, so completion_cost billed at whatever prices
were cached earlier in the process while the assertions read the pinned map.
Clear the cache on both sides of the fixture, matching the local_model_cost_map
fixture in tests/test_litellm/conftest.py.

The anthropic messages streaming tests called GLOBAL_LOGGING_WORKER.flush()
on whatever queue happened to be around. A queue left non-empty by an earlier
test is still bound to that test's loop, so join() either hangs or raises
"bound to a different event loop". Rebind to the running loop before the call
and wait for the captured payload instead of a fixed sleep.
2026-08-22 22:24:05 +00:00
yuneng-jiang
6a0d03914c
test: drop the cwd-relative sys.path.insert calls from the test suite (#37802)
* test: drop the cwd-relative sys.path.insert calls from the test suite

TQ003 stands at 1,077 across 1,058 files, and 1,015 of them are the same shape:
sys.path.insert(0, os.path.abspath("../..")) and its deeper siblings. The
argument resolves against the working directory rather than the file, so from
the repo root, where every job runs pytest, it inserts the directory two levels
above the checkout. It has never pointed at litellm. The package is installed
into the environment anyway, which is what actually makes the import work, and
what the rule's message has said all along.

Removing them leaves 1,634 imports of sys and os with no remaining reference,
and those go too, except where another test module imports the name back out of
the file. The rest of TQ003 is 62 call sites that resolve against __file__ or a
variable, which are a different question and are left alone.

Collection is identical either way: 45,871 tests and the same 51 pre-existing
collection errors before and after, and ruff reports no new undefined name.

* test: drop the duplicate imports the sys.path sweep exposed to F811

* test(pre-call-utils): restore the os import the new bedrock tests need
2026-08-22 09:25:58 -07:00
mateo-berri
092449b32f Merge branch 'litellm_internal_staging' of https://github.com/BerriAI/litellm into litellm_fix_oauth_credential_forwarding 2026-08-22 08:45:05 -07:00
mateo-berri
349e7e6990 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_oauth_credential_forwarding
# Conflicts:
#	tests/test_litellm/test_main.py
2026-08-22 08:44:58 -07:00
yuneng-jiang
0c97eea660
test(cost-calc): stop 182 global writes leaking out of the cost-calc suites (#37815)
* test(cost-calc): stop 182 global writes leaking out of the cost-calc suites

Across test_cost_calculator.py and llm_cost_calc/test_llm_cost_calc_utils.py,
58 tests opened by setting LITELLM_LOCAL_MODEL_COST_MAP in os.environ and
replacing litellm.model_cost, and none of them put the env var back. The
second file already had a _local_model_cost_map fixture doing it by hand with
a try/finally, so both idioms sat in the same file.

Keep that fixture, give it monkeypatch, and have every one of those tests ask
for it. The margin and discount tests drop their hand-rolled
copy-then-restore in favour of monkeypatch.setattr, which also puts the
global back when an assertion fails part way through.

Both files also drop a sys.path.insert whose argument resolves outside the
repo, so it was never what made the imports work.

TQ003 1077 -> 1075, TQ004 768 -> 693, TQ005 2836 -> 2731, and the budget
ceilings come down with them.

* fix(test): make the streamed-cost tests load the map they assert against

The local_cost_map fixture set LITELLM_LOCAL_MODEL_COST_MAP but never reloaded
litellm.model_cost, and reading the variable is not what loads the map. So the
three streaming-cost tests billed against whatever map the process happened to
be holding, and their hardcoded prices only held when something else had
already swapped in the checked-in one. This branch stops the cost-calc tests
leaking that map, which left test_main billing at the ambient prices instead.

The fixture now loads the map it names, so the prices these tests assert hold
on their own.
2026-08-21 21:00:04 -07:00
yuneng-jiang
b416bdadd3
test(main): pin what a streamed response costs, end to end (#37812)
Rebuilding a streamed response and pricing it is the path a spend row comes
from, and nothing asserted it end to end. Reversing either half of the usage
the provider reported left the file green.

Three cases: the rebuilt response bills the usage the last chunk carried,
streaming and not streaming bill the same usage the same, and a stream that
reported no usage is still billed rather than dropped.

The cost is asserted against the catalog prices the run itself reads, with a
non-zero guard in front of it so an all-zeros lookup cannot satisfy it
vacuously. Pinning the dollar figure as a literal would have made a routine
gpt-4o price update fail a test about usage reconstruction.
2026-08-21 20:15:27 -07:00
mateo-berri
43143933d9 fix: stop provider-scoped headers leaking across fallback hops
completion() aliased the caller's header mapping instead of copying it,
then merged the provider-scoped headers into that same object. The router
shares one header dict across fallback attempts, so the credential written
on an Anthropic attempt was still present when a later Bedrock or Vertex
attempt read the dict, defeating the provider scoping.

Copy the mapping before merging so each attempt sees only its own headers.
2026-08-21 19:58:19 -07:00
mateo-berri
e16cf5ed3f Fall back to the GPT version rule when the cost map carries no breakpoint flag
A proxy on the default remote cost map never produced a prompt cache
breakpoint: the published map has the gpt-5.6 entries without
supports_prompt_cache_breakpoint, so the model-map gate returned False
for every listed model and only LITELLM_LOCAL_MODEL_COST_MAP=True (the
repo .env, hence the passing unit tests) made the feature work. The hook
now honors the flag when the entry carries one, True or False, and
otherwise applies the GPT-5.6+ version rule to the model name, so a map
that lags the flag still gets the OpenAI dialect. The model-map tests
pin litellm.model_cost to the bundled backup map and a new test drives
the hook against an unflagged gpt-5.6 entry.

completion() and acompletion() take base_url as an alias for api_base
that only lands on api_base after the cache control hook ran, so a
GPT-5.6 call at a non-OpenAI gateway given through base_url still got
the dialect. Both seed calls and the unstamped request-params read now
look at base_url too.

ResponsesAPIRequestUtils.merge_prompt_management_input reshaped hook
output in place, retyping text parts to input_text on the caller's own
message objects. The merge now shapes a copy of each message as it
emits it, so the identity-based merge keeps working on the hook's
objects and nothing the hook or the client owns is mutated.
2026-08-20 05:47:19 -07:00
mateo-berri
d9aaa95978 Gate OpenAI prompt cache breakpoints on the real target and carry them through /v1/responses
The cache control hook also runs on litellm.responses() input. On a
GPT-5.6 deployment it wrapped a string-content item into a chat-shaped
{"type": "text"} part, which the Responses API rejects, and it never
marked input_text, input_image or input_file parts, so no breakpoint and
no prompt_cache_options reached the provider. Add the Responses part
types to the eligible block set and translate chat-shaped text parts on
non-assistant items to input_text in
ResponsesAPIRequestUtils.merge_prompt_management_input, which both the
async and the sync prompt management sites go through.

The dialect also fired for any GPT-5.6 name that resolved to provider
openai, including deployments pointed at a custom api_base that does not
understand prompt_cache_breakpoint. Decide it once per request from the
provider, the model map and the resolved api_base (request, then
litellm.api_base, then OPENAI_BASE_URL / OPENAI_API_BASE): only
api.openai.com and *.api.openai.com hosts speak the dialect, a top-level
prompt_cache_options opts a custom target in, and litellm_proxy/ targets
never get it. maybe_seed_default_injection_points takes api_base and
stamps the finished decision on the points as _litellm_openai_dialect so
the sync completion() path, whose hook params do not carry api_base,
honors it; maybe_inject_cache_control takes api_base from the
/v1/messages handler.

Eligibility now comes from a supports_prompt_cache_breakpoint model map
flag on the OpenAI gpt-5.6 entries, exposed through
litellm.utils.supports_prompt_cache_breakpoint, with the GPT version rule
kept only for models the map does not know. The OpenAI dialect no longer
reserves a slot for tool_config points, which OpenAI has no cache block
for, and with_prompt_cache_breakpoint plus the chat bridge helper return
a new block instead of mutating their input.
2026-08-20 05:17:13 -07:00
mateo-berri
ec35098108 fix(main): forward store and prompt_cache_key on the MCP gateway early-return 2026-08-18 13:57:07 -07:00
mateo-berri
0e0768df9f Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_pr_33195_head
# Conflicts:
#	litellm/main.py
#	litellm/utils.py
#	tests/test_litellm/test_main.py
2026-08-18 13:03:04 -07:00
Fahima Mokhtari
e1ef7775bd
fix(main): an explicit provider outranks a known OpenAI model name (#36800)
* fix(main): an explicit provider outranks a known OpenAI model name

completion() picks the OpenAI handler whenever `model in
litellm.open_ai_chat_completion_models`, and that clause is evaluated before the
gemini and vertex_ai branches. get_llm_provider() already resolves those names
to "openai", so the clause only adds anything when the provider is something
else, and then it silently overrides it: the config built for the requested
provider is handed to the OpenAI handler.

For gemini that is fatal. VertexGeminiConfig.transform_request raises
NotImplementedError by design, since Vertex builds its request in its own
handler, so `gemini/gpt-4o` dies in async_transform_request before anything is
sent. register_model() reaches the same state without an odd model id: an entry
claiming litellm_provider "openai" adds its name to
open_ai_chat_completion_models, so one mislabelled pricing entry reroutes every
later call to that model in the process.

The name clause now applies only when no other provider was resolved.

* test(main): move the routing regression into the mapped test file

CLAUDE.md asks bug fixes to extend the mapped test file, so these belong in
tests/test_litellm/test_main.py rather than a module of their own.

They also no longer swap out the provider handler objects. Both Gemini cases
inject an HTTPHandler whose post() answers like generativelanguage does, then
assert the URL the request went to and read the reply back; the OpenAI case
injects an OpenAI client and patches its own raw-response create. That asserts
the endpoint the call reaches instead of which attribute the test replaced, and
matches the neighbouring tests in the file.
2026-08-14 11:39:02 -07:00