Commit graph

39832 commits

Author SHA1 Message Date
Ishaan Jaffer
097bab9441
test(ocr): assert rust path resolves key via secret manager 2026-06-22 21:05:16 -07:00
Ishaan Jaffer
a21a80751d
ocr: resolve mistral key via get_secret_str before the rust path (secret-manager parity) 2026-06-22 21:05:16 -07:00
Ishaan Jaffer
53f059a1d3
test(interactions): add budget_exceeded to expected status enum (Google updated the published spec) 2026-06-22 20:58:49 -07:00
Ishaan Jaffer
ca904f5607
test(ocr): assert rust_ocr returns the raw bridge dict 2026-06-22 20:37:21 -07:00
Ishaan Jaffer
17791a2174
ocr: wrap rust bridge dict into OCRResponse at the call site 2026-06-22 20:37:21 -07:00
Ishaan Jaffer
718b1d550c
ocr: make rust_bridge a leaf (return raw dict, no litellm import) so the CodeQL autofix stops re-breaking it 2026-06-22 20:37:21 -07:00
ishaan-berri
c9b712c823
Potential fix for pull request finding 'CodeQL / Cyclic import'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-06-22 20:26:59 -07:00
ishaan-berri
bbbb9ededb
Potential fix for pull request finding 'CodeQL / Cyclic import'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-06-22 20:26:48 -07:00
Ishaan Jaffer
fe1cbe505d
ci: re-trigger checks 2026-06-22 20:14:54 -07:00
Ishaan Jaffer
d6b9928dc1
ci: re-trigger checks 2026-06-22 20:10:01 -07:00
Ishaan Jaffer
a4b4965888
ocr: modernize rust_bridge typing (PEP 604, drop typing.Any/Dict) to satisfy strict-rule gate 2026-06-22 18:56:24 -07:00
Ishaan Jaffer
88a9e6ed57
ocr: guard OCRResponse under TYPE_CHECKING so the annotation resolves 2026-06-22 18:49:03 -07:00
Ishaan Jaffer
3188f93be0
ocr: lazily import rust bridge inside ocr() to break the import cycle the CodeQL autofix mangled 2026-06-22 18:49:03 -07:00
ishaan-berri
54577a691e
Potential fix for pull request finding 'CodeQL / Module-level cyclic import'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-06-22 18:36:24 -07:00
ishaan-berri
8325d52451
Potential fix for pull request finding 'CodeQL / Module-level cyclic import'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-06-22 18:36:10 -07:00
Ishaan Jaffer
6fee78a760
ci(rust): build with --locked to enforce the lockfile 2026-06-22 18:28:06 -07:00
Ishaan Jaffer
c6aa0607fe
rust: commit Cargo.lock for reproducible builds 2026-06-22 18:28:06 -07:00
Ishaan Jaffer
100861fba1
rust: stop ignoring Cargo.lock 2026-06-22 18:28:06 -07:00
Ishaan Jaffer
d2339a9f3b
test(ocr): cover Rust OCR routing + toggle 2026-06-22 18:25:06 -07:00
Ishaan Jaffer
d7ce63f20f
litellm: export use_litellm_rust() 2026-06-22 18:25:06 -07:00
Ishaan Jaffer
b370c1c191
ocr: route mistral to Rust when enabled; keep bare-str file rejection 2026-06-22 18:25:06 -07:00
Ishaan Jaffer
c697154bf3
ocr: add minimal Rust bridge (use_litellm_rust + rust_ocr) 2026-06-22 18:25:06 -07:00
Ishaan Jaffer
8feefec4f4
rust(bridge): end-to-end ocr() + gil_stats(), GIL released for HTTP 2026-06-22 18:25:06 -07:00
Ishaan Jaffer
9240cad089
rust(bridge): add GIL release accounting 2026-06-22 18:25:06 -07:00
Ishaan Jaffer
0ce378d358
rust(bridge): depend on litellm-core 2026-06-22 18:25:06 -07:00
Ishaan Jaffer
055bef8482
rust(providers): end-to-end run_ocr orchestrator with shared client + timeout 2026-06-22 18:25:05 -07:00
Ishaan Jaffer
ba8c541804
rust(mistral): add complete_url + resolve_api_key helpers 2026-06-22 18:25:05 -07:00
Ishaan Jaffer
74b1e6d691
rust(providers): depend on reqwest 2026-06-22 18:25:05 -07:00
Ishaan Jaffer
b86d3168d6
rust: add reqwest (rustls-tls) workspace dependency 2026-06-22 18:25:05 -07:00
Ishaan Jaffer
6b83353639
rust(core): add Auth/Http/Network error variants 2026-06-22 18:25:05 -07:00
Ishaan Jaffer
fea5204a2f
Simplify rust ocr entrypoint 2026-06-22 17:53:18 -07:00
Ishaan Jaffer
6a135105eb
address greptile rust ocr feedback 2026-06-22 17:38:24 -07:00
Ishaan Jaffer
70afb75ff1
Add litellm rust workspace with mistral ocr bridge 2026-06-22 17:07:00 -07:00
yuneng-jiang
dcf1b445e6
Merge pull request #30907 from BerriAI/litellm_internal_staging
Some checks failed
CodeQL / Analyze (actions) (push) Has been cancelled
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (python) (push) Has been cancelled
CodSpeed Benchmarks / benchmarks (push) Has been cancelled
Helm unit test / unit-test (push) Has been cancelled
Scorecard supply-chain security / Scorecard analysis (push) Has been cancelled
GitHub Actions Security Analysis / zizmor (push) Has been cancelled
chore(ci): promote internal staging to main
2026-06-20 18:20:54 -07:00
yuneng-jiang
df5619f9b2
chore: update Next.js build artifacts (2026-06-21 00:50 UTC, node v20.20.2) (#30906) 2026-06-20 18:09:54 -07:00
ryan-crabbe-berri
8125ddd2f8
fix(ui): source api-keys identity from useAuthorized to stop "User ID is not set" (#30903)
The migrated /ui/api-keys route gates rendering on useAuthorized() but read userID from the AuthContext (useAuth), which hydrates asynchronously. On a hard refresh or deep link the route could render UserDashboard before AuthContext had populated userID, so UserDashboard hit its `userID == null` guard and showed "User ID is not set". The legacy index page avoided this by gating on AuthContext's own authLoading; the migration switched the gate to useAuthorized without aligning the identity source.

Read identity from useAuthorized (a synchronous cookie decode) so userID is populated whenever the route is authorized. useAuth is kept only for the backfill setters UserDashboard still expects, until the planned AuthContext consolidation removes them.

Refs LIT-3687
2026-06-20 17:48:35 -07:00
Krrish Dholakia
53593f697d
feat(sandbox): e2b code execution primitive (#30898)
* feat(sandbox): add e2b code execution primitive

Add a provider-agnostic code execution primitive that runs model-generated
code in an isolated sandbox and returns the output, with e2b as the first
backend over raw httpx (no SDK dependency).

Public API: litellm.acode_interpreter_tool (ephemeral create -> run -> delete)
plus the low-level lifecycle litellm.acreate_sandbox / arun_code /
adelete_sandbox. Each is @client-decorated so operations are logged like
litellm.asearch. Backends implement BaseSandboxConfig; resolved via
ProviderConfigManager.get_provider_sandbox_config.

* fix(sandbox): address review feedback and CI gates

- document e2b provider in provider_endpoints_support.json and add a sandbox endpoint definition
- regenerate dashboard CallTypes after the sandbox call-type additions
- guard explicit timeout=0 instead of coercing it to the default
- require a ContainerHandle access token before running code; reject bare-id runs
- return False on a 404 delete now that the shared http handler raises for status
- skip non-JSON NDJSON lines and cap streamed output to bound memory
- move the real-network integration tests out of tests/test_litellm into tests/integration/sandbox

* fix(sandbox): satisfy strict ruff gate and scope star-exports

- modernize annotations in the new sandbox modules to PEP 585/604 (list/dict,
  X | None) and drop the now-unnecessary quoted forward refs so the strict-rule
  budget delta for UP006/UP037/UP045 returns to zero
- add __all__ to litellm/sandbox/main.py so 'import *' only re-exports the four
  public entrypoints instead of leaking module-level imports

* fix(sandbox): drop quotes on sandbox config return annotation

utils.py uses 'from __future__ import annotations', so the quoted forward ref
tripped UP037; the unquoted union is lazily evaluated and keeps the strict-rule
delta at zero

* chore(sandbox): re-trigger automated review after addressing feedback
2026-06-20 16:30:01 -07:00
Mateo Wang
b16cfd7de9
test: point router/completion/triton tests at the local fake OpenAI endpoint (#30900)
* test: point router/completion/triton tests at the local fake OpenAI endpoint

The shared Railway-hosted mock (exampleopenaiendpoint-production.up.railway.app)
takes down unrelated CI jobs whenever it is unreachable. #30695 moved the mounted
proxy configs onto a job-local fake server but left these in-Python api_base
literals pointing at the dead host, so litellm_router_testing, local_testing_part1,
local_testing_part2 and llm_translation_testing still fail with a 404
"Application not found" when Railway is down

Resolve the api_base from FAKE_OPENAI_API_BASE (default http://127.0.0.1:8190)
through a shared helper, auto-start the canned server from the local_testing and
llm_translation conftests when nothing is already serving, and extend the server
with a Triton embeddings route and a slow-endpoint delay so the triton and
latency-timeout tests run fully offline. The deliberately broken fallback URL is
left as-is so fallback handling still has a failing upstream

* fix: ignore non-loopback FAKE_OPENAI_API_BASE so the local mock is used in CI

* fix: drop 0.0.0.0 from loopback hosts, an unreliable client connect target

* fix(tests): keep fake OpenAI mock alive across xdist workers

ensure_fake_openai_endpoint registered atexit on the worker that spawned
the subprocess, so under -n 4 the first worker to drain its queue would
terminate the shared mock while siblings were still hitting it. Detach
the child via start_new_session and drop the per-worker teardown; reuse
on /health handles re-runs and CI containers clean up themselves
2026-06-20 16:20:35 -07:00
Shivam Rawat
c7efa77de3
fix(watsonx): wrap string embedding input in array for WatsonX API (#30897)
* fix(watsonx): wrap string embedding input in array for WatsonX API

WatsonX text/embeddings expects inputs as []string; OpenAI clients often send a single string.

Co-authored-by: Cursor <cursoragent@cursor.com>

* style(watsonx): format watsonx embed transformation for black

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(watsonx): avoid UP006 in embed transformation strict lint gate

Use list[str] and branch-based input normalization instead of List and cast
so the watsonx embedding change does not add strict ruff UP006 violations.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Shivam Rawat <shivamrawat@Shivams-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-20 15:50:44 -07:00
yuneng-jiang
46d2861001
chore(ci): bump deps (#30899)
* bump: version 0.1.42 → 0.1.43

* bump: version 1.89.0 → 1.90.0

* adding uv lock
2026-06-20 15:40:03 -07:00
yuneng-jiang
140ca3012a
chore: update Next.js build artifacts (2026-06-20 21:24 UTC, node v20.20.2) (#30894) 2026-06-20 15:07:32 -07:00
Mateo Wang
a7b0b0ba09
feat: add lint-gate target and truncation-proof summary to the strict ruff gate (#30877)
* feat: add CI-parity mode and truncation-proof summary to strict ruff gate

* refactor: tolerant worktree cleanup and concrete GateInputs types

* fix: clean up temp dir when git worktree add fails

* fix: align lint-gate with CI by dropping unused --ci-parity path

The lint-gate Makefile target invoked ruff_strict_gate.py with --ci-parity,
which counted violations on a throwaway merge of base into HEAD against base
counts at the base tip. CI in test-linting.yml runs the same script without
--ci-parity on a PR-head checkout, taking the gather_fast path that counts on
the live tree against base counts at the merge-base. A local pass could
therefore disagree with CI.

Drop --ci-parity from the Makefile and remove the now-unused gather_ci_parity
branch and flag so there is one code path that both local and CI exercise.
The docstring claim that CI runs against the synthetic merge ref was also
wrong; the workflow checks out github.event.pull_request.head.sha.

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-06-20 11:46:01 -07:00
Mateo Wang
15aa40b36e
test(ui): isolate OldTeams delete-warning tests from leaked mock (#30871)
Some checks are pending
GitHub Actions Security Analysis / zizmor (push) Waiting to run
The deprecated OldTeams component takes only accessToken, userID, userRole
and premiumUser; it ignores the teams prop these tests passed and instead
populates its table from the mocked teamListCall. The delete-warning block
never set teamListCall, and vi.clearAllMocks clears call history but not
implementations, so the table rendered the "Legacy Team" (keys.length 2)
left behind by the previous block's last test. Both delete tests therefore
ran against that leaked team: the keys-present case passed only because the
leaked count happened to be 2, and the no-keys case rendered the same warning
it asserted should be absent, so it failed.

Seed the team through the channel the component actually reads (teamListCall)
and drop the props it never consumes, so each test renders exactly the team
it declares. The keys-present case now uses a distinctive count so it can no
longer pass on a coincidental leak
2026-06-20 09:10:45 -07:00
Yassin Kortam
9c3ad1b094
feat(caching): add valkey-semantic cache backend and fix semantic cache scope keys (#30675)
Adds a "valkey-semantic" cache type so semantic prompt caching can run
against Valkey clusters (for example AWS ElastiCache for Valkey) using the
valkey-search module.

The existing "redis-semantic" backend cannot drive valkey-search. RedisVL
gates the connection on a RediSearch module version that valkey-search does
not report, and its SemanticCache index declares the prompt as a TEXT field,
which valkey-search does not implement. ValkeySemanticCache therefore talks to
valkey-search directly over redis-py: it builds a vector index from the field
types valkey-search supports (TAG for caller scope, VECTOR for the prompt
embedding) and runs KNN queries for retrieval. Prompt extraction, embedding
generation, and cached-response parsing are reused from RedisSemanticCache
since those are backend agnostic. The redis dependency is imported lazily in
the cache dispatch so importing litellm without redis installed still works.

It also fixes semantic-cache scope keys so similarity matching works across
reworded prompts. get_cache_key() hashed messages / prompt / input into the
litellm_cache_key that every semantic backend filters its KNN search on, so a
paraphrase landed in a different bucket and never matched, even far above the
similarity threshold. For semantic cache types the prompt-bearing params are
now excluded from the scope key and the server-set tenant identity
(user_api_key, team, org) is appended instead, restoring embedding matching
within a tenant while keeping cache entries scoped to the authenticated
key / team / org. The three semantic backends share this key, so the same
change fixes redis-semantic and qdrant-semantic.

Connections resolve from VALKEY_HOST / VALKEY_PORT / VALKEY_PASSWORD, falling
back to REDIS_* for drop-in compatibility, and passwordless clusters (IAM or
no-auth) are supported.

Resolves #29121
Fixes #29086
2026-06-19 17:09:17 -07:00
yuneng-jiang
ea17236a1e
fix(ui): warn that team models are deleted in the delete-team modal (#29990)
The delete-team confirmation modal warned that a team's keys would be
deleted but said nothing about models. #29977 made team deletion also
delete the team's BYOK models, so the modal copy was understating what
gets removed.

The warning banner now mentions models alongside keys, and the
always-shown confirmation message does too so a team that has models but
no keys (the banner only renders when keys exist) still gets warned.
2026-06-19 16:30:46 -07:00
yuneng-jiang
60dc8420ed
fix(ui): repoint dead usage guide link to cost tracking docs (#30859)
The "View Usage Guide" button on the legacy Usage page (shown when
DISABLE_EXPENSIVE_DB_QUERIES is set, i.e. SpendLogs has 1M+ rows) linked
to docs/proxy/spending_monitoring, which was removed from the docs and
now returns 404. Point it at docs/proxy/cost_tracking, which is live.

Fixes LIT-2724
2026-06-19 15:31:48 -07:00
Yassin Kortam
4847fa5dd5
fix(proxy): record partial spend on the failure row for interrupted streams (#30788)
A streaming request that breaks mid-flight, for example on a mid-stream read
timeout, still bills the provider for the chunks already delivered, yet the proxy
recorded that interrupted request as a zero-spend failure. An earlier revision
logged the recovered partial usage through the success path, which mislabeled a
failed request as a success and produced a misleading spend row

This recovers the partial usage where the failure is actually logged. The
streaming handler assembles the usage from the chunks seen so far and stashes it,
with its cost, on the logging object before firing the failure handlers. The
proxy failure hook lifts that usage and cost onto request_data before the
non-serialisable logging object is popped, and the spend-log writer records the
real partial spend on the failure row instead of a hardcoded zero;
get_logging_payload honors the recovered usage for the token columns and
_failure_handler_helper_fn preserves the recovered cost so the non-DB failure
loggers stay consistent

A request that recovers via a successful fallback is unaffected: the failure hook
only fires when the whole request fails, so the fallback's combined-usage success
row stays the single source of truth and there is no double counting

Resolves LIT-3825

Co-authored-by: veria-ai[bot] <224490171+veria-ai[bot]@users.noreply.github.com>
2026-06-19 12:03:15 -07:00
Yassin Kortam
bd74c62ff1
fix(passthrough): recover output tokens for interrupted anthropic streams (#30787) 2026-06-19 12:03:02 -07:00
yucheng-berri
1f9323792c
fix(otel): one v2 logger owns the global provider; scope tenant OTLP creds per exporter (#30590)
* fix(otel): one v2 logger owns the global provider; scope tenant creds per exporter

The proxy published the OTel global TracerProvider before callbacks were
initialized, so no preset logger existed yet and a second generic logger was
built that won the global provider. Server spans then exported through a
different provider than the preset's gen-ai spans, orphaning the LLM span on
the preset backend. Publish after callback init and reuse the already-built
logger instead.

Separately, per-request tenant OTLP credentials were stamped onto every OTLP
exporter, leaking one backend's key onto a co-configured backend. Tag each
exporter with the preset that contributed it and apply dynamic credentials
only to the matching owner.

* fix(otel): satisfy Any-discipline on changed lines

Type the logger-selection parameter as Sequence[object] (isinstance narrows
it), cast the list[Any] global at the single call site, and pass model_copy a
typed dict[str, str] update so no changed line carries an Any value.

* fix(otel): annotate the untyped-global boundary with any-ok

select_global_otel_v2_logger consumes litellm._in_memory_loggers, a shared
List[Any] global this change does not own. A cast doesn't satisfy the
Any-discipline checker (it inspects the inner expression), and re-annotating the
global is out of scope, so mark the single boundary line any-ok.

* test(otel): cover the startup global-provider publish via injectable helper

The publish step lived inline in proxy_startup_event (a FastAPI lifespan unit
tests do not execute), so its lines were uncovered though the selection logic
was tested. Extract publish_global_otel_v2_provider, which selects the single v2
logger and publishes its provider through an injected setter, and unit-test that
the published provider is the selected logger's. proxy_server delegates to it.

* refactor(otel): select global provider from the registered owner, not a list scan

The startup publish picked the global TracerProvider by scanning
_in_memory_loggers for the first OpenTelemetryV2, re-deriving an answer the
factory already settled: the first logger built registers itself as
proxy_server.open_telemetry_logger, and every other v2 path (guardrail, identity
seeding, phase spans) routes through that owner via _registered_v2_logger. Pass
that owner into select_global_otel_v2_logger so the global provider reuses the
same logger instead of an independent, order-dependent guess; the list scan
remains the SDK-path fallback. The owner is injected at the proxy call site to
keep the helper free of hidden global reads.

* refactor(otel): type ExporterSpec.owner as an ExporterOwner enum

The owner field carried free-form strings that had to match preset callback
names. Introduce a str-based ExporterOwner enum (values equal to the callback
names, so per-request credential routing's owner==callback_name comparison still
holds) and have each preset tag its exporter with the enum member.

* refactor(otel): rename ExporterOwner.ARIZE to ARIZE_AX

Distinguish the hosted Arize AX backend from Arize Phoenix at the member level
while keeping the value 'arize' (the public callback name routing compares
against). Add a comment noting AX and Phoenix are separate backends.
2026-06-19 11:15:29 -07:00
Mateo Wang
1bd603d1ac
chore(typing): add boto3/botocore stubs so basedpyright resolves the AWS SDK (#30815)
Some checks are pending
GitHub Actions Security Analysis / zizmor (push) Waiting to run
2026-06-19 08:24:49 -07:00