* test(tracing): pin HTTP request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): generate existing HTTP request models from Rust schemas
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): bind trace query params to generated request models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(rust): raise the native wheel size gate to 48 MB
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(tracing): read the trace list clock without a thread-pool dispatch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(traces): share the trace page-size bounds between schema and reader
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): type trace request queries against the generated OpenAPI schema
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs: encode the trace contract boundary in AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): format trace request aliases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): track ClickHouse migrations in a checksummed ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(clickhouse): harden migration execution and reuse retention SQL
* refactor(rust): track ClickHouse migrations in a checksummed ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): keep ClickHouse retention TTLs in the current policy list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): own trace reads behind a cached TraceStore port
Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.
Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.
Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): seed large profiles with server-side copies and long sessions
Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chores
* style(lens): float the investigation setup badge on the tab edge
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(lens): restructure the trace drawer and polish its layout
Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.
Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* typesafety
* wip
* fmt
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat: seed Lens dev with configurable load profiles
* feat: run Lens UI live through the dev launcher
* fix: verify Lens UI startup before seeding
* fix(ui): render run timestamps on one compact line
The agent runs table printed the long locale form with timezone, which
wrapped to two lines per row. Use a fixed-width 24h form with
milliseconds that matches the timeline axis, keep the long form in the
hover title, and show the timezone once in the column header
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): use the root query client for the Lens demo
Drop the demo's nested QueryClient. Cache keys are already partitioned by
scope, and the root client now skips retries on 4xx ApiErrors, which covers
the demo's not-in-demo and read-only rejections and live 4xx alike.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: drop unused synthetic_spend reference from query help
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(traces): drop the deeplite fixtures and map every SDK to a logo
The deeplite captures predate the example repository and carried
synthetic spend rows, which leaked a fixture-only column into the
query help SQL and pinned tests to its shape. Replace them with the
SDK captures in the ClickHouse round trip and query API tests, and
derive the seed tenant lookup from the capture metadata
The runs table only knew the two Anthropic framework slugs. Register
the slugs the normalizer emits for LangChain, LangGraph, Deep Agents,
CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands,
Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic
agent glyph when a run has no known framework
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): inject the Lens sample through data sources instead of demo checks
Components no longer ask whether they are in the demo. The traces source
carries live and handoff, navigation state comes from a URL or memory
route, and the preview action comes from context instead of onDemo props
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): infinite scroll for the agent runs list
Replace the Load more button with a sentinel that fetches the next
cursor page as the list nears its end. Placeholder rows hold the tail
while more runs exist, and a failed page stops auto-loading until Retry.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): keep the whole Lens view in the URL
Lens navigation now lives entirely in query params through nuqs: the
sample session (demo=true, with a Demo data switch in the header), the
open run (trace, trace_ref), the selected step, view and detail section
(span, view, span_tab) and the list filters and range (q, agent, status,
hours). Any Lens view is a shareable link and the back button walks runs
RunView takes its selection injected: the drawer feeds it URL state and
the investigations evidence sheet keeps a local one, so a finding's
original run never writes step ids into the URL. Leaving the sample
session clears every Lens key except the tab so sample ids never point
at live data
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* style(traces): cargo fmt captures tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(traces): render curated SQL examples from shared files
* fix(traces): filter trace-summary example after grouping so boundary-spanning traces keep full totals
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(traces): order failed-spans example by timestamp so newest failures survive the limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* wip
* wip
* wip
* chore(trace): checkpoint ongoing Rust migration
* refactor(trace): group Python bridge under trace package
* refactor(traces): read span conventions through a Convention trait
Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.
The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.
* feat(trace): export Rust-owned wire schemas and enforce contract bounds
* fix(trace): bound quoted counts in ClickHouse wire schemas
* feat(trace): generate Python wire contracts with datamodel-code-generator
* test(trace): validate migrated callers and generated contracts at the native boundary
* refactor(traces): rename normalization convention to format
* fix(traces): reconcile spend evidence and preserve unknown costs
* feat(traces): normalize additional telemetry formats
* test(traces): cover captured normalization fixtures
* refactor(traces): isolate SDK normalization rules
* feat(tracing): seed all trace exports for local dashboard
* fix(clickhouse): preserve custom LiteLLM request metadata
* docs(traces): define normalization module boundaries
* docs(traces): define resolution and OTLP boundaries
* fix(ui): normalize nullable trace message names
* refactor(traces): split resolver modules and cover resolution behavior
* test(traces): replace normalization snapshots with behavior assertions
* fix(ui): align dashboard API contracts with generated types
* refactor(traces): type normalization and storage boundaries
* fix(traces): seed captured SDK spend and preserve provider identities
* wip
* test(traces): verify guide discovery and content ordering
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(bedrock): send grok chat completions through runtime openai path
Unspecified bedrock grok was rewritten to Converse. Chat completions now hit bedrock-runtime /openai/v1/chat/completions, and converse/ still uses Converse
* feat(bedrock): serve gpt-oss and gpt-5.6 chat completions on runtime's native openai path
* fix(bedrock): route gpt-oss response_format to Converse and decide the route once from the raw request
* fix(bedrock): serve region-path and GovCloud gpt-oss ids on native Chat Completions
The cost-map parity tests require every regional variant of a flagged id to carry the same supports_ flags, so the six us-gov gpt-oss entries now carry the native-route flags too. A region path in the model name (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model: the route is looked up on the id after the path, the path's region picks the endpoint and the SigV4 scope, an explicit aws_region_name still wins, and the body carries the bare id AWS expects
* fix(bedrock): keep params AWS refuses natively off the chat completions route
Drop the params each family 400s or 503s on runtime Chat Completions from the native config's supported list (GPT-5.6 penalties, stop, and logprobs, Grok penalties, gpt-oss logit_bias) so drop_params drops them as Converse did, gate legacy functions on GPT-5.6 the same way as tools, and send an Anthropic-style thinking block to Converse, the only route that forwards it
* fix(bedrock): keep schema-less json_object on Converse for the chat completions models
* fix(bedrock): keep every json_object response_format on Converse for the chat completions models
* fix(rust): declare the bedrock runtime chat completions flags on ModelInfo
* fix(bedrock): opt into the native chat completions route through supported_endpoints
* docs(cost-map): describe the bedrock native chat completions capability flags
* revert: docs(cost-map): describe the bedrock native chat completions capability flags
This reverts commit 4101c0ceb2.
cost-map-guard runs main's schema generator under pull_request_target and compares
its output to the PR's committed schema, so a PR that changes the generator's output
cannot pass that required check until the generator change lands on main first. The
descriptions move to a follow-up that lands the generator change ahead of the schema
* test(bedrock): move the native chat completions tests under tests/unit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): drop reasoning_effort none for grok on the native chat completions route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): keep converse extension params on the converse route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): inline http image urls and keep stop on converse for native chat completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(bedrock): share the sync remote media inliner
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(image-handling): infer the image mime type when the server sends a generic content type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): stop sending aws_bedrock_project_id as OpenAI-Project on the runtime chat completions route
* feat(bedrock): make native chat completions an opt-in bedrock/chat_completions/ route
Bare Bedrock OpenAI and Grok model ids stay on Converse as on main. The
bedrock/chat_completions/<model> prefix opts a deployment into bedrock-runtime's
/openai/v1/chat/completions, and a request carrying a Converse-only param
still falls back to Converse. The cost map no longer decides the route.
* fix(bedrock): keep chat_completions/-prefixed deployments on the native Responses surface
* fix(bedrock): keep provider response headers on the runtime chat completions route
* feat(bedrock): serve gpt-5.6 and newer on runtime chat completions by default
Unprefixed bedrock/<gpt-5.6+> models whose cost-map row lists /v1/chat/completions
now route to the native OpenAI-compatible endpoint; converse/ pins Converse and
chat_completions/ still opts gpt-oss and Grok in. Guardrails, application inference
profile ARNs, and tools with reasoning keep falling back to Converse per request.
Hoist the remote-media url comprehension into a single-clause helper.
* fix(bedrock): refuse temperature and top_p natively on GPT 5.6 and newer like Converse does
AWS answers temperature and top_p with a 400 on the native Chat Completions endpoint for the GPT 5.6+ models, the same models whose Converse route already dropped both under drop_params via supports_sampling_params: false. The native config now honors that price-map flag, the gpt-6 and gpt-6.1 rows carry it, and the gpt-6 family joins gpt-5 in refusing frequency_penalty, presence_penalty, logprobs, and top_logprobs before the request reaches AWS.
* fix(bedrock): refuse GPT sampling and logprob params natively only while reasoning is on
On bedrock-runtime's native chat completions endpoint, GPT-5.x and GPT-6.x
accept temperature, top_p, frequency_penalty, presence_penalty, logprobs,
and top_logprobs once reasoning_effort is "none", and refuse them with any
other effort or when the effort is unset. The previous commit refused the
sampling params unconditionally from the cost map's supports_sampling_params
flag, which lost the reasoning-off case and never covered the penalties or
logprobs. The refusal now keys on the model being a GPT id and reasoning
being active, raises a 400 UnsupportedParamsError naming the params unless
drop_params drops them, and lets everything through under "none". Grok and
gpt-oss keep their unconditional family refusals.
* refactor(bedrock): keep the Converse route-prefix strip inside the bedrock llms module
* fix(bedrock): forward a non-string reasoning_effort on the native route instead of crashing
A list or dict reasoning_effort hit a frozenset membership test in
without_refused_reasoning_effort and raised TypeError, which the proxy
surfaced as a 500 APIConnectionError with no upstream call. The value is
now left alone unless it is a string Bedrock's native endpoint refuses,
so AWS answers the malformed value with its own 400 like it does for an int
* fix(bedrock): route overlong GPT version digits to Converse and send native chat completions to the runtime endpoint
A model id with more than 4300 version digits raised ValueError in the route check; the digits are now bounded so such ids fall back to Converse. The native chat completions URL now follows Converse's precedence: aws_bedrock_runtime_endpoint (or AWS_BEDROCK_RUNTIME_ENDPOINT) wins over api_base, so a deployment that sets both keeps sending to the same host
* fix(bedrock): route model_id overrides to Converse and never send an empty bearer natively
A deployment whose litellm_params carry model_id (an application inference profile or provisioned throughput ARN) went to the native Chat Completions route with the base model in the URL and model_id left in the body. It now takes Converse like the bedrock/arn:... model form, which encodes the override into the request URL
A blank api_key on a SigV4 deployment became an Authorization header reading Bearer with nothing after it on the native route, since the OpenAI-like header builder writes any non-None key and the signer keeps a non-AWS4 Authorization header. validate_environment now resolves the key through bedrock_bearer_token, so a blank key is signed with SigV4 the way Converse signs it
* test(bedrock): audit the native GPT chat completions route on the integration rig
Adds the /audit cells for the runtime chat completions route: the scripted Bedrock runtime peer, the happy and fallback wire tests, the sad-path and regex worst-case tests, the chaos burst tests, the Messages adapter tests, and the Responses native-route tests. Tests only, no product diff.
* test(bedrock): harden the runtime chat completions audit cells
The chaos peer's shared counter and process now come from the same spawn context, since a fork-context Value handed to a spawn-context process raises on Linux. The peer-kill test waits for the first six answers to reach the client before killing the peer instead of counting accepted requests. The Responses wire tests look the spend row up under both the ciphertext id the caller received and the issued id behind it, matching the chaos file's rule for the pre-encryption row
* fix(bedrock): refuse or drop a non-string reasoning_effort before the native chat completions call
A reasoning_effort sent as an int, a list, or an object on a GPT 5.6+ deployment the native
route serves now answers 400 from litellm before any wire request, naming the type and the
drop_params way out, and is dropped under drop_params so AWS applies its default effort, the
way Converse dropped it on main. The tip since a0cef91f0b forwarded it for AWS to refuse
* test(bedrock): pin router retries off and give the chaos bursts config deployments on an owned proxy
* test(bedrock): wait for the replacement worker before tearing down the sigkill chaos proxy
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo <mateo@berri.ai>
* refactor(proxy): extract shared spend log read policy
* test(proxy): use named bindings for spend scope regression
* test(proxy): reuse existing spend log query harness
* test(proxy): cover spend log permission lookup adoption
* chore(proxy): relocate existing spend query baseline
* refactor(proxy): make scope query returns explicit
* refactor(proxy): inject deferred log permission lookup
* test(proxy): cover teamless management compatibility lookup
* refactor(proxy): compose user and team log grants
* refactor(proxy): share generic authorization composition
* refactor(proxy): compose trace read permissions
* refactor(proxy): centralize spend and trace authorization
* refactor(proxy): strengthen spend and trace scope types
* refactor(proxy): flatten log read scope into owned logs
Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.
A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): run spend scope tests through one SQLite emulator
Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.
load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(proxy): unify log and trace ownership permissions
* test(tracing): align fixtures with ownership read scopes
* refactor(tracing): align query scopes with row ownership
* refactor(spend): make ownership SQL predicates explicit
* test(spend): validate ownership SQL against PostgreSQL
* docs(traces): drop key-row visibility from query help guide
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend): reach the empty-memberships branch in team lookup test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate dashboard API types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(traces): type the ClickHouse query help response
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(traces): include agent names and frameworks in named contract round trips
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(traces): cover native query help validation in the storage adapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* wip
* wip
* test(traces): separate root status from diagnostic error counts
* test(traces): cover normalization precedence and fallbacks
* chore(cache): remove stray comments from trace PR
* test(traces): name lens test for shared query path
* fix(traces): place query implementation before test module
* test(traces): use unified read scope in migration tests
* ci(rust): allow feature checks to finish
* ci(mcp): allow dependency resolution to finish
* fix(traces): preserve key visibility and safe spend attribution
* feat(traces): add Framework column to otel_traces
* feat(traces): pass span events to normalizers and add framework field
* feat(traces): add Claude Code and Agent SDK span normalizer
* feat(traces): decode events before normalizing and apply tool span names
* feat(traces): list distinct frameworks per trace
* feat(traces): return span framework in trace spans query
* test(traces): add scrubbed Claude Agent SDK OTLP fixtures
* test(traces): cover Claude Agent SDK normalization from real exports
* test(traces): assert trace list frameworks stay scoped per trace
* feat(tracing): validate framework in native normalized spans
* feat(tracing): add framework to Span and frameworks to TraceSummary
* feat(tracing): store normalized framework on span rows
* feat(tracing): surface span framework and trace frameworks
* test(tracing): cover framework aggregation in trace summaries
* test(tracing): decode Claude Agent SDK rows with framework and tool args
* chore(ui): regenerate API types for trace frameworks
* feat(ui): add trace framework registry for Claude Agent SDK and Claude Code
* feat(ui): show SDK logo and label in the runs list Agent column
* feat(ui): show SDK logo and label in the run header
* test(ui): cover SDK label and logo in the runs list
* test(ui): cover SDK label and logo in the run header
* feat(tracing): show the agent's final answer as claude agent span output
* feat(tracing): name claude code agents after their otel service
* test(tracing): cover claude code agent naming from the service
* fix(tracing): mark the span row framework field read-only
* test(tracing): scrub host os details from the claude sdk fixture
* test(tracing): scrub host os details from the detailed claude sdk fixture
* fix(ui): hide the decorative sdk logo from screen readers
* feat(ui): show the agent name with the sdk logo in the runs list
* feat(ui): show the agent name with the sdk logo in the run header
* test(ui): cover agent names beside the sdk logo in the runs list
* test(ui): cover the agent name in the run header
* fix(lens): use recorded agent identities across framework traces
* fix(lens): tighten agent identity and bound trace lookups
* style(tracing): wrap framework agent identity test case
* fix(tracing): use ClickHouse URL for reads by default
* fix(tracing): unify ClickHouse storage configuration
* fix(tracing): update dashboard setup copy for one URL
* test(tracing): make tests/unit/tracing a package
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(config): drop legacy string tracing store variant
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tracing): own ClickHouse defaults in constants and reject unset env references
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tracing): use raw regex patterns in config tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tracing): read ClickHouse env defaults when tracing config resolves
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): split audit log query guard to fit condition-chain budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(tracing): add SQL queries and schema-aware query help
* test(tracing): verify help requests and sync API types
* refactor(tracing): render query help with Askama
* refactor(tracing): use jinja extension for query guide
* fix(tracing): preserve query help when discovery fails
* feat(tracing): enforce team SQL scope with managed ClickHouse readers
* test(tracing): verify reads with one ClickHouse URL
* fix(tracing): revoke rotated trace reader credentials
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): streamline query help catalog assembly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(tracing): run query help discovery sequentially
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tracing): update reader setup request expectations
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): embed migration folders with a shared migrate! macro
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): enable syn proc-macro feature for litellm-migrate-macros
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): reject signed versions and symlinks in migrate!
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.
* feat(lens): add trace analysis and standalone worker
* fix(lens): clarify review limits and finalize main integration
* fix(lens): simplify worker setup and show the next check
* fix(lens): simplify analyzer setup and resolve integration failures
* fix(lens): preserve durations and evidence from later trace reads
* fix(lens): trust server context for internal analysis exclusion
* fix(lens): pin reviewed analyzer image and verify request inclusion
* test(lens): select time units before entering custom duration
* test(lens): allow the standalone analyzer lifetime HTTP client
* test(lens): run analyzer tests in active proxy coverage shard
* feat(ui): adopt the new LiteLLM logo and monogram
Swap the bundled admin UI logos for the new brand assets: the primary
logo in blue for light mode and white for dark mode, and the monogram for
the collapsed sidebar, favicons, and the built-in guardrail cards.
/get_image gains a variant=monogram query parameter so the collapsed
sidebar can request the monogram while admin-configured UI_LOGO_PATH /
UI_LOGO_PATH_DARK logos still take precedence. The bundled light logo
moves from JPEG to a transparent PNG.
* test(ui): query collapsed sidebar logos by role to stay within the lint budget
* fix: point remaining logo consumers at the new bundled assets
The Rust gateway UI served /get_image from the removed litellm_logo.jpg,
the non-root get_image tests pinned logo.jpg, and a cookbook script read
litellm/proxy/logo.jpg. Point them at the monogram and logo.png.
* fix(mcp): serve the BYOK OAuth page logo from /get_image
The page pointed at /ui/assets/logos/litellm_logo.jpg, which the rebrand
removes from the dashboard sources, so the next UI build would drop it.
/get_image?variant=monogram is always served by the proxy and follows any
admin-configured logo.
* fix(ui): invert the LiteLLM monogram on dark guardrail cards
The blue monogram has a transparent train cut-out, so on a dark card it
read as a muddy blue block. Inverting it yields the brand's white mark,
which the logo guidelines prescribe for dark backgrounds.
* fix(gateway-ui): serve theme and variant aware logos from the dashboard export
The Rust gateway served one monogram for every /get_image request, and the
committed export lacked it, so /get_image returned 404 until the next UI
release build. Pick the full or monogram logo in light or dark from the
query, ship those assets in the dashboard's public dir and the committed
export, and drop two comments that restated asserted paths.
* fix(traces): correct ClickHouse rollup partitioning, dedupe keys, and retention changes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(traces): pin spend dedupe timestamps within one second
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* wip
* feat(traces): establish shared Rust storage foundation
* fix(traces): escape ClickHouse text parameters
* test(traces): exercise response cap with bounded strings
* fix(traces): remove unnecessary lint expectation
* fix(traces): encode ClickHouse timestamp units in Rust
* test(traces): mark exception match as a regex
* refactor(traces): execute schema setup in Rust
* refactor(traces): use shared logging execution wrapper
* docs(traces): replace foundation README with boundary rules
* fix(traces): use current bridge execution facade
* fix(traces): account for protocol cast in lint budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost_calculator): add cost_per_second for chat per-second pricing
Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins
Move Bedrock commitment rows to cost_per_second so they bill once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop legacy per-second fields from chat paths
Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost_calculator): recognize output-only per-second rates
Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): extract litellm-host-native as the shared Rust host driver
Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): interrupt the machine when the in-process stream consumer fails
Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): separate the machine contract from coroutine execution
* auth update
* refactor(rust): use standard flow control for host requests
* style(rust): keep host driver imports formatted
* chores
* mostly relocation
* refactor(rust): separate interceptors from queued observers
* refactor(rust): centralize legacy callback mappings and lifecycle
* docs: define Python host boundaries and migration plan
* refactor: enforce Python host and bridge boundaries
* refactor(rust): separate operations from callback composition
* refactor(rust): compose SDK policy through call hooks
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>