Commit graph

443 commits

Author SHA1 Message Date
yujonglee
1bc85fad40
refactor(python-bridge): ship a signature base and read the resolved call (#45450)
* refactor(python-bridge): ship a signature base and read the resolved call

NativeCall carries base (positionals by name plus signature defaults) instead
of the fully bound dict. resolved lays kwargs over base, which is what bound
held, so every pre-hook read and every route host keeps seeing the same values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(bedrock): build the transcription NativeCall with an empty base

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(dispatch): read the resolved call instead of bound

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(host-python): rename effective to effective_py_args and note the shallow copy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 17:20:29 -07:00
yujonglee
abee1c14e7
refactor(python-bridge): derive NativeCall extraction (#45449)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 17:20:28 -07:00
devin-ai-integration[bot]
3822947b0d
fix(rust): add the inline-tools-2026-09-15 beta to AnthropicBeta (#45439)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 15:25:17 -07:00
yujonglee
0a98c7be10
refactor(rust): derive strum VariantArray and string conversions (#45434)
* refactor(rust): derive strum VariantArray and string conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): use strum conversions directly with explicit spellings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): use rstest values for key management systems

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 21:24:49 +00:00
yujonglee
0721cffab2
refactor(python-bridge): take NativeCall directly and fold routes into per-route folders (#45413)
* refactor(python-bridge): take NativeCall directly and fold routes into per-route folders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(python-bridge): use rstest for updated tests and keep embedding's stub parameter name

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 13:48:28 -07:00
yujonglee
9505af0639
feat(rust): add the Anthropic beta header policy (#45419)
* feat(rust): add the Anthropic beta header policy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): resolve known betas held as Other in AnthropicBeta::on

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): use named rstest cases for Other beta resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): make the rstest named-case rule explicit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): move BetaPolicy tests to the public API test crate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 20:22:58 +00:00
nate-berri
e843cb7aca
fix(rust): path-qualify attribute aliases so rust-analyzer resolves them (#45416)
rust-analyzer cannot resolve the bare alias names emitted by
macro_rules_attribute::apply, so every aliased type was invisible to it
(no go-to-definition or find-references). Referencing the aliases through
crate:: resolves them via the re-export that attribute_alias! generates.

Co-authored-by: Nate Armstrong <narmstrong@Nates-MBP.localdomain>
2026-10-08 11:51:58 -07:00
yujonglee
35b0992533
feat(rust): add standalone typed LLM wire contracts (#45188)
* refactor(rust): add typed LLM wire contracts to llms-types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): drop blank lines left in new wire types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): complete additive wire type contracts

* feat(rust): complete additive nested wire contracts

* feat(rust): give Messages content blocks concrete per-block contracts

Replace the shared all-optional ContentBlockPayload with one struct per
documented block and nested tagged unions for server tool results, so
required fields from the official Messages reference are enforced.

Make CustomTool a plain struct with required name and input_schema, add
MessagesToolParam plus the toolset and undated tool-search tags, model
MiniMax media sources as a tagged union, require documented fields on
chat audio, image URLs, and logprobs, and rename the new block enum to
MessagesContentPart so it no longer shadows the streaming struct.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(rust): tighten new wire contracts against the official references

Give each Messages builtin tool a concrete contract with its documented
required fields, fixed names, and typed allowed callers in place of the
shared all-optional ToolDefinition bag. Make MCP tool result content and
fallback triggers optional as the request params document, and narrow
tool search references, web fetch content, and bash results to their
documented block types.

Keep unmodeled Responses output items and new usage iteration and stop
detail types as Recognized values instead of rejecting the whole payload,
accept find_in_page and optional open_page URLs, preserve JSON Schema
property order, and move prompt_cache_breakpoint to Chat Completions
content parts where OpenAI documents it.

Drop OCR table and key/value types that only matched Azure Document
Intelligence and would reject other providers' normalized passthrough,
along with unsourced bounding box and applied edit fields.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(rust): type every documented Responses output item and close wire gaps

Add the 18 remaining documented Responses output item variants with
concrete nested actions, outputs, and safety checks, and add the
documented namespace, async, caller, and image generation fields to the
existing items.

Restrict MCP tool result content to a string or text blocks, model the
Messages diagnostics cache-miss union and its request param with null
distinct from absent, and accept draft-07 tuple items and prefixItems in
JSON Schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(rust): keep Messages contract assertions in test bodies

Compare built-in tool decodes against fully built expected values and
split the required-field rejections per type, so no case carries its
assertions in a callback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(rust): cite the Mistral OCR reference beside the OCR wire types

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 16:27:12 +00:00
tin-berri
057034d0d3
fix(lens): refresh runs until gateway costs are complete (#45105) 2026-10-08 00:24:41 -07:00
ishaan-berri
03ba79d261
feat(lens): show end-user feedback on traces, stored in ClickHouse (#45171)
* feat(lens): add lens_feedback ClickHouse table for human trace scores

ReplacingMergeTree(UpdatedAt, IsDeleted) keyed like agent_traces_by_key so
feedback joins traces in sort order, with a bloom filter on TraceId, a score
CHECK, and the same retention as traces

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add scoped ClickHouse reads for trace feedback

feedback_target resolves a visible trace's team, key and trace_ref before a
write; feedback lists the latest live entry per author; feedback_summary
returns count, average and lowest score for a batch of traces

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens): pin the Claude session to trace id hash shared with feedback

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): expose feedback queries to the Python trace bridge

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add typed feedback models and ClickHouse feedback store

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add /lens/feedback API for 0-10 human scores on traces

PUT saves or replaces the caller's score and comment, GET lists every
entry, DELETE removes the caller's, POST /summary feeds the trace list.
Callers can target a trace_id or a Claude Code session_id

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add typed feedback calls to the traces API

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): flag rated traces and filter the list by feedback

A Feedback column shows each run's average score and rating count, marked
red when any score is 4 or below, and a filter narrows to rated or
low-score runs through the URL

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): show and edit human feedback inside a trace

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): let app keys post end-user feedback on their own traces

Any key can now PUT/DELETE feedback on traces its team or key sent, naming
the end user in an optional user field. Reading feedback stays with Lens
admins

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): show end-user feedback first in a run and flag low-scored rows

Replace the edit popover and feedback filter with a read-only view: the
list shows each run's score in a fixed-width badge and tints the row red
when a user scored it 4 or below, and opening a run shows what users said
with their score right under the header

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-07 20:14:30 -07:00
ishaan-berri
b85756104f
feat(lens): show agent, user and slack thread first in the run header (#45261)
* feat(lens): add who started the run to RunSource

* feat(lens): carry agent.source.user on trace span rows

* feat(lens): resolve the run source user from the span row

* feat(lens): read agent.source.user in trace_spans query

* feat(lens): read agent.source.user in trace_page_spans query

* feat(lens): read agent.source.user in trace_span_batch query

* feat(lens): read agent.source.user in trace_list_span_batch query

* feat(lens): decode the source user from clickhouse span rows

* test(lens): add source user to trace cache test rows

* test(lens): add source user to trace cache read fixtures

* test(lens): add source user to trace cache snapshot fixtures

* test(lens): add source user to capture fixture rows

* test(lens): round-trip source user on span row contract

* test(lens): cover run source user resolution

* chore(lens): regenerate python trace types with source user

* chore(lens): regenerate trace json schema with source user

* chore(lens): regenerate trace page json schema with source user

* chore(ui): add source user to api types

* feat(lens): add run user chip and slack thread chip

* feat(lens): put agent, user and thread on the first row of the run header

* test(lens): cover the run header identity row and compact totals

* feat(lens): start demo runs from a slack thread
2026-10-07 20:01:55 -07:00
ishaan-berri
b970e412d9
perf(lens): faster trace opens and list pages at scale (#45228)
* perf(lens): read trace spans in larger batches

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* perf(lens): evaluate the trace list page once per identity lookup

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): queue trace reads instead of rejecting past 8 in flight

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 01:44:19 +00:00
moe-berri
fdb1ce8d78
fix(lens): preserve numeric tags and report OTLP error codes (#45206) 2026-10-08 00:31:48 +00:00
yujonglee
e638540de9
refactor(rust-bridge): share field and response marshaling (#45187)
* refactor(rust-bridge): share field and response marshaling

* test(rust-bridge): isolate response factory test modules per case

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust-bridge): document shared marshaling helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 17:22:23 -07:00
moe-berri
e9cfba2c17
feat(lens): isolate ingestion and investigations in a Rust service (#45148)
* feat(lens): isolate trace storage and investigation in a Rust service

* fix(lens): include Rust sources in the image build context

* feat(lens): wire service setup, scoped delivery receipts and lease attempts

* fix(lens): complete service routing and reject stale investigation results

* fix(lens): retry key propagation and validate isolated Compose setup

* fix(lens): seed through isolated ingestion and preserve upstream queue fixes

* chore: sync schema.prisma copies from root

* fix(lens): bind nullable due timestamps as text for Prisma

* chore(ui): remove stale lint suppressions

* fix(lens): address CI failures and review findings

* refactor(lens): remove retired Python worker and run evaluations in Rust

* fix(lens): reuse control connections and satisfy review checks

* test(lens): install and upgrade both Helm charts on Kubernetes

* test(lens): run connection reuse coverage as an integration test

* fix(ui): upgrade Next.js to 16.3.8 security release

* fix(lens): fence stale attempts and preserve reviewed evidence

* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"

This reverts commit 2f79a51b25.

* fix(lens): stop failed investigations and stream history excerpts

* test(lens): cover model tool and result contracts

* test(lens): fix retired routes and reuse installation build artifacts

* test(lens): use portable grep in Helm installation smoke

* test(lens): wait for migrations before forwarding Helm services

* fix(lens): keep failed evidence reads retryable

* fix(lens): preserve sandbox output during process exit

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-10-07 17:18:58 -07:00
ishaan-berri
bf9bd35469
feat(lens): scope traces to one agent with a header picker (#45202)
* feat(lens): add trace_agents rollup query

* feat(lens): register the trace_agents read query

* feat(lens): add trace_agents params and row types

* feat(lens): dispatch the trace_agents query

* feat(lens): export trace_agents wire schemas

* test(lens): pin the trace_agents query name

* test(lens): cover trace_agents scope, window and failure counts

* feat(lens): cap the agent list size

* feat(lens): declare the trace_agents bridge query

* feat(lens): read trace agents from clickhouse storage

* feat(lens): add trace agent response models

* feat(lens): list agents within the reader's trace scope

* feat(lens): add GET /v1/traces/agents

* chore(lens): regenerate trace models with trace_agents

* chore(lens): regenerate trace types with trace_agents

* chore(lens): regenerate read query name schema

* chore(lens): add trace agent row schema

* chore(lens): add trace agents params schema

* test(lens): cover the trace agents route

* test(lens): cover agent listing scope and timestamps

* chore(ui): regenerate api types with trace agents route

* feat(lens): derive trace agent types from the schema

* feat(lens): fetch the agent list from the traces api

* feat(lens): roll up demo runs into agents

* feat(lens): serve the agent list in demo data

* feat(lens): load agents seen in the last two weeks

* feat(lens): remember the selected agent per browser

* feat(lens): add the agent picker

* feat(lens): wire agent selection into the lens header

* feat(lens): show the agent picker next to the lens title

* refactor(lens): drop the toolbar agent filter in favor of the header picker

* test(lens): remove tests for the toolbar agent filter

* test(lens): cover agent resolution and rollup

* test(lens): cover scoping, switching and remembering the agent

* test(lens): read only the create request in guided setup

* test(lens): assert the toolbar agent filter is gone

* test(lens): read stubbed requests without their abort signal
2026-10-07 23:48:04 +00:00
ishaan-berri
df23f11e9b
feat(lens): link traces to the conversation that started them (#45169)
* feat(lens): carry lens.source attributes on trace span rows

* feat(lens): read lens.source attributes in trace_spans query

* feat(lens): read lens.source attributes in trace_page_spans query

* feat(lens): read lens.source attributes in trace_span_batch query

* feat(lens): read lens.source attributes in trace_list_span_batch query

* feat(lens): decode lens.source columns from clickhouse span rows

* feat(lens): add run source to the trace summary contract

* feat(lens): export RunSource from litellm-traces

* feat(lens): resolve the https run source from the root span

* test(lens): cover run source resolution and https guard

* test(lens): add source fields to capture fixture rows

* test(lens): round-trip source fields on span row contract

* test(lens): add source fields to trace cache test rows

* test(lens): add source fields to trace cache read fixtures

* test(lens): add source fields to trace cache snapshot fixtures

* chore(lens): regenerate python trace types with run source

* chore(lens): regenerate trace json schema with run source

* chore(lens): regenerate trace page json schema with run source

* chore(ui): regenerate api types with trace run source

* feat(lens): add run source link with hover card

* test(lens): cover run source app detection and url guard

* feat(lens): show the run source next to the trace name

* feat(lens): label the run source link, e.g. Slack thread

* test(lens): assert run source link labels

* feat(lens): move the run source link into the trace stats row

* feat(lens): add a typed run source, e.g. slack or teams

* feat(lens): export RunSourceType

* feat(lens): carry the source type on trace span rows

* feat(lens): resolve the run source type, defaulting to custom

* feat(lens): read agent.source attributes in trace_spans query

* feat(lens): read agent.source attributes in trace_page_spans query

* feat(lens): read agent.source attributes in trace_span_batch query

* feat(lens): read agent.source attributes in trace_list_span_batch query

* feat(lens): decode the source type from clickhouse span rows

* test(lens): cover run source type parsing

* test(lens): add source type to capture fixture rows

* test(lens): round-trip source type on span row contract

* test(lens): add source type to trace cache test rows

* test(lens): add source type to trace cache read fixtures

* test(lens): add source type to trace cache snapshot fixtures

* chore(lens): regenerate python trace types with source type

* chore(lens): regenerate trace json schema with source type

* chore(lens): regenerate trace page json schema with source type

* chore(ui): regenerate api types with run source type

* feat(lens): pick the run source logo and label from its type

* test(lens): cover run source labels by type and slack logo

* feat(lens): show the run source as a Source stat with the app name

* test(lens): assert run source app names

* feat(lens): place the Source stat before Duration
2026-10-07 22:32:37 +00:00
yujonglee
2aaa0b5d5c
refactor(rust-bridge): unify native call inputs and Messages settings (#45126)
* refactor(rust): separate Messages settings from capability inputs

* refactor(rust-bridge): unify Messages OCR and Responses call inputs

* refactor(rust-bridge): share NativeCall across inference entrypoints

* fix(rust-bridge): preserve Responses URL aliases and public test inputs
2026-10-07 15:05:43 -07:00
devin-ai-integration[bot]
46d440ae96
fix(model_prices): consolidate claude-haiku-5-5 over-100k pricing and capability flags (#45151)
* feat(types): declare above_100k_tokens price fields on ModelInfoBase

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 186f6c81a0)

* fix(router): mirror above_100k pricing fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 75c9c3ff1e)

* chore(ui): regenerate schema.d.ts for above_100k pricing fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 53163d2184)

* refactor(types): mark above_100k ModelInfoBase fields ReadOnly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit febed3199a)

* fix(model_prices): bill claude-haiku-5-5 long prompts on every provider and in batch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 218c00cf5a)

* fix(cost): pass *_above_Nk_tokens_batches rates through get_model_info

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 37b6f2fba0)

* fix(model_prices): allow disabling thinking and forced tool use on claude-haiku-5-5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 6e91e0ed36)

* test(model_prices): cite the vendor source for claude-haiku-5-5 capability flags

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 1f97bb27b4)

* test(model_prices): type and tidy the claude-haiku-5-5 config tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 3d2526035f)

* feat(bedrock): add claude haiku 5.5 over 100k token tier

Price-Sync: litellm-providers
(cherry picked from commit ee3822c29c)

* feat(model_prices): add openrouter claude-haiku-5.5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): dedupe above_100k pricing keys from text merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): bill vertex claude-haiku-5-5 prompts over 100k tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): add adaptive thinking and cache minimum to openrouter claude-haiku-5.5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-07 14:14:47 -07:00
devin-ai-integration[bot]
a9b9700790
perf(lens): bound single trace reads by the sampled start time (#45088)
* perf(lens): bound single trace reads by the sampled start time

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): allow unused query fixture field in load tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): pass start_time in every lens content and evidence test

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-07 12:50:35 -07:00
devin-ai-integration[bot]
8f85de740f
perf(lens): prune ClickHouse partitions when sampling and sample in one pass (#45087)
* perf(lens): prune ClickHouse partitions when sampling and sample in one pass

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(lens): allow unused query fixture field in load tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): shrink the sample page when a response exceeds the read limit

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(lens): qualify request sample window columns

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-07 12:24:25 -07:00
devin-ai-integration[bot]
77fc3315e5
fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts (#43878)
* fix(caching): keep tool calls and tool results in semantic cache prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): keep semantic tool prompt helpers within lint budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep structured function_call_output text in semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): split Responses text-field collection to stay within complexity budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): tag each tool result with the position of the call it answers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): encode tool result position and output together so tool text cannot forge result tags

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): expect encoded tool result record in qdrant semantic prompt parity case

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): cover tool result arrangements, SDK clients, concurrency and qdrant outage for semantic cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(caching): embed every semantic cache prompt field except volatile ones

Replace the per-shape allowlist in the Python and Rust semantic cache prompt
walkers with one include-by-default walker. Plain text keeps its old
concatenation; any other block or message is embedded as compact JSON with
call ids mapped to ordinals, cache_control dropped, and signatures, encrypted
content and base64 data replaced with a short sha256 digest.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts"

This reverts commit 86c82b949b.

* Revert "test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check"

This reverts commit 39efb5d9da.

* Revert "feat(caching): embed every semantic cache prompt field except volatile ones"

This reverts commit 5aed3ab3de.

* refactor(caching): rename get_str_from_messages_with_tools to get_semantic_cache_prompt_from_messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): split semantic cache prompt extraction by API format

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): drop TypeIs guard and register Responses prompt walker with the recursion check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): pick the Responses text field without a Final inside a loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): walk semantic cache prompts as plain dicts, dumping pydantic items once up front

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): write the semantic cache prompt builders as plain loops

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): drop formatting-only churn from the redis semantic cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(integration): drop the caching group wiring that main already carries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): recurse into tool_result content in the semantic cache prompt helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): check the max_messages cap on the shared exact-cache proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): read list-form function_call_output text in semantic cache prompts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(caching): extract nested Responses input lookup to keep walker under complexity limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "refactor(caching): extract nested Responses input lookup to keep walker under complexity limit"

This reverts commit 0665296bf1.

* style(caching): suppress C901 on the Responses input walker instead of splitting it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 17:32:39 +00:00
ishaan-berri
d8bc2b78e4
fix(lens): show the first user message as the run input (#44958)
* fix(traces): use the first user message for the run input preview

* test(traces): cover first user message as the input preview

* test(traces): check fixture previews against the first user message

* fix(lens): show the whole input preview on one line in the runs table

* test(lens): cover multi-line input previews in the runs table
2026-10-06 16:30:52 -07:00
ishaan-berri
5fb23eabd8
feat(lens): add datasets built from real traces (#44765)
* feat(lens): add dataset case and size limits

* feat(lens): add dataset, case and build models

* feat(lens): build dataset cases from traces, findings and text

* feat(lens): store dataset revisions insert-only

* feat(lens): add dataset routes for build, save, export and eval cases

* feat(lens): mount the dataset router before lens routes

* feat(lens): add LiteLLM_LensDataset table

* feat(lens): add LiteLLM_LensDataset table to proxy schema

* feat(lens): add LiteLLM_LensDataset table to extras schema

* feat(lens): add migration that creates the dataset table

* test(lens): cover dataset case building, dedupe and limits

* test(lens): cover dataset revisions, conflicts and eval cases

* chore(ui): regenerate API types for lens datasets

* feat(lens): add dataset UI types

* feat(lens): add datasets API client

* feat(lens): add dataset query and mutation hooks

* feat(lens): add case selection and expected edit logic

* test(lens): cover case selection and expected edits

* feat(lens): add the add to dataset dialog

* test(lens): cover saving picked cases from the dialog

* feat(lens): add datasets list

* feat(lens): add dataset detail with revisions and export

* test(lens): cover editing, revisions and export in datasets tab

* feat(lens): expose datasets on the lens API

* feat(lens): add in-memory datasets for demo mode

* feat(lens): wire demo datasets into the demo lens API

* feat(lens): add optional lens API hook

* feat(lens): add optional onboarding hook

* feat(lens): add datasets tab and dataset routing

* feat(lens): show the datasets tab

* feat(lens): add to dataset from the trace header

* feat(lens): add a single turn to a dataset from a step

* feat(lens): add finding evidence to a dataset

* docs(lens): add datasets screenshots for the PR

* test(lens): cover dataset revision storage against Postgres

* test(lens): cover dataset trace paging, findings and route errors

* test(lens): cover dataset build fallbacks and no_content skips

* fix(lens): register the dataset table for postgres span names

* feat(lens): add invalid skip reason for unparseable case lines

* fix(lens): keep valid JSONL cases, provenance and size limits; stop reading past the case cap

* test(lens): cover malformed JSONL, re-import provenance, size fields and early cap

* chore(ui): regenerate API types for the invalid skip reason

* feat(lens): show text for the invalid skip reason

* feat(lens): render datasets in the same inspector table as investigations

* test(lens): open a dataset by clicking its table row

* docs(lens): update the datasets list screenshot

* style(lens): format the datasets table

* fix(lens): normalize JSON text span input and output into messages when building cases

* test(lens): cover JSON text span normalization for dataset cases

* Revert "fix(lens): normalize JSON text span input and output into messages when building cases"

This reverts commit 29c3e34962fa12427dbd516c08f2599db997a905.

* fix(traces): normalize agent assistant summaries into UI messages

* feat(lens): add dataset case view helpers built on the trace parsers

* test(lens): cover dataset case view helpers

* feat(lens): show dataset cases in an inspector table

* feat(lens): open a dataset case in a side panel with trace message cards

* feat(lens): rebuild the dataset page header and layout

* feat(lens): keep the open dataset case in the URL

* test(lens): drive dataset edits through the case table and panel

* docs(lens): update dataset view screenshots

* Revert "test(lens): cover JSON text span normalization for dataset cases"

This reverts commit 30229ae8ae90265c09d755c43c1f1eeaa66050c7.

* fix(lens): import StateMessage from shared in dataset detail

* fix(lens): import StateMessage from shared in datasets list

* fix(lens): give the datasets migration a unique timestamp after review checkpoints
2026-10-06 23:04:43 +00:00
moe-berri
d181bc7b80
fix(lens): refresh open traces without claiming session completion (#44900)
* fix(lens): refresh open traces without claiming session completion

* fix(lens): cancel paused refreshes and distinguish refresh failures

* style(lens): format live refresh regressions

* fix(lens): preserve manual reads when pausing live updates

* fix(tracing): preserve optional provider evidence in fixture replay

* fix(tracing): keep copied provider identities consistent

* fix(lens): refresh resumed native sessions and retain paging

* fix(lens): serialize conversation paging with refresh

* fix(lens): refresh recorded content with trace details

* fix(lens): refresh content using resolved trace references

* fix(lens): preserve content through refresh failures

* fix(lens): serialize conversation paging with content refresh
2026-10-06 15:31:45 -07:00
moe-berri
a3e74a5223
fix(tracing): preserve optional provider evidence in fixture replay (#44930)
* fix(tracing): preserve optional provider evidence in fixture replay

* fix(tracing): keep copied provider identities consistent
2026-10-06 14:00:40 -07:00
moe-berri
5086fb3038
fix(tracing): retain native logs and headless tool results (#44894)
* fix(tracing): retain native logs and headless tool results

* fix(tracing): bound uncorrelated log annotations after decoding

* fix(tracing): canonicalize absent native log context
2026-10-06 13:51:38 -07:00
ishaan-berri
5ed7ec8511
fix(lens): price agent traces by joining gen_ai.response.id to spend logs (#44738)
* fix(lens): price agent traces by joining gen_ai.response.id to spend logs

Trace spend now joins each model call to spend_logs on one key: the
span's response id (gen_ai.response.id or the id the normalizers read
from OpenInference/LangChain output) against spend_logs.response_id or
the upstream id embedded in a managed resp_ id. The litellm.call_id and
traceparent transport join paths and the per-row ownership gate are
removed; the spend SQL still restricts rows to what the reader can see.

A run with some unpriced calls now reports the sum of its priced calls
plus priced_calls, instead of an unknown total.

* test(lens): cover response id spend join and partial trace totals

* chore(lens): regenerate trace types for priced_calls

* feat(lens): show partial run cost as a lower bound with priced call count

* fix(lens): treat litellm.call_id as the same assigned call id for spend joins

The id LiteLLM assigned to a call is either the response id it returned
(gen_ai.response.id -> spend_logs.response_id) or its gateway call id
(litellm.call_id -> spend_logs.litellm_call_id). Both are exact ids the
gateway mints and logs, so the join stays one rule. Transport span
matching and the ownership gate stay removed.

* test(lens): cover litellm.call_id spend joins and restore captured totals

* feat(lens): link each priced model call to its spend log

Spans gain spend_log_request_id, the spend_logs.request_id the call was
priced from, and spend_match, which says whether a model call matched or
why not (no assigned id on the span, no spend log with that id, or an
ambiguous match). A model call span is priced from the same ids as the
run total, so its cost and the total agree.

* test(lens): cover spend log links on model call spans

* chore(lens): regenerate trace types for spend log links

* feat(lens): open the matched spend log from an LLM step

An LLM step's header now shows a Spend log chip with the matched
request id and cost; clicking it opens the request log drawer over the
run, fetched by the exact spend_logs.request_id instead of the span's
own response id. Unpriced steps say why (no assigned id on the span, or
no spend log with it). Tree rows show each model call's cost, and a
partial run cost shows its priced call count inline.

* test(lens): cover the spend log link and unmatched cost reasons

* feat(lens): show the spend log link as a bordered LiteLLM Spend Log button

* feat(lens): add a back link from the spend log drawer to the agent trace

* feat(lens): label the spend log back link Back to Lens trace with the Lens icon

* fix(lens): ignore assigned ids that name no spend log when pricing a call

An id that names no row no longer vetoes the call, so a span carrying
both a response id and a call id still prices from a spend row logged
before litellm_call_id existed. An id naming two or more rows makes the
call ambiguous, and the match reason comes from the same per-id result,
so a single matched row with no cost is reported as matched.

* fix(lens): price a trace only from spend logs in its own team

A reader with several teams could see the same assigned id in another
team's spend log; only rows from the trace's team now price it. The
user and key ownership gate stays removed.

* perf(lens): resolve each model call's spend once per trace

Model call matches are computed once when the trace is resolved and
looked up by span index, instead of scanning the model call list for
every span and walking the graph again for spans, agents and the run
total.

* fix(lens): hide a step's Cost fact only when its spend log link shows the cost

* chore(lens): drop narrative doc comments from the spend join

* fix(lens): price a model call only when its ids agree on one spend log per span

* fix(lens): keep pricing spend logs written before litellm_call_id by their request id

* fix(lens): price every attempt a model call's ids name when they agree
2026-10-06 19:55:26 +00:00
devin-ai-integration[bot]
b6d587d23e
refactor(rust_bridge): remove rule-gated native secret-manager selection (#44906)
* refactor(rust_bridge): remove rule-gated native secret-manager selection

Mirror the cache treatment: SecretManagerRule/SecretManagerContext and the resolve_native_* binding plumbing are gone. The Rust bridge now selects the native backend from the explicitly configured client (capture_secret_manager / _SecretManagerRuntime.from_client), and get_secret_from_manager is the plain Python handler path.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(deps): bump sharp to 0.35.5 for GHSA-wq5f-xc86-pv6w

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 19:14:55 +00:00
devin-ai-integration[bot]
837c6a7481
fix(security): remove the publicly known master key from the repo (#44718)
* fix(security): hash the publicly known master key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: replace weak master key examples and regenerate artifacts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: replace weak key fixtures with generated test keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: generate master keys for proxy startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve lens dev key entropy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore proxy key compatibility in scrub examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: scrub merged SSO fixture and refresh dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ci): stabilize test keys and metadata collection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: drop the rebuilt dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:55:24 -07:00
devin-ai-integration[bot]
878ba39e7c
refactor(rust): extract inference-testing crate (#44873)
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.

Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 16:33:06 +00:00
devin-ai-integration[bot]
de74f81c69
chore(rust): prune inference deps and rewrite layering docs (#44836)
* chore(rust): prune unused inference crate dependencies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): rewrite inference layering docs for the split format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:27 -07:00
devin-ai-integration[bot]
ae35c9d775
refactor(rust): extract inference-ocr crate (#44832)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
6a55e0a8aa
refactor(rust): extract inference-chat crate (#44827)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
d7f5b40ab4
refactor(rust): extract inference-messages crate (#44818)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:26 -07:00
devin-ai-integration[bot]
785cb12f37
refactor(rust): extract inference-responses crate (#44811)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:25 -07:00
devin-ai-integration[bot]
63babf23e6
refactor(rust): extract inference-transcription crate (#44809)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
1a9b533e71
refactor(rust): rename litellm-core to litellm-inference (#44802)
* refactor(rust): rename litellm-core to litellm-inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): expose inference base API for format crates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): move shared inference test helpers behind test-support

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 07:17:24 -07:00
devin-ai-integration[bot]
5f9eff55f8
refactor(cache): remove dead Cache._native_cache runtime path (#44785)
The native response-cache runtime attached through Cache._native_cache is
unreachable since the V2 cache replaced it. Drop the Python branches and
wrapper, the _ResponseCacheRuntime pyclass and its backend/activation/
semantic modules, the python-bridge deps only they used, and the tests and
fixtures dedicated to that path.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:58:19 -07:00
moe-berri
2190003bb8
fix(lens): reconstruct native coding agent conversations (#44711)
* fix(lens): reconstruct native coding agent conversations

* fix(lens): display timestamps used for conversation ordering

* fix(lens): preserve coding trace identity and message provenance

* fix(lens): complete capture checks after trace pagination

* fix(lens): show loaded replies during trace pagination
2026-10-05 17:51:05 -07:00
moe-berri
c29b42a32b
fix(lens): preserve span timestamps in investigation evidence (#44702)
* fix(lens): preserve span timestamps in investigation evidence

* style(lens): format chronology regression test
2026-10-05 16:42:42 -07:00
devin-ai-integration[bot]
b9251dafad
refactor(rust): derive string enum serde through strum and serde_with (#44675)
* docs(rust): document string enum serde conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): derive string enum serde through strum and serde_with

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:34:34 -07:00
devin-ai-integration[bot]
74ad633e53
refactor(rust): rename litellm-framing crate to litellm-framer (#44683)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:32:00 +00:00
devin-ai-integration[bot]
45e7be1abc
feat(tracing)!: return only data from SQL queries (#44609)
* feat(tracing): add generated SQL response contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): return data-only trace SQL responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): format trace response schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 12:11:59 -07:00
devin-ai-integration[bot]
9b6a6a0b71
refactor(tracing): generate existing HTTP request models from Rust schemas (#44591)
* test(tracing): pin HTTP request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): generate existing HTTP request models from Rust schemas

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): bind trace query params to generated request models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 48 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(tracing): read the trace list clock without a thread-pool dispatch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(traces): share the trace page-size bounds between schema and reader

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): type trace request queries against the generated OpenAPI schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: encode the trace contract boundary in AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format trace request aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 17:37:52 +00:00
devin-ai-integration[bot]
fdc0e7dc2b
refactor(rust): track ClickHouse migrations in a checksummed ledger (#44580)
* refactor(rust): track ClickHouse migrations in a checksummed ledger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(clickhouse): harden migration execution and reuse retention SQL

* refactor(rust): track ClickHouse migrations in a checksummed ledger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(rust): keep ClickHouse retention TTLs in the current policy list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 16:04:44 +00:00
devin-ai-integration[bot]
461a58c40a
refactor(lens): storage-independent trace reads, shared keyset pager, typed read failures (#44422)
* feat(lens): own trace reads behind a cached TraceStore port

Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.

Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.

Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(lens): seed large profiles with server-side copies and long sessions

Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chores

* style(lens): float the investigation setup badge on the tab edge

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens): restructure the trace drawer and polish its layout

Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.

Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* typesafety

* wip

* fmt

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 12:39:47 -07:00
devin-ai-integration[bot]
62fb808d4b
feat: improve Lens dev seeding and live UI (#44468)
* feat: seed Lens dev with configurable load profiles

* feat: run Lens UI live through the dev launcher

* fix: verify Lens UI startup before seeding

* fix(ui): render run timestamps on one compact line

The agent runs table printed the long locale form with timezone, which
wrapped to two lines per row. Use a fixed-width 24h form with
milliseconds that matches the timeline axis, keep the long form in the
hover title, and show the timezone once in the column header

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): use the root query client for the Lens demo

Drop the demo's nested QueryClient. Cache keys are already partitioned by
scope, and the root client now skips retries on 4xx ApiErrors, which covers
the demo's not-in-demo and read-only rejections and live 4xx alike.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: drop unused synthetic_spend reference from query help

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(traces): drop the deeplite fixtures and map every SDK to a logo

The deeplite captures predate the example repository and carried
synthetic spend rows, which leaked a fixture-only column into the
query help SQL and pinned tests to its shape. Replace them with the
SDK captures in the ClickHouse round trip and query API tests, and
derive the seed tenant lookup from the capture metadata

The runs table only knew the two Anthropic framework slugs. Register
the slugs the normalizer emits for LangChain, LangGraph, Deep Agents,
CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands,
Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic
agent glyph when a run has no known framework

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): inject the Lens sample through data sources instead of demo checks

Components no longer ask whether they are in the demo. The traces source
carries live and handoff, navigation state comes from a URL or memory
route, and the preview action comes from context instead of onDemo props

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): infinite scroll for the agent runs list

Replace the Load more button with a sentinel that fetches the next
cursor page as the list nears its end. Placeholder rows hold the tail
while more runs exist, and a failed page stops auto-loading until Retry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): keep the whole Lens view in the URL

Lens navigation now lives entirely in query params through nuqs: the
sample session (demo=true, with a Demo data switch in the header), the
open run (trace, trace_ref), the selected step, view and detail section
(span, view, span_tab) and the list filters and range (q, agent, status,
hours). Any Lens view is a shareable link and the back button walks runs

RunView takes its selection injected: the drawer feeds it URL state and
the investigations evidence sheet keeps a local one, so a finding's
original run never writes step ids into the URL. Leaving the sample
session clears every Lens key except the tab so sample ids never point
at live data

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* style(traces): cargo fmt captures tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 01:03:11 +00:00
devin-ai-integration[bot]
21ecd0af55
fix(traces): reject conflicting spend aliases and unrelated HTTP siblings (#44456)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-10-03 23:43:10 +00:00
yujonglee
c7e60f03de
fix(tracing): preserve spend identity and gateway correlation (#44421)
* fix(tracing): preserve spend identity and gateway correlation

* test(tracing): refresh real SDK spend captures

* fix(tracing): resolve complete gateway attempt costs across SDKs

* docs: add trace cost screenshot for PR 44421

* update fixtures

* wip

* docs: remove trace cost screenshot from PR evidence
2026-10-03 16:00:49 -07:00