* refactor(python-bridge): ship a signature base and read the resolved call
NativeCall carries base (positionals by name plus signature defaults) instead
of the fully bound dict. resolved lays kwargs over base, which is what bound
held, so every pre-hook read and every route host keeps seeing the same values.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(bedrock): build the transcription NativeCall with an empty base
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(dispatch): read the resolved call instead of bound
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(host-python): rename effective to effective_py_args and note the shallow copy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): derive strum VariantArray and string conversions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): use strum conversions directly with explicit spellings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): use rstest values for key management systems
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(python-bridge): take NativeCall directly and fold routes into per-route folders
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(python-bridge): use rstest for updated tests and keep embedding's stub parameter name
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(rust): add the Anthropic beta header policy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): resolve known betas held as Other in AnthropicBeta::on
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): use named rstest cases for Other beta resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): make the rstest named-case rule explicit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): move BetaPolicy tests to the public API test crate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
rust-analyzer cannot resolve the bare alias names emitted by
macro_rules_attribute::apply, so every aliased type was invisible to it
(no go-to-definition or find-references). Referencing the aliases through
crate:: resolves them via the re-export that attribute_alias! generates.
Co-authored-by: Nate Armstrong <narmstrong@Nates-MBP.localdomain>
* refactor(rust): add typed LLM wire contracts to llms-types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(rust): drop blank lines left in new wire types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): complete additive wire type contracts
* feat(rust): complete additive nested wire contracts
* feat(rust): give Messages content blocks concrete per-block contracts
Replace the shared all-optional ContentBlockPayload with one struct per
documented block and nested tagged unions for server tool results, so
required fields from the official Messages reference are enforced.
Make CustomTool a plain struct with required name and input_schema, add
MessagesToolParam plus the toolset and undated tool-search tags, model
MiniMax media sources as a tagged union, require documented fields on
chat audio, image URLs, and logprobs, and rename the new block enum to
MessagesContentPart so it no longer shadows the streaming struct.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(rust): tighten new wire contracts against the official references
Give each Messages builtin tool a concrete contract with its documented
required fields, fixed names, and typed allowed callers in place of the
shared all-optional ToolDefinition bag. Make MCP tool result content and
fallback triggers optional as the request params document, and narrow
tool search references, web fetch content, and bash results to their
documented block types.
Keep unmodeled Responses output items and new usage iteration and stop
detail types as Recognized values instead of rejecting the whole payload,
accept find_in_page and optional open_page URLs, preserve JSON Schema
property order, and move prompt_cache_breakpoint to Chat Completions
content parts where OpenAI documents it.
Drop OCR table and key/value types that only matched Azure Document
Intelligence and would reject other providers' normalized passthrough,
along with unsourced bounding box and applied edit fields.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(rust): type every documented Responses output item and close wire gaps
Add the 18 remaining documented Responses output item variants with
concrete nested actions, outputs, and safety checks, and add the
documented namespace, async, caller, and image generation fields to the
existing items.
Restrict MCP tool result content to a string or text blocks, model the
Messages diagnostics cache-miss union and its request param with null
distinct from absent, and accept draft-07 tuple items and prefixItems in
JSON Schema.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(rust): keep Messages contract assertions in test bodies
Compare built-in tool decodes against fully built expected values and
split the required-field rejections per type, so no case carries its
assertions in a callback.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(rust): cite the Mistral OCR reference beside the OCR wire types
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): add lens_feedback ClickHouse table for human trace scores
ReplacingMergeTree(UpdatedAt, IsDeleted) keyed like agent_traces_by_key so
feedback joins traces in sort order, with a bloom filter on TraceId, a score
CHECK, and the same retention as traces
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): add scoped ClickHouse reads for trace feedback
feedback_target resolves a visible trace's team, key and trace_ref before a
write; feedback lists the latest live entry per author; feedback_summary
returns count, average and lowest score for a batch of traces
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(lens): pin the Claude session to trace id hash shared with feedback
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): expose feedback queries to the Python trace bridge
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): add typed feedback models and ClickHouse feedback store
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): add /lens/feedback API for 0-10 human scores on traces
PUT saves or replaces the caller's score and comment, GET lists every
entry, DELETE removes the caller's, POST /summary feeds the trace list.
Callers can target a trace_id or a Claude Code session_id
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): add typed feedback calls to the traces API
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): flag rated traces and filter the list by feedback
A Feedback column shows each run's average score and rating count, marked
red when any score is 4 or below, and a filter narrows to rated or
low-score runs through the URL
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): show and edit human feedback inside a trace
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): let app keys post end-user feedback on their own traces
Any key can now PUT/DELETE feedback on traces its team or key sent, naming
the end user in an optional user field. Reading feedback stays with Lens
admins
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens-ui): show end-user feedback first in a run and flag low-scored rows
Replace the edit popover and feedback filter with a read-only view: the
list shows each run's score in a fixed-width badge and tints the row red
when a user scored it 4 or below, and opening a run shows what users said
with their score right under the header
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(lens): add who started the run to RunSource
* feat(lens): carry agent.source.user on trace span rows
* feat(lens): resolve the run source user from the span row
* feat(lens): read agent.source.user in trace_spans query
* feat(lens): read agent.source.user in trace_page_spans query
* feat(lens): read agent.source.user in trace_span_batch query
* feat(lens): read agent.source.user in trace_list_span_batch query
* feat(lens): decode the source user from clickhouse span rows
* test(lens): add source user to trace cache test rows
* test(lens): add source user to trace cache read fixtures
* test(lens): add source user to trace cache snapshot fixtures
* test(lens): add source user to capture fixture rows
* test(lens): round-trip source user on span row contract
* test(lens): cover run source user resolution
* chore(lens): regenerate python trace types with source user
* chore(lens): regenerate trace json schema with source user
* chore(lens): regenerate trace page json schema with source user
* chore(ui): add source user to api types
* feat(lens): add run user chip and slack thread chip
* feat(lens): put agent, user and thread on the first row of the run header
* test(lens): cover the run header identity row and compact totals
* feat(lens): start demo runs from a slack thread
* refactor(rust-bridge): share field and response marshaling
* test(rust-bridge): isolate response factory test modules per case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust-bridge): document shared marshaling helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): isolate trace storage and investigation in a Rust service
* fix(lens): include Rust sources in the image build context
* feat(lens): wire service setup, scoped delivery receipts and lease attempts
* fix(lens): complete service routing and reject stale investigation results
* fix(lens): retry key propagation and validate isolated Compose setup
* fix(lens): seed through isolated ingestion and preserve upstream queue fixes
* chore: sync schema.prisma copies from root
* fix(lens): bind nullable due timestamps as text for Prisma
* chore(ui): remove stale lint suppressions
* fix(lens): address CI failures and review findings
* refactor(lens): remove retired Python worker and run evaluations in Rust
* fix(lens): reuse control connections and satisfy review checks
* test(lens): install and upgrade both Helm charts on Kubernetes
* test(lens): run connection reuse coverage as an integration test
* fix(ui): upgrade Next.js to 16.3.8 security release
* fix(lens): fence stale attempts and preserve reviewed evidence
* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"
This reverts commit 2f79a51b25.
* fix(lens): stop failed investigations and stream history excerpts
* test(lens): cover model tool and result contracts
* test(lens): fix retired routes and reuse installation build artifacts
* test(lens): use portable grep in Helm installation smoke
* test(lens): wait for migrations before forwarding Helm services
* fix(lens): keep failed evidence reads retryable
* fix(lens): preserve sandbox output during process exit
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(lens): add trace_agents rollup query
* feat(lens): register the trace_agents read query
* feat(lens): add trace_agents params and row types
* feat(lens): dispatch the trace_agents query
* feat(lens): export trace_agents wire schemas
* test(lens): pin the trace_agents query name
* test(lens): cover trace_agents scope, window and failure counts
* feat(lens): cap the agent list size
* feat(lens): declare the trace_agents bridge query
* feat(lens): read trace agents from clickhouse storage
* feat(lens): add trace agent response models
* feat(lens): list agents within the reader's trace scope
* feat(lens): add GET /v1/traces/agents
* chore(lens): regenerate trace models with trace_agents
* chore(lens): regenerate trace types with trace_agents
* chore(lens): regenerate read query name schema
* chore(lens): add trace agent row schema
* chore(lens): add trace agents params schema
* test(lens): cover the trace agents route
* test(lens): cover agent listing scope and timestamps
* chore(ui): regenerate api types with trace agents route
* feat(lens): derive trace agent types from the schema
* feat(lens): fetch the agent list from the traces api
* feat(lens): roll up demo runs into agents
* feat(lens): serve the agent list in demo data
* feat(lens): load agents seen in the last two weeks
* feat(lens): remember the selected agent per browser
* feat(lens): add the agent picker
* feat(lens): wire agent selection into the lens header
* feat(lens): show the agent picker next to the lens title
* refactor(lens): drop the toolbar agent filter in favor of the header picker
* test(lens): remove tests for the toolbar agent filter
* test(lens): cover agent resolution and rollup
* test(lens): cover scoping, switching and remembering the agent
* test(lens): read only the create request in guided setup
* test(lens): assert the toolbar agent filter is gone
* test(lens): read stubbed requests without their abort signal
* feat(lens): carry lens.source attributes on trace span rows
* feat(lens): read lens.source attributes in trace_spans query
* feat(lens): read lens.source attributes in trace_page_spans query
* feat(lens): read lens.source attributes in trace_span_batch query
* feat(lens): read lens.source attributes in trace_list_span_batch query
* feat(lens): decode lens.source columns from clickhouse span rows
* feat(lens): add run source to the trace summary contract
* feat(lens): export RunSource from litellm-traces
* feat(lens): resolve the https run source from the root span
* test(lens): cover run source resolution and https guard
* test(lens): add source fields to capture fixture rows
* test(lens): round-trip source fields on span row contract
* test(lens): add source fields to trace cache test rows
* test(lens): add source fields to trace cache read fixtures
* test(lens): add source fields to trace cache snapshot fixtures
* chore(lens): regenerate python trace types with run source
* chore(lens): regenerate trace json schema with run source
* chore(lens): regenerate trace page json schema with run source
* chore(ui): regenerate api types with trace run source
* feat(lens): add run source link with hover card
* test(lens): cover run source app detection and url guard
* feat(lens): show the run source next to the trace name
* feat(lens): label the run source link, e.g. Slack thread
* test(lens): assert run source link labels
* feat(lens): move the run source link into the trace stats row
* feat(lens): add a typed run source, e.g. slack or teams
* feat(lens): export RunSourceType
* feat(lens): carry the source type on trace span rows
* feat(lens): resolve the run source type, defaulting to custom
* feat(lens): read agent.source attributes in trace_spans query
* feat(lens): read agent.source attributes in trace_page_spans query
* feat(lens): read agent.source attributes in trace_span_batch query
* feat(lens): read agent.source attributes in trace_list_span_batch query
* feat(lens): decode the source type from clickhouse span rows
* test(lens): cover run source type parsing
* test(lens): add source type to capture fixture rows
* test(lens): round-trip source type on span row contract
* test(lens): add source type to trace cache test rows
* test(lens): add source type to trace cache read fixtures
* test(lens): add source type to trace cache snapshot fixtures
* chore(lens): regenerate python trace types with source type
* chore(lens): regenerate trace json schema with source type
* chore(lens): regenerate trace page json schema with source type
* chore(ui): regenerate api types with run source type
* feat(lens): pick the run source logo and label from its type
* test(lens): cover run source labels by type and slack logo
* feat(lens): show the run source as a Source stat with the app name
* test(lens): assert run source app names
* feat(lens): place the Source stat before Duration
* refactor(rust): separate Messages settings from capability inputs
* refactor(rust-bridge): unify Messages OCR and Responses call inputs
* refactor(rust-bridge): share NativeCall across inference entrypoints
* fix(rust-bridge): preserve Responses URL aliases and public test inputs
* feat(types): declare above_100k_tokens price fields on ModelInfoBase
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 186f6c81a0)
* fix(router): mirror above_100k pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 75c9c3ff1e)
* chore(ui): regenerate schema.d.ts for above_100k pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 53163d2184)
* refactor(types): mark above_100k ModelInfoBase fields ReadOnly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit febed3199a)
* fix(model_prices): bill claude-haiku-5-5 long prompts on every provider and in batch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 218c00cf5a)
* fix(cost): pass *_above_Nk_tokens_batches rates through get_model_info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 37b6f2fba0)
* fix(model_prices): allow disabling thinking and forced tool use on claude-haiku-5-5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 6e91e0ed36)
* test(model_prices): cite the vendor source for claude-haiku-5-5 capability flags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 1f97bb27b4)
* test(model_prices): type and tidy the claude-haiku-5-5 config tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 3d2526035f)
* feat(bedrock): add claude haiku 5.5 over 100k token tier
Price-Sync: litellm-providers
(cherry picked from commit ee3822c29c)
* feat(model_prices): add openrouter claude-haiku-5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): dedupe above_100k pricing keys from text merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): bill vertex claude-haiku-5-5 prompts over 100k tokens
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): add adaptive thinking and cache minimum to openrouter claude-haiku-5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
* perf(lens): bound single trace reads by the sampled start time
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): allow unused query fixture field in load tests
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* test(lens): pass start_time in every lens content and evidence test
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
* fix(caching): keep tool calls and tool results in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): keep semantic tool prompt helpers within lint budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): keep structured function_call_output text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split Responses text-field collection to stay within complexity budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): tag each tool result with the position of the call it answers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): encode tool result position and output together so tool text cannot forge result tags
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): expect encoded tool result record in qdrant semantic prompt parity case
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): cover tool result arrangements, SDK clients, concurrency and qdrant outage for semantic cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(caching): embed every semantic cache prompt field except volatile ones
Replace the per-shape allowlist in the Python and Rust semantic cache prompt
walkers with one include-by-default walker. Plain text keeps its old
concatenation; any other block or message is embedded as compact JSON with
call ids mapped to ordinals, cache_control dropped, and signatures, encrypted
content and base64 data replaced with a short sha256 digest.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "test(rust): expect structured JSON for unknown fields in redis and valkey semantic prompts"
This reverts commit 86c82b949b.
* Revert "test(code-quality): allow the bounded semantic cache prompt walkers in the recursion check"
This reverts commit 39efb5d9da.
* Revert "feat(caching): embed every semantic cache prompt field except volatile ones"
This reverts commit 5aed3ab3de.
* refactor(caching): rename get_str_from_messages_with_tools to get_semantic_cache_prompt_from_messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): split semantic cache prompt extraction by API format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): drop TypeIs guard and register Responses prompt walker with the recursion check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): pick the Responses text field without a Final inside a loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): walk semantic cache prompts as plain dicts, dumping pydantic items once up front
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): write the semantic cache prompt builders as plain loops
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip the cache past max_messages and keep tool_result text in semantic prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): drop formatting-only churn from the redis semantic cache tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(integration): drop the caching group wiring that main already carries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): recurse into tool_result content in the semantic cache prompt helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): check the max_messages cap on the shared exact-cache proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): read list-form function_call_output text in semantic cache prompts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(caching): extract nested Responses input lookup to keep walker under complexity limit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Revert "refactor(caching): extract nested Responses input lookup to keep walker under complexity limit"
This reverts commit 0665296bf1.
* style(caching): suppress C901 on the Responses input walker instead of splitting it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(traces): use the first user message for the run input preview
* test(traces): cover first user message as the input preview
* test(traces): check fixture previews against the first user message
* fix(lens): show the whole input preview on one line in the runs table
* test(lens): cover multi-line input previews in the runs table
* feat(lens): add dataset case and size limits
* feat(lens): add dataset, case and build models
* feat(lens): build dataset cases from traces, findings and text
* feat(lens): store dataset revisions insert-only
* feat(lens): add dataset routes for build, save, export and eval cases
* feat(lens): mount the dataset router before lens routes
* feat(lens): add LiteLLM_LensDataset table
* feat(lens): add LiteLLM_LensDataset table to proxy schema
* feat(lens): add LiteLLM_LensDataset table to extras schema
* feat(lens): add migration that creates the dataset table
* test(lens): cover dataset case building, dedupe and limits
* test(lens): cover dataset revisions, conflicts and eval cases
* chore(ui): regenerate API types for lens datasets
* feat(lens): add dataset UI types
* feat(lens): add datasets API client
* feat(lens): add dataset query and mutation hooks
* feat(lens): add case selection and expected edit logic
* test(lens): cover case selection and expected edits
* feat(lens): add the add to dataset dialog
* test(lens): cover saving picked cases from the dialog
* feat(lens): add datasets list
* feat(lens): add dataset detail with revisions and export
* test(lens): cover editing, revisions and export in datasets tab
* feat(lens): expose datasets on the lens API
* feat(lens): add in-memory datasets for demo mode
* feat(lens): wire demo datasets into the demo lens API
* feat(lens): add optional lens API hook
* feat(lens): add optional onboarding hook
* feat(lens): add datasets tab and dataset routing
* feat(lens): show the datasets tab
* feat(lens): add to dataset from the trace header
* feat(lens): add a single turn to a dataset from a step
* feat(lens): add finding evidence to a dataset
* docs(lens): add datasets screenshots for the PR
* test(lens): cover dataset revision storage against Postgres
* test(lens): cover dataset trace paging, findings and route errors
* test(lens): cover dataset build fallbacks and no_content skips
* fix(lens): register the dataset table for postgres span names
* feat(lens): add invalid skip reason for unparseable case lines
* fix(lens): keep valid JSONL cases, provenance and size limits; stop reading past the case cap
* test(lens): cover malformed JSONL, re-import provenance, size fields and early cap
* chore(ui): regenerate API types for the invalid skip reason
* feat(lens): show text for the invalid skip reason
* feat(lens): render datasets in the same inspector table as investigations
* test(lens): open a dataset by clicking its table row
* docs(lens): update the datasets list screenshot
* style(lens): format the datasets table
* fix(lens): normalize JSON text span input and output into messages when building cases
* test(lens): cover JSON text span normalization for dataset cases
* Revert "fix(lens): normalize JSON text span input and output into messages when building cases"
This reverts commit 29c3e34962fa12427dbd516c08f2599db997a905.
* fix(traces): normalize agent assistant summaries into UI messages
* feat(lens): add dataset case view helpers built on the trace parsers
* test(lens): cover dataset case view helpers
* feat(lens): show dataset cases in an inspector table
* feat(lens): open a dataset case in a side panel with trace message cards
* feat(lens): rebuild the dataset page header and layout
* feat(lens): keep the open dataset case in the URL
* test(lens): drive dataset edits through the case table and panel
* docs(lens): update dataset view screenshots
* Revert "test(lens): cover JSON text span normalization for dataset cases"
This reverts commit 30229ae8ae90265c09d755c43c1f1eeaa66050c7.
* fix(lens): import StateMessage from shared in dataset detail
* fix(lens): import StateMessage from shared in datasets list
* fix(lens): give the datasets migration a unique timestamp after review checkpoints
* fix(lens): price agent traces by joining gen_ai.response.id to spend logs
Trace spend now joins each model call to spend_logs on one key: the
span's response id (gen_ai.response.id or the id the normalizers read
from OpenInference/LangChain output) against spend_logs.response_id or
the upstream id embedded in a managed resp_ id. The litellm.call_id and
traceparent transport join paths and the per-row ownership gate are
removed; the spend SQL still restricts rows to what the reader can see.
A run with some unpriced calls now reports the sum of its priced calls
plus priced_calls, instead of an unknown total.
* test(lens): cover response id spend join and partial trace totals
* chore(lens): regenerate trace types for priced_calls
* feat(lens): show partial run cost as a lower bound with priced call count
* fix(lens): treat litellm.call_id as the same assigned call id for spend joins
The id LiteLLM assigned to a call is either the response id it returned
(gen_ai.response.id -> spend_logs.response_id) or its gateway call id
(litellm.call_id -> spend_logs.litellm_call_id). Both are exact ids the
gateway mints and logs, so the join stays one rule. Transport span
matching and the ownership gate stay removed.
* test(lens): cover litellm.call_id spend joins and restore captured totals
* feat(lens): link each priced model call to its spend log
Spans gain spend_log_request_id, the spend_logs.request_id the call was
priced from, and spend_match, which says whether a model call matched or
why not (no assigned id on the span, no spend log with that id, or an
ambiguous match). A model call span is priced from the same ids as the
run total, so its cost and the total agree.
* test(lens): cover spend log links on model call spans
* chore(lens): regenerate trace types for spend log links
* feat(lens): open the matched spend log from an LLM step
An LLM step's header now shows a Spend log chip with the matched
request id and cost; clicking it opens the request log drawer over the
run, fetched by the exact spend_logs.request_id instead of the span's
own response id. Unpriced steps say why (no assigned id on the span, or
no spend log with it). Tree rows show each model call's cost, and a
partial run cost shows its priced call count inline.
* test(lens): cover the spend log link and unmatched cost reasons
* feat(lens): show the spend log link as a bordered LiteLLM Spend Log button
* feat(lens): add a back link from the spend log drawer to the agent trace
* feat(lens): label the spend log back link Back to Lens trace with the Lens icon
* fix(lens): ignore assigned ids that name no spend log when pricing a call
An id that names no row no longer vetoes the call, so a span carrying
both a response id and a call id still prices from a spend row logged
before litellm_call_id existed. An id naming two or more rows makes the
call ambiguous, and the match reason comes from the same per-id result,
so a single matched row with no cost is reported as matched.
* fix(lens): price a trace only from spend logs in its own team
A reader with several teams could see the same assigned id in another
team's spend log; only rows from the trace's team now price it. The
user and key ownership gate stays removed.
* perf(lens): resolve each model call's spend once per trace
Model call matches are computed once when the trace is resolved and
looked up by span index, instead of scanning the model call list for
every span and walking the graph again for spans, agents and the run
total.
* fix(lens): hide a step's Cost fact only when its spend log link shows the cost
* chore(lens): drop narrative doc comments from the spend join
* fix(lens): price a model call only when its ids agree on one spend log per span
* fix(lens): keep pricing spend logs written before litellm_call_id by their request id
* fix(lens): price every attempt a model call's ids name when they agree
* refactor(rust_bridge): remove rule-gated native secret-manager selection
Mirror the cache treatment: SecretManagerRule/SecretManagerContext and the resolve_native_* binding plumbing are gone. The Rust bridge now selects the native backend from the explicitly configured client (capture_secret_manager / _SecretManagerRuntime.from_client), and get_secret_from_manager is the plain Python handler path.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(deps): bump sharp to 0.35.5 for GHSA-wq5f-xc86-pv6w
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.
Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(rust): prune unused inference crate dependencies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): rewrite inference layering docs for the split format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): note transcription and RouteError alias exceptions in inference AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): rename litellm-core to litellm-inference
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): expose inference base API for format crates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): move shared inference test helpers behind test-support
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The native response-cache runtime attached through Cache._native_cache is
unreachable since the V2 cache replaced it. Drop the Python branches and
wrapper, the _ResponseCacheRuntime pyclass and its backend/activation/
semantic modules, the python-bridge deps only they used, and the tests and
fixtures dedicated to that path.
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(tracing): pin HTTP request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): generate existing HTTP request models from Rust schemas
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): bind trace query params to generated request models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(rust): raise the native wheel size gate to 48 MB
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(tracing): read the trace list clock without a thread-pool dispatch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(traces): share the trace page-size bounds between schema and reader
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): type trace request queries against the generated OpenAPI schema
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs: encode the trace contract boundary in AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): format trace request aliases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): track ClickHouse migrations in a checksummed ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(clickhouse): harden migration execution and reuse retention SQL
* refactor(rust): track ClickHouse migrations in a checksummed ledger
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(rust): keep ClickHouse retention TTLs in the current policy list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(lens): own trace reads behind a cached TraceStore port
Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.
Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.
Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(lens): seed large profiles with server-side copies and long sessions
Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chores
* style(lens): float the investigation setup badge on the tab edge
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(lens): restructure the trace drawer and polish its layout
Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.
Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* typesafety
* wip
* fmt
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat: seed Lens dev with configurable load profiles
* feat: run Lens UI live through the dev launcher
* fix: verify Lens UI startup before seeding
* fix(ui): render run timestamps on one compact line
The agent runs table printed the long locale form with timezone, which
wrapped to two lines per row. Use a fixed-width 24h form with
milliseconds that matches the timeline axis, keep the long form in the
hover title, and show the timezone once in the column header
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): use the root query client for the Lens demo
Drop the demo's nested QueryClient. Cache keys are already partitioned by
scope, and the root client now skips retries on 4xx ApiErrors, which covers
the demo's not-in-demo and read-only rejections and live 4xx alike.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: drop unused synthetic_spend reference from query help
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(traces): drop the deeplite fixtures and map every SDK to a logo
The deeplite captures predate the example repository and carried
synthetic spend rows, which leaked a fixture-only column into the
query help SQL and pinned tests to its shape. Replace them with the
SDK captures in the ClickHouse round trip and query API tests, and
derive the seed tenant lookup from the capture metadata
The runs table only knew the two Anthropic framework slugs. Register
the slugs the normalizer emits for LangChain, LangGraph, Deep Agents,
CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands,
Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic
agent glyph when a run has no known framework
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(ui): inject the Lens sample through data sources instead of demo checks
Components no longer ask whether they are in the demo. The traces source
carries live and handoff, navigation state comes from a URL or memory
route, and the preview action comes from context instead of onDemo props
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): infinite scroll for the agent runs list
Replace the Load more button with a sentinel that fetches the next
cursor page as the list nears its end. Placeholder rows hold the tail
while more runs exist, and a failed page stops auto-loading until Retry.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(ui): keep the whole Lens view in the URL
Lens navigation now lives entirely in query params through nuqs: the
sample session (demo=true, with a Demo data switch in the header), the
open run (trace, trace_ref), the selected step, view and detail section
(span, view, span_tab) and the list filters and range (q, agent, status,
hours). Any Lens view is a shareable link and the back button walks runs
RunView takes its selection injected: the drawer feeds it URL state and
the investigations evidence sheet keeps a local one, so a finding's
original run never writes step ids into the URL. Leaving the sample
session clears every Lens key except the tab so sample ids never point
at live data
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* style(traces): cargo fmt captures tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>