- Sanitize error messages: generic 'An internal error occurred' sent to client,
full exception logged server-side via verbose_proxy_logger
- Defense-in-depth: _process_tool_call validates fn_name against role-based
allowlist before dispatch (even though LLM only receives allowed tools)
- Revert tsconfig.json jsx back to 'preserve' (Next.js recommended default)
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
- Restrict team/tag tools to admin-only users (non-admins only get get_usage_data)
- Constrain ChatMessage.role to Literal['user', 'assistant'] to prevent system prompt injection
- Add test for base tools restriction (non-admin gets 1 tool, admin gets 3)
- Issues 3 (unused imports) and 4 (inline datetime) were already fixed in prior commit
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
- Backend emits tool_call events with tool_name, label, args, and status
- Frontend shows each tool call as a step with ✓/spinner/✗ indicator
- Tool call steps show icon, label, date range, and filters
- AI responses rendered with ReactMarkdown (bold, lists, tables, code)
- Cursor-like UX: Thinking → tool calls → Analyzing → streamed answer
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
- AI agent now has 3 tools: get_usage_data, get_team_usage_data, get_tag_usage_data
- Stream status events (Thinking... Fetching... Analyzing...) to UI
- Frontend shows spinner + status text during tool execution
- Better system prompt guiding tool selection
- Entity summariser for team/tag data with ranked breakdowns
- 13 backend tests, 34 frontend tests passing
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Backend:
- New /usage/ai/chat SSE streaming endpoint
- AI agent has get_usage_data tool that queries /user/daily/activity/aggregated
- Follows same architecture as policy AI suggest (litellm.acompletion + tools)
- Non-admin users are restricted to their own data
- 12 backend unit tests
Frontend:
- Panel now calls /usage/ai/chat backend endpoint via SSE
- Removed direct OpenAI client calls from frontend
- Added usageAiChatStream networking function following enrichPolicyTemplateStream pattern
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
- Replace UsageAIChatModal with UsageAIChatPanel
- Panel slides in from right side, usage page stays visible
- Full-height panel with header, model selector, chat area, and input
- Smooth CSS transition for open/close animation
- Update tests for new panel component (34 tests passing)
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
- Create UsageAIChatModal component with streaming chat interface
- Integrate with existing model hub for model selection
- Pass usage data context (spend, models, providers, keys) to AI
- Add Ask AI button next to Export Data button in global view
- Add tests for the new component and integration
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(router): preserve _hidden_params in FallbackStreamWrapper so x-litellm-overhead-duration-ms is emitted for streaming requests
* test(router): add regression test for FallbackStreamWrapper _hidden_params preservation
Set prometheus_emit_stream_label: true in litellm_settings to emit a
stream label (True/False/None) on litellm_proxy_total_requests_metric.
Opt-in to avoid breaking cardinality on existing deployments.
* staged first pass
* black
* Update litellm/proxy/health_check.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* simpler
* restore cached logo
* fix tests for perform_health_check max_concurrency arg
* implement pr suggestion
* and the helm chart
* add configureable resources and probes to the deployment in the helm chart
* more helm chart unittests
* move some background healthcheck loggin to debug
---------
Co-authored-by: Sean Glover <sglover@athenahealth.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
# Please enter a commit message to explain why this merge is necessary,
# especially if it merges an updated upstream into a topic branch.
#
# Lines starting with '#' will be ignored, and an empty message aborts
# the commit.
Non-owner Internal Users could see and interact with the "Edit Settings"
button in the key Settings tab for keys they don't own. The button was
gated by `rolesWithWriteAccess.includes(userRole)` (role-only check)
instead of `canModifyKey` (ownership-aware), unlike the Regenerate and
Delete buttons which already used the correct check.
Replace the condition with `canModifyKey` so the Edit Settings button
follows the same proxy-admin / team-admin / key-owner logic as the
other action buttons. Add tests covering all permission paths.
- Wrap onboarding page with QueryClientProvider to prevent runtime crash
(mirrors the same pattern used in LoginPage)
- Stage deletion of stale litellm/ui/litellm-dashboard/src/app/onboarding/page.tsx
committed at the wrong path
- Rename all 16 test names to start with "should" per AGENTS.md convention
* feat(realtime): add guardrail hook for voice transcription in Realtime API
Adds a new `realtime_input_transcription` guardrail event hook that fires
after Whisper transcription completes, before the LLM generates a response.
When a guardrail blocks, a synthetic warning is sent to the client and
`response.create` is never forwarded — the LLM never responds.
Also rewrites `create_response: true` → `false` in client `session.update`
so the proxy controls when responses are triggered.
* feat(realtime): speak guardrail block message as audio via TTS
Instead of sending synthetic text events when a guardrail blocks,
send response.create with forced instructions so OpenAI's TTS speaks
the warning message — user hears the block instead of just seeing text.
* fix(realtime): speak exact content filter error message via TTS
Extract the human-readable error string from HTTPException.detail
so the spoken warning says e.g. "Content blocked: keyword 'system update'
detected" instead of the raw str(e) repr.
* fix(realtime): reliably enforce create_response=false for guardrails
- Proxy now injects session.update with create_response=false immediately
on session.created (when guardrails are active), instead of rewriting
the client's session.update — works regardless of what the client sends
- Add response.cancel before the warning response.create to kill any
in-flight LLM response that snuck through before the guardrail fired
* refactor(realtime): call apply_guardrail directly, remove dedicated hook method
The async_realtime_input_transcription_hook in CustomGuardrail and
ContentFilterGuardrail was just a thin wrapper that called apply_guardrail —
the same interface used by /chat and /messages. Remove the wrapper and call
apply_guardrail directly from run_realtime_guardrails, keeping the pattern
consistent across all endpoints.
* docs: add Realtime API guardrails tutorial and flow diagram
* fix: address Greptile review comments
- Forward user_api_key_dict through realtime_api/main.py (_arealtime) so
it actually reaches RealTimeStreaming instead of always being None
- Run guardrail interception in provider_config path too (e.g. Gemini),
not only the OpenAI direct path
- Narrow exception catch to HTTPException/ValueError only; re-raise
unexpected errors so programming bugs surface in logs rather than
silently appearing as guardrail blocks
- Update tests: mock apply_guardrail directly (hook method was removed),
replace session.update client-rewrite test with session.created
injection test matching the new server-side approach
* fix: address latest Greptile review comments
- Remove fastapi import from SDK-layer file; check for status_code/detail
attrs instead to identify guardrail-block exceptions vs programming errors
- Add store_message() before continue in transcription interception so
transcription events are logged in the non-provider_config path
- Inject create_response=false on session.created in provider_config path
(Gemini etc.) to match the OpenAI path — prevents LLM auto-responding
before guardrail runs on VAD-detected turns