From 32a52594075fe7a5c7e011c53492f700443a6bde Mon Sep 17 00:00:00 2001 From: yuneng-jiang Date: Wed, 12 Aug 2026 13:39:58 -0700 Subject: [PATCH] test(e2e-ui): verify UI mutations against the API instead of trusting the toast (#36632) * test(e2e-ui): cover the Playground, Logs and Usage manual-QA flows These three pages carried no e2e coverage, so the manual QA checklist was the only thing standing behind them. Playground: sends a chat from the UI for both configured models, and for both virtual-key sources (the logged-in session, and a key pasted into the panel). This is the only spec that drives the dashboard's own LLM call path rather than an admin CRUD endpoint. Logs: a request the proxy actually served appears in the table, its drawer expands to the real request and response bodies, both copy to the clipboard, the Input card collapses, the JSON view exposes Request/Response, and the End User filter narrows the table to one customer. Usage: traffic billed to a virtual key reaches Top Virtual Keys, the card toggles between table and chart, and the key opens its key-info panel. Router settings: the existing spec proved the UI can record a fallback; the new one proves the fallback is honoured, by pointing a model at an unreachable upstream and asserting the reply comes back anyway. It asserts the un-fallen-back call fails first, so a quietly-working primary cannot fake a pass. Supporting changes: - helpers/traffic.ts generates the traffic these pages render, rather than seeding rows no code produced. Its two wait helpers exist because the Logs and Usage pages read different stores: spend logs are flushed on a timer, and the Usage page reads a background rollup *and* fetches once on mount, so waiting on the DOM there can never converge. - helpers/playground.ts holds the playground controls, now shared with the fallback spec. Everything is scoped to the visible copy of the config panel, which is rendered twice for the docked and collapsed layouts. - run_e2e.sh gains E2E_KEEP_ALIVE=1, which brings the stack up and blocks so a spec can be re-run against it without paying for a UI rebuild each iteration. Verified with the full suite on a fresh stack: 89 passed, 0 failed, 5 skipped. * test(e2e-ui): cover listing and calling MCP tools Covers the two MCP manual-QA items the create-only spec cannot reach: opening a server's tool list, and calling a tool and seeing its result. Both need an MCP server that actually answers, so this points at DeepWiki's public MCP server -- Streamable HTTP, auth None, so there is no credential to hold and nothing to leak from a public repo. The call is made by the proxy, not the browser; nothing in the e2e chart restricts that egress. The external dependency is real and is left visible: an upstream outage turns these red rather than auto-skipping, because a spec that skips itself on connection trouble also skips when the proxy's MCP client is what broke. E2E_SKIP_EXTERNAL_MCP=1 is the explicit opt-out. Not yet executed against a live stack. * test(e2e-ui): verify key mutations round-trip instead of trusting the toast The recurring customer report is a form that says "Saved!" and then either no-ops or clobbers an unrelated field. A toast-only assertion passes in both cases, and outside three specs that is all this suite checks. Adds helpers/roundTrip.ts, factoring out the idiom clearCustomPricing, credentials and routerSettings already use: capture the outgoing request body, then read the resource back through the management API. Applies it to the keys spec: - create: the key is readable from /key/list and owns a team_id, rather than trusting a table row rendered from the create response the UI already held - update limits: TPM/RPM are on the wire AND persisted, and the key's models and team are unchanged -- bumping one field wiping another is the reported failure mode (PR #34452), not a hypothetical - delete: the key is gone from /key/list, not merely toasted as deleted - regenerate: the stored token actually changed /key/list shape is per KeyListResponseObject in litellm/proxy/_types.py. Not yet executed: ports 4000/8090 are held by a parallel run. * test(e2e-ui): let the local harness run on non-default ports Two checkouts cannot run run_e2e.sh at the same time: it hardcodes 4000/5432/ 8090, so the second aborts on "port 4000 is in use" and the only way forward is to stop someone else's stack. PROXY_PORT / POSTGRES_PORT / MOCK_LLM_PORT now override those, defaulting to the historical values so an unset environment behaves exactly as before -- CI, the CircleCI job and the chart's sidecar all keep working untouched. Two details that would otherwise make a relocated stack fail confusingly: - the suite resolves its target from E2E_UI_BASE_URL, which defaults to :4000 independently, so the run would build and boot correctly and then test whatever was on the default port. run_e2e.sh now derives it. - the mock server binds its port in server.py, so moving it needs MOCK_LLM_PORT there too. Its HOST stays loopback-only: 127.0.0.1:8090 from inside the proxy's own pod is the contract the e2e chart's sidecar is written against. * test(e2e-ui): cover MCP server edit and delete, verified via the API mcpServers.spec.ts only ever creates a server, and creation is the one MCP operation nobody has complained about. The reports are all on the other side: an alias rename that needs three or four saves to take, a delete that needs two attempts. Both produce a success toast on the failing attempt, so a toast-only assertion cannot tell them from working software. Rename asserts the new alias and the target server_id are on the PUT, then polls /v1/mcp/server until the stored alias matches -- one save has to be enough. Delete asserts the server is really gone from the list. Points at an unreachable URL: these exercise litellm's persistence, never the upstream, so a live MCP server would add a network dependency for nothing. mcpTools.spec.ts is where a real upstream is needed. Both pass against a local stack, as do the mcpTools specs from 5e189e9b1a. Neither reproduced the reported failures on this build -- they guard, they did not catch. * test(e2e-ui): verify team create, invite and delete against the API Three team mutations stopped at a toast, and one of those toasts is matched as loosely as /success/i -- almost any notification satisfied it. - create: the team is readable from /team/list and kept the models chosen in the modal, rather than trusting the UI's own "Team created" - invite: the invited address really appears in members_with_roles, which is the point of the flow - delete: the team is gone from /team/list. The existing assertion was that the row vanished, which is the client dropping it from local state and happens whether or not the delete reached the database. Shapes read off a live proxy: /team/list is a bare array; /team/info nests the record under team_info. All 6 tests pass locally. * test(e2e-ui): verify team-admin member and key mutations against the API The team-admin flows stopped at a success toast. A member add that lands on the wrong team, a remove that takes out the wrong row, and a key that comes back unscoped all produce the same toast as the working case, so the existing assertions could not tell them apart. Each mutation now pins what went on the wire and reads the result back: member add/remove assert team_id and the member identifier on the request, then poll /team/info's roster; the team key asserts team_id on /key/generate and reads /key/list back to confirm the key is owned by the admin's own team rather than orphaned. * test(e2e-ui): verify model add and limit edits against the stored deployment The Models specs checked the rendered result: the TPM/RPM edit asserted the new numbers were visible in view mode, and the two add flows asserted a row showed up in the table. Both render from state the UI already holds, so a save the backend dropped and a save it kept look the same. Each mutation now pins the request and reads the deployment back. The limits edit also asserts the fields it did not touch -- upstream model and team ownership -- are unchanged, because handleModelUpdate rebuilds and PATCHes the whole litellm_params blob, which is how an unrelated field gets clobbered by a save that reports success. The two add flows assert model_name, the routed model and custom_llm_provider on the wire and in storage; a deployment that loses its provider looks correct in the table and is unroutable. The Team-BYOK test is unchanged -- it is skipped without a license, so any change to it would be unverified. * test(e2e-ui): delete the MCP servers these specs create MCP servers outlive the test that made them, the MCP page contacts every server it lists, and most of the ones these specs create point at an unreachable host. They accumulate, and each one makes navigateToPage's networkidle wait a little slower to settle. Measured on a local stack: with eleven leaked servers the whole MCP suite failed on a 30s navigation timeout, including specs that leaked nothing. Deleting the leftovers made all five pass. With per-test cleanup added, a run from a clean slate leaves zero behind and takes 20s instead of 1m24s. mcpServers.spec.ts carried a note that no teardown was needed because the runner brings up a fresh database each time. That holds for CI and is why this went unnoticed; it does not hold for a local stack that is reused. * test(e2e-ui): say why the team-model setup call failed The setup that creates a team-scoped model asserted a bare `ok()`, so a failure read "expected true, received false" and pointed at the UI. The call is enterprise-gated -- creating a model with model_info.team_id returns 403 without LITELLM_LICENSE -- and that is invisible from the old message. It now carries the status and body, which names the cause immediately. The api_base also pointed at the mock's default port rather than the one the harness started; nothing in the test calls the model, but the two should not disagree. * test(e2e-ui): add a model through the UI and serve traffic with it Every existing Add Model test stops at "the row appears in the table", which a deployment that cannot serve a single request also does. The manual-QA item this replaces is the whole loop: fill the form, pass Test Connect, add it, confirm it works. The new test ends by calling the model it just created. That is the only assertion that rules out a dropped api_base, a mangled provider prefix, or a name the router never registers -- all of which look identical in the UI. No provider credential is involved. OpenAI-Compatible is the provider whose form exposes API Base, so the deployment points at the harness's own mock LLM. The mock speaks the OpenAI wire format, so Test Connect performs a real completion against a real endpoint and really succeeds. Also adds teardown for the deployment it creates. A local run throws its database away, but the deployed stack does not, and a leaked deployment shows up in every later Models table and /v2/model/info readback. Both new assertions were mutation-tested: pointing the traffic poll at a name that was never created fails the test, and the wire assertion fails when the typed name is not what reaches /model/new. * test(e2e-ui): print the proxy log when the proxy dies on its own In E2E_KEEP_ALIVE=1 mode the harness blocks until the proxy pid goes away, then printed a bare "Proxy exited." and fell straight into cleanup, which rm -f's the log. The proxy has now exited by itself twice, minutes after a run had finished, leaving nothing to look at. Both startup failure paths already tail -n 100 the log before giving up, so this was the one death that stayed silent Dump the same 100 lines before exiting. A normal Ctrl-C teardown still deletes the log and prints nothing, which is why INT and TERM now exit instead of running cleanup and falling back into the wait loop: under the single trap a SIGTERM deleted the log, resumed the loop, and would then report "tail: no such file", besides running cleanup twice * test(e2e-ui): split the log-drawer copy assertions off the expand test The copy assertions need `navigator.clipboard`, which the browser only exposes in a secure context. Locally the suite runs against http://127.0.0.1 and localhost is trustworthy, so it is there. In CI the run pod is pointed at a plain-HTTP cluster DNS name, where it is undefined -- and InputCard.handleCopy calls writeText unguarded, so the click throws before MessageManager.success and no toast ever renders. That failed all three attempts of litellm-e2e-ui build 10. Measured rather than inferred: on http://127.0.0.1:4100 isSecureContext/typeof navigator.clipboard are true/"object", and on a DNS name resolving to that same address they are false/"undefined", which reproduces the CI failure exactly. Splitting keeps the drawer-rendering coverage running everywhere and confines the skip to the part the browser has actually switched off. The copy assertions still run in full wherever the origin is trustworthy. The underlying product behaviour is left alone deliberately: any deployment served over plain HTTP on a hostname has a copy button that throws and gives no feedback, and that deserves its own fix rather than being papered over from a test. * test(e2e-ui): cut the added comments back to what the code cannot say itself Greptile flagged the helper commentary, and it was right: CLAUDE.md says not to write comments unless they explain very complex business logic, and much of what was added here narrated ordinary test setup and motivation instead. Trims 310 comment lines across the 14 files this branch touched. Kept only the notes that record something unrecoverable from the code: why the request listener is armed before the click, why a locator walks up the DOM, why an assertion exists beyond the toast. Pre-existing comments are left alone. No behaviour change. The only non-comment hunk is a prettier reformat. --- .../e2e/ui/fixtures/mock_llm_server/server.py | 11 +- tests/e2e/ui/helpers/mcp.ts | 65 ++++++ tests/e2e/ui/helpers/playground.ts | 46 ++++ tests/e2e/ui/helpers/roundTrip.ts | 28 +++ tests/e2e/ui/helpers/traffic.ts | 125 ++++++++++ tests/e2e/ui/run_e2e.sh | 115 +++++++-- tests/e2e/ui/tests/logs/logs.spec.ts | 221 ++++++++++++++++++ tests/e2e/ui/tests/mcp/mcpServerEdit.spec.ts | 92 ++++++++ tests/e2e/ui/tests/mcp/mcpServers.spec.ts | 13 +- tests/e2e/ui/tests/mcp/mcpTools.spec.ts | 80 +++++++ .../e2e/ui/tests/modelsPage/addModel.spec.ts | 198 +++++++++++++--- .../ui/tests/playground/playground.spec.ts | 50 ++++ tests/e2e/ui/tests/proxy-admin/keys.spec.ts | 63 ++++- tests/e2e/ui/tests/proxy-admin/teams.spec.ts | 39 +++- .../ui/tests/settings/routerSettings.spec.ts | 110 ++++++++- .../e2e/ui/tests/team-admin/teamAdmin.spec.ts | 75 +++++- tests/e2e/ui/tests/usage/usagePage.spec.ts | 74 ++++++ 17 files changed, 1335 insertions(+), 70 deletions(-) create mode 100644 tests/e2e/ui/helpers/mcp.ts create mode 100644 tests/e2e/ui/helpers/playground.ts create mode 100644 tests/e2e/ui/helpers/roundTrip.ts create mode 100644 tests/e2e/ui/helpers/traffic.ts create mode 100644 tests/e2e/ui/tests/logs/logs.spec.ts create mode 100644 tests/e2e/ui/tests/mcp/mcpServerEdit.spec.ts create mode 100644 tests/e2e/ui/tests/mcp/mcpTools.spec.ts create mode 100644 tests/e2e/ui/tests/playground/playground.spec.ts create mode 100644 tests/e2e/ui/tests/usage/usagePage.spec.ts diff --git a/tests/e2e/ui/fixtures/mock_llm_server/server.py b/tests/e2e/ui/fixtures/mock_llm_server/server.py index 8e92065c696..82c90a9dd64 100644 --- a/tests/e2e/ui/fixtures/mock_llm_server/server.py +++ b/tests/e2e/ui/fixtures/mock_llm_server/server.py @@ -3,6 +3,7 @@ Mock LLM server for UI e2e tests. Responds to OpenAI-format endpoints with canned responses. """ +import os import time import json import uuid @@ -117,4 +118,12 @@ async def embeddings(request: Request): if __name__ == "__main__": - uvicorn.run(app, host="127.0.0.1", port=8090) + # The port is overridable so two checkouts can run the harness at the same + # time; the default keeps every existing caller (run_e2e.sh, the CircleCI + # job, the e2e chart's sidecar) working untouched. + # + # The HOST is deliberately NOT configurable. Binding loopback is what makes + # this reachable at 127.0.0.1:8090 from inside the proxy's own pod, which is + # the contract the deployed config.yml and the e2e values file are written + # against. + uvicorn.run(app, host="127.0.0.1", port=int(os.environ.get("MOCK_LLM_PORT", "8090"))) diff --git a/tests/e2e/ui/helpers/mcp.ts b/tests/e2e/ui/helpers/mcp.ts new file mode 100644 index 00000000000..b41aec59ded --- /dev/null +++ b/tests/e2e/ui/helpers/mcp.ts @@ -0,0 +1,65 @@ +import { expect, Page as PwPage } from "@playwright/test"; +import { navigateToPage } from "./navigation"; +import { Page } from "../fixtures/pages"; +import { masterKey } from "./traffic"; + +/** Creates an MCP server through the UI's discovery to custom-form flow and returns its name. */ +export async function createMcpServer(page: PwPage, url: string): Promise { + await navigateToPage(page, Page.McpServers); + + await page.getByRole("button", { name: /Add New MCP Server/i }).click(); + const discovery = page.getByRole("dialog").filter({ hasText: "Add MCP Server" }); + await expect(discovery).toBeVisible({ timeout: 5_000 }); + await discovery.getByRole("button", { name: /Custom Server/i }).click(); + + const formModal = page.locator(".ant-modal:visible").filter({ hasText: "MCP Server Name" }); + await expect(formModal).toBeVisible({ timeout: 5_000 }); + + // validateMCPServerName rejects spaces and hyphens; the worker index avoids a same-millisecond collision. + const name = `e2e_mcp_${process.env.TEST_WORKER_INDEX ?? "0"}_${Date.now()}`; + await formModal.locator('input[id="server_name"]').fill(name); + + const transportField = formModal.locator(".ant-form-item", { hasText: "Transport Type" }); + await transportField.locator(".ant-select").click(); + await page.locator(".ant-select-dropdown:visible").getByText("Streamable HTTP").click(); + + await formModal.locator('input[id="url"]').fill(url); + + // The auth_type Form.Item has no label prop, so anchor on the enclosing Collapse panel. + const authSection = formModal.locator(".ant-collapse-item", { hasText: /^Authentication/ }); + await authSection.locator(".ant-form-item").first().locator(".ant-select").click(); + await page.locator(".ant-select-dropdown:visible").getByText("None", { exact: true }).click(); + + await formModal.getByRole("button", { name: /^Add MCP Server$/ }).click(); + await expect(page.getByText("MCP Server created successfully").first()).toBeVisible({ timeout: 15_000 }); + + const card = page.getByTestId("mcp-servers-grid").getByText(name).first(); + await expect(card).toBeVisible({ timeout: 10_000 }); + return name; +} + +/** + * Deletes every server carrying `serverName`. Leaked servers break unrelated MCP specs: the page + * reaches out to each one it lists, so unreachable leftovers stall networkidle until it times out. + * Errors are swallowed because this runs from afterEach. + */ +export async function deleteMcpServerByName(page: PwPage, serverName: string): Promise { + const headers = { Authorization: `Bearer ${masterKey()}` }; + try { + const res = await page.request.get("/v1/mcp/server", { headers }); + if (!res.ok()) return; + const servers = (await res.json()) as { server_id: string; server_name?: string }[]; + for (const server of servers.filter((candidate) => candidate.server_name === serverName)) { + await page.request.delete(`/v1/mcp/server/${server.server_id}`, { headers }); + } + } catch { + // best effort, see above + } +} + +/** Opens a server from the grid and switches to its MCP Tools tab. */ +export async function openMcpToolsTab(page: PwPage, serverName: string): Promise { + await page.getByTestId("mcp-servers-grid").getByText(serverName).first().click(); + await expect(page.getByRole("button", { name: /Back to All Servers/i })).toBeVisible({ timeout: 10_000 }); + await page.getByRole("tab", { name: "MCP Tools" }).click(); +} diff --git a/tests/e2e/ui/helpers/playground.ts b/tests/e2e/ui/helpers/playground.ts new file mode 100644 index 00000000000..39aae8398a5 --- /dev/null +++ b/tests/e2e/ui/helpers/playground.ts @@ -0,0 +1,46 @@ +import { expect, type Locator, type Page as PlaywrightPage } from "@playwright/test"; +import { navigateToPage, dismissFeedbackPopup } from "./navigation"; +import { Page } from "../fixtures/pages"; + +/** Controls for the Test Key / Playground page, shared with the router-fallback specs. */ + +/** + * The configuration panel is rendered twice, docked and overlay, with one visible at a time. + * Every control is narrowed to the visible copy or it trips strict mode against its hidden twin. + */ +export const onlyVisible = (locator: Locator): Locator => locator.filter({ visible: true }).first(); + +/** The model dropdown, addressed by the placeholder it shows before selection. */ +export const modelSelect = (page: PlaywrightPage): Locator => + onlyVisible(page.locator('.ant-select:has(.ant-select-selection-placeholder:text-is("Select a Model"))')); + +/** Send button is icon-only (an up-arrow), so there is no accessible name. */ +export const sendButton = (page: PlaywrightPage): Locator => onlyVisible(page.locator("button:has(.anticon-arrow-up)")); + +/** The Virtual Key Source dropdown, addressed by its currently selected label. */ +export const keySourceSelect = (page: PlaywrightPage, current: string): Locator => + onlyVisible(page.locator(`.ant-select:has(.ant-select-selection-item[title="${current}"])`)); + +export async function openPlayground(page: PlaywrightPage): Promise { + await navigateToPage(page, Page.LlmPlayground); + await dismissFeedbackPopup(page); + await expect(onlyVisible(page.getByText("Virtual Key Source"))).toBeVisible({ + timeout: 20_000, + }); +} + +export async function selectModel(page: PlaywrightPage, model: string): Promise { + const select = modelSelect(page); + await select.click(); + // Virtualized: options outside the rendered window are absent from the DOM, so search first. + await select.locator("input.ant-select-selection-search-input").fill(model); + // antd portals its dropdown to the body; options carry the value as `title`. + await onlyVisible(page.locator(`.ant-select-item-option[title="${model}"]`)).click({ timeout: 15_000 }); +} + +export async function sendMessage(page: PlaywrightPage, message: string): Promise { + const input = onlyVisible(page.getByPlaceholder("Type your message", { exact: false })); + await expect(input).toBeVisible({ timeout: 15_000 }); + await input.fill(message); + await sendButton(page).click(); +} diff --git a/tests/e2e/ui/helpers/roundTrip.ts b/tests/e2e/ui/helpers/roundTrip.ts new file mode 100644 index 00000000000..8d6e264e622 --- /dev/null +++ b/tests/e2e/ui/helpers/roundTrip.ts @@ -0,0 +1,28 @@ +import { expect, Page } from "@playwright/test"; +import { masterKey } from "./traffic"; + +/** + * Runs `action` and returns the parsed body of the first matching request. + * + * `action` is a callback so the listener is armed before the click; awaiting the + * click first lets the request go by, and the test then hangs until timeout. + */ +export async function captureRequestBody( + page: Page, + match: { method: string; urlIncludes: string }, + action: () => Promise, +): Promise> { + const pending = page.waitForRequest((req) => req.method() === match.method && req.url().includes(match.urlIncludes)); + await action(); + const request = await pending; + return JSON.parse(request.postData() ?? "{}") as Record; +} + +/** Reads an endpoint as the master key, so a failure is bad data and not an expired UI token. */ +export async function readBack(page: Page, endpoint: string): Promise { + const res = await page.request.get(endpoint, { + headers: { Authorization: `Bearer ${masterKey()}` }, + }); + expect(res.ok(), `GET ${endpoint}`).toBe(true); + return (await res.json()) as T; +} diff --git a/tests/e2e/ui/helpers/traffic.ts b/tests/e2e/ui/helpers/traffic.ts new file mode 100644 index 00000000000..a2fc9463c94 --- /dev/null +++ b/tests/e2e/ui/helpers/traffic.ts @@ -0,0 +1,125 @@ +import { APIRequestContext, expect } from "@playwright/test"; + +/** Model names served by fixtures/config.yml, both backed by the mock LLM server. */ +export const CHAT_MODEL_A = "fake-openai-gpt-4"; +export const CHAT_MODEL_B = "fake-anthropic-claude"; + +/** The only completion text fixtures/mock_llm_server/server.py ever returns. */ +export const MOCK_RESPONSE_TEXT = "This is a mock response."; + +export const masterKey = (): string => process.env.LITELLM_MASTER_KEY || "sk-1234"; + +const rootPath = (): string => process.env.SERVER_ROOT_PATH ?? ""; + +interface ChatOptions { + model: string; + prompt: string; + apiKey?: string; + /** Sent as `user`, which lands in the spend log's end_user column. */ + endUser?: string; +} + +/** POST /v1/chat/completions and return the completion id (the Logs Request ID). */ +export async function sendChatCompletion(request: APIRequestContext, opts: ChatOptions): Promise { + const res = await request.post(`${rootPath()}/v1/chat/completions`, { + headers: { + Authorization: `Bearer ${opts.apiKey ?? masterKey()}`, + "Content-Type": "application/json", + }, + data: { + model: opts.model, + messages: [{ role: "user", content: opts.prompt }], + ...(opts.endUser ? { user: opts.endUser } : {}), + }, + }); + expect(res.ok(), `chat completion for ${opts.model} failed (${res.status()}): ${await res.text()}`).toBe(true); + const body = await res.json(); + expect(body.choices?.[0]?.message?.content).toContain(MOCK_RESPONSE_TEXT); + return body.id as string; +} + +/** `key` is the sk- value to authenticate with; `token` is its hash, which spend aggregates are keyed by. */ +export async function createVirtualKey( + request: APIRequestContext, + data: Record = {}, +): Promise<{ key: string; token: string; alias?: string }> { + const res = await request.post(`${rootPath()}/key/generate`, { + headers: { + Authorization: `Bearer ${masterKey()}`, + "Content-Type": "application/json", + }, + data, + }); + expect(res.ok(), `key generate failed (${res.status()}): ${await res.text()}`).toBe(true); + const body = await res.json(); + return { + key: body.key as string, + token: (body.token ?? body.token_id) as string, + alias: body.key_alias as string | undefined, + }; +} + +/** Spend logs are flushed on a timer, so an assertion straight after a completion races the writer. */ +export async function waitForSpendLog( + request: APIRequestContext, + requestId: string, + timeoutMs = 60_000, +): Promise { + const deadline = Date.now() + timeoutMs; + let lastStatus = 0; + while (Date.now() < deadline) { + const res = await request.get(`${rootPath()}/spend/logs?request_id=${encodeURIComponent(requestId)}`, { + headers: { Authorization: `Bearer ${masterKey()}` }, + }); + lastStatus = res.status(); + if (res.ok()) { + const body = await res.json(); + const rows = Array.isArray(body) ? body : (body?.data ?? []); + if (rows.length > 0) { + return; + } + } + await new Promise((r) => setTimeout(r, 2_000)); + } + throw new Error(`spend log for request ${requestId} never appeared (last /spend/logs status ${lastStatus})`); +} + +const isoDay = (d: Date): string => d.toISOString().slice(0, 10); + +/** + * The Usage page reads /user/daily/activity, a rollup written by a background job, and fetches it once + * on mount. Navigating before the rollup lands leaves a stale render that never refreshes. + */ +export async function waitForKeyInDailyActivity( + request: APIRequestContext, + keyToken: string, + timeoutMs = 120_000, +): Promise { + const now = new Date(); + const start = new Date(now); + start.setDate(start.getDate() - 7); + const query = `start_date=${isoDay(start)}&end_date=${isoDay(now)}`; + + const deadline = Date.now() + timeoutMs; + let lastStatus = 0; + while (Date.now() < deadline) { + const res = await request.get(`${rootPath()}/user/daily/activity?${query}`, { + headers: { Authorization: `Bearer ${masterKey()}` }, + }); + lastStatus = res.status(); + if (res.ok()) { + const body = await res.json(); + const seen = (body?.results ?? []).some( + (day: { breakdown?: { api_keys?: Record } }) => keyToken in (day.breakdown?.api_keys ?? {}), + ); + if (seen) { + return; + } + } + await new Promise((r) => setTimeout(r, 3_000)); + } + throw new Error( + `key ${keyToken} never appeared in /user/daily/activity (last status ${lastStatus}); ` + + "the daily spend rollup may not be running", + ); +} diff --git a/tests/e2e/ui/run_e2e.sh b/tests/e2e/ui/run_e2e.sh index 858eb401c8e..67e3225f668 100755 --- a/tests/e2e/ui/run_e2e.sh +++ b/tests/e2e/ui/run_e2e.sh @@ -12,6 +12,10 @@ set -euo pipefail # ./run_e2e.sh --repeat-each=5 # Run each test 5 times # ./run_e2e.sh --headed # Run with browser visible # +# Ports default to 4000 / 5432 / 8090 and can be moved when another checkout +# already holds them: +# PROXY_PORT=4100 POSTGRES_PORT=5532 MOCK_LLM_PORT=8190 ./run_e2e.sh +# # In CI (CI=true), expects: # - PostgreSQL already running on 127.0.0.1:5432 # - DATABASE_URL already set @@ -28,12 +32,50 @@ MOCK_PID="" PROXY_PID="" PROXY_LOG="" +# Ports, overridable so two checkouts can run this harness at the same time -- +# otherwise a second run aborts on "port 4000 is in use" and the only way out is +# to stop someone else's stack. Defaults are the historical values, so an unset +# environment behaves exactly as before (CI, the CircleCI job and the docs all +# assume 4000/5432/8090). +PROXY_PORT="${PROXY_PORT:-4000}" +POSTGRES_PORT="${POSTGRES_PORT:-5432}" +MOCK_LLM_PORT="${MOCK_LLM_PORT:-8090}" +export MOCK_LLM_PORT + # --- Ensure common tool paths are available (local dev only) --- if [ "$IS_CI" = "false" ]; then for p in /usr/local/bin /opt/homebrew/bin "$HOME/.local/bin" /opt/homebrew/opt/postgresql@14/bin /opt/homebrew/opt/libpq/bin; do [ -d "$p" ] && export PATH="$p:$PATH" done - [ -s "$HOME/.nvm/nvm.sh" ] && source "$HOME/.nvm/nvm.sh" + # Sourcing nvm only makes `nvm` available -- it leaves you on whatever the + # default alias points at, which is frequently an older Node than the + # dashboard's engines allow. `npm install` then fails EBADENGINE, npm exits + # non-zero, and because the install below is `--silent ... || true` the error + # is swallowed and the run dies later with the far less obvious + # "sh: next: command not found". + # + # So select a Node that satisfies ui/litellm-dashboard's engines.node, and if + # none is available say so here rather than 200 lines downstream. + if [ -s "$HOME/.nvm/nvm.sh" ]; then + # shellcheck disable=SC1091 + source "$HOME/.nvm/nvm.sh" + required_major="$(sed -nE 's/.*"node"[[:space:]]*:[[:space:]]*">=?([0-9]+).*/\1/p' \ + "$DASHBOARD_DIR/package.json" 2>/dev/null | head -1)" + if [ -n "$required_major" ]; then + current_major="$(node --version 2>/dev/null | sed -E 's/^v([0-9]+).*/\1/')" + if [ -z "$current_major" ] || [ "$current_major" -lt "$required_major" ]; then + echo "Node $(node --version 2>/dev/null || echo 'not found') is below the dashboard's required v${required_major}; selecting a newer one via nvm" + nvm use "$required_major" >/dev/null 2>&1 || nvm use --lts >/dev/null 2>&1 || true + current_major="$(node --version 2>/dev/null | sed -E 's/^v([0-9]+).*/\1/')" + if [ -z "$current_major" ] || [ "$current_major" -lt "$required_major" ]; then + echo "Error: ui/litellm-dashboard requires Node >= v${required_major}, and no such version is installed." + echo " Install one with: nvm install ${required_major}" + exit 1 + fi + fi + echo "Using Node $(node --version) / npm $(npm --version)" + fi + fi fi # --- Cleanup on exit --- @@ -47,7 +89,11 @@ cleanup() { fi echo "Done." } -trap cleanup EXIT INT TERM +on_signal() { + exit 130 +} +trap cleanup EXIT +trap on_signal INT TERM # --- Pre-flight checks --- for cmd in python3 npx uv; do @@ -59,9 +105,14 @@ if [ "$IS_CI" = "false" ]; then for cmd in docker psql; do command -v "$cmd" >/dev/null 2>&1 || { echo "Error: $cmd not found."; exit 1; } done - for port in 4000 5432 8090; do - if lsof -ti ":$port" >/dev/null 2>&1; then - echo "Error: port $port is in use" + # Only a LISTENER conflicts with us. Without -sTCP:LISTEN this also matches + # ESTABLISHED sockets, so an unrelated *outbound* connection from this machine + # to someone else's :5432 (a psql session, a running app, a Prisma engine + # talking to a remote database) aborts the run with "port 5432 is in use" + # while nothing is actually bound locally. + for port in "$PROXY_PORT" "$POSTGRES_PORT" "$MOCK_LLM_PORT"; do + if lsof -nP -iTCP:"$port" -sTCP:LISTEN >/dev/null 2>&1; then + echo "Error: port $port is in use (override with PROXY_PORT / POSTGRES_PORT / MOCK_LLM_PORT)" exit 1 fi done @@ -69,12 +120,12 @@ if [ "$IS_CI" = "false" ]; then export POSTGRES_USER="e2euser" export POSTGRES_PASSWORD="$(openssl rand -hex 32)" export POSTGRES_DB="litellm_e2e" - export DATABASE_URL="postgresql://${POSTGRES_USER}:${POSTGRES_PASSWORD}@127.0.0.1:5432/${POSTGRES_DB}" + export DATABASE_URL="postgresql://${POSTGRES_USER}:${POSTGRES_PASSWORD}@127.0.0.1:${POSTGRES_PORT}/${POSTGRES_DB}" echo "=== Starting PostgreSQL ===" docker run -d --rm --name "$CONTAINER_NAME" \ -e POSTGRES_USER -e POSTGRES_PASSWORD -e POSTGRES_DB \ - -p 127.0.0.1:5432:5432 \ + -p "127.0.0.1:${POSTGRES_PORT}:5432" \ postgres:16 echo "Waiting for PostgreSQL..." @@ -91,8 +142,13 @@ fi # --- Credentials --- export LITELLM_MASTER_KEY="sk-1234" -export MOCK_LLM_URL="http://127.0.0.1:8090/v1" +export MOCK_LLM_URL="http://127.0.0.1:${MOCK_LLM_PORT}/v1" export DISABLE_SCHEMA_UPDATE="true" +# The suite resolves its target from E2E_UI_BASE_URL (constants.ts), which +# otherwise defaults to :4000 -- so without this a relocated stack would be +# built and booted correctly and then tested against whatever happens to be +# listening on the default port. +export E2E_UI_BASE_URL="${E2E_UI_BASE_URL:-http://127.0.0.1:${PROXY_PORT}}" # Ensure the proxy serves UI at /ui (not behind a subpath) export SERVER_ROOT_PATH="" # Boot with an external logout URL so proxyLogoutUrl.spec.ts can assert the @@ -108,7 +164,11 @@ export LITELLM_LICENSE="${LITELLM_LICENSE:-}" # --- Rebuild UI from source --- echo "=== Building UI from source ===" cd "$DASHBOARD_DIR" -npm install --silent 2>/dev/null || true +# NOT silenced, and NOT `|| true`. Swallowing this is what turns a one-line +# EBADENGINE ("dashboard requires node >=24, you have v20") into the +# considerably less helpful "sh: next: command not found" from the build below, +# because the deps that provide `next` were never installed. +npm install npm run build # Copy the fresh build to the proxy's static UI directory cp -r "$DASHBOARD_DIR/out/" "$REPO_ROOT/litellm/proxy/_experimental/out/" @@ -139,7 +199,7 @@ uv run --no-sync python "$SCRIPT_DIR/fixtures/mock_llm_server/server.py" & MOCK_PID=$! for i in $(seq 1 15); do - if curl -sf http://127.0.0.1:8090/health >/dev/null 2>&1; then break; fi + if curl -sf http://127.0.0.1:${MOCK_LLM_PORT}/health >/dev/null 2>&1; then break; fi sleep 1 done @@ -149,7 +209,7 @@ cd "$REPO_ROOT" PROXY_LOG="${TMPDIR:-/tmp}/litellm-e2e-proxy-$$.log" uv run --no-sync python -m litellm.proxy.proxy_cli \ --config "$SCRIPT_DIR/fixtures/config.yml" \ - --port 4000 >"$PROXY_LOG" 2>&1 & + --port "$PROXY_PORT" >"$PROXY_LOG" 2>&1 & PROXY_PID=$! echo "Waiting for proxy (logs: $PROXY_LOG)..." @@ -160,7 +220,7 @@ for i in $(seq 1 180); do tail -n 100 "$PROXY_LOG" exit 1 fi - HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" http://127.0.0.1:4000/health -H "Authorization: Bearer $LITELLM_MASTER_KEY" 2>/dev/null || true) + HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" http://127.0.0.1:${PROXY_PORT}/health -H "Authorization: Bearer $LITELLM_MASTER_KEY" 2>/dev/null || true) if [ "$HTTP_CODE" = "200" ]; then PROXY_READY=1 break @@ -188,9 +248,38 @@ PGPASSWORD="$DB_PASS" psql -h "$DB_HOST" -p "$DB_PORT" -U "$DB_USER" -d "$DB_NAM # --- Playwright --- echo "=== Installing Playwright dependencies ===" cd "$SCRIPT_DIR" -npm install --silent 2>/dev/null || true +# Same reasoning as the dashboard install above: a failure here means the suite +# has no @playwright/test, and the run should say that rather than fail later. +npm install npx playwright install chromium --with-deps 2>/dev/null || npx playwright install chromium +# Authoring a new spec means running it over and over against a stack that is +# already up -- rebuilding the UI and re-seeding for every iteration costs +# minutes each time. E2E_KEEP_ALIVE brings the stack up, then blocks, so you can +# run `npx playwright test ` yourself from another shell against it. +# Ctrl-C here tears everything down through the usual trap. +if [ "${E2E_KEEP_ALIVE:-0}" = "1" ]; then + cat < + +Press Ctrl-C to tear the stack down. +EOF + while kill -0 "$PROXY_PID" 2>/dev/null; do + sleep 5 + done + echo "Error: proxy process exited unexpectedly. Proxy output:" + tail -n 100 "$PROXY_LOG" + exit 1 +fi + echo "=== Running Playwright tests ===" npx playwright test --config playwright.config.ts "$@" EXIT_CODE=$? diff --git a/tests/e2e/ui/tests/logs/logs.spec.ts b/tests/e2e/ui/tests/logs/logs.spec.ts new file mode 100644 index 00000000000..fc5cce53511 --- /dev/null +++ b/tests/e2e/ui/tests/logs/logs.spec.ts @@ -0,0 +1,221 @@ +import { test, expect, type Locator, type Page as PlaywrightPage } from "@playwright/test"; +import { ADMIN_STORAGE_PATH } from "../../constants"; +import { navigateToPage, dismissFeedbackPopup } from "../../helpers/navigation"; +import { Page } from "../../fixtures/pages"; +import { CHAT_MODEL_A, MOCK_RESPONSE_TEXT, sendChatCompletion, waitForSpendLog } from "../../helpers/traffic"; + +/** + * Anchored to traffic this spec generates itself, with a unique prompt and end user per run, so it + * neither depends on seeded spend rows nor collides with other specs under parallelism. + */ + +const uniqueSuffix = (): string => `${Date.now()}-${Math.random().toString(36).slice(2, 8)}`; + +/** + * Walking up from the label is the only stable handle: the header carries no role, test id or class, + * and its copy button is icon-only with a hover-only tooltip. + */ +const sectionHeader = (drawer: Locator, label: "Input" | "Output"): Locator => + drawer.getByText(label, { exact: true }).locator("xpath=../../.."); + +/** Every tab stays mounted, so the DOM holds four tables at once; scope to the visible one. */ +const requestLogsRows = (page: PlaywrightPage): Locator => + page.locator("table").filter({ visible: true }).first().locator("tbody tr"); + +const visibleTestId = (page: PlaywrightPage, id: string): Locator => page.getByTestId(id).filter({ visible: true }); + +/** Open the Logs page and filter the table down to a single request id. */ +async function openLogsForRequest(page: PlaywrightPage, requestId: string): Promise { + await navigateToPage(page, Page.Logs); + await dismissFeedbackPopup(page); + + const search = visibleTestId(page, "datatable-search"); + await expect(search).toBeVisible({ timeout: 20_000 }); + await search.fill(requestId); + + const row = requestLogsRows(page).filter({ hasText: requestId }); + await expect(row, `no logs row for request ${requestId}`).toHaveCount(1, { + timeout: 30_000, + }); + return row; +} + +test.describe("Logs page", () => { + test.use({ + storageState: ADMIN_STORAGE_PATH, + // The copy buttons go through navigator.clipboard, which rejects without these. + permissions: ["clipboard-read", "clipboard-write"], + }); + + test("a served request expands to its request and response", async ({ page, request }) => { + const prompt = `logs-detail-prompt-${uniqueSuffix()}`; + const requestId = await sendChatCompletion(request, { + model: CHAT_MODEL_A, + prompt, + }); + await waitForSpendLog(request, requestId); + + const row = await openLogsForRequest(page, requestId); + + // Expand: clicking the row opens the detail drawer for that request. + await row.click(); + const drawer = page.locator(".ant-drawer-content").first(); + await expect(drawer).toBeVisible({ timeout: 20_000 }); + await expect(drawer.getByText("Request & Response")).toBeVisible({ + timeout: 20_000, + }); + + // The prompt we sent and the mock server's reply are both rendered. + await expect(drawer.getByText(prompt, { exact: false })).toBeVisible({ + timeout: 20_000, + }); + await expect(drawer.getByText(MOCK_RESPONSE_TEXT, { exact: false }).first()).toBeVisible({ timeout: 20_000 }); + }); + + // Split out because only the copy path needs a secure context; folding it in would + // take the drawer-rendering coverage down with it. + test("the drawer copies the request and the response to the clipboard", async ({ page, request }) => { + // `navigator.clipboard` is undefined outside a secure context, and handleCopy calls + // writeText unguarded, so on plain HTTP served from a hostname the click throws and no + // toast renders. Skipped rather than weakened so the product gap stays visible. + await page.goto("/ui"); + const isSecure = await page.evaluate(() => window.isSecureContext); + test.skip(!isSecure, "origin is not a secure context, so navigator.clipboard is unavailable"); + + const prompt = `logs-copy-prompt-${uniqueSuffix()}`; + const requestId = await sendChatCompletion(request, { + model: CHAT_MODEL_A, + prompt, + }); + await waitForSpendLog(request, requestId); + + const row = await openLogsForRequest(page, requestId); + await row.click(); + const drawer = page.locator(".ant-drawer-content").first(); + await expect(drawer).toBeVisible({ timeout: 20_000 }); + + // Copy request: the Input card's copy button puts the prompt on the clipboard. + await sectionHeader(drawer, "Input").getByRole("button").click(); + await expect(page.getByText("Input copied")).toBeVisible({ + timeout: 10_000, + }); + expect(await page.evaluate(() => navigator.clipboard.readText())).toContain(prompt); + + // Copy response: the Output card's copy button puts the completion on it. + await sectionHeader(drawer, "Output").getByRole("button").click(); + await expect(page.getByText("Output copied")).toBeVisible({ + timeout: 10_000, + }); + expect(await page.evaluate(() => navigator.clipboard.readText())).toContain(MOCK_RESPONSE_TEXT); + }); + + test("the Input card collapses and expands", async ({ page, request }) => { + const prompt = `logs-collapse-prompt-${uniqueSuffix()}`; + const requestId = await sendChatCompletion(request, { + model: CHAT_MODEL_A, + prompt, + }); + await waitForSpendLog(request, requestId); + + const row = await openLogsForRequest(page, requestId); + await row.click(); + + const drawer = page.locator(".ant-drawer-content").first(); + await expect(drawer.getByText("Request & Response")).toBeVisible({ + timeout: 20_000, + }); + + // The body collapses via `max-height: 0; overflow: hidden`, which zeroes its own bounding + // box, so the wrapper reads as hidden while the clipped text node inside it does not. + const header = sectionHeader(drawer, "Input"); + const body = header.locator("xpath=following-sibling::div[1]"); + await expect(header.locator(".anticon-up")).toBeVisible(); + await expect(body).toBeVisible(); + + await header.click(); + await expect(header.locator(".anticon-down")).toBeVisible({ + timeout: 10_000, + }); + await expect(body).toBeHidden({ timeout: 10_000 }); + + await header.click(); + await expect(header.locator(".anticon-up")).toBeVisible({ + timeout: 10_000, + }); + await expect(body).toBeVisible({ timeout: 10_000 }); + await expect(drawer.getByText(prompt, { exact: false })).toBeVisible({ + timeout: 10_000, + }); + }); + + test("the JSON view exposes Request and Response tabs", async ({ page, request }) => { + const prompt = `logs-json-prompt-${uniqueSuffix()}`; + const requestId = await sendChatCompletion(request, { + model: CHAT_MODEL_A, + prompt, + }); + await waitForSpendLog(request, requestId); + + const row = await openLogsForRequest(page, requestId); + await row.click(); + + const drawer = page.locator(".ant-drawer-content").first(); + await expect(drawer.getByText("Request & Response")).toBeVisible({ + timeout: 20_000, + }); + + // antd Radio.Button hides the under its