mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-03 02:21:44 +00:00
* feat: shared resilient-fetch (retries + circuit breaker)
Add a small, runtime-agnostic resilience layer in gitnexus-shared and
migrate every backend HTTP outbound call (CLI, MCP, wiki LLM, web → backend)
through it.
Helpers (gitnexus-shared/src/integrations/):
- retry.ts — withRetry(fn, opts) with caller-supplied
retryability classification and full-jitter
exponential backoff.
- circuit-breaker.ts — closed/open/half-open per-process breaker with
injectable clock, plus a keyed registry so
callers targeting the same endpoint share state.
- resilient-fetch.ts — composed wrapper: retries 5xx + 429 + retryable
network throws, treats AbortSignal.timeout()
and 4xx (other than 429) as terminal, honors
Retry-After (capped at 30s), throws
CircuitOpenError when the breaker opens.
Migrations (no behaviour regression — all existing tests pass):
- gitnexus/src/core/embeddings/http-client.ts (covers analyze + MCP
query path) — replaces inline linear-backoff retry.
- gitnexus/src/core/wiki/llm-client.ts — preserves Azure content-filter
branch; resilientFetch handles 5xx/429.
- gitnexus-web/src/services/backend-client.ts (fetchWithTimeout helper)
— small retry budget (2 attempts, 250–1500 ms) so a dead local
backend still fails fast for the user.
- gitnexus-web/src/core/llm/settings-service.ts (OpenRouter model list).
Deliberately not migrated:
- gitnexus-web/src/services/backend-client.ts streamJob() — Server-Sent
Events stream; the existing reconnect-with-Last-Event-ID logic is
not unary-fetch shaped.
- gitnexus-web/src/components/SettingsPanel.tsx checkOllamaStatus() —
one-shot health probe; retrying delays the "Ollama not running"
error rather than improving UX.
41 new helper tests cover backoff math, breaker state transitions,
Retry-After parsing (delta-seconds + HTTP-date), 401/422 terminal
classification, and breaker fail-fast on three exhausted retry batches.
* fix(review): apply autofix feedback
Address Claude's two MEDIUM blocking findings on PR #1448 plus the
CodeQL SSRF false-positive flag.
- backend-client `fetchWithTimeout` now uses `AbortSignal.timeout()`
merged with the caller's signal via `AbortSignal.any()`. Timer-fired
aborts surface as `DOMException(name='TimeoutError')` so
resilientFetch routes them through the terminal-network branch
(no retry, no breaker hit), instead of incrementing the breaker
for user-side network slowness.
- Method-aware retry budget in `fetchWithTimeout`: idempotent verbs
(GET/HEAD/OPTIONS) keep the 2-attempt budget; POST/PATCH/PUT/DELETE
default to single-attempt so a 5xx on `startAnalyze` cannot start
a duplicate job. New `forceRetry` parameter for callers that
know-idempotent mutations (e.g. DELETE of a known-deleted resource).
- `resilient-fetch.ts` carries a documented suppression for CodeQL
js/server-side-request-forgery on the inner fetch call. Every
concrete caller passes a hardcoded URL constant or a value from
configuration (env vars, saved settings); user request input never
flows into the URL parameter.
- New test file `backend-client-retry.test.ts` covers all three
paths: GET retries on 503, POST does not retry, timeout does not
increment the breaker.
* fix(resilient-fetch): address Codex adversarial findings
Closes the three blocking issues from Codex's review on PR #1448.
U1 — Add `recordNeutral()` to CircuitBreaker.
Third outcome path that's an explicit no-op for state and the
consecutive-failure counter. Distinct from `recordSuccess` (closes
the breaker) and `recordFailure` (may open it). Used for outcomes
that are neither evidence of backend health nor evidence of
backend failure.
U2 — Route terminal-client / terminal-network through `recordNeutral`.
Previously a 401 or local timeout called `recordSuccess`, which
reset `consecutiveFailures` to 0. A 5xx → 401 → 5xx → 401 → 5xx
sequence would NEVER trip the breaker because each 4xx in between
erased the running count. Also classify external `AbortError` as
terminal-network (was retryable-network), so caller-driven
cancellation no longer retries against an already-aborted signal
or counts toward breaker failures on exhaustion.
U3 — Per-origin breaker key in web `fetchWithTimeout`.
Was hardcoded to `'web-backend'` even though `_backendUrl` is
mutable via `setBackendUrl`. Switching backend URLs after a
circuit tripped on host-A would strand the user during the full
cooldown. Key is now `web-backend:<origin>`, so each backend URL
gets its own breaker state.
Tests: +5 recordNeutral, +4 resilient-fetch (interleaved 4xx/5xx,
external AbortError, prior-state preservation), +1 web switch-backend
regression. All 70 gitnexus integration tests + 15 web tests green.
* fix(resilient-fetch): tolerate header-less fetch mocks on 429
`classifyOutcome` called `resp.headers.get('Retry-After')` directly,
which crashed when a test stubs `fetch` with a plain object like
`{ ok: false, status: 429 }` (no `headers` field). Real `Response`
always has Headers, so this surfaces only in test setups, but the
helper has no business assuming caller-side correctness on this — the
defensive guard is cheap and a missing `Retry-After` falls through to
exponential-backoff retry like any 429 without the header.
Surfaced by `gitnexus/test/unit/http-embedder.test.ts > retries on
rate limit`, which the embeddings migration exercises against a
plain-object 429 stub. Locked in with a new
`classifies 429 from a header-less fetch mock without throwing` case.
* fix(review): apply autofix feedback
Closes findings from the third multi-agent review pass on PR #1448.
#1 (P1) callLLM had no per-attempt timeout
Wiki LLM calls passed no `signal` to resilientFetch; each of three
retry attempts could hang indefinitely on a frozen TCP connection.
Add `signal: AbortSignal.timeout(60_000)` so the per-attempt budget
matches what http-client.ts and backend-client.ts already provide.
#2 (P2) drop dead `lastRetryableResp` post-loop fallback
Variable was set in one switch arm but only read in unreachable code
after the loop. The retry loop always returns/throws on every
iteration. Keep only the defensive `throw` so TypeScript's
control-flow analysis still sees `Promise<Response>` as the return.
#5 (P2) gate test-only exports behind a subpath
`__resetBreakerRegistry__` and `classifyOutcome` were reachable from
the main `gitnexus-shared` barrel — production code calling
`__resetBreakerRegistry__` from a tool implementation would silently
nuke every circuit breaker process-wide. Move to a new
`gitnexus-shared/test-helpers` subpath export. Production callers
see the cleaner public API; tests import via the explicit
`gitnexus-shared/test-helpers` path.
#6 (P2) exhaustiveness guard on Outcome switch
Add a `default: const _: never = outcome` arm so a future sixth
`Outcome.kind` won't compile silently — it'll surface at the switch
site rather than fall through to a retry/no-retry default.
#9 (P3) document cumulative wall-clock budget
Add a "Cumulative wall-clock budget" paragraph to resilientFetch's
JSDoc explaining the worst-case total wait (`maxAttempts × (per-attempt
timeout + capDelayMs)` ≈ 60s with defaults) and pointing callers at
outer `AbortSignal.timeout()` when they want a tighter bound.
Deferred to follow-up PRs (per review's Auto-resolve recommendation):
- #3 idempotency knob to shared API (forceRetry into ResilientFetchOptions)
- #4 publish.ts migration to resilientFetch
- #7 parseRetryAfter past-HTTP-date / negative-seconds asymmetry
- #8 recordNeutral counter time-decay (documented breaker semantic)
* fix(circuit-breaker): gate half-open to a single in-flight probe
Closes the Codex adversarial-review finding on PR #1448 that flagged a
recovery-time thundering herd: when cooldown expired, every concurrent
caller transitioned the breaker to half-open and probed the still-
recovering dependency in lockstep, defeating the breaker's "fail fast"
promise.
U1 — probe-permit gate in CircuitBreaker.check()
Added a `probeInFlight: boolean` field. After cooldown expires, the
first `check()` admits the probe and consumes the permit; subsequent
callers throw `CircuitOpenError` with a configurable
`halfOpenRetryAfterMs` (default 1000ms) until the probe resolves.
Critical design point: `recordNeutral` now RELEASES the permit but
does NOT transition state. Without that split, a single `TimeoutError`
from per-attempt `AbortSignal.timeout` (which routes through neutral
classification) would permanently park the breaker in half-open. By
separating permit-release from state-resolution, we keep the
"neutral doesn't claim health" semantic without creating that wedge.
Other changes:
- `halfOpenRetryAfterMs` is now a constructor option for consumers
with long-running protected ops (LLM streaming, large uploads).
- `getState()` is documented as a pure read; the implicit
Open -> Half-Open transition lives in `check()` only, so tests
that inspect state never inadvertently consume a probe permit.
- `isProbeInFlight()` test-only accessor for assertion clarity.
- JSDoc on `check()` records the JS event-loop atomicity dependency
and the load-bearing `try/finally` pairing invariant.
U2 — End-to-end concurrency regression through resilientFetch
Three new scenarios in resilient-fetch.test.ts (26 -> 29):
- 3 concurrent calls + probe gets 200 -> 1 hits fetch, 2 throw
CircuitOpenError, breaker closes.
- 3 concurrent calls + probe gets 503 -> ResilientFetchExhaustedError
on probe; concurrent callers see halfOpenRetryAfterMs (1000ms);
fresh caller after probe resolves sees the FULL new cooldown
(10000ms), not the probe-in-flight default.
- Probe cancelled mid-flight via AbortError -> permit released,
state stays half-open, next caller becomes the new probe and
succeeds.
Plus 9 new circuit-breaker unit tests (16 -> 25) covering the permit
gate, recordNeutral-releases-permit semantic, fresh-cooldown distinction,
default vs configurable halfOpenRetryAfterMs, getState() purity, and
the three-probes-via-neutrals chain.
Total integration test count: 70 -> 82. All 106 gitnexus + 15 web
tests pass; both packages typecheck.
Maintainer decisions (deferred per plan 003 Open Questions):
- Plan 002's deferral judgement was reversed on Codex's argument
without new measurement / incident data. The reversal is defensible
on principle (Hystrix / Resilience4j alignment) but lacks workload-
driven evidence.
- Probe-blocked callers throw silently (no log / event hook). R4's
"no new public API" prevents adding observability; loosen if a
debug log on probe-blocked is wanted.
* refactor(embeddings): replace bespoke HF breaker with shared CircuitBreaker
Deleted the local `HfDownloadCircuitBreaker` class and the manual
retry loop in `withHfDownloadRetry`. Both are now backed by the
shared `gitnexus-shared` primitives:
- `hfDownloadCircuit` is `new CircuitBreaker({ failureThreshold,
cooldownMs, key: 'hf-download' })` — same state machine as before
PLUS the single-permit half-open gate that prevents recovery-time
stampedes when CLI + MCP embedders concurrently re-load the model.
- `withHfDownloadRetry` delegates the loop to `withRetry` from the
shared package. Per-attempt timeout (`withDownloadTimeout`),
network-vs-non-network classification, circuit recording, and the
`onRetry` callback wire through `withRetry`'s `isRetryable`
callback.
Behaviour preserved:
- Pre-flight `CIRCUIT_OPEN_TAG` rejection when the breaker is open.
- Mid-loop `CIRCUIT_OPEN_TAG` "opened after N consecutive failures"
when a network error trips the threshold.
- Non-network errors (e.g. CUDA unavailable) bypass retry and go
through `recordNeutral` instead of resetting the breaker's
failure-count progress.
- `onRetry(attempt+1, max, err)` fires only when there's a next
attempt, matching the prior semantic.
Generic CircuitBreaker gained two inspection accessors:
- `getOpenedAt(): number | null`
- `getCooldownMs(): number`
Used by `withHfDownloadRetry` to compute `secsUntilReset` without
consuming a probe permit (which `check()` would do).
Test consolidation: the 7 bespoke `HfDownloadCircuitBreaker`
state-machine tests in hf-env.test.ts were 1:1 duplicates of
existing tests in `circuit-breaker.test.ts` and were deleted.
Remaining 42 hf-env tests all pass; full integration sweep (148
gitnexus + 15 web) green.
395 lines
14 KiB
TypeScript
395 lines
14 KiB
TypeScript
import { describe, it, expect, beforeEach } from 'vitest';
|
|
import { CircuitBreaker, CircuitOpenError, getBreaker } from 'gitnexus-shared';
|
|
import { __resetBreakerRegistry__ } from 'gitnexus-shared/test-helpers';
|
|
|
|
describe('CircuitBreaker', () => {
|
|
beforeEach(() => __resetBreakerRegistry__());
|
|
|
|
function makeClock(start = 1_700_000_000_000) {
|
|
let t = start;
|
|
return {
|
|
now: () => t,
|
|
advance: (ms: number) => {
|
|
t += ms;
|
|
},
|
|
};
|
|
}
|
|
|
|
it('runs through check/recordSuccess in closed state', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 3, cooldownMs: 30_000, now: clock.now });
|
|
expect(b.getState()).toBe('closed');
|
|
b.check(); // does not throw
|
|
b.recordSuccess();
|
|
expect(b.getState()).toBe('closed');
|
|
expect(b.getConsecutiveFailures()).toBe(0);
|
|
});
|
|
|
|
it('stays closed below the failure threshold', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 3, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
b.recordFailure();
|
|
expect(b.getState()).toBe('closed');
|
|
expect(b.getConsecutiveFailures()).toBe(2);
|
|
});
|
|
|
|
it('opens after failureThreshold consecutive failures and check throws', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 3, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
b.recordFailure();
|
|
b.recordFailure();
|
|
expect(b.getState()).toBe('open');
|
|
expect(() => b.check()).toThrow(CircuitOpenError);
|
|
});
|
|
|
|
it('CircuitOpenError.retryAfterMs decreases as time advances', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
let caught: CircuitOpenError | null = null;
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught?.retryAfterMs).toBe(30_000);
|
|
|
|
clock.advance(10_000);
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught?.retryAfterMs).toBe(20_000);
|
|
});
|
|
|
|
it('transitions Open -> Half-Open after cooldown elapses (via check)', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
expect(b.getState()).toBe('open');
|
|
clock.advance(31_000);
|
|
b.check(); // should not throw
|
|
// After check, internal state is half-open (next call probes).
|
|
expect(b.getConsecutiveFailures()).toBe(1); // unchanged until next outcome
|
|
});
|
|
|
|
it('half-open + recordSuccess -> closed and counter reset', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
b.check();
|
|
b.recordSuccess();
|
|
expect(b.getState()).toBe('closed');
|
|
expect(b.getConsecutiveFailures()).toBe(0);
|
|
});
|
|
|
|
it('half-open + recordFailure -> open with fresh openedAt', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
const firstOpen = clock.now();
|
|
clock.advance(31_000); // cooldown expired
|
|
b.check(); // half-open
|
|
b.recordFailure();
|
|
// Open with fresh timestamp — full cooldown again.
|
|
let caught: CircuitOpenError | null = null;
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught).toBeInstanceOf(CircuitOpenError);
|
|
expect(caught?.retryAfterMs).toBe(30_000);
|
|
// Sanity: not the original openedAt (would be negative remaining).
|
|
expect(clock.now()).toBeGreaterThan(firstOpen);
|
|
});
|
|
|
|
it('recordSuccess from closed state with prior partial failures resets counter', () => {
|
|
const b = new CircuitBreaker({ failureThreshold: 5 });
|
|
b.recordFailure();
|
|
b.recordFailure();
|
|
expect(b.getConsecutiveFailures()).toBe(2);
|
|
b.recordSuccess();
|
|
expect(b.getConsecutiveFailures()).toBe(0);
|
|
expect(b.getState()).toBe('closed');
|
|
});
|
|
|
|
describe('recordNeutral (U1)', () => {
|
|
it('is a no-op from closed state with zero prior failures', () => {
|
|
const b = new CircuitBreaker({ failureThreshold: 3 });
|
|
b.recordNeutral();
|
|
expect(b.getState()).toBe('closed');
|
|
expect(b.getConsecutiveFailures()).toBe(0);
|
|
});
|
|
|
|
it('preserves partial-failure progress (does not reset counter)', () => {
|
|
const b = new CircuitBreaker({ failureThreshold: 3 });
|
|
b.recordFailure();
|
|
b.recordFailure();
|
|
b.recordNeutral();
|
|
expect(b.getConsecutiveFailures()).toBe(2);
|
|
expect(b.getState()).toBe('closed');
|
|
// Real third failure still trips the breaker — neutrals didn't
|
|
// erase the running count toward the threshold.
|
|
b.recordFailure();
|
|
expect(b.getState()).toBe('open');
|
|
});
|
|
|
|
it('does not reset openedAt or transition out of open state', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
expect(b.getState()).toBe('open');
|
|
b.recordNeutral();
|
|
// Still open; cooldown clock unchanged.
|
|
expect(() => b.check()).toThrow(CircuitOpenError);
|
|
});
|
|
|
|
it('leaves half-open state alone (next true outcome decides)', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
b.check(); // half-open
|
|
b.recordNeutral();
|
|
// Still half-open; a subsequent recordFailure flips to open.
|
|
b.recordFailure();
|
|
let caught: CircuitOpenError | null = null;
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught).toBeInstanceOf(CircuitOpenError);
|
|
});
|
|
|
|
it('integration: 2 failures + 5 neutrals + 1 failure → opens on third real failure', () => {
|
|
const b = new CircuitBreaker({ failureThreshold: 3 });
|
|
b.recordFailure();
|
|
b.recordFailure();
|
|
for (let i = 0; i < 5; i++) b.recordNeutral();
|
|
expect(b.getConsecutiveFailures()).toBe(2);
|
|
expect(b.getState()).toBe('closed');
|
|
b.recordFailure();
|
|
expect(b.getState()).toBe('open');
|
|
});
|
|
});
|
|
|
|
describe('half-open probe permit gate (U1)', () => {
|
|
it('admits exactly one caller after cooldown; subsequent check() throws halfOpenRetryAfterMs', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({
|
|
failureThreshold: 1,
|
|
cooldownMs: 30_000,
|
|
halfOpenRetryAfterMs: 1_000,
|
|
now: clock.now,
|
|
});
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
|
|
// Caller A: gets the probe permit.
|
|
b.check();
|
|
expect(b.isProbeInFlight()).toBe(true);
|
|
|
|
// Caller B: blocked.
|
|
let caught: CircuitOpenError | null = null;
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught).toBeInstanceOf(CircuitOpenError);
|
|
expect(caught?.retryAfterMs).toBe(1_000);
|
|
|
|
// Caller A's recordSuccess clears the breaker.
|
|
b.recordSuccess();
|
|
expect(b.isProbeInFlight()).toBe(false);
|
|
expect(b.getState()).toBe('closed');
|
|
|
|
// Caller C: succeeds in closed state.
|
|
b.check();
|
|
expect(b.getState()).toBe('closed');
|
|
});
|
|
|
|
it('recordFailure on probe re-opens with fresh cooldown (NOT halfOpenRetryAfterMs)', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({
|
|
failureThreshold: 1,
|
|
cooldownMs: 30_000,
|
|
halfOpenRetryAfterMs: 1_000,
|
|
now: clock.now,
|
|
});
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
|
|
b.check(); // A: probe
|
|
expect(() => b.check()).toThrow(CircuitOpenError); // B: blocked
|
|
|
|
b.recordFailure(); // A reports failure → reopens with fresh openedAt
|
|
|
|
// C: should see the fresh cooldown remaining, not the probe-in-flight 1s default.
|
|
let caught: CircuitOpenError | null = null;
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught).toBeInstanceOf(CircuitOpenError);
|
|
// Fresh openedAt = current clock; cooldown is 30s; retryAfter ≈ 30s.
|
|
expect(caught?.retryAfterMs).toBe(30_000);
|
|
});
|
|
|
|
it('recordNeutral releases the probe permit but leaves state half-open', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
|
|
b.check(); // A: probe
|
|
expect(b.isProbeInFlight()).toBe(true);
|
|
|
|
b.recordNeutral(); // A: neutral — permit released, state untouched
|
|
expect(b.isProbeInFlight()).toBe(false);
|
|
expect(b.getState()).toBe('half-open');
|
|
|
|
// B: succeeds (becomes the new probe), no longer blocked.
|
|
b.check();
|
|
expect(b.isProbeInFlight()).toBe(true);
|
|
|
|
// B's recordSuccess clears the breaker.
|
|
b.recordSuccess();
|
|
expect(b.getState()).toBe('closed');
|
|
});
|
|
|
|
it('three sequential probes via neutrals: A → A.neutral → B → B.neutral → C', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
const initialFailures = b.getConsecutiveFailures();
|
|
clock.advance(31_000);
|
|
|
|
for (let i = 0; i < 3; i++) {
|
|
b.check();
|
|
b.recordNeutral();
|
|
}
|
|
// Counter unchanged; state still half-open; permit released.
|
|
expect(b.getConsecutiveFailures()).toBe(initialFailures);
|
|
expect(b.getState()).toBe('half-open');
|
|
expect(b.isProbeInFlight()).toBe(false);
|
|
});
|
|
|
|
it('5 same-tick sequential callers: exactly one passes, the other 4 throw', () => {
|
|
// `check()` is synchronous — these calls execute on a single
|
|
// microtask in declaration order. The first mutates probeInFlight
|
|
// = true; the next four observe the mutation and throw. This
|
|
// tests mutation ordering, not true concurrency (the actual
|
|
// interleaved-async-microtask scenario lives in U2).
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
|
|
const results: Array<'pass' | 'throw'> = [];
|
|
for (let i = 0; i < 5; i++) {
|
|
try {
|
|
b.check();
|
|
results.push('pass');
|
|
} catch {
|
|
results.push('throw');
|
|
}
|
|
}
|
|
expect(results.filter((r) => r === 'pass').length).toBe(1);
|
|
expect(results.filter((r) => r === 'throw').length).toBe(4);
|
|
});
|
|
|
|
it('probe permit consumed; clock advances another full cooldown without record*; still throws', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
|
|
b.check(); // probe permit consumed
|
|
clock.advance(60_000); // another full cooldown elapses, no record*
|
|
|
|
// Half-open semantics: wait for an outcome, not a timer. The
|
|
// permit-consumed state doesn't auto-resolve on time.
|
|
expect(() => b.check()).toThrow(CircuitOpenError);
|
|
});
|
|
|
|
it('halfOpenRetryAfterMs default is 1000 when not configured', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
b.check();
|
|
|
|
let caught: CircuitOpenError | null = null;
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught?.retryAfterMs).toBe(1_000);
|
|
});
|
|
|
|
it('halfOpenRetryAfterMs is configurable for long-running protected ops', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({
|
|
failureThreshold: 1,
|
|
cooldownMs: 30_000,
|
|
halfOpenRetryAfterMs: 10_000, // LLM-streaming-friendly
|
|
now: clock.now,
|
|
});
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
b.check();
|
|
|
|
let caught: CircuitOpenError | null = null;
|
|
try {
|
|
b.check();
|
|
} catch (err) {
|
|
caught = err as CircuitOpenError;
|
|
}
|
|
expect(caught?.retryAfterMs).toBe(10_000);
|
|
});
|
|
|
|
it('getState() is a pure read — does not consume the probe permit', () => {
|
|
const clock = makeClock();
|
|
const b = new CircuitBreaker({ failureThreshold: 1, cooldownMs: 30_000, now: clock.now });
|
|
b.recordFailure();
|
|
clock.advance(31_000);
|
|
|
|
// Test calls getState() to inspect — must not consume the permit.
|
|
expect(b.getState()).toBe('half-open');
|
|
expect(b.isProbeInFlight()).toBe(false);
|
|
// First check() still gets the permit.
|
|
b.check();
|
|
expect(b.isProbeInFlight()).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('getBreaker registry', () => {
|
|
it('returns the same instance for the same key', () => {
|
|
const a = getBreaker('endpoint-a');
|
|
const b = getBreaker('endpoint-a');
|
|
expect(a).toBe(b);
|
|
});
|
|
|
|
it('returns different instances for different keys', () => {
|
|
const a = getBreaker('endpoint-a');
|
|
const b = getBreaker('endpoint-b');
|
|
expect(a).not.toBe(b);
|
|
});
|
|
|
|
it('__resetBreakerRegistry__ clears all instances', () => {
|
|
const a = getBreaker('endpoint-a');
|
|
__resetBreakerRegistry__();
|
|
const a2 = getBreaker('endpoint-a');
|
|
expect(a2).not.toBe(a);
|
|
});
|
|
});
|
|
});
|