mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-10-10 03:27:59 +00:00
* fix(test): widen worker pool retry timeout to prevent flake under load The "replaces a timed-out worker" test used 150ms idle timeout (600ms retry), which is too tight when CPU is contended during parallel test runs. Increase to 500ms (2s retry) — the test exercises the retry mechanism, not tight timing. Closes #1323 * fix(pool): wait for replacement worker to come online before dispatching Root cause: replaceWorker() spawned a new Worker but returned immediately without waiting for the thread to start. The subsequent runWorker() call started the idle timer and posted the sub-batch while the thread was still booting. Under CPU contention, thread startup latency consumed most of the retry timeout budget, causing the flake. Wait for the 'online' event before assigning the replacement worker. This ensures the idle timeout measures actual processing time, not thread startup overhead. Reverts the test timeout widening (500ms→150ms) since the root cause is now addressed. No production performance regression was found — the 30s default timeout is unaffected. Only the tight test timeouts were sensitive to startup latency. * fix(pool): harden replacement worker startup with three-event helper Address review feedback on the waitForWorkerOnline implementation: 1. Add waitForWorkerOnline helper that listens for 'online', 'error', and 'exit' events with proper cleanup after settlement. Prevents the dispatch promise from hanging if a replacement worker crashes before coming online (e.g. OOM, native addon failure). 2. Wrap replaceWorker call site in try/catch that routes failures through fail() — prevents unhandled promise rejections in the async setTimeout callback. 3. Re-check stopped flag after awaiting replacement startup — prevents injecting a live worker into a pool that was stopped by a concurrent failure during the await window. Terminates the orphaned replacement. 4. Add integration test for replacement worker crash during startup: worker throws on second load (marker-file gated), verifying the pool rejects the dispatch instead of hanging. * fix(pool): preserve original error in replacement worker catch The bare catch{} discarded the original error from waitForWorkerOnline, causing the startup-crash test regex to miss. Bind the error and include its message in the re-thrown Error. |
||
|---|---|---|
| .. | ||
| fixtures | ||
| helpers | ||
| integration | ||
| unit | ||
| utils | ||