11 KiB
Testing — GitNexus
How we structure tests and which commands to run locally and in CI.
Packages
| Package | Path | Runner | Notes |
|---|---|---|---|
| CLI + MCP core | gitnexus/ |
Vitest | Primary test surface in CI |
| Web UI | gitnexus-web/ |
Vitest | Unit/component tests |
| Web UI E2E | gitnexus-web/ |
Playwright | Run when changing UI flows |
Test lanes
gitnexus/ commands
From gitnexus/:
| Command | What it runs | When to use |
|---|---|---|
npm test |
Full suite (all 3 vitest projects) | Before opening a PR |
npm run typecheck:tests |
TypeScript checks for source, tests, and helpers | Before opening a PR |
npm run test:unit |
Unit tests only (test/unit/) |
Tight development loop |
npm run test:integration |
Integration tests (test/integration/) |
After changing pipelines, DB, workers |
npm run test:coverage |
Full suite + v8 coverage with thresholds | Checking coverage impact |
npm run test:parity |
Scope-resolution parity for all migrated languages | After changing resolver or scope code |
npm run test:cross-platform |
Platform-sensitive subset only | Debugging a Windows/macOS issue |
npm run test:watch |
Vitest in watch mode | Active development |
Vitest transpiles TypeScript without checking types. Run npm run typecheck:tests
in gitnexus/ to check test code and helpers against the production types. This
uses tsc --noEmit -p tsconfig.test.json; parser input files under test/fixtures/
are excluded because they are sample source code, not part of the test program.
gitnexus-web/ commands
From gitnexus-web/:
| Command | What it runs | When to use |
|---|---|---|
npm test |
Unit/component tests (vitest) | After changing web code |
npm run test:coverage |
Unit tests + coverage | Checking coverage impact |
npm run test:e2e |
Playwright browser tests | After changing UI flows (requires gitnexus serve + npm run dev) |
Before opening a PR
# gitnexus-shared/dist must exist first. `npm install` / `npm run build` in
# gitnexus/ compiles it via parent `lib/tsc.js` (do not npm ci gitnexus-shared).
cd gitnexus && npx tsc --noEmit && npm run typecheck:tests && npm test
cd ../gitnexus-web && npx tsc -b --noEmit && npm test
Pre-commit hook
A husky pre-commit hook (.husky/pre-commit) runs automatically on every git commit:
- Formatting —
lint-stagedruns prettier on staged files gitnexus-web/files staged →tsc -b --noEmitgitnexus/files staged →tsc --noEmit
Tests do not run in the pre-commit hook — they run in CI (ci-tests.yml) only.
Skip with git commit --no-verify (use sparingly).
Vitest projects
gitnexus/vitest.config.ts defines three projects for safety isolation:
| Project | Files | Parallelism | Purpose |
|---|---|---|---|
lbug-db |
Native LadybugDB integration tests (explicit list) | Sequential | Prevents file-lock conflicts from native mmap addon |
cli-e2e |
skills-e2e.test.ts |
Sequential | CLI process spawning requires serial execution |
default |
Everything else | Parallel | Fast execution for pure logic and parser tests |
When adding a new test that uses native LadybugDB (@ladybugdb/core), add it to the lbug-db project's explicit include list and the default project's exclude list.
Test categories
- Unit — Pure logic, parsers, graph/query helpers; fast; no network.
- Integration — Real combinations (filesystem, MCP wiring, larger pipelines) as already organized under
gitnexus/test/integration. - Resolver / parity — Language-specific call-resolution tests in
test/integration/resolvers/. - E2E (web) — Critical user paths only; prefer
data-testidattributes for stable selectors. Tests run against real backend (gitnexus serve) and Vite dev server.
Scope-resolution tests
Every language resolves calls and inheritance through the scope-resolution pipeline — the legacy call-resolution DAG and the per-language REGISTRY_PRIMARY_<LANG> flag were removed in RING4-1 (#942). Each language's resolver test lives at test/integration/resolvers/<slug>.test.ts and runs once, on the single scope-resolution path, as part of the normal tests job (vitest test/**/*.test.ts).
Adding a language: register its ScopeResolver in scope-resolution/pipeline/registry.ts (SCOPE_RESOLVERS) and add the resolver test file — no workflow or config edit needed.
Cross-platform testing
Windows and macOS CI runs only the platform-sensitive test subset (~50 files out of 373). The full suite runs on Ubuntu.
The subset is defined in gitnexus/scripts/cross-platform-tests.ts and includes:
- Platform-specific logic — tests with
process.platformguards, path.sep behavior, EPERM/EBUSY error classification - Native LadybugDB — all
lbug-*integration tests (N-API addon with known platform-varying behavior) - Process spawning / CLI — tests using real
child_process.spawn, shell quoting, CLI invocations - Worker threads — tests spawning real
worker_threads - Native addon loading — tree-sitter grammar loading smoke tests
- Filesystem behavior — CRLF handling, directory walking, symlinks
When adding a platform-sensitive test, add it to the appropriate section in scripts/cross-platform-tests.ts.
Confirming no tests are orphaned
Every test file matches one of the three vitest projects. To verify:
cd gitnexus
npx vitest list 2>/dev/null | wc -l # should match total test count
To check the cross-platform list is up to date, run npm run test:cross-platform — it fails fast if any listed file is missing.
CI integration
GitHub Actions (.github/workflows/ci.yml) orchestrate:
| Workflow | Jobs | Purpose |
|---|---|---|
ci-quality.yml |
format, lint, typecheck, typecheck-web, workflow-convention | Code quality gates |
ci-tests.yml |
ubuntu/coverage, cross-platform (Win/Mac), packaged-install-smoke | Full suite + coverage on Ubuntu; platform-sensitive subset on Win/Mac |
ci-scope-parity.yml |
discover, parity | Scope-resolution parity for all migrated languages |
ci-e2e.yml |
e2e (chromium) | Playwright E2E, gated on gitnexus-web/** changes |
The CI Gate job in ci.yml requires the quality and test workflows to pass.
The browser E2E workflow must pass or be skipped because no web files changed.
Branch protection also requires six platform check names from the former
three-shard matrix. These names remain as aggregate gates: all native shards
and the every test executed audit must succeed before any of them passes.
Failed, cancelled, skipped, or missing dependency results fail these gates.
The actual native tests run in the current Windows/macOS shard matrix.
The typecheck job runs both the production compiler check and
npm run typecheck:tests. A type error in either check fails the job and the CI gate.
Complete execution, including platform and benchmark tests
The required every test executed job reconciles execution receipts from Ubuntu
coverage, every Windows/macOS shard, the serial benchmark run, and the real Python
workflow preflight. It requires a recorded pass for every collected test. A skip
on Linux is satisfied only by a pass of that exact test in another required job.
Test identities include the file, suite/title and source location; ambiguous
parameterized cases must have unique titles. Missing receipts, missing test files,
unhandled runner errors, failed hooks, failed assertions, and tests with no pass
all fail the gate. Web tests are checked separately with the same rules.
The locked pytest suite and Linux/Windows containment jobs also upload JUnit receipts. The same gate checks every Python test file was collected and every case passed in at least one job, while preserving failures from any job. The Linux containment job supplies Bubblewrap, the pinned CLI, and built Vitest dependencies for tests that cannot run in the basic Python job. Python results are included in the combined PR report.
The PR report shows the reconciled result as Unverified. Zero means every test has execution evidence; individual OS logs still show tests that require another OS as skipped. Failed executions remain failures even if another job passes.
npm run test:benchmarks discovers all tests gated by GITNEXUS_BENCH and runs
them serially. Keep timing measurements out of parallel coverage workers. The
eval-tests job installs the locked Python dependencies and runs the workflow
preflight Vitest tests as well as pytest; those tests must exercise real Python
validation, never a stubbed success.
Regression testing
Re-run the full relevant suite when:
- Prompt or agent-behavior documentation changes (if tests encode behavior)
- Model or embedding-related code paths change
- Graph schema, query contracts, or MCP tool shapes change
- Dependencies with parsing or runtime impact upgrade
User acceptance / beta (optional)
For staged releases or UI betas: deploy to a staging environment, collect structured feedback, watch errors and latency, then iterate before a wider release.