mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-08-28 05:25:25 +00:00
* feat(lbug): add ensureEmbeddingRowDmlSafe VECTOR gate for embedding-row DML LadybugDB refuses every mutation of a table carrying an HNSW index while the VECTOR extension is not loaded on that connection: DELETE and CREATE raise a Binder exception, DROP TABLE is refused while the index references it, and SET segfaults the process. Dropping the index is not an available recovery either — CALL DROP_VECTOR_INDEX is itself a VECTOR-extension function and is undefined in exactly that state. Add a single primitive that loads VECTOR under the analyze install policy and, only when that fails, reads CALL SHOW_INDEXES (which works without the extension) to decide whether an index actually exists to trip over. No call sites yet. Refs #2623 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(lbug): pin the #2623 VECTOR gate for embedding-row DML Three cases: no index + VECTOR unavailable stays safe (no needless escalation); index present + VECTOR unavailable is reported blocked AND the raw deleteNodesForFiles genuinely throws 'extension is not loaded' (proving the hazard is real, not theoretical); index present + VECTOR loadable is safe, the delete works, and the HNSW index survives — the invariant run-analyze relies on when it keeps the index across a surgical incremental run. Refs #2623 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(analyze): load VECTOR before the incremental writeback touches embedding rows Incremental analyze died on every content change once a repo had built code_embedding_idx: Analysis failed: Binder exception: Trying to delete from an index on table CodeEmbedding but its extension is not loaded. The surgical writeback's first statement is deleteNodesForFiles' CodeEmbedding join-delete, but nothing on that path loaded VECTOR until Phase 4 — so the engine refused the delete. This is an ordering defect, not an environment one: it reproduces on machines where VECTOR loads fine. The dirty-flag recovery then forced a full rebuild on the next run, which is why it read as 'just slow'. Call ensureEmbeddingRowDmlSafe() once, before the escalation gate and before any row is touched — the same 'index lifecycle before row DML' seam dropSearchFTSIndexes occupies for FTS (#2589). Unconditional, because a DB carrying the index from an earlier --embeddings run hits the same wall on a plain incremental run. When VECTOR truly cannot load the table is immutable (the index cannot be dropped without the extension either), so the run falls through to the existing wipe-and-COPY escalation with a message naming cause, consequence and remedy. Fixes #2623 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(analyze): pin the #2623 VECTOR-before-embedding-DML ordering end-to-end Sibling of the #2589 FTS drop-before-delete suite, same shape: drive the real runFullAnalysis incremental path over a real git repo and a real LadybugDB, seed real embedding rows, build the HNSW index, then assert the index state at the exact moment deleteNodesForFiles is invoked. Both cases were confirmed to discriminate — with the run-analyze change reverted they fail with the reported 'Trying to delete from an index on table CodeEmbedding but its extension is not loaded', and pass with it: - surgical path: the run completes, the index is still present AND extension_loaded at delete time, exactly one row per nodeId survives, and the untouched file's rows are preserved - blocked path: with GITNEXUS_LBUG_EXTENSION_INSTALL=never the run escalates to a full DB write and says so, instead of crashing Also applies prettier's reindent to the run-analyze log ternary. Refs #2623 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(lbug): cite the pinned LadybugDB version in the #2623 probe note The probe matrix behind ensureEmbeddingRowDmlSafe was first recorded on 0.18.0, but gitnexus/package-lock.json pins 0.18.2 (#2587). Re-ran every case on 0.18.2: refused DELETE, refused CREATE, SIGSEGV on SET, DROP_VECTOR_INDEX undefined, DROP TABLE refused, SHOW_INDEXES readable with extension_loaded intact. Identical on both, so the design is unchanged — only the citation was wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(analyze): preserve embeddings across the VECTOR-blocked rebuild, and check the catalog before loading Three follow-ups from reviewing the fix itself. 1. Data loss on the blocked path. Escalating wipes the DB files, and Phase 3.5 restores embedding rows from cachedEmbeddings — which deriveEmbeddingMode only populates when meta.stats.embeddings > 0. A DB holding embedding rows that its meta does not account for therefore had every vector destroyed silently by a rebuild it never asked for. Probe on a 3-file repo: 3 rows before, 0 after, no warning. Read the rows before escalating (a plain MATCH, no extension needed) so the existing restore has something to restore, and say so in the log. The blocked-path test now asserts the seeded rows survive exactly once, and that assertion fails without this rescue. 2. Catalog before extension. ensureEmbeddingRowDmlSafe loaded VECTOR first and only read SHOW_INDEXES on failure, so every incremental analyze on a machine without VECTOR paid a bounded out-of-process INSTALL attempt plus an 'extension unavailable' warning — including repos that never built an embedding index and can never hit this bug. One local catalog read settles that case first; the load is attempted only when an index actually gates DML, or when the catalog cannot be read. 3. Dead branch. targetConn is always the module singleton there, so the isSharedSingletonConn ternary could never take its second arm. Collapsed to withConnLock. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(doctor): live-probe the VECTOR extension instead of printing the static platform capability Review finding on #2624 (MEDIUM), and exactly what #2623's reporter hit: doctor printed 'VECTOR index: available' — derived from a static platform check — while every incremental analyze on the same machine was dying on an unloaded VECTOR extension. The FTS line was switched to a live LOAD probe for the identical contradiction under #2374; VECTOR now gets the same treatment. probeVectorExtensionLoad shares the FTS probe's implementation (bounded, offline-safe, never runs the installer) and doctor's semantic-mode line now follows the probe, not the platform: without a loadable extension the vector index can be neither built nor queried, so search really is on exact scan. The load-error classifier's remedies are label-parameterized so the VECTOR row stops dispensing FTS-specific advice — 'run analyze --repair-fts' repairs FTS indexes only and was actively wrong for a missing vector extension. Default label stays 'FTS'; every existing caller and pinned remedy string is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(lbug): remove the stale Windows VECTOR gate — the extension ships for win_amd64 The codebase categorically refused VECTOR on Windows (platform !== 'win32' in isVectorExtensionSupportedByPlatform, plus a hard early-return in loadVectorExtension) on the strength of an early-era report that in-process INSTALL VECTOR could SIGSEGV (#1365). That belief is stale, verified directly: - the extension server hosts win_amd64 VECTOR artifacts for every 0.18.x extension version — v0.18.0 and v0.18.1 both serve a real 14 MB PE32+ DLL (curl-probed; 'file' confirms PE32+ x86-64) - the pinned 0.18.2 core resolves its extension directory to 0.18.1 (strace-verified LOAD open()), so the pinned version's Windows artifact exists too - INSTALL now runs in a spawned child (installDuckDbExtensionOutOfProcess), so even a crashing installer kills only the child and degrades to unavailable — the original hazard cannot reach the parent process any more Windows now takes the same runtime path as every other OS: try LOAD, install out-of-process when policy allows, degrade to exact scan when it truly fails. The MCP semantic-search lane loses its static platform gate too — it always attempts the vector index and falls back to the exact scan on runtime failure, with a once-per-backend diagnostic naming the real error instead of a platform-policy message. isVectorExtensionSupportedByPlatform is deleted; getRuntimeCapabilities reports the platform capability as available everywhere and defers machine truth to the live probe. Windows CI is the enforcement: the vector suites skip visibly only when the extension genuinely cannot load, so green Windows lanes now actually exercise VECTOR instead of silently skipping by policy. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(lbug): pin the catalog-read-failure fallback in ensureEmbeddingRowDmlSafe Review finding on #2624 (LOW): the one branch where the gate cannot cheaply prove safety — SHOW_INDEXES itself erroring — was exercised only by inference. Force it with a Connection.prototype.query spy over the real DB: the catalog read fails, and the gate must fall through to actually attempting the extension load (asserted via the recorded statement stream) rather than guessing, returning true here because the extension is loadable. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(mcp): load VECTOR on the pool's shared Database so the semantic vector lane actually works Review finding on #2624 (MEDIUM): extension load scope is per-Database (probe-verified — LOAD on one connection enables QUERY_VECTOR_INDEX on every connection of the same Database), and the pool pre-warm loaded only FTS. So LocalBackend's vector lane has ALWAYS raised 'Catalog exception: function QUERY_VECTOR_INDEX is not defined' through the pool and silently fallen back to the exact scan — repos above the 10k exact-scan cap got empty semantic results. The serve path was unaffected (the embedding pipeline loads the extension itself). Mirror the FTS line at BOTH load sites — doInitLbug's pre-warm and initLbugWithDb's external-Database adoption — under the same load-only contract (the read pool never triggers a network install), tracked by a new SharedDB.vectorLoaded flag reset where ftsLoaded resets. The new pool test is discriminating and deliberately closes the writable core adapter before the pool opens: a shared/injected Database would inherit the VECTOR load from test seeding and pass either way, so the case forces the pool onto its OWN fresh read-only Database where only the pre-warm can make the lane legal. Verified: fails at the pre-fix tree with the exact Catalog exception, passes with the fix. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci: run the #2623 ordering suite on Windows/macOS and pre-install VECTOR alongside FTS Two review findings on #2624, both landing in existing seams: - scripts/cross-platform-tests.ts gains incremental-vector-extension-ordering .test.ts: the win32 VECTOR gate is gone in this PR, so the #2623 drop-ordering + blocked-path escalation must be proven on the windows-latest native addon, not just Ubuntu. (The review's claim that lbug-delete-nodes-for-files.test.ts was also missing was wrong — it has been on the roster since #2409.) - scripts/ensure-fts.ts now pre-installs VECTOR under the same best-effort auto-policy contract, so every sharded CI process LOADs from ~/.lbdb instead of racing its own bounded out-of-process INSTALL; the workflow's extension cache already covers it (path is the whole extension dir — key kept for cache continuity). The cross-platform job sets GITNEXUS_REQUIRE_VECTOR=1 beside GITNEXUS_REQUIRE_FTS so a genuinely unavailable VECTOR is a loud failure, never a silent skip. Windows/macOS cannot be executed locally; the PR's CI lanes are the proof for this commit. Linux smoke: ensure-fts.ts reports both extensions ready; all 79 roster entries resolve. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(pool): register loadVectorExtension in the pool unit-suite mocks The pool adapter's new loadVectorExtension import surfaced in four suites that mock lbug-adapter.js with explicit factories (vitest fails loudly on a missing mocked export). Register the export in each — resolving false where the suite's world assumes no vector, true where it mirrors FTS — and extend lbug-pool-fts-load.test.ts, the suite that owns pre-warm extension loading, with the vector pair: successful load cached per shared Database, failed load retried on the next open, both pinned to policy load-only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(analyze): use POSIX literals for graph paths in the #2623 ordering suite First Windows CI run of this suite (it joined the cross-platform roster this PR) failed with 'Parser exception: Invalid input <MATCH (n:Function) WHERE n.filePath = '>' — path.join produces backslashes on Windows, and a backslash inside the seed helper's single-quoted Cypher literal breaks the parser. The graph stores repo-relative filePaths with forward slashes on every OS, so graph-side paths are POSIX literals now (the incremental-orchestration convention); path.join stays only for real filesystem access. The same Windows lane also proved the substance this suite exists for: lbug-vector-extension passed 7/7 on windows-latest — the extension installed, loaded, and built a real HNSW index there — and the pool vector-lane and DML gate suites passed too. This commit fixes the harness, not the fix. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Gergo Magyar <abhigyan1.patwari@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
645 lines
29 KiB
YAML
645 lines
29 KiB
YAML
name: Tests
|
|
|
|
on:
|
|
workflow_call:
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
jobs:
|
|
# Ubuntu full-suite coverage, sharded. Each shard writes a vitest blob report
|
|
# (carrying its slice of V8 coverage) with thresholds forced OFF — a single
|
|
# shard's partial coverage can't meet the gate. The coverage-merge job below
|
|
# reduces the blobs and enforces the real thresholds on the combined coverage.
|
|
# FTS self-installs per shard (test/helpers/fts-availability.ts), so sharding
|
|
# the full suite across fresh runners is safe. Shard count: shard-plan.cov_total.
|
|
tests:
|
|
name: ubuntu / coverage ${{ matrix.shard }}/${{ needs.shard-plan.outputs.cov_total }}
|
|
needs: shard-plan
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 25
|
|
strategy:
|
|
fail-fast: false
|
|
matrix:
|
|
shard: ${{ fromJSON(needs.shard-plan.outputs.cov_shards) }}
|
|
# Fail loudly (don't silently skip) if the FTS extension is unavailable, so
|
|
# FTS-dependent lbug integration suites are guaranteed to run in CI.
|
|
env:
|
|
GITNEXUS_REQUIRE_FTS: '1'
|
|
steps:
|
|
# persist-credentials: false — runs tests + uploads a blob artifact; the
|
|
# default-persisted token must not be capturable through it (zizmor
|
|
# credential-persistence / artipacked audit). The job never pushes.
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: ./.github/actions/setup-gitnexus
|
|
with:
|
|
build: 'true'
|
|
# Warm-cache the FTS extension (same per-OS key as the cross-platform job)
|
|
# and install it up front, so every coverage shard has FTS in ~/.lbdb before
|
|
# any test module loads. The file-path FTS gate (extension-binary-real)
|
|
# resolves the extension at module load and can't self-install, so sharding
|
|
# could otherwise drop it into a shard with no installer sibling.
|
|
- name: Cache LadybugDB FTS extension
|
|
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v5
|
|
with:
|
|
path: ~/.lbdb/extension
|
|
key: lbug-fts-${{ runner.os }}-${{ hashFiles('gitnexus/package-lock.json') }}
|
|
- name: Ensure FTS + VECTOR extensions installed
|
|
run: npx tsx scripts/ensure-fts.ts
|
|
working-directory: gitnexus
|
|
- name: Run sharded tests with coverage (blob)
|
|
# Shard via env var (not `${{ }}` inlined into the shell) so it isn't a
|
|
# template-injection sink; shell: bash makes "$SHARD" expand uniformly.
|
|
# Thresholds forced to 0 — the merge job enforces the real gate on the
|
|
# MERGED coverage; a single shard's partial coverage would always fail.
|
|
shell: bash
|
|
env:
|
|
SHARD: ${{ matrix.shard }}/${{ needs.shard-plan.outputs.cov_total }}
|
|
run: >-
|
|
npx vitest run
|
|
--shard="$SHARD"
|
|
--reporter=default
|
|
--reporter=blob
|
|
--coverage
|
|
--coverage.thresholds.lines=0
|
|
--coverage.thresholds.functions=0
|
|
--coverage.thresholds.branches=0
|
|
--coverage.thresholds.statements=0
|
|
working-directory: gitnexus
|
|
- name: Upload coverage blob
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: coverage-blob-${{ matrix.shard }}
|
|
path: gitnexus/.vitest-reports/
|
|
# .vitest-reports is a dotdir; upload-artifact excludes hidden files by
|
|
# default, which would upload an empty artifact and break the merge.
|
|
include-hidden-files: true
|
|
retention-days: 5
|
|
|
|
# Merge the sharded coverage blobs into one report and enforce the real
|
|
# thresholds on the combined ('new') coverage — `vitest --mergeReports` re-runs
|
|
# nothing, it just reduces the stored blobs. Also emits the merged
|
|
# test-results.json and runs the (unsharded) web + docker suites, so the
|
|
# `test-reports` artifact keeps the exact shape ci-report.yml consumes for its
|
|
# base-branch ('baseline') vs new coverage delta.
|
|
coverage-merge:
|
|
name: ubuntu / coverage merge
|
|
needs: tests
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 15
|
|
env:
|
|
GITNEXUS_REQUIRE_FTS: '1'
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: ./.github/actions/setup-gitnexus
|
|
with:
|
|
build: 'true'
|
|
- name: Download coverage blobs
|
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
|
|
with:
|
|
pattern: coverage-blob-*
|
|
path: gitnexus/.vitest-reports
|
|
merge-multiple: true
|
|
- name: Merge coverage + enforce thresholds
|
|
run: >-
|
|
npx vitest --mergeReports
|
|
--reporter=default
|
|
--reporter=json
|
|
--outputFile=test-results.json
|
|
--coverage
|
|
--coverage.reporter=json-summary
|
|
--coverage.reporter=json
|
|
--coverage.reporter=text
|
|
--coverage.thresholdAutoUpdate=false
|
|
working-directory: gitnexus
|
|
# gitnexus-shared already built by setup-gitnexus above
|
|
- name: Install gitnexus-web dependencies
|
|
run: npm ci
|
|
working-directory: gitnexus-web
|
|
- name: Run gitnexus-web unit tests
|
|
run: >-
|
|
npx vitest run
|
|
--reporter=default
|
|
--reporter=json
|
|
--outputFile=web-test-results.json
|
|
working-directory: gitnexus-web
|
|
- name: Run docker-server integration tests
|
|
run: node --test docker-server.test.mjs
|
|
- name: Upload test reports
|
|
if: always()
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
with:
|
|
name: test-reports
|
|
path: |
|
|
gitnexus/coverage/coverage-summary.json
|
|
gitnexus/coverage/coverage-final.json
|
|
gitnexus/test-results.json
|
|
gitnexus-web/web-test-results.json
|
|
retention-days: 5
|
|
|
|
# Single source of truth for the platform-sensitive shard count. TOTAL below
|
|
# generates both the shard index list (the matrix) and the /N denominator (job
|
|
# name + --shard arg), so they can't drift — bump the shard count by editing
|
|
# TOTAL alone. Checkout-free (ubuntu ships jq), so no credential surface.
|
|
shard-plan:
|
|
runs-on: ubuntu-latest
|
|
outputs:
|
|
shards: ${{ steps.gen.outputs.shards }}
|
|
total: ${{ steps.gen.outputs.total }}
|
|
cov_shards: ${{ steps.gen.outputs.cov_shards }}
|
|
cov_total: ${{ steps.gen.outputs.cov_total }}
|
|
steps:
|
|
- id: gen
|
|
run: |
|
|
TOTAL=3 # cross-platform (windows/macOS) shards per OS
|
|
COV_TOTAL=3 # ubuntu coverage shards (merged before thresholds)
|
|
if [ "$TOTAL" -lt 1 ] || [ "$COV_TOTAL" -lt 1 ]; then
|
|
echo "shard totals must be >= 1" >&2; exit 1
|
|
fi
|
|
{
|
|
echo "shards=$(jq -nc --argjson n "$TOTAL" '[range(1; $n + 1)]')"
|
|
echo "total=$TOTAL"
|
|
echo "cov_shards=$(jq -nc --argjson n "$COV_TOTAL" '[range(1; $n + 1)]')"
|
|
echo "cov_total=$COV_TOTAL"
|
|
} >> "$GITHUB_OUTPUT"
|
|
|
|
# Platform-sensitive subset only — the full suite runs on Ubuntu above.
|
|
# See gitnexus/scripts/cross-platform-tests.ts for the file list and
|
|
# rationale for each included test.
|
|
cross-platform:
|
|
name: ${{ matrix.os }} (platform-sensitive) ${{ matrix.shard }}/${{ needs.shard-plan.outputs.total }}
|
|
needs: shard-plan
|
|
strategy:
|
|
fail-fast: false
|
|
matrix:
|
|
# Ubuntu already covered by the coverage job above
|
|
os: [windows-latest, macos-latest]
|
|
# Shard the fixed file list across N runners per OS (N = TOTAL in the
|
|
# shard-plan job). The suite is dominated by ~50 CLI/worker process
|
|
# spawns and Windows is ~5x slower than macOS at those, so the unsharded
|
|
# run crept past the 15-min watchdog in run-cross-platform.ts. vitest
|
|
# shards by file COUNT, not runtime, so the heaviest spawn suites can
|
|
# cluster on one shard. The busiest Windows shard has grown to the old
|
|
# 15-minute watchdog (14m57s on the v1.6.10-rc.19 green run, one
|
|
# observed timeout since — #2449), so the job env below raises the
|
|
# per-shard watchdog to 20 minutes, still bounded by timeout-minutes.
|
|
# Shard indices come from the shard-plan job (single source of truth):
|
|
# its TOTAL drives this list and the /N in the job name + --shard arg.
|
|
shard: ${{ fromJSON(needs.shard-plan.outputs.shards) }}
|
|
runs-on: ${{ matrix.os }}
|
|
timeout-minutes: 25
|
|
# Same guarantee on the platform-sensitive runners: FTS-dependent suites in
|
|
# the cross-platform subset must run, not silently skip.
|
|
#
|
|
# GITNEXUS_E2E_CLI=dist: the e2e suites spawn the CLI ~50 times; each spawn via
|
|
# `node --import tsx src/cli/index.ts` re-transpiles the whole CLI, and Windows
|
|
# is ~5x slower at process startup. `build: true` below produces a fresh dist
|
|
# before tests, so opting these runners into the built CLI removes that
|
|
# per-spawn transpile (see test/helpers/cli-entry.ts). Deliberately scoped to
|
|
# THIS job: the Ubuntu coverage job leaves it unset, so it keeps exercising the
|
|
# tsx-on-source path in CI (both entry points stay covered).
|
|
env:
|
|
GITNEXUS_REQUIRE_FTS: '1'
|
|
# #2623: the win32 VECTOR gate is gone, so the vector suites genuinely
|
|
# run here — require the extension so an unavailable VECTOR is a loud
|
|
# failure, never a silent skip (same contract as GITNEXUS_REQUIRE_FTS).
|
|
GITNEXUS_REQUIRE_VECTOR: '1'
|
|
GITNEXUS_E2E_CLI: dist
|
|
# #2449: hosted Windows runners intermittently push the busiest shard past
|
|
# the default 15-minute watchdog. 20 minutes restores real headroom while
|
|
# the 25-minute job timeout above still bounds a genuine hang.
|
|
GITNEXUS_CROSS_PLATFORM_TIMEOUT_MINUTES: '20'
|
|
steps:
|
|
# persist-credentials: false — runs tests only, never pushes (zizmor
|
|
# credential-persistence / artipacked audit).
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: ./.github/actions/setup-gitnexus
|
|
with:
|
|
build: 'true'
|
|
# Warm-cache the installed LadybugDB FTS + VECTOR extensions
|
|
# (~/.lbdb/extension) per OS + lockfile so a warm run skips the network
|
|
# install entirely, and the parallel shards share one download across
|
|
# runs. Pure reliability/speed: on a cache miss the tests self-install on
|
|
# demand (see test/helpers/fts-availability.ts), so a miss just falls
|
|
# back to install — never a correctness dependency. Keyed by lockfile
|
|
# hash so a LadybugDB version bump re-installs; per-OS because the
|
|
# extensions are native binaries. (Key name kept as lbug-fts for cache
|
|
# continuity — the path covers every extension in the shared home.)
|
|
- name: Cache LadybugDB FTS extension
|
|
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v5
|
|
with:
|
|
path: ~/.lbdb/extension
|
|
key: lbug-fts-${{ runner.os }}-${{ hashFiles('gitnexus/package-lock.json') }}
|
|
- name: Ensure FTS + VECTOR extensions installed
|
|
run: npx tsx scripts/ensure-fts.ts
|
|
working-directory: gitnexus
|
|
- name: Run platform-sensitive tests
|
|
# Pass the shard through an env var (not `${{ }}` inlined into the shell)
|
|
# so it isn't a template-injection sink (zizmor). shell: bash makes the
|
|
# `"$SHARD"` expansion uniform across the windows + macOS matrix (the
|
|
# default run shell is pwsh on Windows, where `$SHARD` would be empty).
|
|
shell: bash
|
|
env:
|
|
SHARD: ${{ matrix.shard }}/${{ needs.shard-plan.outputs.total }}
|
|
run: npx tsx scripts/run-cross-platform.ts --shard="$SHARD"
|
|
working-directory: gitnexus
|
|
|
|
# Tree-sitter ABI gate (#1922). Two halves, both blocking:
|
|
# 1. Static, offline: assert every grammar's compiled ABI loads on the
|
|
# pinned runtime (check-tree-sitter-upgrade-readiness.py --assert-current).
|
|
# 2. Dynamic: run the parser-loader ABI load-smoke on the OS matrix so an
|
|
# ABI-incompatible committed vendor prebuilt (e.g. Swift's — the static
|
|
# check introspects source, not the shipped .node) fails on the platform
|
|
# it ships to.
|
|
abi-assert:
|
|
name: tree-sitter ABI (${{ matrix.os }})
|
|
strategy:
|
|
fail-fast: false
|
|
matrix:
|
|
os: [ubuntu-latest, windows-latest, macos-latest]
|
|
runs-on: ${{ matrix.os }}
|
|
timeout-minutes: 20
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: ./.github/actions/setup-gitnexus
|
|
with:
|
|
build: 'true'
|
|
|
|
- name: Assert installed + vendored grammar ABIs (static)
|
|
shell: bash
|
|
run: python3 .github/scripts/check-tree-sitter-upgrade-readiness.py --assert-current
|
|
|
|
- name: Run parser-loader ABI load-smoke (dynamic)
|
|
run: npx vitest run test/unit/parser-loader-abi.test.ts
|
|
working-directory: gitnexus
|
|
|
|
# End-to-end smoke test for the #1728 packaging fix: pack the published
|
|
# tarball, install it globally into a temp prefix, and assert no junction
|
|
# creation (the EPERM root cause) plus working CLI plus vendor cleanliness
|
|
# (#836). Runs on windows-latest because that is the platform the fix
|
|
# targets; the in-repo `npm ci` job above only exercises the dev-tree path
|
|
# and skips the tarball reify step where the historical EPERM occurred.
|
|
packaged-install-smoke:
|
|
name: packaged install smoke (${{ matrix.os }})
|
|
strategy:
|
|
fail-fast: false
|
|
matrix:
|
|
os: [windows-latest, ubuntu-latest]
|
|
runs-on: ${{ matrix.os }}
|
|
timeout-minutes: 15
|
|
steps:
|
|
# persist-credentials: false — this job runs npm pack + npm install -g
|
|
# from a tarball and never pushes back; the token in .git/config would
|
|
# be at risk of leaking through any future artifact-upload step
|
|
# (zizmor artipacked audit). Disable upfront.
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: ./.github/actions/setup-gitnexus
|
|
with:
|
|
build: 'true'
|
|
|
|
- name: Pack gitnexus tarball
|
|
shell: bash
|
|
run: npm pack
|
|
working-directory: gitnexus
|
|
|
|
- name: Install gitnexus tarball into isolated prefix
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
PREFIX="$RUNNER_TEMP/gitnexus-smoke"
|
|
mkdir -p "$PREFIX"
|
|
TARBALL=$(find . -maxdepth 1 -name 'gitnexus-*.tgz' -print -quit)
|
|
if [ -z "$TARBALL" ]; then
|
|
echo "ERROR: no gitnexus-*.tgz tarball found in $(pwd)" >&2
|
|
exit 1
|
|
fi
|
|
echo "Installing $TARBALL into $PREFIX"
|
|
npm install -g --prefix "$PREFIX" "./$TARBALL" --no-audit --no-fund
|
|
echo "PREFIX=$PREFIX" >> "$GITHUB_ENV"
|
|
working-directory: gitnexus
|
|
|
|
- name: Assert no junctions or vendor build artifacts
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
# Locate the installed gitnexus package across npm prefix layouts
|
|
# (lib/node_modules on POSIX, node_modules on Windows).
|
|
for candidate in "$PREFIX/lib/node_modules/gitnexus" "$PREFIX/node_modules/gitnexus"; do
|
|
if [ -d "$candidate" ]; then
|
|
INSTALLED="$candidate"
|
|
break
|
|
fi
|
|
done
|
|
if [ -z "${INSTALLED:-}" ]; then
|
|
echo "ERROR: installed gitnexus package not found under $PREFIX" >&2
|
|
ls -la "$PREFIX" || true
|
|
exit 1
|
|
fi
|
|
echo "Installed package at: $INSTALLED"
|
|
|
|
# #836 invariant: no node_modules/ or build/ under any vendor/*.
|
|
BAD=$(find "$INSTALLED/vendor" \( -name node_modules -o -name build \) -print 2>/dev/null || true)
|
|
if [ -n "$BAD" ]; then
|
|
echo "ERROR: vendor tree contains forbidden build artifacts (#836):" >&2
|
|
echo "$BAD" >&2
|
|
exit 1
|
|
fi
|
|
|
|
# #1728 invariant: materialized grammar dirs are real directories,
|
|
# not junctions/symlinks (which is what the EPERM regression created).
|
|
for name in tree-sitter-dart tree-sitter-proto tree-sitter-swift; do
|
|
entry="$INSTALLED/node_modules/$name"
|
|
if [ ! -e "$entry" ]; then
|
|
echo "WARN: $name not materialized (toolchain/prebuild may be unavailable on $RUNNER_OS)"
|
|
continue
|
|
fi
|
|
if [ -L "$entry" ]; then
|
|
echo "ERROR: $entry is a symlink/junction — #1728 regression" >&2
|
|
exit 1
|
|
fi
|
|
if [ ! -d "$entry" ]; then
|
|
echo "ERROR: $entry is not a directory" >&2
|
|
exit 1
|
|
fi
|
|
done
|
|
|
|
- name: Assert gitnexus --version works
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "$RUNNER_OS" = "Windows" ]; then
|
|
"$PREFIX/gitnexus.cmd" --version
|
|
else
|
|
"$PREFIX/bin/gitnexus" --version
|
|
fi
|
|
|
|
# Node engines-floor gate (#2372). A module that statically names an API
|
|
# newer than the supported floor (e.g. `module.registerHooks`, added in
|
|
# 22.15) fails to LINK on the floor — a class vitest/tsx transforms
|
|
# structurally mask, and the default `node-version: 22` (resolves to latest)
|
|
# never hits. Build the dist on 22.x, then import-link every module R1 names
|
|
# as a load surface on the pinned engines floor (22.18.0, per package.json
|
|
# `engines: ^22.18.0 || >=24.11.0`) so a regression fails here instead of
|
|
# shipping to users on the minimum supported Node.
|
|
node-floor-compat:
|
|
name: node floor compat (22.18)
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 15
|
|
steps:
|
|
# persist-credentials: false — builds and import-links only, never pushes
|
|
# (zizmor credential-persistence / artipacked audit).
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
|
|
with:
|
|
node-version: '22'
|
|
cache: npm
|
|
cache-dependency-path: gitnexus/package-lock.json
|
|
- name: Build gitnexus-shared
|
|
run: npm ci && npm run build
|
|
working-directory: gitnexus-shared
|
|
- name: Install and build gitnexus
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
npm ci
|
|
npm run build
|
|
working-directory: gitnexus
|
|
# Switch to the engines-floor Node AFTER building — native deps built on
|
|
# 22.x load across the whole 22.x ABI line, and nothing installs after this
|
|
# (so no package-manager cache is needed).
|
|
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
|
|
with:
|
|
node-version: '22.18.0'
|
|
package-manager-cache: false
|
|
- name: Import-link the built dist on Node 22.18
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
node --version
|
|
node --version | grep -q '^v22\.18\.' || { echo "expected Node 22.18.x" >&2; exit 1; }
|
|
for m in \
|
|
core/embeddings/runtime-install \
|
|
core/embeddings/onnxruntime-node-resolver \
|
|
core/embeddings/onnxruntime-common-resolver \
|
|
cli/embeddings \
|
|
cli/analyze \
|
|
cli/doctor \
|
|
mcp/core/embedder; do
|
|
echo "import dist/$m.js"
|
|
node --input-type=module -e "await import('./dist/$m.js')"
|
|
done
|
|
working-directory: gitnexus
|
|
|
|
# ── Dedicated benchmark gate ─────────────────────────────────────
|
|
# The cross-language `*-pipeline-benchmark.test.ts` suites are gated behind
|
|
# GITNEXUS_BENCH (they generate synthetic codebases at scale), so the main
|
|
# coverage job above SKIPS them — their O(n^2) scaling guards never ran in CI.
|
|
# Run them here with GITNEXUS_BENCH=1, alongside the Python scope-capture and
|
|
# import-resolution fingerprint + scaling guards (PR #1918 P2a).
|
|
#
|
|
# `--no-file-parallelism` is REQUIRED: these suites measure wall-clock and peak
|
|
# heap, so parallel forks both skew the timings and OOM the worker pool — they
|
|
# must run one file at a time.
|
|
#
|
|
# go-pipeline-benchmark.test.ts is deliberately NOT included: its
|
|
# worker-pool (#1848) suite spins a real worker pool that exits unexpectedly
|
|
# under vitest's fork pool (reproduced in validation), which would make this
|
|
# gate flaky. Go is already guarded by its non-gated O(n^2) tripwire (runs in
|
|
# the main coverage job) plus its golden capture-parity test.
|
|
benchmarks:
|
|
name: benchmarks (GITNEXUS_BENCH)
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 25
|
|
steps:
|
|
# persist-credentials: false — this job only runs npm + vitest benchmarks
|
|
# and never pushes; the default-persisted token in .git/config would be at
|
|
# risk of leaking through an artifact upload (zizmor credential-persistence
|
|
# / artipacked audit). Mirrors the packaged-install-smoke job below.
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: ./.github/actions/setup-gitnexus
|
|
with:
|
|
build: 'true'
|
|
|
|
- name: Python scope-capture + import-resolution fingerprint / scaling guards
|
|
run: |
|
|
node --import tsx bench/python-scope/measure.mjs --check
|
|
node --import tsx bench/python-scope/import-target-fingerprint.mjs --check
|
|
working-directory: gitnexus
|
|
|
|
- name: Cross-language scope-capture fingerprint + scaling guards
|
|
# Build-free: asserts emit<Lang>ScopeCaptures output is unchanged
|
|
# (fingerprint) and stays linear (scaling < 1.5) for go/csharp/rust/php/
|
|
# ruby/cobol. Catches an O(n^2) re-regression without the worker pool.
|
|
run: node --import tsx bench/scope-capture/measure.mjs --check
|
|
working-directory: gitnexus
|
|
|
|
- name: CFG construction time / disk / memory guards (#2081 M1)
|
|
# Build-free: asserts collectFunctionCfgs output is unchanged
|
|
# (fingerprint) and that wall-time, cfgSideChannel disk bytes, AND
|
|
# retained heap all stay sub-quadratic for the straight-line /
|
|
# many-functions / branchy scenarios. Catches an O(n^2) re-regression in
|
|
# the per-function CFG builder (e.g. an extendBlock concat chain) and a
|
|
# memory/disk blow-up. --expose-gc enables the retained-heap measurement.
|
|
run: node --expose-gc --import tsx bench/cfg/measure.mjs --check
|
|
working-directory: gitnexus
|
|
|
|
- name: Emit-persistence throughput / byte-identity guards (#2203)
|
|
# Build-free: asserts streamAllCSVsToDisk output is byte-identical
|
|
# (order-independent CSV-line fingerprint — the #2203 U2/U3 emit
|
|
# optimisations must not change graph content) and that emit wall-time
|
|
# stays linear in node+edge count. The LadybugDB COPY half needs a real
|
|
# DB, so its timing lives in the runtime PROF_LBUG_LOAD breakdown.
|
|
run: node --import tsx bench/emit-persistence/measure.mjs --check
|
|
working-directory: gitnexus
|
|
|
|
- name: Streaming PDG-emit byte-identity / bounded-RSS guards (#2202)
|
|
# Build-free: asserts the streaming PdgEmitSink emits a CSV row SET
|
|
# byte-identical to the whole-graph streamAllCSVsToDisk emit, AND that
|
|
# the in-memory graph retains zero BasicBlock nodes (the O(chunk) peak-RSS
|
|
# bound that unblocks full-kernel-scale repos). Fails on fingerprint drift
|
|
# or any resident BasicBlock.
|
|
run: node --import tsx bench/emit-persistence/measure-streaming.mjs --check
|
|
working-directory: gitnexus
|
|
|
|
- name: Cross-language pipeline benchmarks (GITNEXUS_BENCH, serial)
|
|
env:
|
|
GITNEXUS_BENCH: '1'
|
|
run: >-
|
|
npx vitest run --no-file-parallelism
|
|
test/integration/cobol-pipeline-benchmark.test.ts
|
|
test/integration/csharp-pipeline-benchmark.test.ts
|
|
test/integration/rust-pipeline-benchmark.test.ts
|
|
test/integration/php-pipeline-benchmark.test.ts
|
|
test/integration/ruby-pipeline-benchmark.test.ts
|
|
working-directory: gitnexus
|
|
|
|
# Locked eval suite. setup-uv and uv itself are immutable so CI exercises
|
|
# exactly the dependency graph developers run from eval/uv.lock.
|
|
eval-tests:
|
|
name: eval / locked pytest
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 15
|
|
steps:
|
|
# persist-credentials: false — runs tests only, never pushes.
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: astral-sh/setup-uv@11f9893b081a58869d3b5fccaea48c9e9e46f990 # v8.3.2
|
|
with:
|
|
version: '0.11.23'
|
|
python-version: '3.13'
|
|
enable-cache: true
|
|
cache-dependency-glob: eval/uv.lock
|
|
- run: uv run --locked --extra dev python -m pytest tests -q
|
|
working-directory: eval
|
|
|
|
# Native Linux ownership and Bubblewrap boundary. The environment flag makes
|
|
# the real namespace test mandatory; a missing/blocked bwrap is a failure.
|
|
eval-containment-linux:
|
|
name: eval / containment (ubuntu)
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 20
|
|
env:
|
|
GITNEXUS_REQUIRE_BWRAP_CANARY: '1'
|
|
GITNEXUS_REQUIRE_CLAUDE_CANARY: '1'
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
|
|
with:
|
|
node-version: '22.18.0'
|
|
cache: npm
|
|
cache-dependency-path: |
|
|
gitnexus/package-lock.json
|
|
gitnexus-shared/package-lock.json
|
|
- uses: astral-sh/setup-uv@11f9893b081a58869d3b5fccaea48c9e9e46f990 # v8.3.2
|
|
with:
|
|
version: '0.11.23'
|
|
python-version: '3.13'
|
|
enable-cache: true
|
|
cache-dependency-glob: eval/uv.lock
|
|
- name: Install sandbox runtime and pinned Claude CLI
|
|
run: |
|
|
set -euo pipefail
|
|
sudo apt-get update
|
|
sudo apt-get install --yes --no-install-recommends bubblewrap socat
|
|
apparmor_userns=/proc/sys/kernel/apparmor_restrict_unprivileged_userns
|
|
if [[ -r "${apparmor_userns}" ]] && [[ "$(<"${apparmor_userns}")" == '1' ]]; then
|
|
sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
|
|
fi
|
|
canary_runtime="${RUNNER_TEMP}/claude-canary"
|
|
install -d -m 0700 "${canary_runtime}"
|
|
install -m 0600 \
|
|
.github/claude-canary-runtime/package.json \
|
|
"${canary_runtime}/package.json"
|
|
install -m 0600 \
|
|
.github/claude-canary-runtime/package-lock.json \
|
|
"${canary_runtime}/package-lock.json"
|
|
npm ci \
|
|
--prefix "${canary_runtime}" \
|
|
--ignore-scripts=false \
|
|
--audit=false \
|
|
--fund=false
|
|
node -e \
|
|
"const p=require(process.argv[1]); if(p.version!=='2.1.214') process.exit(1)" \
|
|
"${canary_runtime}/node_modules/@anthropic-ai/claude-code/package.json"
|
|
test "$("${canary_runtime}/node_modules/@anthropic-ai/claude-code-linux-x64/claude" --version)" = \
|
|
'2.1.214 (Claude Code)'
|
|
- name: Build pinned shared runtime
|
|
run: |
|
|
npm ci
|
|
npm run build
|
|
working-directory: gitnexus-shared
|
|
- name: Install and build pinned GitNexus runtime
|
|
run: |
|
|
npm ci
|
|
npm run build
|
|
working-directory: gitnexus
|
|
- name: Prove process-tree and sandbox containment
|
|
env:
|
|
CLAUDE_CANARY_BIN: ${{ runner.temp }}/claude-canary/node_modules/@anthropic-ai/claude-code-linux-x64/claude
|
|
run: >-
|
|
uv run --locked --extra dev python -m pytest
|
|
tests/test_process_control.py
|
|
tests/test_proposer_sandbox.py
|
|
tests/test_workflow_bench_sessions.py
|
|
tests/test_ce_plugin_runtime.py -q
|
|
working-directory: eval
|
|
|
|
# Native Windows Job Object canary. POSIX-only tests skip by platform, while
|
|
# the grandchild delayed-write test must execute and pass on this runner.
|
|
eval-containment-windows:
|
|
name: eval / containment (windows)
|
|
runs-on: windows-latest
|
|
timeout-minutes: 15
|
|
steps:
|
|
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
|
|
with:
|
|
persist-credentials: false
|
|
- uses: astral-sh/setup-uv@11f9893b081a58869d3b5fccaea48c9e9e46f990 # v8.3.2
|
|
with:
|
|
version: '0.11.23'
|
|
python-version: '3.13'
|
|
enable-cache: true
|
|
cache-dependency-glob: eval/uv.lock
|
|
- name: Prove Windows process-tree ownership
|
|
run: >-
|
|
uv run --locked --extra dev python -m pytest
|
|
tests/test_process_control.py -q
|
|
working-directory: eval
|