GitNexus/GUARDRAILS.md
John R. Eakin c68d7975e6
docs: agent development framework, GitHub templates, eval refactor (#479)
* ci: E2E workflow, web typecheck job, pre-commit hook, test suite

CI:
- ci.yml consolidated to reference ci-tests.yml
- ci-quality.yml: add typecheck-web job for gitnexus-web/
- ci-e2e.yml: E2E workflow with dorny/paths-filter (web changes only)
- ci-report.yml: remove dead integration-reports references
- CI gate allows skipped E2E status
- .gitignore: playwright artifacts, eval test artifacts

Pre-commit hook:
- .githooks/pre-commit: typecheck + unit tests for both packages
- Activated via git config core.hooksPath in prepare script

Test infrastructure:
- Vitest + React Testing Library: 58 unit tests
  (graph, server-connection, mermaid, settings, constants, utils, paths)
- Playwright E2E: 5 tests + manual recording harness
- vitest.config from vitest/config, engines.node >= 20
- Playwright artifacts retain-on-failure
- wait-on in devDependencies
- vitest/coverage-v8 aligned with vitest 4.x

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: update gitnexus-web package-lock.json

Reflects devDependency additions (vitest, playwright, wait-on,
@testing-library, etc.) from package.json changes in this PR.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): add missing process-list-loaded testid, increase CI timeouts

- Add data-testid="process-list-loaded" to ProcessesPanel (E2E tests
  were waiting for an element that didn't exist)
- Increase server connect timeouts from 5s to 10s for slower CI

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): run gitnexus-web unit tests in CI, remove unused variable

- Add gitnexus-web npm ci + vitest run to ci-tests.yml so web unit
  tests are gated by the CI status check (were only running locally)
- Remove unused IS_PLAYWRIGHT_AUTOMATION variable from E2E spec

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): add process-row testid, wait for networkidle on page load

- Add data-testid="process-row" to ProcessItem component (E2E tests
  referenced it but it didn't exist in the source)
- Use waitUntil: 'networkidle' on page.goto to ensure Vite dev server
  is fully ready before interacting (fixes first-test timeout in CI)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): add process-view-button and process-highlight-button testids

E2E tests referenced these data-testid attributes but they didn't
exist in ProcessItem. All 6 E2E testids now have matching source
elements: status-ready, process-list-loaded, process-row,
process-view-button, process-highlight-button, server-url-input.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): remove networkidle — Vite HMR WebSocket prevents it from resolving

networkidle waits for zero network activity for 500ms, but Vite's HMR
WebSocket stays open permanently, causing page.goto to timeout at 60s
on all tests after the first. The explicit toBeVisible waits on UI
elements are sufficient and deterministic.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): wait for Server button visibility, add CI retry, all 5 tests pass locally

Root cause: test 1 clicked the Server button before React hydrated,
so the tab content never rendered and the input wasn't found.

Fixes:
- Wait for Server button toBeVisible before clicking
- Increase input wait to 15s
- Remove networkidle (Vite HMR WebSocket prevents it from resolving)
- Add retries: 1 in CI for transient cold-start flakiness

Verified locally: all 5 E2E tests pass, 198 unit tests pass, typecheck clean.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): tolerate LadybugDB native crash during analyze step

gitnexus analyze can crash with "double free or corruption" (known
issue #273) during the LadybugDB native addon shutdown. The index is
usually written successfully before the crash. The workflow now:
1. Allows analyze to exit non-zero with a warning
2. Verifies .gitnexus index was actually created
3. Only fails if no index exists (real failure)

All tests verified locally: 198 unit, 5 E2E pass, typecheck clean.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): fix shell quoting in analyze step, simplify to || true

The previous echo string had special characters that broke bash
quoting in GitHub Actions. Simplified to: analyze || true, then
check if .gitnexus exists.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add agent development framework, GitHub templates, eval refactor

Agent framework (layered docs for AI-assisted contributions):
- AGENTS.md: canonical instructions, impact analysis, MCP tools
- CLAUDE.md: Claude Code-specific deltas and hooks
- GUARDRAILS.md: safety boundaries, non-negotiables, escalation
- ARCHITECTURE.md: monorepo layout, data flow map
- TESTING.md: test structure, commands, categories
- RUNBOOK.md: copy-paste operations for dev/CI/MCP
- llms.txt: minimal LLM context pointer

Editor integration:
- .cursor/index.mdc + rules/100-monorepo.mdc

GitHub templates:
- PR template with areas-touched checkboxes
- Bug report + feature request issue forms

Eval harness:
- Refactored mcp_bridge, tool_registry, constants
- Error sanitization utilities
- Property-based tests via Hypothesis

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(eval): use format_exception instead of format_exc in sanitize_exception

format_exc() returns the currently handled exception traceback, which
may be unrelated if called outside an active except block. Using
format_exception(type(exc), exc, exc.__traceback__) reliably captures
the passed exception's traceback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update CONTRIBUTING.md and TESTING.md for current CI/hook setup

- CONTRIBUTING.md: add gitnexus-web typecheck command, pre-commit hook
  checklist item
- TESTING.md: add gitnexus-web typecheck command, pre-commit hook
  section (husky), update CI integration to list actual workflow files
  (ci-quality, ci-tests, ci-e2e)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update testing docs to reflect CI/E2E changes from PR #486

- AGENTS.md: update test counts (CLI ~2000 unit, ~1850 integration),
  add gitnexus-web testing section (198 unit, 5 E2E with commands)
- RUNBOOK.md: fix Node requirement to >=20, fix E2E local repro command
- TESTING.md: E2E uses data-testid selectors + real servers, not mocks
- .cursor/rules/100-monorepo.mdc: add web test/E2E commands

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: address context engineering review — deduplicate tokens, expand Cursor rules

- Remove ~100-line gitnexus:start block from CLAUDE.md (was duplicated from AGENTS.md)
- Fix gitnexus:start block inlined inside AGENTS.md Reference Docs bullet (doubled)
- Replace CLAUDE.md scope table with pointer to AGENTS.md (single source of truth)
- Expand .cursor/index.mdc with 5 non-negotiable safety rules for always-on context
- Add .cursor/rules/200-eval.mdc with Python/eval commands (glob-scoped to eval/**)
- Improve llms.txt with priority annotations and descriptions
- Bump version headers to 1.2.0, last-reviewed to 2026-03-24

Saves ~1,400 tokens/session with zero information loss.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
2026-03-25 06:48:41 +00:00

4.8 KiB
Raw Blame History

Guardrails — GitNexus (repo + agents)

Rules for human contributors and AI agents working on this codebase or publishing artifacts. These complement AGENTS.md / CLAUDE.md (which focus on GitNexus-in-GitNexus workflows).

Scope (typical agent session)

When automating changes in this repository, treat scope as least privilege:

  • Read: Source, tests, docs, public config as needed for the task.
  • Write: Only files required for the requested fix or feature; avoid unrelated formatting or refactors.
  • Execute: Tests, typecheck, and documented CLI commands; do not run destructive commands on user data outside the repo without explicit approval.
  • Off-limits: Other peoples machines, production deployments you dont own, and credentials you didnt receive permission to use.

Adjust explicitly if the maintainer defines a different scope for a task.


Non-negotiables

  1. Never commit secrets — API keys, tokens, .env with real values, private URLs, or session cookies. Use .env.example with placeholders only.
  2. Never rename symbols with blind find-and-replace when working in a GitNexus-indexed project — use the rename MCP tool with dry_run: true first, then review graph vs text_search edits. (There is no separate gitnexus rename CLI; renaming goes through MCP or editor integration.)
  3. Run impact analysis before editing shared symbols — use impact (upstream) for functions/classes/methods others call; do not ignore HIGH / CRITICAL risk without maintainer sign-off.
  4. Prefer detect_changes before commit — confirm diffs map to expected symbols/processes when the graph is available.
  5. Preserve embeddings — if .gitnexus/meta.json shows embeddings, run npx gitnexus analyze --embeddings when refreshing the index; plain analyze can drop them.

Signs (recurring failure patterns)

Use this format: Trigger → Instruction → Reason.
Append new Signs here when the same mistake repeats (e.g. CI broken twice the same way).

Sign: Stale graph after edits

  • Trigger: MCP or resources warn the index is behind HEAD, or code search doesnt match latest commit.
  • Instruction: Run npx gitnexus analyze from the repo root (plus --embeddings if the project used them).
  • Reason: Tools query LadybugDB built at last analyze; git changes are invisible until re-indexed.

Sign: Embeddings vanished after analyze

  • Trigger: Semantic search quality drops; stats.embeddings in .gitnexus/meta.json is 0 after a refresh.
  • Instruction: Re-run npx gitnexus analyze --embeddings and confirm meta.json reflects stored embeddings.
  • Reason: Embedding generation is opt-in; analyze without the flag does not preserve prior vectors.

Sign: MCP lists no repos

  • Trigger: MCP stderr says no indexed repos.
  • Instruction: Run npx gitnexus analyze in the target repository; verify npx gitnexus list shows it.
  • Reason: The MCP server discovers repos via ~/.gitnexus/registry.json, populated by analyze.

Sign: Wrong repo in multi-repo setups

  • Trigger: Query/impact results clearly belong to another project.
  • Instruction: Call list_repos, then pass repo on subsequent tools (or use per-workspace MCP config).
  • Reason: Default target may be ambiguous when multiple repos are registered.

Sign: LadybugDB lock / “database busy”

  • Trigger: Errors opening .gitnexus/lbug while MCP and analyze both run.
  • Instruction: Stop overlapping processes; one writer at a time. Retry analyze or restart MCP.
  • Reason: Embedded DB expects single-process ownership of the store.

Publishing & supply chain

  • npm: Do not publish from unreviewed automation; follow maintainer release process. Bump version intentionally; tag releases to match package.json.
  • Dependencies: Prefer minimal, auditable changes to package.json; run tests and CI after lockfile updates.
  • License: This project ships under PolyForm Noncommercial 1.0.0 — do not relicense or imply a different license in docs or metadata without maintainer approval.

Escalation

Stop and ask a human maintainer when:

  • Impact analysis shows HIGH / CRITICAL risk and the task still requires the change.
  • You need to alter CI, release, or security-sensitive config.
  • Requirements conflict (e.g. “speed up analyze” vs “must keep all embeddings on huge repo”).
  • You are unsure whether data loss is acceptable (clean, forced migrations, schema changes).