mirror of
https://github.com/abhigyanpatwari/GitNexus.git
synced 2026-08-28 05:25:25 +00:00
feat(wiki): allow explicit HTTP LLM hosts (#2491)
* feat(wiki): allow explicit HTTP LLM hosts Keep wiki LLM HTTP endpoints fail-closed by default while adding a narrow exact-host opt-in for LAN/self-hosted models. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(wiki): simplify insecure LLM flag name Rename the wiki HTTP opt-in flag to --allow-insecure-connection per review feedback. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(wiki): simplify insecure connection env Rename the wiki HTTP allowlist environment variable and align validation errors with the CLI flag naming. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
This commit is contained in:
parent
3dd553b345
commit
a333d94a00
10 changed files with 264 additions and 117 deletions
186
README.md
186
README.md
|
|
@ -82,15 +82,15 @@ That's it. `analyze` indexes the codebase, installs agent skills, registers Clau
|
|||
|
||||
## Two Ways to Use GitNexus
|
||||
|
||||
| | **CLI + MCP** (recommended) | **Web UI** |
|
||||
| ----------- | ---------------------------------------------------------------------- | --------------------------------------------------------------------- |
|
||||
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
|
||||
| **For** | Daily development with Cursor, Claude Code, Antigravity, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
|
||||
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
|
||||
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
|
||||
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
|
||||
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
|
||||
| **Privacy** | Everything local, no network | Everything in-browser, no server |
|
||||
| | **CLI + MCP** (recommended) | **Web UI** |
|
||||
| ----------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------------------- |
|
||||
| **What** | Index repos locally, connect AI agents via MCP | Visual graph explorer + AI chat in browser |
|
||||
| **For** | Daily development with Cursor, Claude Code, Antigravity, Codex, Windsurf, OpenCode | Quick exploration, demos, one-off analysis |
|
||||
| **Scale** | Full repos, any size | Limited by browser memory (~5k files), or unlimited via backend mode |
|
||||
| **Install** | `npm install -g gitnexus` | No install — [gitnexus.vercel.app](https://gitnexus.vercel.app) |
|
||||
| **Storage** | LadybugDB native (fast, persistent) | LadybugDB WASM (in-memory, per session) |
|
||||
| **Parsing** | Tree-sitter native bindings | Tree-sitter WASM |
|
||||
| **Privacy** | Everything local, no network | Everything in-browser, no server |
|
||||
|
||||
> **Bridge mode:** `gitnexus serve` connects the two — the web UI auto-detects the local server and can browse all your CLI-indexed repos without re-uploading or re-indexing.
|
||||
|
||||
|
|
@ -137,49 +137,49 @@ flowchart TB
|
|||
|
||||
### 17 MCP tools (15 per-repo + 2 group)
|
||||
|
||||
| Tool | What It Does |
|
||||
| ---------------- | --------------------------------------------------------------------- |
|
||||
| `list_repos` | Discover all indexed repositories (paginated — `limit`/`offset`) |
|
||||
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) |
|
||||
| `context` | 360-degree symbol view — categorized refs, process participation |
|
||||
| `impact` | Blast radius analysis with depth grouping and confidence |
|
||||
| `trace` | Shortest directed path between two symbols (call + class-member edges)|
|
||||
| `detect_changes` | Git-diff impact — maps changed lines to affected processes |
|
||||
| `check` | Read-only structural checks against the indexed graph |
|
||||
| `rename` | Multi-file coordinated rename with graph + text search |
|
||||
| `cypher` | Raw Cypher graph queries |
|
||||
| `route_map` | API route map — which components fetch which endpoints, and handlers |
|
||||
| `tool_map` | MCP/RPC tool definitions — where they're defined and handled |
|
||||
| `shape_check` | Validate API response shapes against consumers' property accesses |
|
||||
| `api_impact` | Pre-change impact report for an API route handler |
|
||||
| `explain` | Explain persisted taint findings (source→sink flows, `--pdg` indexes) |
|
||||
| `pdg_query` | Query control/data dependence at statement level (`--pdg` indexes) |
|
||||
| `group_list` | List configured repository groups |
|
||||
| `group_sync` | Rebuild a group's Contract Registry and cross-repo links |
|
||||
| Tool | What It Does |
|
||||
| ---------------- | ---------------------------------------------------------------------- |
|
||||
| `list_repos` | Discover all indexed repositories (paginated — `limit`/`offset`) |
|
||||
| `query` | Process-grouped hybrid search (BM25 + semantic + RRF) |
|
||||
| `context` | 360-degree symbol view — categorized refs, process participation |
|
||||
| `impact` | Blast radius analysis with depth grouping and confidence |
|
||||
| `trace` | Shortest directed path between two symbols (call + class-member edges) |
|
||||
| `detect_changes` | Git-diff impact — maps changed lines to affected processes |
|
||||
| `check` | Read-only structural checks against the indexed graph |
|
||||
| `rename` | Multi-file coordinated rename with graph + text search |
|
||||
| `cypher` | Raw Cypher graph queries |
|
||||
| `route_map` | API route map — which components fetch which endpoints, and handlers |
|
||||
| `tool_map` | MCP/RPC tool definitions — where they're defined and handled |
|
||||
| `shape_check` | Validate API response shapes against consumers' property accesses |
|
||||
| `api_impact` | Pre-change impact report for an API route handler |
|
||||
| `explain` | Explain persisted taint findings (source→sink flows, `--pdg` indexes) |
|
||||
| `pdg_query` | Query control/data dependence at statement level (`--pdg` indexes) |
|
||||
| `group_list` | List configured repository groups |
|
||||
| `group_sync` | Rebuild a group's Contract Registry and cross-repo links |
|
||||
|
||||
> Per-repo tools take an optional `repo` parameter (omit it when only one repo is indexed) and an optional `branch` for indexes pinned with `gitnexus analyze --branch`. Omitting `branch` queries the workspace index, which follows your checked-out working tree — switching branches and re-running `gitnexus analyze` updates it incrementally. `explain` and `pdg_query` need an index built with `gitnexus analyze --pdg`.
|
||||
|
||||
### Resources for instant context
|
||||
|
||||
| Resource | Purpose |
|
||||
| ---------------------------------------- | ---------------------------------------------------- |
|
||||
| `gitnexus://repos` | List all indexed repositories (read this first) |
|
||||
| `gitnexus://setup` | Setup and usage guidance for agents |
|
||||
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
|
||||
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
|
||||
| `gitnexus://repo/{name}/cluster/{name}` | Cluster members and details |
|
||||
| `gitnexus://repo/{name}/processes` | All execution flows |
|
||||
| `gitnexus://repo/{name}/process/{name}` | Full process trace with steps |
|
||||
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher queries |
|
||||
| `gitnexus://group/{name}/contracts` | A group's extracted contracts and cross-links |
|
||||
| `gitnexus://group/{name}/status` | Staleness of repos in a group |
|
||||
| Resource | Purpose |
|
||||
| --------------------------------------- | ---------------------------------------------------- |
|
||||
| `gitnexus://repos` | List all indexed repositories (read this first) |
|
||||
| `gitnexus://setup` | Setup and usage guidance for agents |
|
||||
| `gitnexus://repo/{name}/context` | Codebase stats, staleness check, and available tools |
|
||||
| `gitnexus://repo/{name}/clusters` | All functional clusters with cohesion scores |
|
||||
| `gitnexus://repo/{name}/cluster/{name}` | Cluster members and details |
|
||||
| `gitnexus://repo/{name}/processes` | All execution flows |
|
||||
| `gitnexus://repo/{name}/process/{name}` | Full process trace with steps |
|
||||
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher queries |
|
||||
| `gitnexus://group/{name}/contracts` | A group's extracted contracts and cross-links |
|
||||
| `gitnexus://group/{name}/status` | Staleness of repos in a group |
|
||||
|
||||
### 2 MCP prompts for guided workflows
|
||||
|
||||
| Prompt | What It Does |
|
||||
| --------------- | -------------------------------------------------------------------------- |
|
||||
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
|
||||
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
|
||||
| Prompt | What It Does |
|
||||
| --------------- | ------------------------------------------------------------------------- |
|
||||
| `detect_impact` | Pre-commit change analysis — scope, affected processes, risk level |
|
||||
| `generate_map` | Architecture documentation from the knowledge graph with mermaid diagrams |
|
||||
|
||||
### 6 agent skills installed to `.claude/skills/` automatically
|
||||
|
||||
|
|
@ -196,20 +196,21 @@ flowchart TB
|
|||
|
||||
`gitnexus setup` auto-detects your editors and writes the correct global MCP config. Run it once. To configure only selected integrations, pass `--coding-agent`/`-c` with a comma-separated list, e.g. `gitnexus setup -c cursor,codex`.
|
||||
|
||||
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|
||||
| ------------------------ | --- | ------ | ---------------------------------------------------------------------------------------- | ------------ |
|
||||
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
|
||||
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
|
||||
| Editor | MCP | Skills | Hooks (auto-augment) | Support |
|
||||
| ------------------------ | --- | ------ | ----------------------------------------------------------------------------------------------------------------- | ------------ |
|
||||
| **Claude Code** | Yes | Yes | Yes (PreToolUse + PostToolUse) | **Full** |
|
||||
| **Cursor** | Yes | Yes | Yes (postToolUse, [manual install](gitnexus-cursor-integration/README.md#hook-install)) | **Full** |
|
||||
| **Antigravity** (Google) | Yes | Yes | Yes (AfterTool, [Gemini CLI hooks schema](https://geminicli.com/docs/hooks/reference/))[¹](#fn-antigravity-hooks) | **Full** |
|
||||
| **Codex** | Yes | Yes | Yes (PreToolUse + PostToolUse, [Codex hooks](https://developers.openai.com/codex/hooks)) | **Full** |
|
||||
| **OpenCode** | Yes | Yes | — | MCP + Skills |
|
||||
| **CodeBuddy** (Tencent) | Yes | Yes | — | MCP + Skills |
|
||||
| **Qoder** (Alibaba) | Yes | Yes | — | MCP + Skills |
|
||||
| **Windsurf** | Yes | — | — | MCP |
|
||||
| **Codex** | Yes | Yes | Yes (PreToolUse + PostToolUse, [Codex hooks](https://developers.openai.com/codex/hooks)) | **Full** |
|
||||
| **OpenCode** | Yes | Yes | — | MCP + Skills |
|
||||
| **CodeBuddy** (Tencent) | Yes | Yes | — | MCP + Skills |
|
||||
| **Qoder** (Alibaba) | Yes | Yes | — | MCP + Skills |
|
||||
| **Windsurf** | Yes | — | — | MCP |
|
||||
|
||||
> **Claude Code** and **Codex** get the deepest integration: MCP tools + agent skills + PreToolUse hooks that enrich searches with graph context + PostToolUse hooks that detect a stale index after commits and prompt the agent to reindex.
|
||||
|
||||
<a id="fn-antigravity-hooks"></a>
|
||||
|
||||
> ¹ **Antigravity hooks** follow the [Gemini CLI hooks reference](https://geminicli.com/docs/hooks/reference/) (Antigravity 2.0 is the documented successor to Gemini CLI). Augmentation runs in `AfterTool` because `BeforeTool` has no context-injection channel in the Gemini contract — the agent sees graph context appended to the tool result via `hookSpecificOutput.additionalContext`. Stale-index hints land in the same channel after a successful `git commit/merge/rebase/cherry-pick/pull`. The schema may evolve if Antigravity-specific hook docs diverge from Gemini CLI's; the implementation will track those changes.
|
||||
|
||||
<details>
|
||||
|
|
@ -446,7 +447,7 @@ Commit a `.gitnexusrc` JSON file at the repo root to preconfigure recurring `ana
|
|||
"skipContextFiles": true, // alias of skipAgentsMd: keep your own AGENTS.md/CLAUDE.md
|
||||
"skipSkills": true, // don't install standard .claude/skills/gitnexus-* skills
|
||||
"embeddings": true, // generate embeddings by default
|
||||
"workerTimeout": 60
|
||||
"workerTimeout": 60,
|
||||
}
|
||||
```
|
||||
|
||||
|
|
@ -470,32 +471,32 @@ Notes:
|
|||
|
||||
Most `analyze` knobs are also CLI flags (`--workers`, `--worker-timeout`, `--max-file-size`, `--verbose`). Use the env-var form when you'd otherwise repeat the same flag every run, or when invoking GitNexus from a long-running host (MCP server, eval-server, CI shell) that already manages its own environment. CLI flags take precedence over env vars; env vars take precedence over built-in defaults.
|
||||
|
||||
| Variable | Default | Effect | Tune when… |
|
||||
| -------------------------------------- | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size (must be ≥ 1). Equivalent to `--workers <n>`. The worker pool is the sole parse path — there is no sequential parser, so `0` is rejected with an actionable error (the pool self-heals via quarantine + respawn). | Constrained containers (cgroup CPU limits) or CI runners with explicit quotas. To narrow down a worker crash set `1` for a single-worker pool — not `0`. |
|
||||
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
|
||||
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
|
||||
| `GITNEXUS_AUTH_TOKEN` | unset | Bearer token required when `eval-server` binds beyond loopback. May also be read from `.env.local` or `.env`; shell values take precedence. | Exposing the evaluation HTTP tools to a container, VM, or LAN. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED` | unset | When `1`, emits `[deferred-profile]` timing/progress logs for the post-chunk deferred resolution band (imports → heritage → buildHeritageMap → legacy call resolution). Implied by `GITNEXUS_VERBOSE`. | Diagnosing analyze stalls in "Resolving calls (all chunks)" on large Java/Kotlin repos (issue #1741) without the full verbose ingestion noise. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED_SLOW_MS` | `3000` (verbose) / `5000` | Per-file threshold in ms above which `processCallsFromExtracted` emits a `slow file …` log line. Parsed via `Number()`: accepts integers (`5000`), scientific notation (`2.5e3`), decimals (`.5`), and hex (`0x10`). Non-finite or non-positive values fall back to the default. | Hunting a few outlier files dominating the deferred call-resolution stage; lower to surface more, raise to focus only on the worst. |
|
||||
| `PROF_LBUG_LOAD` | unset | When `1`, emits one `[lbug-load prof]` summary line per `loadGraphToLbug` call breaking the graph-DB persistence wall into stages (`csv-emit` / `copy-nodes` / `copy-rels` / `fallback` / `total`) plus node & edge counts. Zero-cost when unset. | Attributing large-repo analyze wall time across CSV generation vs. LadybugDB `COPY` (issue #2203) — the analyze "emit" timing is the scope-resolution bucket, not this DB-write path. |
|
||||
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
|
||||
| `GITNEXUS_FTS_STEMMER` | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` for matching repository comments. Re-run `gitnexus analyze --repair-fts` after changing it. | Keyword search quality is poor for non-English comments or identifiers under English stemming. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold in bytes. Equivalent to `--wal-checkpoint-threshold <bytes>`. `-1` keeps LadybugDB's stock threshold (~16 MiB). Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. | You need a larger or smaller WAL auto-checkpoint threshold for your analyze workload. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD`| `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
|
||||
| `GITNEXUS_WORKER_SHUTDOWN_DRAIN_MS` | `30000` | Max wait at pool shutdown for a retired worker still inside native code. The worker is terminated at its next JS-safe point instead of mid-native-call (which aborts the whole process with `Napi::Error`, #2432); on expiry it is left running, unref'd, and terminated when it surfaces. | Shutdown latency matters more than draining a wedged worker (lower), or a legitimately-slow native grammar needs longer to surface (raise). |
|
||||
| `GITNEXUS_CPP_CAPTURE_BUDGET_MS` | `20000` | Per-file wall-clock budget for C++ capture extraction. On breach the file keeps the captures accumulated so far and logs a warning — the worker returns to JS instead of stalling in native-heavy loops (#2432). `0` expires immediately. | Pathological generated C++ that still exceeds the budget after the indexed lookups; raise for completeness, lower to fail-fast. |
|
||||
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Chunk boundary used for cache-key composition and dispatch. Smaller = finer-grained cache hits but more dispatch overhead. | Tuning incremental-analyze cache behavior on monorepos. |
|
||||
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
|
||||
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips the vendored grammar materialize for `tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`, and `tree-sitter-kotlin` at install time (and the Dart/Proto source builds). Those four won't be parsed; the install still succeeds. | Installing on a host without a C++ toolchain or where the vendored prebuilds don't match; willing to skip Dart/Proto/Swift/Kotlin parsing. |
|
||||
| `GITNEXUS_MCP_READ_ONLY` | unset | Set to `1` to expose only proven single-repository read tools and resources; `0` disables the policy and any other value fails startup. | The MCP server runs in an environment where graph mutation, raw Cypher, and cross-repository group routing must be unavailable. |
|
||||
| `GITNEXUS_MCP_ALLOWED_REPOS` | unset | Comma-separated allowlist of canonical indexed repository names or absolute paths. Invalid, ambiguous, or blank entries fail startup. | One MCP process must expose only a bounded subset of the repositories in the global registry. |
|
||||
| `GITNEXUS_MCP_DEFAULT_REPO` | unset | Canonical indexed repository name or absolute path used when a tool or resource omits its repository. Must belong to the allowlist when one is set. | Several repositories are available but unqualified MCP calls should resolve deterministically. |
|
||||
| `GITNEXUS_MCP_DEFAULT_MAX_TOKENS` | unset | Default positive-integer response budget for MCP `query`, `context`, and `impact`, estimated at four UTF-8 bytes per token. Explicit `maxTokens` wins. | Long MCP responses consume too much model context and callers cannot reliably add a per-request budget. |
|
||||
| Variable | Default | Effect | Tune when… |
|
||||
| ----------------------------------------------- | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_POOL_SIZE` | `cores - 1`, capped at 16 | Parse worker pool size (must be ≥ 1). Equivalent to `--workers <n>`. The worker pool is the sole parse path — there is no sequential parser, so `0` is rejected with an actionable error (the pool self-heals via quarantine + respawn). | Constrained containers (cgroup CPU limits) or CI runners with explicit quotas. To narrow down a worker crash set `1` for a single-worker pool — not `0`. |
|
||||
| `GITNEXUS_PARSE_CHUNK_CONCURRENCY` | `2` | Number of chunks whose file contents may be read into memory in parallel while the pool dispatches the current chunk. Worker dispatch itself stays serial. | Repos large enough to chunk (multi-MB total source) where disk I/O is a measurable fraction of analyze wall-clock. |
|
||||
| `GITNEXUS_VERBOSE` | unset | When `1`, enables verbose ingestion logs (skipped-file warnings, per-chunk throughput, parse-cache stats). Equivalent to `--verbose`. | Debugging an analyze that "completed" but seems to have missed files; tuning `--workers` / chunk concurrency against observable throughput. |
|
||||
| `GITNEXUS_AUTH_TOKEN` | unset | Bearer token required when `eval-server` binds beyond loopback. May also be read from `.env.local` or `.env`; shell values take precedence. | Exposing the evaluation HTTP tools to a container, VM, or LAN. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED` | unset | When `1`, emits `[deferred-profile]` timing/progress logs for the post-chunk deferred resolution band (imports → heritage → buildHeritageMap → legacy call resolution). Implied by `GITNEXUS_VERBOSE`. | Diagnosing analyze stalls in "Resolving calls (all chunks)" on large Java/Kotlin repos (issue #1741) without the full verbose ingestion noise. |
|
||||
| `GITNEXUS_PROFILE_DEFERRED_SLOW_MS` | `3000` (verbose) / `5000` | Per-file threshold in ms above which `processCallsFromExtracted` emits a `slow file …` log line. Parsed via `Number()`: accepts integers (`5000`), scientific notation (`2.5e3`), decimals (`.5`), and hex (`0x10`). Non-finite or non-positive values fall back to the default. | Hunting a few outlier files dominating the deferred call-resolution stage; lower to surface more, raise to focus only on the worst. |
|
||||
| `PROF_LBUG_LOAD` | unset | When `1`, emits one `[lbug-load prof]` summary line per `loadGraphToLbug` call breaking the graph-DB persistence wall into stages (`csv-emit` / `copy-nodes` / `copy-rels` / `fallback` / `total`) plus node & edge counts. Zero-cost when unset. | Attributing large-repo analyze wall time across CSV generation vs. LadybugDB `COPY` (issue #2203) — the analyze "emit" timing is the scope-resolution bucket, not this DB-write path. |
|
||||
| `GITNEXUS_MAX_FILE_SIZE` | `512` (KB) | Walker skip threshold in KB. Hard cap is `32768` (tree-sitter buffer ceiling). Equivalent to `--max-file-size <kb>`. | Indexing repos with intentionally-large source files (generated parsers, vendored bundles) that should still be parsed. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_TIMEOUT_MS` | `30000` | Worker idle timeout in milliseconds before retry/fallback. Equivalent to `--worker-timeout <seconds>` × 1000. | Slow-parsing files (large minified JS, deeply-nested TS types) that legitimately need more than 30s. |
|
||||
| `GITNEXUS_FTS_STEMMER` | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` for matching repository comments. Re-run `gitnexus analyze --repair-fts` after changing it. | Keyword search quality is poor for non-English comments or identifiers under English stemming. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold in bytes. Equivalent to `--wal-checkpoint-threshold <bytes>`. `-1` keeps LadybugDB's stock threshold (~16 MiB). Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. | You need a larger or smaller WAL auto-checkpoint threshold for your analyze workload. |
|
||||
| `GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES` | `8388608` (8 MB) | Per-job byte budget the pool will send to a worker in one `postMessage`. | Very large individual files; mostly diagnostic — bumping past 8 MB risks structured-clone memory pressure. |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per worker slot before the slot is dropped from the active rotation. Bounds respawn loops on a chronically-crashing slot. | Hosts where a flaky worker should retry more (raise) or fail-fast (lower) before the slot is dropped. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Combined with `timeoutBackoffFactor`, prevents exponentially-growing retries from stalling for hours. | Slow files that legitimately need long total retry windows; lower to fail-fast on stalls. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, every subsequent dispatch rejects until a fresh pool is created. | Hosts where a SIGSEGV-prone native grammar should trip the breaker sooner; CI runners that should fail loudly. |
|
||||
| `GITNEXUS_WORKER_SHUTDOWN_DRAIN_MS` | `30000` | Max wait at pool shutdown for a retired worker still inside native code. The worker is terminated at its next JS-safe point instead of mid-native-call (which aborts the whole process with `Napi::Error`, #2432); on expiry it is left running, unref'd, and terminated when it surfaces. | Shutdown latency matters more than draining a wedged worker (lower), or a legitimately-slow native grammar needs longer to surface (raise). |
|
||||
| `GITNEXUS_CPP_CAPTURE_BUDGET_MS` | `20000` | Per-file wall-clock budget for C++ capture extraction. On breach the file keeps the captures accumulated so far and logs a warning — the worker returns to JS instead of stalling in native-heavy loops (#2432). `0` expires immediately. | Pathological generated C++ that still exceeds the budget after the indexed lookups; raise for completeness, lower to fail-fast. |
|
||||
| `GITNEXUS_CHUNK_BYTE_BUDGET` | `2097152` (2 MB) | Chunk boundary used for cache-key composition and dispatch. Smaller = finer-grained cache hits but more dispatch overhead. | Tuning incremental-analyze cache behavior on monorepos. |
|
||||
| `GITNEXUS_NO_GITIGNORE` | unset | When set, skips `.gitignore` parsing. `.gitnexusignore` is still honored. | Indexing a repo whose `.gitignore` excludes files you actually want indexed (e.g., generated code committed for cross-repo lookup). |
|
||||
| `GITNEXUS_SKIP_OPTIONAL_GRAMMARS` | unset | When `=1` strictly, skips the vendored grammar materialize for `tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`, and `tree-sitter-kotlin` at install time (and the Dart/Proto source builds). Those four won't be parsed; the install still succeeds. | Installing on a host without a C++ toolchain or where the vendored prebuilds don't match; willing to skip Dart/Proto/Swift/Kotlin parsing. |
|
||||
| `GITNEXUS_MCP_READ_ONLY` | unset | Set to `1` to expose only proven single-repository read tools and resources; `0` disables the policy and any other value fails startup. | The MCP server runs in an environment where graph mutation, raw Cypher, and cross-repository group routing must be unavailable. |
|
||||
| `GITNEXUS_MCP_ALLOWED_REPOS` | unset | Comma-separated allowlist of canonical indexed repository names or absolute paths. Invalid, ambiguous, or blank entries fail startup. | One MCP process must expose only a bounded subset of the repositories in the global registry. |
|
||||
| `GITNEXUS_MCP_DEFAULT_REPO` | unset | Canonical indexed repository name or absolute path used when a tool or resource omits its repository. Must belong to the allowlist when one is set. | Several repositories are available but unqualified MCP calls should resolve deterministically. |
|
||||
| `GITNEXUS_MCP_DEFAULT_MAX_TOKENS` | unset | Default positive-integer response budget for MCP `query`, `context`, and `impact`, estimated at four UTF-8 bytes per token. Explicit `maxTokens` wins. | Long MCP responses consume too much model context and callers cannot reliably add a per-request budget. |
|
||||
|
||||
</details>
|
||||
|
||||
|
|
@ -739,10 +740,17 @@ gitnexus wiki --force
|
|||
gitnexus wiki --timeout <seconds> # LLM request timeout in seconds (default: disabled)
|
||||
gitnexus wiki --retries <n> # Max LLM retry attempts per request (default: 3)
|
||||
|
||||
# Allow a specific LAN/self-hosted HTTP LLM host (HTTPS is preferred for remote endpoints)
|
||||
gitnexus wiki --base-url http://llama-box.local:8080/v1 --allow-insecure-connection llama-box.local
|
||||
# Or set a comma-separated host allowlist:
|
||||
GITNEXUS_ALLOW_INSECURE_CONNECTION=llama-box.local,192.168.1.23
|
||||
|
||||
# Change the output language
|
||||
gitnexus wiki --lang <lang> # e.g. english, chinese, spanish, japanese
|
||||
```
|
||||
|
||||
For safety, `http://` LLM base URLs are allowed by default only for loopback hosts (`localhost`, `127.0.0.1`, `::1`). `--allow-insecure-connection` and `GITNEXUS_ALLOW_INSECURE_CONNECTION` accept exact hostnames or IP addresses only; do not include schemes, ports, paths, credentials, or wildcards.
|
||||
|
||||
The wiki generator reads the indexed graph structure, groups files into modules via LLM, generates per-module documentation pages, and creates an overview page — all with cross-references to the knowledge graph.
|
||||
|
||||
## Web UI (browser-based)
|
||||
|
|
@ -781,10 +789,10 @@ This starts the server on `http://localhost:4747` and the web UI on `http://loca
|
|||
|
||||
The official setup ships **two signed images**, published identically to **GitHub Container Registry** (GHCR) and **Docker Hub** — same build, same digest, same Cosign signature:
|
||||
|
||||
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
|
||||
| ----------------------------------------------------------------------- | ---------------------------------------------- | ------------------------------- |
|
||||
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
|
||||
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
|
||||
| Purpose | GHCR (default in `docker-compose.yaml`) | Docker Hub mirror |
|
||||
| ---------------------------------------------------------------------- | --------------------------------------------- | ------------------------------ |
|
||||
| CLI / `gitnexus serve` backend (HTTP API on port `4747`, MCP, indexer) | `ghcr.io/abhigyanpatwari/gitnexus:latest` | `akonlabs/gitnexus:latest` |
|
||||
| Static web UI (port `4173`) | `ghcr.io/abhigyanpatwari/gitnexus-web:latest` | `akonlabs/gitnexus-web:latest` |
|
||||
|
||||
A named volume (`gitnexus-data`) persists the global registry, indexes, and cloned repos at `/data/gitnexus` inside the server container. To make repos on your host machine indexable, set `WORKSPACE_DIR` before bringing the stack up:
|
||||
|
||||
|
|
@ -914,11 +922,11 @@ Enterprise includes:
|
|||
|
||||
Built by the community — not officially maintained, but worth checking out.
|
||||
|
||||
| Project | Author | Description |
|
||||
| ------------------------------------------------------------------------------ | ------------------------------------------------------- | ------------------------------------------------------------------------ |
|
||||
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
|
||||
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
|
||||
| [KiloCode MCP workflow](Documentation/kilo-code-mcp.md) | [@oktanishq](https://github.com/oktanishq) | Guide to connect GitNexus MCP to Kilo Code and verify tools. |
|
||||
| Project | Author | Description |
|
||||
| ----------------------------------------------------------------------------- | ------------------------------------------------------ | ----------------------------------------------------------------------- |
|
||||
| [pi-gitnexus](https://github.com/tintinweb/pi-gitnexus) | [@tintinweb](https://github.com/tintinweb) | GitNexus plugin for [pi](https://pi.dev) — `pi install npm:pi-gitnexus` |
|
||||
| [gitnexus-stable-ops](https://github.com/ShunsukeHayashi/gitnexus-stable-ops) | [@ShunsukeHayashi](https://github.com/ShunsukeHayashi) | Stable ops & deployment workflows (Miyabi ecosystem) |
|
||||
| [KiloCode MCP workflow](Documentation/kilo-code-mcp.md) | [@oktanishq](https://github.com/oktanishq) | Guide to connect GitNexus MCP to Kilo Code and verify tools. |
|
||||
|
||||
> Have a project built on GitNexus? Open a PR to add it here!
|
||||
|
||||
|
|
|
|||
|
|
@ -249,6 +249,8 @@ gitnexus clean # Delete index for current repo
|
|||
gitnexus clean --all --force # Delete all indexes
|
||||
gitnexus wiki [path] # Generate LLM-powered docs from knowledge graph
|
||||
gitnexus wiki --model <model> # Wiki with custom LLM model (default: minimax/minimax-m2.5)
|
||||
gitnexus wiki --base-url http://llama-box.local:8080/v1 --allow-insecure-connection llama-box.local
|
||||
# Allow an exact LAN/self-hosted HTTP LLM host; env: GITNEXUS_ALLOW_INSECURE_CONNECTION
|
||||
gitnexus doctor # Show runtime platform capabilities and embedding configuration
|
||||
|
||||
# Direct graph queries — the same tools the MCP server exposes, no MCP daemon needed
|
||||
|
|
@ -461,14 +463,14 @@ GitNexus uses optional DuckDB extensions for BM25 and vector search. The `gitnex
|
|||
|
||||
Configure the behavior with these environment variables:
|
||||
|
||||
| Variable | Values | Default | Effect |
|
||||
| -------------------------------------------- | ---------------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL` | `auto`, `load-only`, `never` | `auto` | `auto` runs one bounded install if LOAD fails — a plain `INSTALL`, escalating to `FORCE INSTALL` only when the LOAD error shows the present extension file is broken. `load-only` only uses already-installed extensions (recommended for offline / firewalled environments). `never` skips optional extensions entirely. |
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS` | positive integer | `15000` | Wall-clock budget for the out-of-process extension-install child before it is killed. |
|
||||
| `GITNEXUS_FTS_STEMMER` | supported LadybugDB stemmer | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` when that better matches repository comments and identifiers. Re-run `gitnexus analyze --repair-fts` after changing it. |
|
||||
| `GITNEXUS_FTS_CJK_SEGMENTATION` | `none`, `bigram` | `none` | `bigram` inserts overlapping character-bigram boundaries into Chinese/Japanese Han-ideograph spans in `content`/`description` before FTS indexing, so LadybugDB's space-only tokenizer can see sub-phrase word boundaries. Scoped to CJK Unified Ideographs only — Japanese Hiragana/Katakana and Korean Hangul are not currently segmented. Unlike `GITNEXUS_FTS_STEMMER`, this rewrites stored text — enabling it on an already-indexed repo requires a full `gitnexus analyze --force`; neither `--repair-fts` nor a plain incremental `analyze` applies it to previously-indexed files. Set the same value wherever `analyze` and search-serving processes (CLI query, MCP server, web server) run. |
|
||||
| `GITNEXUS_COMMUNITY_ENGINE` | `graphology`, `icebug`, `auto` | `graphology` | Community-detection engine used during analyze. `graphology` uses the bundled default path. `icebug` and `auto` currently behave identically: both try the experimental Icebug CSR path and fall back to Graphology if the optional native module is unavailable or incompatible. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | integer `>= -1` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold during analyze (bytes). Auto-checkpoint remains enabled; `-1` keeps Ladybug's stock ~16 MiB. Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. |
|
||||
| Variable | Values | Default | Effect |
|
||||
| -------------------------------------------- | ------------------------------ | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL` | `auto`, `load-only`, `never` | `auto` | `auto` runs one bounded install if LOAD fails — a plain `INSTALL`, escalating to `FORCE INSTALL` only when the LOAD error shows the present extension file is broken. `load-only` only uses already-installed extensions (recommended for offline / firewalled environments). `never` skips optional extensions entirely. |
|
||||
| `GITNEXUS_LBUG_EXTENSION_INSTALL_TIMEOUT_MS` | positive integer | `15000` | Wall-clock budget for the out-of-process extension-install child before it is killed. |
|
||||
| `GITNEXUS_FTS_STEMMER` | supported LadybugDB stemmer | `porter` | Stemmer used when rebuilding BM25/FTS indexes. Use `none` for CJK-heavy repositories, or a language stemmer such as `german`, `french`, or `spanish` when that better matches repository comments and identifiers. Re-run `gitnexus analyze --repair-fts` after changing it. |
|
||||
| `GITNEXUS_FTS_CJK_SEGMENTATION` | `none`, `bigram` | `none` | `bigram` inserts overlapping character-bigram boundaries into Chinese/Japanese Han-ideograph spans in `content`/`description` before FTS indexing, so LadybugDB's space-only tokenizer can see sub-phrase word boundaries. Scoped to CJK Unified Ideographs only — Japanese Hiragana/Katakana and Korean Hangul are not currently segmented. Unlike `GITNEXUS_FTS_STEMMER`, this rewrites stored text — enabling it on an already-indexed repo requires a full `gitnexus analyze --force`; neither `--repair-fts` nor a plain incremental `analyze` applies it to previously-indexed files. Set the same value wherever `analyze` and search-serving processes (CLI query, MCP server, web server) run. |
|
||||
| `GITNEXUS_COMMUNITY_ENGINE` | `graphology`, `icebug`, `auto` | `graphology` | Community-detection engine used during analyze. `graphology` uses the bundled default path. `icebug` and `auto` currently behave identically: both try the experimental Icebug CSR path and fall back to Graphology if the optional native module is unavailable or incompatible. |
|
||||
| `GITNEXUS_WAL_CHECKPOINT_THRESHOLD` | integer `>= -1` | `67108864` (64 MiB) | LadybugDB WAL auto-checkpoint threshold during analyze (bytes). Auto-checkpoint remains enabled; `-1` keeps Ladybug's stock ~16 MiB. Larger thresholds reduce checkpoint frequency but increase the WAL size at rotation time — choose a smaller value on disk-constrained environments. |
|
||||
|
||||
```bash
|
||||
# Offline/airgapped: never reach the network for extensions
|
||||
|
|
@ -533,13 +535,13 @@ For repositories with very large source files, `GITNEXUS_WORKER_SUB_BATCH_MAX_BY
|
|||
|
||||
Three env vars expose the pool's resilience layers (respawn budget, cumulative-timeout cap, circuit breaker). Defaults are tuned for typical repos; bump them when an analyze legitimately needs more retries, or lower them to fail-fast on a known-bad shape.
|
||||
|
||||
| Variable | Default | Effect |
|
||||
| ----------------------------------------------- | ----------------------- | --------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per slot before the slot is dropped from the active rotation. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
|
||||
| Variable | Default | Effect |
|
||||
| ----------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT` | `3` | Max replacement spawns per slot before the slot is dropped from the active rotation. |
|
||||
| `GITNEXUS_WORKER_MAX_CUMULATIVE_TIMEOUT_MS` | `5 × subBatchTimeoutMs` | Total retry wall-time budget per job before quarantining. Bounds exponentially-growing retry waits. |
|
||||
| `GITNEXUS_WORKER_CONSECUTIVE_FAILURE_THRESHOLD` | `max(3, poolSize)` | Per-slot consecutive deaths before the pool's circuit breaker trips. After tripping, dispatches require a fresh pool. |
|
||||
| `GITNEXUS_WORKER_SHUTDOWN_DRAIN_MS` | `30000` | Max wait at pool shutdown for a retired worker still inside native code — terminated at its next JS-safe point instead of mid-native-call, which would abort the process (`Napi::Error`, #2432). |
|
||||
| `GITNEXUS_CPP_CAPTURE_BUDGET_MS` | `20000` | Per-file wall-clock budget for C++ capture extraction; on breach the file keeps partial captures with a warning (#2432). `0` expires immediately. |
|
||||
| `GITNEXUS_CPP_CAPTURE_BUDGET_MS` | `20000` | Per-file wall-clock budget for C++ capture extraction; on breach the file keeps partial captures with a warning (#2432). `0` expires immediately. |
|
||||
|
||||
### Graph cleanup tuning
|
||||
|
||||
|
|
|
|||
|
|
@ -96,6 +96,7 @@ const OPTION_DESCRIPTION_KEYS = {
|
|||
'wiki|--concurrency <n>': 'help.option.wiki.concurrency',
|
||||
'wiki|--timeout <seconds>': 'help.option.wiki.timeout',
|
||||
'wiki|--retries <n>': 'help.option.wiki.retries',
|
||||
'wiki|--allow-insecure-connection <host>': 'help.option.wiki.allowInsecureConnection',
|
||||
'wiki|--gist': 'help.option.wiki.gist',
|
||||
'wiki|-v, --verbose': 'help.option.verbose',
|
||||
'wiki|--review': 'help.option.wiki.review',
|
||||
|
|
|
|||
|
|
@ -235,6 +235,8 @@ export const en = {
|
|||
'help.option.wiki.concurrency': 'Parallel LLM calls (default: 3)',
|
||||
'help.option.wiki.timeout': 'LLM request timeout in seconds (default: disabled)',
|
||||
'help.option.wiki.retries': 'Max LLM retry attempts per request (default: 3)',
|
||||
'help.option.wiki.allowInsecureConnection':
|
||||
'Allow exact host(s) for http:// LLM base URLs (comma-separated; HTTPS is preferred)',
|
||||
'help.option.wiki.gist': 'Publish wiki as a public GitHub Gist after generation',
|
||||
'help.option.wiki.review':
|
||||
'Stop after grouping to review module structure before generating pages',
|
||||
|
|
|
|||
|
|
@ -222,6 +222,8 @@ export const zhCN = {
|
|||
'help.option.wiki.concurrency': '并行 LLM 调用数(默认:3)',
|
||||
'help.option.wiki.timeout': 'LLM 请求超时时间(秒,默认:禁用)',
|
||||
'help.option.wiki.retries': '每个请求的最大 LLM 重试次数(默认:3)',
|
||||
'help.option.wiki.allowInsecureConnection':
|
||||
'允许 http:// LLM base URL 使用的精确主机(逗号分隔;推荐使用 HTTPS)',
|
||||
'help.option.wiki.gist': '生成后发布 Wiki 为公开 GitHub Gist',
|
||||
'help.option.wiki.review': '分组后停止,以便在生成页面前审查模块结构',
|
||||
'help.option.wiki.lang': '生成文档的输出语言(如 english、chinese、spanish、japanese)',
|
||||
|
|
|
|||
|
|
@ -304,6 +304,10 @@ program
|
|||
.option('--concurrency <n>', 'Parallel LLM calls (default: 3)', '3')
|
||||
.option('--timeout <seconds>', 'LLM request timeout in seconds (default: disabled)')
|
||||
.option('--retries <n>', 'Max LLM retry attempts per request (default: 3)')
|
||||
.option(
|
||||
'--allow-insecure-connection <host>',
|
||||
'Allow exact host(s) for http:// LLM base URLs (comma-separated; HTTPS is preferred)',
|
||||
)
|
||||
.option('--gist', 'Publish wiki as a public GitHub Gist after generation')
|
||||
.option('-v, --verbose', 'Enable verbose output (show LLM commands and responses)')
|
||||
.option('--review', 'Stop after grouping to review module structure before generating pages')
|
||||
|
|
|
|||
|
|
@ -17,7 +17,11 @@ import {
|
|||
saveCLIConfig,
|
||||
} from '../storage/repo-manager.js';
|
||||
import { WikiGenerator, type WikiOptions } from '../core/wiki/generator.js';
|
||||
import { resolveLLMConfig, type LLMProvider } from '../core/wiki/llm-client.js';
|
||||
import {
|
||||
parseLLMAllowedInsecureHttpHosts,
|
||||
resolveLLMConfig,
|
||||
type LLMProvider,
|
||||
} from '../core/wiki/llm-client.js';
|
||||
import { detectCursorCLI } from '../core/wiki/cursor-client.js';
|
||||
import { detectLocalCLI } from '../core/wiki/local-cli-client.js';
|
||||
import { logger } from '../core/logger.js';
|
||||
|
|
@ -37,6 +41,7 @@ export interface WikiCommandOptions {
|
|||
timeout?: string;
|
||||
retries?: string;
|
||||
lang?: string;
|
||||
allowInsecureConnection?: string;
|
||||
}
|
||||
|
||||
function parsePositiveIntegerOption(
|
||||
|
|
@ -185,9 +190,14 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
|
||||
let timeoutSeconds: number | undefined;
|
||||
let retries: number | undefined;
|
||||
let allowedInsecureHttpHosts: string[] | undefined;
|
||||
try {
|
||||
timeoutSeconds = parsePositiveIntegerOption(options?.timeout, '--timeout', 1000);
|
||||
retries = parsePositiveIntegerOption(options?.retries, '--retries');
|
||||
allowedInsecureHttpHosts =
|
||||
options?.allowInsecureConnection === undefined
|
||||
? undefined
|
||||
: parseLLMAllowedInsecureHttpHosts(options.allowInsecureConnection);
|
||||
} catch (error) {
|
||||
console.log(` Error: ${(error as Error).message}\n`);
|
||||
process.exitCode = 1;
|
||||
|
|
@ -245,6 +255,7 @@ const wikiCommandImpl = async (inputPath?: string, options?: WikiCommandOptions)
|
|||
provider: options?.provider,
|
||||
apiVersion: options?.apiVersion,
|
||||
isReasoningModel: options?.reasoningModel,
|
||||
allowedInsecureHttpHosts,
|
||||
});
|
||||
|
||||
// Run interactive setup if no saved config and no CLI flags provided
|
||||
|
|
|
|||
|
|
@ -35,6 +35,8 @@ export interface LLMConfig {
|
|||
requestTimeoutMs?: number;
|
||||
/** Max fetch attempts before giving up (default: 3). */
|
||||
maxAttempts?: number;
|
||||
/** Exact hostnames allowed for explicit http:// LLM endpoints. */
|
||||
allowedInsecureHttpHosts?: readonly string[];
|
||||
}
|
||||
|
||||
export interface LLMResponse {
|
||||
|
|
@ -94,6 +96,9 @@ export async function resolveLLMConfig(overrides?: Partial<LLMConfig>): Promise<
|
|||
apiVersion:
|
||||
overrides?.apiVersion || process.env.GITNEXUS_AZURE_API_VERSION || savedConfig.apiVersion,
|
||||
isReasoningModel: overrides?.isReasoningModel ?? savedConfig.isReasoningModel,
|
||||
allowedInsecureHttpHosts:
|
||||
overrides?.allowedInsecureHttpHosts ??
|
||||
parseLLMAllowedInsecureHttpHosts(process.env[LLM_ALLOW_INSECURE_CONNECTION_ENV]),
|
||||
};
|
||||
}
|
||||
|
||||
|
|
@ -117,6 +122,38 @@ function isTimeoutLikeError(err: unknown): boolean {
|
|||
return /time(d)?\s*out|timeout/i.test(err.message);
|
||||
}
|
||||
|
||||
export const LLM_ALLOW_INSECURE_CONNECTION_ENV = 'GITNEXUS_ALLOW_INSECURE_CONNECTION';
|
||||
|
||||
function normalizeAllowedInsecureHttpHost(host: string): string {
|
||||
const trimmed = host.trim().toLowerCase();
|
||||
const fail = () => {
|
||||
throw new Error(
|
||||
`--allow-insecure-connection / ${LLM_ALLOW_INSECURE_CONNECTION_ENV} entries must be exact hostnames or IP addresses`,
|
||||
);
|
||||
};
|
||||
if (!trimmed || /[/@?#]/.test(trimmed)) fail();
|
||||
|
||||
if (trimmed.startsWith('[')) {
|
||||
if (!trimmed.endsWith(']')) fail();
|
||||
const normalized = trimmed.slice(1, -1);
|
||||
if (!normalized || /[\[\]]/.test(normalized)) fail();
|
||||
return normalized;
|
||||
}
|
||||
|
||||
if (/[\[\]]/.test(trimmed)) fail();
|
||||
if ((trimmed.match(/:/g)?.length ?? 0) === 1) {
|
||||
// URL.hostname never includes the port, so accepting "host:port" would
|
||||
// create a confusing no-op allowlist entry.
|
||||
fail();
|
||||
}
|
||||
return trimmed;
|
||||
}
|
||||
|
||||
export function parseLLMAllowedInsecureHttpHosts(value: string | undefined): string[] {
|
||||
if (value === undefined || value.trim() === '') return [];
|
||||
return [...new Set(value.split(',').map(normalizeAllowedInsecureHttpHost))];
|
||||
}
|
||||
|
||||
/**
|
||||
* Validate that a base URL supplied for LLM API calls is a safe HTTP/HTTPS
|
||||
* endpoint (CWE-918 / CodeQL js/http-to-file-access).
|
||||
|
|
@ -124,15 +161,22 @@ function isTimeoutLikeError(err: unknown): boolean {
|
|||
* Allowed:
|
||||
* - https:// with any hostname (public LLM APIs, Azure, OpenRouter, …)
|
||||
* - http:// restricted to localhost / 127.0.0.1 (local servers: Ollama, LiteLLM, …)
|
||||
* - http:// to exact hosts explicitly allowlisted for LAN/self-hosted LLMs
|
||||
*
|
||||
* Rejected:
|
||||
* - file://, data:, javascript:, and any other non-HTTP scheme
|
||||
* - http:// aimed at non-loopback hosts (avoids SSRF against internal networks)
|
||||
* - http:// aimed at non-loopback hosts unless explicitly allowlisted
|
||||
* (avoids SSRF against internal networks by default)
|
||||
*
|
||||
* Throws with a descriptive message on validation failure so callers surface a
|
||||
* clear error rather than an opaque network error.
|
||||
*/
|
||||
export function validateLLMBaseUrl(baseUrl: string): void {
|
||||
export function validateLLMBaseUrl(
|
||||
baseUrl: string,
|
||||
allowedInsecureHttpHosts: readonly string[] = parseLLMAllowedInsecureHttpHosts(
|
||||
process.env[LLM_ALLOW_INSECURE_CONNECTION_ENV],
|
||||
),
|
||||
): void {
|
||||
let parsed: URL;
|
||||
try {
|
||||
parsed = new URL(baseUrl);
|
||||
|
|
@ -150,10 +194,12 @@ export function validateLLMBaseUrl(baseUrl: string): void {
|
|||
// Node's URL parser preserves IPv6 brackets in hostname (e.g. "[::1]"),
|
||||
// so strip them before comparing to bare address literals.
|
||||
const host = parsed.hostname.toLowerCase().replace(/^\[|\]$/g, '');
|
||||
if (host !== 'localhost' && host !== '127.0.0.1' && host !== '::1') {
|
||||
const allowedHosts = new Set(allowedInsecureHttpHosts.map(normalizeAllowedInsecureHttpHost));
|
||||
if (host !== 'localhost' && host !== '127.0.0.1' && host !== '::1' && !allowedHosts.has(host)) {
|
||||
// Use parsed.origin (scheme+host+port, no credentials) instead of the full URL.
|
||||
throw new Error(
|
||||
`Insecure http:// LLM base URLs are only allowed for localhost/127.0.0.1. ` +
|
||||
`Insecure http:// LLM base URLs are only allowed for localhost/127.0.0.1 ` +
|
||||
`or hosts listed by --allow-insecure-connection / ${LLM_ALLOW_INSECURE_CONNECTION_ENV}. ` +
|
||||
`Use https:// for remote endpoints (got ${parsed.origin})`,
|
||||
);
|
||||
}
|
||||
|
|
@ -212,7 +258,7 @@ export async function callLLM(
|
|||
options?: CallLLMOptions,
|
||||
): Promise<LLMResponse> {
|
||||
// Validate base URL before any fetch (CodeQL js/http-to-file-access)
|
||||
validateLLMBaseUrl(config.baseUrl);
|
||||
validateLLMBaseUrl(config.baseUrl, config.allowedInsecureHttpHosts);
|
||||
|
||||
const messages: Array<{ role: string; content: string }> = [];
|
||||
if (systemPrompt) {
|
||||
|
|
|
|||
|
|
@ -579,6 +579,17 @@ describe('wikiCommand --timeout mapping', () => {
|
|||
|
||||
async function loadWikiCommandHarness() {
|
||||
let capturedConfig: Record<string, unknown> | undefined;
|
||||
const resolveLLMConfig = vi.fn().mockImplementation((overrides = {}) =>
|
||||
Promise.resolve({
|
||||
apiKey: 'sk-test',
|
||||
baseUrl: 'https://api.openai.com/v1',
|
||||
model: 'gpt-4o',
|
||||
maxTokens: 16_384,
|
||||
temperature: 0,
|
||||
provider: 'openai',
|
||||
...overrides,
|
||||
}),
|
||||
);
|
||||
const generatorCtor = vi
|
||||
.fn()
|
||||
.mockImplementation(function (_repoPath, _storagePath, _lbugPath, config) {
|
||||
|
|
@ -609,14 +620,7 @@ describe('wikiCommand --timeout mapping', () => {
|
|||
const actual = await importOriginal<typeof import('../../src/core/wiki/llm-client.js')>();
|
||||
return {
|
||||
...actual,
|
||||
resolveLLMConfig: vi.fn().mockResolvedValue({
|
||||
apiKey: 'sk-test',
|
||||
baseUrl: 'https://api.openai.com/v1',
|
||||
model: 'gpt-4o',
|
||||
maxTokens: 16_384,
|
||||
temperature: 0,
|
||||
provider: 'openai',
|
||||
}),
|
||||
resolveLLMConfig,
|
||||
};
|
||||
});
|
||||
vi.doMock('../../src/core/wiki/generator.js', () => ({
|
||||
|
|
@ -642,6 +646,7 @@ describe('wikiCommand --timeout mapping', () => {
|
|||
generatorCtor,
|
||||
consoleSpy,
|
||||
getCapturedConfig: () => capturedConfig,
|
||||
resolveLLMConfig,
|
||||
};
|
||||
}
|
||||
|
||||
|
|
@ -671,6 +676,25 @@ describe('wikiCommand --timeout mapping', () => {
|
|||
expect(harness.generatorCtor).toHaveBeenCalledTimes(1);
|
||||
expect(harness.getCapturedConfig()?.maxAttempts).toBe(5);
|
||||
});
|
||||
|
||||
it('maps --allow-insecure-connection to allowedInsecureHttpHosts', async () => {
|
||||
const harness = await loadWikiCommandHarness();
|
||||
|
||||
await harness.wikiCommand('/tmp/repo', {
|
||||
allowInsecureConnection: 'llama-box.local,192.168.1.23,llama-box.local',
|
||||
});
|
||||
|
||||
expect(harness.resolveLLMConfig).toHaveBeenCalledWith(
|
||||
expect.objectContaining({
|
||||
allowedInsecureHttpHosts: ['llama-box.local', '192.168.1.23'],
|
||||
}),
|
||||
);
|
||||
expect(harness.generatorCtor).toHaveBeenCalledTimes(1);
|
||||
expect(harness.getCapturedConfig()?.allowedInsecureHttpHosts).toEqual([
|
||||
'llama-box.local',
|
||||
'192.168.1.23',
|
||||
]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('wikiCommand timeout messaging', () => {
|
||||
|
|
|
|||
|
|
@ -2,9 +2,12 @@ import { describe, it, expect, vi, afterEach } from 'vitest';
|
|||
|
||||
// Import the function we'll add in the next step
|
||||
import {
|
||||
LLM_ALLOW_INSECURE_CONNECTION_ENV,
|
||||
isAzureProvider,
|
||||
isReasoningModel,
|
||||
buildRequestUrl,
|
||||
parseLLMAllowedInsecureHttpHosts,
|
||||
resolveLLMConfig,
|
||||
validateLLMBaseUrl,
|
||||
} from '../../src/core/wiki/llm-client.js';
|
||||
|
||||
|
|
@ -470,6 +473,10 @@ describe('readSSEStream — content_filter handling', () => {
|
|||
});
|
||||
|
||||
describe('validateLLMBaseUrl', () => {
|
||||
afterEach(() => {
|
||||
delete process.env[LLM_ALLOW_INSECURE_CONNECTION_ENV];
|
||||
});
|
||||
|
||||
it('allows https:// for any public host', () => {
|
||||
expect(() => validateLLMBaseUrl('https://api.openai.com/v1')).not.toThrow();
|
||||
expect(() => validateLLMBaseUrl('https://openrouter.ai/api/v1')).not.toThrow();
|
||||
|
|
@ -498,6 +505,46 @@ describe('validateLLMBaseUrl', () => {
|
|||
);
|
||||
});
|
||||
|
||||
it('allows explicit http:// hosts only when exactly allowlisted', () => {
|
||||
expect(() =>
|
||||
validateLLMBaseUrl('http://llama-box.local:8080/v1', ['llama-box.local']),
|
||||
).not.toThrow();
|
||||
expect(() =>
|
||||
validateLLMBaseUrl('http://LLAMA-BOX.local:8080/v1', [' llama-box.LOCAL ']),
|
||||
).not.toThrow();
|
||||
expect(() => validateLLMBaseUrl('http://llama-box.local.evil/v1', ['llama-box.local'])).toThrow(
|
||||
'Insecure http://',
|
||||
);
|
||||
expect(() => validateLLMBaseUrl('http://192.168.1.23:8080/v1', ['192.168.1.23'])).not.toThrow();
|
||||
});
|
||||
|
||||
it('parses and validates comma-separated insecure HTTP host allowlists', () => {
|
||||
expect(
|
||||
parseLLMAllowedInsecureHttpHosts(' llama-box.local,192.168.1.23,llama-box.local '),
|
||||
).toEqual(['llama-box.local', '192.168.1.23']);
|
||||
expect(parseLLMAllowedInsecureHttpHosts('[fe80::1]')).toEqual(['fe80::1']);
|
||||
expect(() => parseLLMAllowedInsecureHttpHosts('http://llama-box.local')).toThrow(
|
||||
'exact hostnames or IP addresses',
|
||||
);
|
||||
expect(() => parseLLMAllowedInsecureHttpHosts('llama-box.local/path')).toThrow(
|
||||
'exact hostnames or IP addresses',
|
||||
);
|
||||
expect(() => parseLLMAllowedInsecureHttpHosts('llama-box.local:8080')).toThrow(
|
||||
'exact hostnames or IP addresses',
|
||||
);
|
||||
expect(() => parseLLMAllowedInsecureHttpHosts('[fe80::1]:8080')).toThrow(
|
||||
'exact hostnames or IP addresses',
|
||||
);
|
||||
});
|
||||
|
||||
it('resolveLLMConfig reads insecure HTTP hosts from env when no override is passed', async () => {
|
||||
process.env[LLM_ALLOW_INSECURE_CONNECTION_ENV] = 'llama-box.local,192.168.1.23';
|
||||
|
||||
const config = await resolveLLMConfig();
|
||||
|
||||
expect(config.allowedInsecureHttpHosts).toEqual(['llama-box.local', '192.168.1.23']);
|
||||
});
|
||||
|
||||
it('rejects http:// hostname-spoofing attempts', () => {
|
||||
// Full-hostname comparison prevents prefix/suffix attacks
|
||||
expect(() => validateLLMBaseUrl('http://localhost.evil.com/v1')).toThrow('Insecure http://');
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue