feat(grammars): unify tree-sitter-swift with the vendored-source build pipeline

Swift was the last grammar handled differently — it shipped only upstream
prebuilds, while c/dart/proto/kotlin vendor their grammar source and use a
prefer-prebuild -> source-build-fallback activation script. Vendor swift's
source so all five are handled identically (one uniform build path).

- vendor/tree-sitter-swift: add binding.gyp (win-hardened), bindings/node/
  binding.cc, src/parser.c (ABI-14 default, ~18 MB), src/scanner.c, and
  src/tree_sitter/ headers. The 6/6 prebuilds are retained. The legacy
  parser_abi13.c alternate is intentionally not vendored.
- build-tree-sitter-swift.cjs: rewrite the prebuild probe into the dart-style
  prefer-prebuild then source-build fallback (keeps the GITNEXUS_SKIP gate and
  the never-exit-non-zero postinstall invariant).
- build-tree-sitter-prebuilds.yml: register swift (kind 'vendored'); add its
  package.json to the version-gated pull_request paths and a validate snippet.
- prebuild-coverage guard auto-moves swift into the source-fallback cohort
  (binding.gyp now present); refresh the stale "swift is prebuild-only" comments.
- tests: add build-tree-sitter-swift-probe.test.ts; fix the pre-existing
  build-tree-sitter-kotlin-probe.test.ts breakage (it still asserted the old
  probe strings after kotlin's dart-style conversion); assert swift's vendored
  source in cli-commands.test.ts.
- docs: README / .devcontainer / kotlin vendor README — swift's prebuilds are
  now GitNexus-cross-built from vendored source like the rest, not upstream-only.

Verified: swift source-builds against node-addon-api@8 -> N-API binary -> loads
against the pinned tree-sitter@0.21.1 (ABI 14) -> parses cleanly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Gergo Magyar 2026-06-09 11:40:47 +00:00
parent 50253d5d24
commit 140cd067e2
19 changed files with 542781 additions and 74 deletions

View file

@ -310,7 +310,7 @@ VS Code's Ports panel shows forwarded ports once their listener starts.
- **LadybugDB integration tests may fail in containers** (file-locking, `AGENTS.md` § Testing). Default to `npm run test:unit` inside the container; run integration tests on the host. Tracking issue: documented as a known limitation.
- **Single-writer LadybugDB constraint** (`GUARDRAILS.md` § LadybugDB lock). Don't run `gitnexus analyze` on the host and inside the container against the same `.gitnexus/` directory simultaneously — the second writer will get `database busy`.
- **Native grammar builds add ~30s to first install.** Tree-sitter Dart/Proto compile from vendored source during `gitnexus`'s `postinstall` (a toolchain is needed only if no prebuild matches the host); Swift and Kotlin are vendored with prebuilt `.node` binaries (`node-gyp-build` selects one — no compile). Set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` (in your shell or `remoteEnv`, then rebuild) to skip all four; each loses parsing for the affected language(s), and the install still succeeds.
- **Native grammar builds add ~30s to first install.** Tree-sitter Dart/Proto/Swift/Kotlin are all vendored uniformly: `node-gyp-build` picks a committed GitNexus-built prebuilt `.node` at install time (no compile), and only falls back to compiling from the vendored source during `postinstall` if no prebuild matches the host (then a toolchain is needed). Set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` (in your shell or `remoteEnv`, then rebuild) to skip all four; each loses parsing for the affected language(s), and the install still succeeds.
- **`tree-sitter-kotlin`/`tree-sitter-swift` warnings on install** only appear when no prebuild matches the platform-arch (per `AGENTS.md`); they are non-fatal — parsing for that language is simply unavailable.
- **`.mcp.json` works inside the container**: `npx -y gitnexus@latest mcp` resolves cleanly because npm registry is reachable and the workspace bind mount exposes the same `.mcp.json` the host sees.
- **Husky pre-commit fires inside the container** without extra setup. The root `npm install` (run automatically in `postCreateCommand`) installs the hook via `package.json` `prepare`.

View file

@ -6,12 +6,19 @@ name: Build tree-sitter prebuilds
# risk for any tree-sitter grammar" pipeline.
#
# Grammars covered here (the at-risk set — everything else already ships 6
# upstream prebuilds AND stays dependency-review-tracked, so it is left alone):
# - tree-sitter-dart (vendored; currently `npx node-gyp rebuild` at postinstall)
# - tree-sitter-proto (vendored; currently `npx node-gyp rebuild` at postinstall)
# - tree-sitter-kotlin (third-party optionalDependency, source-only — needs a
# vendor skeleton in place before this workflow targets it)
# (tree-sitter-swift already vendors upstream-shipped prebuilds and needs nothing.)
# upstream prebuilds AND stays dependency-review-tracked, so it is left alone).
# All five are vendored under gitnexus/vendor/; `kind` (below) only picks where
# the build job fetches the C source to compile:
# - tree-sitter-c (vendored prebuild-only; built from the published npm
# package — closes upstream's 4/6 ARM gap #2116 for a
# REQUIRED grammar)
# - tree-sitter-dart (vendored source; built from gitnexus/vendor/)
# - tree-sitter-proto (vendored source; built from gitnexus/vendor/)
# - tree-sitter-kotlin (vendored source; built from the published npm package —
# upstream ships source only)
# - tree-sitter-swift (vendored source; built from gitnexus/vendor/ — its
# prebuilds were originally upstream-shipped, now
# GitNexus-cross-built like the rest for uniformity)
#
# Output: gitnexus/vendor/<grammar>/prebuilds/<platform-arch>/<grammar>.node for
# all 6 targets ({linux,darwin,win32}-{x64,arm64}). tree-sitter grammars are
@ -37,7 +44,7 @@ on:
workflow_dispatch:
inputs:
grammars:
description: 'Comma-separated grammar shortnames to build (dart,proto,kotlin), or "all".'
description: 'Comma-separated grammar shortnames to build (c,dart,proto,kotlin,swift), or "all".'
required: false
type: string
default: 'all'
@ -64,6 +71,7 @@ on:
- 'gitnexus/vendor/tree-sitter-dart/package.json'
- 'gitnexus/vendor/tree-sitter-proto/package.json'
- 'gitnexus/vendor/tree-sitter-kotlin/package.json'
- 'gitnexus/vendor/tree-sitter-swift/package.json'
# Transition window: kotlin's pin still lives here until it is vendored.
- 'gitnexus/package.json'
- 'gitnexus/package-lock.json'
@ -126,6 +134,10 @@ jobs:
dart: { name: 'tree-sitter-dart', kind: 'vendored' },
proto: { name: 'tree-sitter-proto', kind: 'vendored' },
kotlin: { name: 'tree-sitter-kotlin', kind: 'npm' },
// swift is vendored WITH its source (parser.c/scanner.c/binding.gyp),
// so it builds from gitnexus/vendor/ like dart/proto. Its prebuilds
// were originally upstream-shipped; rebuilding them here unifies it.
swift: { name: 'tree-sitter-swift', kind: 'vendored' },
};
const PLATFORMS = [
{ platform_arch: 'linux-x64', os: 'ubuntu-24.04' },
@ -330,6 +342,7 @@ jobs:
dart: "void main() { print(\"hi\"); }",
proto: "syntax = \"proto3\";\nmessage M { int32 id = 1; }",
kotlin: "fun main() { println(\"hi\") }",
swift: "func greet() { print(\"hi\") }",
};
const lang = require("node-gyp-build")(process.cwd());
const Parser = require("tree-sitter");

View file

@ -94,8 +94,9 @@ jobs:
# 1. Static, offline: assert every grammar's compiled ABI loads on the
# pinned runtime (check-tree-sitter-upgrade-readiness.py --assert-current).
# 2. Dynamic: run the parser-loader ABI load-smoke on the OS matrix so an
# ABI-incompatible prebuilt (esp. the binary-only Swift vendor, which the
# static check can't introspect) fails on the platform it ships to.
# ABI-incompatible committed vendor prebuilt (e.g. Swift's — the static
# check introspects source, not the shipped .node) fails on the platform
# it ships to.
abi-assert:
name: tree-sitter ABI (${{ matrix.os }})
strategy:

View file

@ -119,7 +119,7 @@ To configure MCP for your editor, run `npx gitnexus setup` once — or set it up
> **Faster install (no C++ toolchain needed):** set `GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1` before `npm install -g gitnexus` to skip the vendored grammar materialize/build for `tree-sitter-dart`, `tree-sitter-proto`, `tree-sitter-swift`, and `tree-sitter-kotlin` — those four won't be parsed, but install completes in seconds without `python3`/`make`/`g++`. Strict `=1` only — any other value falls through to the rebuild. See the `tree-sitter-kotlin` note below.
>
> **About `tree-sitter-kotlin`:** like Dart/Proto/Swift, Kotlin is a **vendored** grammar (under `gitnexus/vendor/tree-sitter-kotlin`). Upstream `tree-sitter-kotlin` ships **source only** (no prebuilt binaries), so — unlike Swift, whose prebuilds are copied from upstream — GitNexus builds the Kotlin platform prebuilds itself (via the `build-tree-sitter-prebuilds` GitHub Actions workflow) and vendors them. `node-gyp-build` selects the right `.node` at require time, so **no C/C++ toolchain is needed**. If no prebuild matches your platform-arch, only Kotlin (`.kt`/`.kts`) parsing is unavailable; the rest of `gitnexus` is unaffected.
> **About `tree-sitter-kotlin`:** like Dart/Proto/Swift, Kotlin is a **vendored** grammar (under `gitnexus/vendor/tree-sitter-kotlin`). Upstream `tree-sitter-kotlin` ships **source only** (no prebuilt binaries), so GitNexus builds the Kotlin platform prebuilds itself (via the `build-tree-sitter-prebuilds` GitHub Actions workflow) and vendors them — the same uniform pipeline now used for Dart, Proto, and Swift (Swift's prebuilds were originally copied from upstream; they're now GitNexus-cross-built too). `node-gyp-build` selects the right `.node` at require time, so **no C/C++ toolchain is needed**. If no prebuild matches your platform-arch, only Kotlin (`.kt`/`.kts`) parsing is unavailable; the rest of `gitnexus` is unaffected.
### MCP Setup

View file

@ -1,39 +1,71 @@
#!/usr/bin/env node
/**
* Probe tree-sitter-swift prebuild availability at install time.
* Activate the tree-sitter-swift native binding after materialize-vendor-grammars.cjs.
*
* The vendored package ships platform prebuilds; node-gyp-build selects the
* correct binary at require time. This script calls node-gyp-build once
* against the materialized package so a missing-prebuild failure surfaces
* as an install-time warning (with the rest of the gitnexus install
* succeeding) rather than as a runtime error the first time Swift parsing
* is requested. The result is discarded — it does not copy, register, or
* mutate anything; the runtime require() path in parser-loader does the
* actual load. Running this probe here instead of an npm `install` script
* on the vendored package preserves the #836 hygiene (no scripts.install
* inside vendor/).
* Swift is vendored. Unlike its historical prebuild-only form, the grammar
* source (parser.c/scanner.c/binding.gyp + src/) is now ALSO vendored, so this
* script mirrors Dart/Proto/Kotlin/C exactly: prefer a committed prebuild for
* this platform-arch (toolchain-free); otherwise build from the vendored source
* so Swift parsing still works on any host with a toolchain — e.g. CI, where the
* GitNexus-cross-built prebuilds may not yet be vendored. The committed
* prebuilds for every platform-arch are produced by
* .github/workflows/build-tree-sitter-prebuilds.yml.
*
* MUST NEVER throw or exit non-zero — it must never break `gitnexus` install.
*/
const fs = require('fs');
const path = require('path');
const { execSync } = require('child_process');
// Opt-out: Swift is optional, so the env var skips its build entirely (also
// skipped at materialize). Strict `=== '1'` only.
if (process.env.GITNEXUS_SKIP_OPTIONAL_GRAMMARS === '1') {
console.warn('[tree-sitter-swift] Skipping prebuild probe (GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1).');
console.warn(
'[tree-sitter-swift] Skipping build (GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1). Swift parsing will be unavailable until reinstalled without the env var.',
);
process.exit(0);
}
const swiftDir = path.join(__dirname, '..', 'node_modules', 'tree-sitter-swift');
const bindingGyp = path.join(swiftDir, 'binding.gyp');
const bindingNode = path.join(swiftDir, 'build', 'Release', 'tree_sitter_swift_binding.node');
try {
if (!fs.existsSync(path.join(swiftDir, 'bindings', 'node', 'index.js'))) {
if (!fs.existsSync(bindingGyp) || fs.existsSync(bindingNode)) {
process.exit(0);
}
const nodeGypBuild = require('node-gyp-build');
nodeGypBuild(swiftDir);
// Prefer a committed prebuild for this platform-arch (no toolchain needed).
try {
require('node-gyp-build').path(swiftDir);
process.exit(0);
} catch {
// No matching prebuild — fall through to the source build below.
}
try {
require.resolve('node-addon-api');
require.resolve('node-gyp-build');
} catch (resolveErr) {
console.warn(
'[tree-sitter-swift] Skipping build: hoisted build deps not resolvable (%s).',
resolveErr.message,
);
console.warn(
'[tree-sitter-swift] Swift parsing will be unavailable until a prebuild or toolchain is present.',
);
process.exit(0);
}
console.log(
'[tree-sitter-swift] No prebuild for this platform — building native binding from source...',
);
execSync('npx node-gyp rebuild', { cwd: swiftDir, stdio: 'pipe', timeout: 180000 });
console.log('[tree-sitter-swift] Native binding built successfully');
} catch (err) {
console.warn('[tree-sitter-swift] Prebuild probe failed:', err.message);
console.warn('[tree-sitter-swift] Could not build native binding:', err.message);
console.warn(
'[tree-sitter-swift] Swift parsing will be unavailable. Non-Swift functionality is unaffected.',
'[tree-sitter-swift] Swift (.swift) parsing will be unavailable. Non-Swift functionality is unaffected.',
);
process.exit(0);
}

View file

@ -6,20 +6,22 @@ import path from 'node:path';
import { fileURLToPath } from 'node:url';
/**
* Behavioral coverage for the postinstall probe `scripts/build-tree-sitter-kotlin.cjs`.
* Behavioral coverage for the postinstall activation script
* `scripts/build-tree-sitter-kotlin.cjs`.
*
* Kotlin is a vendored grammar (like Swift): the probe calls `node-gyp-build`
* against the materialized package to surface a single install-time warning when
* no prebuild matches this platform-arch, instead of a first-use runtime error.
* Its hard invariant is that it MUST NEVER exit non-zero — it runs in
* `gitnexus`'s postinstall, so a non-zero exit would break `npm install gitnexus`
* for every user. This suite executes the real script bytes across its branches
* and asserts exit code 0 every time.
* Kotlin is a vendored grammar (like Swift/Dart/Proto/C). The script prefers a
* committed prebuild for this platform-arch (toolchain-free); if none matches it
* source-builds from the vendored grammar source. Its hard invariant is that it
* MUST NEVER exit non-zero — it runs in `gitnexus`'s postinstall, so a non-zero
* exit would break `npm install gitnexus` for every user. This suite executes
* the real script bytes across its branches and asserts exit code 0 every time.
*
* The probe is copied into an isolated temp `scripts/` dir so its
* The script is copied into an isolated temp `scripts/` dir so its
* `__dirname`-relative `../node_modules/tree-sitter-kotlin` resolves under our
* control (absent dir, or a present-but-no-prebuild dir) without touching the
* repo's real node_modules.
* control (absent dir, or a present-but-unbuildable dir) without touching the
* repo's real node_modules. The temp dir has no reachable `node-gyp-build` /
* `node-addon-api`, so the source-build path stops at the "hoisted build deps
* not resolvable" guard (still exit 0) instead of invoking a real compile.
*/
const probeSource = readFileSync(
@ -27,13 +29,15 @@ const probeSource = readFileSync(
'utf8',
);
const UNAVAILABLE = 'Kotlin (.kt/.kts) parsing will be unavailable';
// Catch-branch sentinel (only printed when an actual node-gyp build is attempted
// and throws) — must NOT appear on the deps-unavailable guard path.
const CATCH_UNAVAILABLE = 'Kotlin (.kt/.kts) parsing will be unavailable';
let tmpRoot: string;
let scriptPath: string;
beforeAll(() => {
tmpRoot = mkdtempSync(path.join(tmpdir(), 'gn-kotlin-probe-'));
tmpRoot = mkdtempSync(path.join(tmpdir(), 'gn-kotlin-build-'));
mkdirSync(path.join(tmpRoot, 'scripts'), { recursive: true });
scriptPath = path.join(tmpRoot, 'scripts', 'build-tree-sitter-kotlin.cjs');
writeFileSync(scriptPath, probeSource);
@ -53,41 +57,43 @@ function runProbe(overrides: Record<string, string | undefined>) {
if (v === undefined) delete env[k];
else env[k] = v;
}
return spawnSync(process.execPath, [scriptPath], { env, encoding: 'utf8', timeout: 10_000 });
return spawnSync(process.execPath, [scriptPath], { env, encoding: 'utf8', timeout: 30_000 });
}
describe('build-tree-sitter-kotlin.cjs vendored prebuild probe', () => {
describe('build-tree-sitter-kotlin.cjs vendored grammar activation', () => {
it('exits 0 and reports skipping when GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1', () => {
const r = runProbe({ GITNEXUS_SKIP_OPTIONAL_GRAMMARS: '1' });
expect(r.status).toBe(0);
expect(r.signal).toBeNull();
expect(r.stderr).toContain('Skipping prebuild probe');
expect(r.stderr).not.toContain(UNAVAILABLE);
expect(r.stderr).toContain('Skipping build');
expect(r.stderr).not.toContain(CATCH_UNAVAILABLE);
});
it('exits 0 silently when the materialized package is absent', () => {
// No node_modules/tree-sitter-kotlin next to the script — nothing to probe
// (materialize was skipped/failed). Swift-style: silent exit 0.
it('exits 0 silently when the materialized package is absent (no binding.gyp)', () => {
// No node_modules/tree-sitter-kotlin next to the script — materialize was
// skipped/failed, so there is no binding.gyp to build. Silent exit 0.
const r = runProbe({});
expect(r.status).toBe(0);
expect(r.signal).toBeNull();
expect(r.stderr).not.toContain(UNAVAILABLE);
expect(r.stderr).not.toContain(CATCH_UNAVAILABLE);
});
it('warns (and exits 0) when the package is present but no prebuild loads', () => {
// Materialize a package shell (bindings/node/index.js present) with no
// loadable prebuild → node-gyp-build throws → the probe must warn, not exit
// non-zero. (Here the throw is a missing node-gyp-build resolution, an
// equivalent trigger of the catch branch's never-fail guarantee.)
const pkg = path.join(tmpRoot, 'node_modules', 'tree-sitter-kotlin', 'bindings', 'node');
mkdirSync(pkg, { recursive: true });
writeFileSync(path.join(pkg, 'index.js'), '');
it('exits 0 (warning) when the package has a binding.gyp but no prebuild/build deps', () => {
// Materialize a package with binding.gyp present but no prebuild and no
// build/Release/*.node. The script falls through prefer-prebuild to the
// source-build path; in this temp env node-gyp-build/node-addon-api are not
// resolvable, so it stops at the deps guard (or, if they were resolvable,
// the node-gyp build would fail) — either way it warns and exits 0.
const pkg = path.join(tmpRoot, 'node_modules', 'tree-sitter-kotlin');
mkdirSync(path.join(pkg, 'bindings', 'node'), { recursive: true });
writeFileSync(path.join(pkg, 'binding.gyp'), '{ "targets": [] }');
writeFileSync(path.join(pkg, 'bindings', 'node', 'index.js'), '');
try {
const r = runProbe({});
expect(r.status).toBe(0);
expect(r.signal).toBeNull();
expect(r.stderr).toContain('Prebuild probe failed');
expect(r.stderr).toContain(UNAVAILABLE);
expect(r.stderr).toMatch(/hoisted build deps not resolvable|Could not build native binding/);
expect(r.stderr).not.toContain('built successfully');
} finally {
rmSync(path.join(tmpRoot, 'node_modules'), { recursive: true, force: true });
}

View file

@ -0,0 +1,109 @@
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import { spawnSync } from 'node:child_process';
import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import path from 'node:path';
import { fileURLToPath } from 'node:url';
/**
* Behavioral coverage for the postinstall activation script
* `scripts/build-tree-sitter-swift.cjs`.
*
* Swift is a vendored grammar, unified with Kotlin/Dart/Proto/C: the script
* prefers a committed prebuild for this platform-arch (toolchain-free); if none
* matches it source-builds from the vendored grammar source. Its hard invariant
* is that it MUST NEVER exit non-zero — it runs in `gitnexus`'s postinstall, so a
* non-zero exit would break `npm install gitnexus` for every user. This suite
* executes the real script bytes across its branches and asserts exit code 0
* every time (mirrors build-tree-sitter-kotlin-probe.test.ts).
*
* The script is copied into an isolated temp `scripts/` dir so its
* `__dirname`-relative `../node_modules/tree-sitter-swift` resolves under our
* control. The temp dir has no reachable `node-gyp-build` / `node-addon-api`, so
* the source-build path stops at the "hoisted build deps not resolvable" guard
* (still exit 0) instead of invoking a real compile.
*/
const probeSource = readFileSync(
fileURLToPath(new URL('../../scripts/build-tree-sitter-swift.cjs', import.meta.url)),
'utf8',
);
// Catch-branch sentinel (only printed when an actual node-gyp build is attempted
// and throws) — must NOT appear on the deps-unavailable guard path.
const CATCH_UNAVAILABLE = 'Swift (.swift) parsing will be unavailable';
let tmpRoot: string;
let scriptPath: string;
beforeAll(() => {
tmpRoot = mkdtempSync(path.join(tmpdir(), 'gn-swift-build-'));
mkdirSync(path.join(tmpRoot, 'scripts'), { recursive: true });
scriptPath = path.join(tmpRoot, 'scripts', 'build-tree-sitter-swift.cjs');
writeFileSync(scriptPath, probeSource);
});
afterAll(() => {
rmSync(tmpRoot, { recursive: true, force: true });
});
function runProbe(overrides: Record<string, string | undefined>) {
const env: Record<string, string> = {};
for (const [k, v] of Object.entries(process.env)) {
if (v !== undefined) env[k] = v;
}
delete env.GITNEXUS_SKIP_OPTIONAL_GRAMMARS;
for (const [k, v] of Object.entries(overrides)) {
if (v === undefined) delete env[k];
else env[k] = v;
}
return spawnSync(process.execPath, [scriptPath], { env, encoding: 'utf8', timeout: 30_000 });
}
describe('build-tree-sitter-swift.cjs vendored grammar activation', () => {
it('exits 0 and reports skipping when GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1', () => {
const r = runProbe({ GITNEXUS_SKIP_OPTIONAL_GRAMMARS: '1' });
expect(r.status).toBe(0);
expect(r.signal).toBeNull();
expect(r.stderr).toContain('Skipping build');
expect(r.stderr).not.toContain(CATCH_UNAVAILABLE);
});
it('exits 0 silently when the materialized package is absent (no binding.gyp)', () => {
// No node_modules/tree-sitter-swift next to the script — materialize was
// skipped/failed, so there is no binding.gyp to build. Silent exit 0.
const r = runProbe({});
expect(r.status).toBe(0);
expect(r.signal).toBeNull();
expect(r.stderr).not.toContain(CATCH_UNAVAILABLE);
});
it('exits 0 (warning) when the package has a binding.gyp but no prebuild/build deps', () => {
// Materialize a package with binding.gyp present but no prebuild and no
// build/Release/*.node. The script falls through prefer-prebuild to the
// source-build path; in this temp env node-gyp-build/node-addon-api are not
// resolvable, so it stops at the deps guard (or, if they were resolvable,
// the node-gyp build would fail) — either way it warns and exits 0.
const pkg = path.join(tmpRoot, 'node_modules', 'tree-sitter-swift');
mkdirSync(path.join(pkg, 'bindings', 'node'), { recursive: true });
writeFileSync(path.join(pkg, 'binding.gyp'), '{ "targets": [] }');
writeFileSync(path.join(pkg, 'bindings', 'node', 'index.js'), '');
try {
const r = runProbe({});
expect(r.status).toBe(0);
expect(r.signal).toBeNull();
expect(r.stderr).toMatch(/hoisted build deps not resolvable|Could not build native binding/);
expect(r.stderr).not.toContain('built successfully');
} finally {
rmSync(path.join(tmpRoot, 'node_modules'), { recursive: true, force: true });
}
});
it('never exits non-zero across env permutations (postinstall hard invariant)', () => {
for (const overrides of [{ GITNEXUS_SKIP_OPTIONAL_GRAMMARS: '1' }, {}]) {
const r = runProbe(overrides);
expect(r.status).toBe(0);
expect(r.signal).toBeNull();
}
});
});

View file

@ -84,7 +84,7 @@ describe('CLI commands', () => {
expect(pkg.default.files).toContain('vendor');
});
it('keeps vendored Swift runtime with prebuilds and hoisted activation script', async () => {
it('keeps vendored Swift runtime with vendored source + GitNexus-built prebuilds and hoisted activation script', async () => {
const pkg = await import('../../package.json', { with: { type: 'json' } });
const swiftPkg = await import('../../vendor/tree-sitter-swift/package.json', {
with: { type: 'json' },
@ -95,9 +95,22 @@ describe('CLI commands', () => {
expect(pkg.default.dependencies['tree-sitter']).toBe('0.21.1');
expect(pkg.default.scripts.postinstall).toContain('build-tree-sitter-swift.cjs');
expect(swiftPkg.default.version).toBe('0.7.1');
// No scripts.install / dependencies inside vendor/ (#836 / #1728 hygiene).
expect(swiftPkg.default.scripts?.install).toBeUndefined();
expect(swiftPkg.default.dependencies).toBeUndefined();
expect(swiftPkg.default.peerDependencies['tree-sitter']).toContain('^0.21.1');
// Swift is now unified with Dart/Proto/Kotlin/C: the grammar SOURCE is
// vendored so build-tree-sitter-swift.cjs can source-build the binding
// when no committed prebuild matches (e.g. CI before prebuilds land).
const bindingGyp = await fs.readFile(
path.join(REPO_ROOT, 'gitnexus/vendor/tree-sitter-swift/binding.gyp'),
'utf8',
);
expect(bindingGyp).toContain('tree_sitter_swift_binding');
expect(bindingGyp).toContain('src/parser.c');
await expect(
fs.stat(path.join(REPO_ROOT, 'gitnexus/vendor/tree-sitter-swift/src/parser.c')),
).resolves.toBeDefined();
});
it('keeps vendored Kotlin runtime with GitNexus-built prebuilds and hoisted activation script (#2107)', async () => {

View file

@ -20,8 +20,9 @@ import { fileURLToPath } from 'node:url';
*
* Two cohorts:
* 1. VENDORED grammars (gitnexus/vendor/tree-sitter-*) — GitNexus owns these
* prebuilds (cross-built by .github/workflows/build-tree-sitter-prebuilds.yml,
* or copied from upstream for Swift). Every one MUST cover all 6 tuples.
* prebuilds (cross-built by .github/workflows/build-tree-sitter-prebuilds.yml;
* Swift's were originally upstream-shipped, now rebuilt the same way). Each
* one that does NOT also vendor its build source MUST cover all 6 tuples.
* 2. npm-dependency grammars — upstream owns their prebuilds. We assert 6/6
* too, with documented exceptions (see KNOWN_NPM_GAPS).
*/
@ -90,9 +91,10 @@ describe('vendored grammar prebuild coverage (toolchain-free on every supported
// A grammar that vendors its build sources (binding.gyp) can source-build the
// gaps on any toolchain host (e.g. CI), so an incomplete prebuild set is
// tolerated for it — the build-tree-sitter-prebuilds workflow fills the
// prebuilds to make it toolchain-free. A prebuild-only grammar (no source,
// e.g. swift, whose prebuilds come from upstream) MUST ship all six, or it is
// dead on the missing platform.
// prebuilds to make it toolchain-free. Every grammar GitNexus currently
// vendors carries its source (incl. swift, unified with the rest), so the
// strict branch below is defensive: a hypothetical prebuild-only grammar (no
// binding.gyp) MUST ship all six, or it is dead on the missing platform.
const hasSourceFallback = existsSync(path.join(grammarDir, 'binding.gyp'));
it(

View file

@ -18,8 +18,10 @@ prebuilds itself and vendors them here. `node-gyp-build` selects the correct
binary at require time; `build-tree-sitter-kotlin.cjs` probes availability at
install time.
This differs from `tree-sitter-swift`, whose prebuilds are **copied from the
upstream package** (Swift ships them). Kotlin's are **GitNexus-cross-built**.
`tree-sitter-swift` is handled the same way now: its prebuilds were originally
**copied from upstream** (Swift ships them), but it is unified with this pipeline —
its source is vendored and its prebuilds are **GitNexus-cross-built** too, so all
of Dart/Proto/Swift/Kotlin go through one uniform build path.
### Updating this vendor package

View file

@ -1,14 +1,29 @@
## GitNexus vendor notice
This directory is a GitNexus-managed vendored copy of the official
`tree-sitter-swift@0.7.1` npm runtime package, including its official native
prebuilds. GitNexus keeps the top-level `tree-sitter` dependency pinned to
`^0.21.1` until the broader parser runtime upgrade is handled separately.
`tree-sitter-swift@0.7.1` npm runtime package. GitNexus keeps the top-level
`tree-sitter` dependency pinned to `^0.21.1` until the broader parser runtime
upgrade is handled separately.
Unified with the Dart/Proto/Kotlin/C vendored grammars, this copy also vendors
the grammar **source** — `binding.gyp`, `bindings/node/binding.cc`,
`src/parser.c` (the ABI-14 default; ~18 MB, compresses heavily in git),
`src/scanner.c`, and `src/tree_sitter/` — so `gitnexus/scripts/build-tree-sitter-swift.cjs`
can source-build the native binding on any toolchain host when no committed
prebuild matches (e.g. CI before the prebuilds land). Note: upstream
deliberately omits the generated `parser.c` (see the FAQ below); GitNexus
commits it on purpose so the source-build fallback is deterministic and never
needs the tree-sitter CLI at install time. The native `prebuilds/` are
GitNexus-cross-built by `.github/workflows/build-tree-sitter-prebuilds.yml`
(originally upstream-shipped).
When updating this vendor package, replace it from an official
`tree-sitter-swift` npm release, keep the native `prebuilds/` artifacts, update
the `_vendoredBy` provenance fields in `package.json`, and verify the packed
GitNexus tarball can load `tree-sitter-swift`.
`tree-sitter-swift` npm release: refresh `src/parser.c`/`src/scanner.c`/
`src/tree_sitter/`/`binding.gyp`/`bindings/node/binding.cc` (use the ABI-14
`parser.c`, not the legacy `parser_abi13.c`), bump `version` in `package.json`
to retrigger the prebuild workflow, update the `_vendoredBy` provenance, and
verify the packed GitNexus tarball can both load a committed prebuild and
source-build `tree-sitter-swift`.
![Parse rate badge](https://byob.yarr.is/alex-pinkus/tree-sitter-swift/parse_rate)
[![Crates.io badge](https://byob.yarr.is/alex-pinkus/tree-sitter-swift/crates_io_version)](https://crates.io/crates/tree-sitter-swift)

View file

@ -0,0 +1,30 @@
{
"targets": [
{
"target_name": "tree_sitter_swift_binding",
"dependencies": [
"<!(node -p \"require('node-addon-api').targets\"):node_addon_api_except",
],
"include_dirs": [
"src",
],
"sources": [
"bindings/node/binding.cc",
"src/parser.c",
"src/scanner.c"
],
"conditions": [
["OS!='win'", {
"cflags_c": [
"-std=c11",
],
}, { # OS == "win"
"cflags_c": [
"/std:c11",
"/utf-8",
],
}],
],
}
]
}

View file

@ -0,0 +1,20 @@
#include <napi.h>
typedef struct TSLanguage TSLanguage;
extern "C" TSLanguage *tree_sitter_swift();
// "tree-sitter", "language" hashed with BLAKE2
const napi_type_tag LANGUAGE_TYPE_TAG = {
0x8AF2E5212AD58ABF, 0xD5006CAD83ABBA16
};
Napi::Object Init(Napi::Env env, Napi::Object exports) {
exports["name"] = Napi::String::New(env, "swift");
auto language = Napi::External<TSLanguage>::New(env, tree_sitter_swift());
language.TypeTag(&LANGUAGE_TYPE_TAG);
exports["language"] = language;
return exports;
}
NODE_API_MODULE(tree_sitter_swift_binding, Init)

View file

@ -9,7 +9,7 @@
"type": "git",
"url": "git+https://github.com/alex-pinkus/tree-sitter-swift.git"
},
"_vendoredBy": "gitnexus - minimal runtime package copied from official tree-sitter-swift@0.7.1 (gitHead 88bfd19a89be9d0481b14566fb6160cccea2fe0a). Prebuild activation runs via gitnexus/scripts/build-tree-sitter-swift.cjs after materialize-vendor-grammars.cjs (no install script here — avoids #836 / #1728).",
"_vendoredBy": "gitnexus - runtime package derived from official tree-sitter-swift@0.7.1 (gitHead 88bfd19a89be9d0481b14566fb6160cccea2fe0a). Unified with Dart/Proto/Kotlin/C: the grammar source (parser.c/scanner.c/binding.gyp + src/) is ALSO vendored so build-tree-sitter-swift.cjs can source-build the binding on a toolchain host when no prebuild matches (e.g. CI before prebuilds land); src/parser.c is the ABI-14 default (~18 MB on disk, compresses heavily in git — the upstream parser_abi13.c alternate is not vendored). The native prebuilds/ are GitNexus-cross-built by .github/workflows/build-tree-sitter-prebuilds.yml (originally upstream-shipped). Build activation runs via gitnexus/scripts/build-tree-sitter-swift.cjs after materialize-vendor-grammars.cjs (no scripts.install here — avoids #836 / #1728).",
"peerDependencies": {
"tree-sitter": "^0.21.1 || ^0.22.1"
},

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,929 @@
#include "tree_sitter/parser.h"
#include <string.h>
#include <wctype.h>
#define TOKEN_COUNT 33
enum TokenType {
BLOCK_COMMENT,
RAW_STR_PART,
RAW_STR_CONTINUING_INDICATOR,
RAW_STR_END_PART,
IMPLICIT_SEMI,
EXPLICIT_SEMI,
ARROW_OPERATOR,
DOT_OPERATOR,
CONJUNCTION_OPERATOR,
DISJUNCTION_OPERATOR,
NIL_COALESCING_OPERATOR,
EQUAL_SIGN,
EQ_EQ,
PLUS_THEN_WS,
MINUS_THEN_WS,
BANG,
THROWS_KEYWORD,
RETHROWS_KEYWORD,
DEFAULT_KEYWORD,
WHERE_KEYWORD,
ELSE_KEYWORD,
CATCH_KEYWORD,
AS_KEYWORD,
AS_QUEST,
AS_BANG,
ASYNC_KEYWORD,
CUSTOM_OPERATOR,
HASH_SYMBOL,
DIRECTIVE_IF,
DIRECTIVE_ELSEIF,
DIRECTIVE_ELSE,
DIRECTIVE_ENDIF,
FAKE_TRY_BANG
};
#define OPERATOR_COUNT 20
const char* OPERATORS[OPERATOR_COUNT] = {
"->",
".",
"&&",
"||",
"??",
"=",
"==",
"+",
"-",
"!",
"throws",
"rethrows",
"default",
"where",
"else",
"catch",
"as",
"as?",
"as!",
"async"
};
enum IllegalTerminatorGroup {
ALPHANUMERIC,
OPERATOR_SYMBOLS,
OPERATOR_OR_DOT,
NON_WHITESPACE
};
const enum IllegalTerminatorGroup OP_ILLEGAL_TERMINATORS[OPERATOR_COUNT] = {
OPERATOR_SYMBOLS, // ->
OPERATOR_OR_DOT, // .
OPERATOR_SYMBOLS, // &&
OPERATOR_SYMBOLS, // ||
OPERATOR_SYMBOLS, // ??
OPERATOR_SYMBOLS, // =
OPERATOR_SYMBOLS, // ==
NON_WHITESPACE, // +
NON_WHITESPACE, // -
OPERATOR_SYMBOLS, // !
ALPHANUMERIC, // throws
ALPHANUMERIC, // rethrows
ALPHANUMERIC, // default
ALPHANUMERIC, // where
ALPHANUMERIC, // else
ALPHANUMERIC, // catch
ALPHANUMERIC, // as
OPERATOR_SYMBOLS, // as?
OPERATOR_SYMBOLS, // as!
ALPHANUMERIC // async
};
const enum TokenType OP_SYMBOLS[OPERATOR_COUNT] = {
ARROW_OPERATOR,
DOT_OPERATOR,
CONJUNCTION_OPERATOR,
DISJUNCTION_OPERATOR,
NIL_COALESCING_OPERATOR,
EQUAL_SIGN,
EQ_EQ,
PLUS_THEN_WS,
MINUS_THEN_WS,
BANG,
THROWS_KEYWORD,
RETHROWS_KEYWORD,
DEFAULT_KEYWORD,
WHERE_KEYWORD,
ELSE_KEYWORD,
CATCH_KEYWORD,
AS_KEYWORD,
AS_QUEST,
AS_BANG,
ASYNC_KEYWORD
};
const uint64_t OP_SYMBOL_SUPPRESSOR[OPERATOR_COUNT] = {
0, // ARROW_OPERATOR,
0, // DOT_OPERATOR,
0, // CONJUNCTION_OPERATOR,
0, // DISJUNCTION_OPERATOR,
0, // NIL_COALESCING_OPERATOR,
0, // EQUAL_SIGN,
0, // EQ_EQ,
0, // PLUS_THEN_WS,
0, // MINUS_THEN_WS,
1UL << FAKE_TRY_BANG, // BANG,
0, // THROWS_KEYWORD,
0, // RETHROWS_KEYWORD,
0, // DEFAULT_KEYWORD,
0, // WHERE_KEYWORD,
0, // ELSE_KEYWORD,
0, // CATCH_KEYWORD,
0, // AS_KEYWORD,
0, // AS_QUEST,
0, // AS_BANG,
0, // ASYNC_KEYWORD
};
#define RESERVED_OP_COUNT 31
const char* RESERVED_OPS[RESERVED_OP_COUNT] = {
"/",
"=",
"-",
"+",
"!",
"*",
"%",
"<",
">",
"&",
"|",
"^",
"?",
"~",
".",
"..",
"->",
"/*",
"*/",
"+=",
"-=",
"*=",
"/=",
"%=",
">>",
"<<",
"++",
"--",
"===",
"...",
"..<"
};
static bool is_cross_semi_token(enum TokenType op) {
switch(op) {
case ARROW_OPERATOR:
case DOT_OPERATOR:
case CONJUNCTION_OPERATOR:
case DISJUNCTION_OPERATOR:
case NIL_COALESCING_OPERATOR:
case EQUAL_SIGN:
case EQ_EQ:
case PLUS_THEN_WS:
case MINUS_THEN_WS:
case THROWS_KEYWORD:
case RETHROWS_KEYWORD:
case DEFAULT_KEYWORD:
case WHERE_KEYWORD:
case ELSE_KEYWORD:
case CATCH_KEYWORD:
case AS_KEYWORD:
case AS_QUEST:
case AS_BANG:
case ASYNC_KEYWORD:
case CUSTOM_OPERATOR:
return true;
case BANG:
default:
return false;
}
}
#define NON_CONSUMING_CROSS_SEMI_CHAR_COUNT 3
const uint32_t NON_CONSUMING_CROSS_SEMI_CHARS[NON_CONSUMING_CROSS_SEMI_CHAR_COUNT] = { '?', ':', '{' };
/**
* All possible results of having performed some sort of parsing.
*
* A parser can return a result along two dimensions:
* 1. Should the scanner continue trying to find another result?
* 2. Was some result produced by this parsing attempt?
*
* These are flattened into a single enum together. When the function returns one of the `TOKEN_FOUND` cases, it
* will always populate its `symbol_result` field. When it returns one of the `STOP_PARSING` cases, callers should
* immediately return (with the value, if there is one).
*/
enum ParseDirective {
CONTINUE_PARSING_NOTHING_FOUND,
CONTINUE_PARSING_TOKEN_FOUND,
CONTINUE_PARSING_SLASH_CONSUMED,
STOP_PARSING_NOTHING_FOUND,
STOP_PARSING_TOKEN_FOUND,
STOP_PARSING_END_OF_FILE
};
struct ScannerState {
uint32_t ongoing_raw_str_hash_count;
};
void *tree_sitter_swift_external_scanner_create() {
return calloc(0, sizeof(struct ScannerState));
}
void tree_sitter_swift_external_scanner_destroy(void *payload) {
free(payload);
}
void tree_sitter_swift_external_scanner_reset(void *payload) {
struct ScannerState *state = (struct ScannerState *)payload;
state->ongoing_raw_str_hash_count = 0;
}
unsigned tree_sitter_swift_external_scanner_serialize(void *payload, char *buffer) {
struct ScannerState *state = (struct ScannerState *)payload;
uint32_t hash_count = state->ongoing_raw_str_hash_count;
buffer[0] = (hash_count >> 24) & 0xff;
buffer[1] = (hash_count >> 16) & 0xff;
buffer[2] = (hash_count >> 8) & 0xff;
buffer[3] = (hash_count) & 0xff;
return 4;
}
void tree_sitter_swift_external_scanner_deserialize(
void *payload,
const char *buffer,
unsigned length
) {
if (length < 4) {
return;
}
uint32_t hash_count = (
(((uint32_t) buffer[0]) << 24) |
(((uint32_t) buffer[1]) << 16) |
(((uint32_t) buffer[2]) << 8) |
(((uint32_t) buffer[3]))
);
struct ScannerState *state = (struct ScannerState *)payload;
state->ongoing_raw_str_hash_count = hash_count;
}
static void advance(TSLexer *lexer) {
lexer->advance(lexer, false);
}
static bool should_treat_as_wspace(int32_t character) {
return iswspace(character) || (((int32_t) ';') == character);
}
static int32_t encountered_op_count(bool *encountered_operator) {
int32_t encountered = 0;
for (int op_idx = 0; op_idx < OPERATOR_COUNT; op_idx++) {
if (encountered_operator[op_idx]) {
encountered++;
}
}
return encountered;
}
static bool any_reserved_ops(uint8_t *encountered_reserved_ops) {
for (int op_idx = 0; op_idx < RESERVED_OP_COUNT; op_idx++) {
if (encountered_reserved_ops[op_idx] == 2) {
return true;
}
}
return false;
}
static bool is_legal_custom_operator(
int32_t char_idx,
int32_t first_char,
int32_t cur_char
) {
bool is_first_char = !char_idx;
switch (cur_char) {
case '=':
case '-':
case '+':
case '!':
case '%':
case '<':
case '>':
case '&':
case '|':
case '^':
case '?':
case '~':
return true;
case '.':
// Grammar allows `.` for any operator that starts with `.`
return is_first_char || first_char == '.';
case '*':
case '/':
// Not listed in the grammar, but `/*` and `//` can't be the start of an operator since they start comments
return char_idx != 1 || first_char != '/';
default:
if (
(cur_char >= 0x00A1 && cur_char <= 0x00A7) ||
(cur_char == 0x00A9) ||
(cur_char == 0x00AB) ||
(cur_char == 0x00AC) ||
(cur_char == 0x00AE) ||
(cur_char >= 0x00B0 && cur_char <= 0x00B1) ||
(cur_char == 0x00B6) ||
(cur_char == 0x00BB) ||
(cur_char == 0x00BF) ||
(cur_char == 0x00D7) ||
(cur_char == 0x00F7) ||
(cur_char >= 0x2016 && cur_char <= 0x2017) ||
(cur_char >= 0x2020 && cur_char <= 0x2027) ||
(cur_char >= 0x2030 && cur_char <= 0x203E) ||
(cur_char >= 0x2041 && cur_char <= 0x2053) ||
(cur_char >= 0x2055 && cur_char <= 0x205E) ||
(cur_char >= 0x2190 && cur_char <= 0x23FF) ||
(cur_char >= 0x2500 && cur_char <= 0x2775) ||
(cur_char >= 0x2794 && cur_char <= 0x2BFF) ||
(cur_char >= 0x2E00 && cur_char <= 0x2E7F) ||
(cur_char >= 0x3001 && cur_char <= 0x3003) ||
(cur_char >= 0x3008 && cur_char <= 0x3020) ||
(cur_char == 0x3030)
) {
return true;
} else if (
(cur_char >= 0x0300 && cur_char <= 0x036f) ||
(cur_char >= 0x1DC0 && cur_char <= 0x1DFF) ||
(cur_char >= 0x20D0 && cur_char <= 0x20FF) ||
(cur_char >= 0xFE00 && cur_char <= 0xFE0F) ||
(cur_char >= 0xFE20 && cur_char <= 0xFE2F) ||
(cur_char >= 0xE0100 && cur_char <= 0xE01EF)
) {
return !is_first_char;
} else {
return false;
}
}
}
static bool eat_operators(
TSLexer *lexer,
const bool *valid_symbols,
bool mark_end,
const int32_t prior_char,
enum TokenType *symbol_result
) {
bool possible_operators[OPERATOR_COUNT];
uint8_t reserved_operators[RESERVED_OP_COUNT];
for (int op_idx = 0; op_idx < OPERATOR_COUNT; op_idx++) {
possible_operators[op_idx] = valid_symbols[OP_SYMBOLS[op_idx]] && (!prior_char || OPERATORS[op_idx][0] == prior_char);
}
for (int op_idx = 0; op_idx < RESERVED_OP_COUNT; op_idx++) {
reserved_operators[op_idx] = !prior_char || RESERVED_OPS[op_idx][0] == prior_char;
}
bool possible_custom_operator = valid_symbols[CUSTOM_OPERATOR];
int32_t first_char = prior_char ? prior_char : lexer->lookahead;
int32_t last_examined_char = first_char;
int32_t str_idx = prior_char ? 1 : 0;
int32_t full_match = -1;
while(true) {
for (int op_idx = 0; op_idx < OPERATOR_COUNT; op_idx++) {
if (!possible_operators[op_idx]) {
continue;
}
if (OPERATORS[op_idx][str_idx] == '\0') {
// Make sure that the operator is allowed to have the next character as its lookahead.
enum IllegalTerminatorGroup illegal_terminators = OP_ILLEGAL_TERMINATORS[op_idx];
switch (lexer->lookahead) {
// See "Operators":
// https://docs.swift.org/swift-book/ReferenceManual/LexicalStructure.html#ID418
case '/':
case '=':
case '-':
case '+':
case '!':
case '*':
case '%':
case '<':
case '>':
case '&':
case '|':
case '^':
case '?':
case '~':
if (illegal_terminators == OPERATOR_SYMBOLS) {
break;
} // Otherwise, intentionally fall through to the OPERATOR_OR_DOT case
// fall through
case '.':
if (illegal_terminators == OPERATOR_OR_DOT) {
break;
} // Otherwise, fall through to DEFAULT which checks its groups directly
// fall through
default:
if (iswalnum(lexer->lookahead) && illegal_terminators == ALPHANUMERIC) {
break;
}
if (!iswspace(lexer->lookahead) && illegal_terminators == NON_WHITESPACE) {
break;
}
full_match = op_idx;
if (mark_end) {
lexer->mark_end(lexer);
}
}
possible_operators[op_idx] = false;
continue;
}
if (OPERATORS[op_idx][str_idx] != lexer->lookahead) {
possible_operators[op_idx] = false;
continue;
}
}
for (int op_idx = 0; op_idx < RESERVED_OP_COUNT; op_idx++) {
if (!reserved_operators[op_idx]) {
continue;
}
if (RESERVED_OPS[op_idx][str_idx] == '\0') {
reserved_operators[op_idx] = 0;
continue;
}
if (RESERVED_OPS[op_idx][str_idx] != lexer->lookahead) {
reserved_operators[op_idx] = 0;
continue;
}
if (RESERVED_OPS[op_idx][str_idx + 1] == '\0') {
reserved_operators[op_idx] = 2;
continue;
}
}
possible_custom_operator = possible_custom_operator && is_legal_custom_operator(
str_idx,
first_char,
lexer->lookahead
);
uint32_t encountered_ops = encountered_op_count(possible_operators);
if (encountered_ops == 0) {
if (!possible_custom_operator) {
break;
} else if (mark_end && full_match == -1) {
lexer->mark_end(lexer);
}
}
last_examined_char = lexer->lookahead;
lexer->advance(lexer, false);
str_idx += 1;
if (encountered_ops == 0 && !is_legal_custom_operator(
str_idx,
first_char,
lexer->lookahead
)) {
break;
}
}
if (full_match != -1) {
// We have a match -- first see if that match has a symbol that suppresses it. For example, in `try!`, we do not
// want to emit the `!` as a symbol in our scanner, because we want the parser to have the chance to parse it as
// an immediate token.
uint64_t suppressing_symbols = OP_SYMBOL_SUPPRESSOR[full_match];
if (suppressing_symbols) {
for (uint64_t suppressor = 0; suppressor < TOKEN_COUNT; suppressor++) {
if (!(suppressing_symbols & 1 << suppressor)) {
continue;
}
// The suppressing symbol is valid in this position, so skip it.
if (valid_symbols[suppressor]) {
return false;
}
}
}
*symbol_result = OP_SYMBOLS[full_match];
return true;
}
if (possible_custom_operator && !any_reserved_ops(reserved_operators)) {
if ((last_examined_char != '<' || iswspace(lexer->lookahead)) && mark_end) {
lexer->mark_end(lexer);
}
*symbol_result = CUSTOM_OPERATOR;
return true;
}
return false;
}
static enum ParseDirective eat_comment(
TSLexer *lexer,
const bool *valid_symbols,
bool mark_end,
enum TokenType *symbol_result
) {
if (lexer->lookahead != '/') {
return CONTINUE_PARSING_NOTHING_FOUND;
}
advance(lexer);
if (lexer->lookahead != '*') {
return CONTINUE_PARSING_SLASH_CONSUMED;
}
advance(lexer);
bool after_star = false;
unsigned nesting_depth = 1;
for (;;) {
switch (lexer->lookahead) {
case '\0':
return STOP_PARSING_END_OF_FILE;
case '*':
advance(lexer);
after_star = true;
break;
case '/':
if (after_star) {
advance(lexer);
after_star = false;
nesting_depth--;
if (nesting_depth == 0) {
if (mark_end) {
lexer->mark_end(lexer);
}
*symbol_result = BLOCK_COMMENT;
return STOP_PARSING_TOKEN_FOUND;
}
} else {
advance(lexer);
after_star = false;
if (lexer->lookahead == '*') {
nesting_depth++;
advance(lexer);
}
}
break;
default:
advance(lexer);
after_star = false;
break;
}
}
}
static enum ParseDirective eat_whitespace(
TSLexer *lexer,
const bool *valid_symbols,
enum TokenType *symbol_result
) {
enum ParseDirective ws_directive = CONTINUE_PARSING_NOTHING_FOUND;
bool semi_is_valid = valid_symbols[IMPLICIT_SEMI] && valid_symbols[EXPLICIT_SEMI];
uint32_t lookahead;
while (should_treat_as_wspace(lookahead = lexer->lookahead)) {
if (lookahead == ';') {
if (semi_is_valid) {
ws_directive = STOP_PARSING_TOKEN_FOUND;
lexer->advance(lexer, false);
}
break;
}
lexer->advance(lexer, true);
lexer->mark_end(lexer);
if (ws_directive == CONTINUE_PARSING_NOTHING_FOUND && (lookahead == '\n' || lookahead == '\r')) {
ws_directive = CONTINUE_PARSING_TOKEN_FOUND;
}
}
enum ParseDirective any_comment = CONTINUE_PARSING_NOTHING_FOUND;
if (ws_directive == CONTINUE_PARSING_TOKEN_FOUND && lookahead == '/') {
bool has_seen_single_comment = false;
while (lexer->lookahead == '/') {
// It's possible that this is a comment - start an exploratory mission to find out, and if it is, look for what
// comes after it. We care about what comes after it for the purpose of suppressing the newline.
enum TokenType multiline_comment_result;
any_comment = eat_comment(lexer, valid_symbols, /* mark_end */ false, &multiline_comment_result);
if (any_comment == STOP_PARSING_TOKEN_FOUND) {
// This is a multiline comment. This scanner should be parsing those, so we might want to bail out and
// emit it instead. However, we only want to do that if we haven't advanced through a _single_ line
// comment on the way - otherwise that will get lumped into this.
if (!has_seen_single_comment) {
lexer->mark_end(lexer);
*symbol_result = multiline_comment_result;
return STOP_PARSING_TOKEN_FOUND;
}
} else if (any_comment == STOP_PARSING_END_OF_FILE) {
return STOP_PARSING_END_OF_FILE;
} else if (any_comment == CONTINUE_PARSING_SLASH_CONSUMED) {
// We accidentally ate a slash -- we should actually bail out, say we saw nothing, and let the next pass
// take it from after the newline.
return CONTINUE_PARSING_SLASH_CONSUMED;
} else if (lexer->lookahead == '/') {
// There wasn't a multiline comment, which we know means that the comment parser ate its `/` and then
// bailed out. If it had seen anything comment-like after that first `/` it would have continued going
// and eventually had a well-formed comment or an EOF. Thus, if we're currently looking at a `/`, it's
// the second one of those and it means we have a single-line comment.
has_seen_single_comment = true;
while (lexer->lookahead != '\n' && lexer->lookahead != '\0') {
lexer->advance(lexer, true);
}
} else if (iswspace(lexer->lookahead)) {
// We didn't see any type of comment - in fact, we saw an operator that we don't normally treat as an
// operator. Still, this is a reason to stop parsing.
return STOP_PARSING_NOTHING_FOUND;
}
// If we skipped through some comment, we're at whitespace now, so advance.
while(iswspace(lexer->lookahead)) {
any_comment = CONTINUE_PARSING_NOTHING_FOUND; // We're advancing, so clear out the comment
lexer->advance(lexer, true);
}
}
enum TokenType operator_result;
bool saw_operator = eat_operators(
lexer,
valid_symbols,
/* mark_end */ false,
'\0',
&operator_result
);
if (saw_operator) {
// The operator we saw should suppress the newline, so bail out.
return STOP_PARSING_NOTHING_FOUND;
} else {
// Promote the implicit newline to an explicit one so we don't check for operators again.
*symbol_result = IMPLICIT_SEMI;
ws_directive = STOP_PARSING_TOKEN_FOUND;
}
}
// Let's consume operators that can live after a "semicolon" style newline. Before we do that, though, we want to
// check for a set of characters that we do not consume, but that still suppress the semi.
if (ws_directive == CONTINUE_PARSING_TOKEN_FOUND) {
for (int i = 0; i < NON_CONSUMING_CROSS_SEMI_CHAR_COUNT; i++) {
if (NON_CONSUMING_CROSS_SEMI_CHARS[i] == lookahead) {
return CONTINUE_PARSING_NOTHING_FOUND;
}
}
}
if (semi_is_valid && ws_directive != CONTINUE_PARSING_NOTHING_FOUND) {
*symbol_result = lookahead == ';' ? EXPLICIT_SEMI : IMPLICIT_SEMI;
return ws_directive;
}
return CONTINUE_PARSING_NOTHING_FOUND;
}
#define DIRECTIVE_COUNT 4
const char* DIRECTIVES[OPERATOR_COUNT] = {
"if",
"elseif",
"else",
"endif"
};
const enum TokenType DIRECTIVE_SYMBOLS[DIRECTIVE_COUNT] = {
DIRECTIVE_IF,
DIRECTIVE_ELSEIF,
DIRECTIVE_ELSE,
DIRECTIVE_ENDIF
};
static enum TokenType find_possible_compiler_directive(TSLexer *lexer) {
bool possible_directives[DIRECTIVE_COUNT];
for (int dir_idx = 0; dir_idx < DIRECTIVE_COUNT; dir_idx++) {
possible_directives[dir_idx] = true;
}
int32_t str_idx = 0;
int32_t full_match = -1;
while(true) {
for (int dir_idx = 0; dir_idx < DIRECTIVE_COUNT; dir_idx++) {
if (!possible_directives[dir_idx]) {
continue;
}
uint8_t expected_char = DIRECTIVES[dir_idx][str_idx];
if (expected_char == '\0') {
full_match = dir_idx;
lexer->mark_end(lexer);
}
if (expected_char != lexer->lookahead) {
possible_directives[dir_idx] = false;
continue;
}
}
uint8_t match_count = 0;
for (int dir_idx = 0; dir_idx < DIRECTIVE_COUNT; dir_idx += 1) {
if (possible_directives[dir_idx]) {
match_count += 1;
}
}
if (match_count == 0) {
break;
}
lexer->advance(lexer, false);
str_idx += 1;
}
if (full_match == -1) {
// No compiler directive found, so just match the starting symbol
return HASH_SYMBOL;
}
return DIRECTIVE_SYMBOLS[full_match];
}
static bool eat_raw_str_part(
struct ScannerState *state,
TSLexer *lexer,
const bool *valid_symbols,
enum TokenType *symbol_result
) {
uint32_t hash_count = state->ongoing_raw_str_hash_count;
if (!valid_symbols[RAW_STR_PART]) {
return false;
} else if (hash_count == 0) {
// If this is a raw_str_part, it's the first one - look for hashes
while (lexer->lookahead == '#') {
hash_count += 1;
advance(lexer);
}
if (hash_count == 0) {
return false;
}
if (lexer->lookahead == '"') {
advance(lexer);
} else if (hash_count == 1) {
lexer->mark_end(lexer);
*symbol_result = find_possible_compiler_directive(lexer);
return true;
} else {
return false;
}
} else if (valid_symbols[RAW_STR_CONTINUING_INDICATOR]) {
// This is the end of an interpolation - now it's another raw_str_part. This is a synthetic
// marker to tell us that the grammar just consumed a `(` symbol to close a raw
// interpolation (since we don't want to fire on every `(` in existence). We don't have
// anything to do except continue.
} else {
return false;
}
// We're in a state where anything other than `hash_count` hash symbols in a row should be eaten
// and is part of a string.
// The last character _before_ the hashes will tell us what happens next.
// Matters are also complicated by the fact that we don't want to consume every character we
// visit; if we see a `\#(`, for instance, with the appropriate number of hash symbols, we want
// to end our parsing _before_ that sequence. This allows highlighting tools to treat that as a
// separate token.
while (lexer->lookahead != '\0') {
uint8_t last_char = '\0';
lexer->mark_end(lexer); // We always want to parse thru the start of the string so far
// Advance through anything that isn't a hash symbol, because we want to count those.
while (lexer->lookahead != '#' && lexer->lookahead != '\0') {
last_char = lexer->lookahead;
advance(lexer);
if (last_char != '\\' || lexer->lookahead == '\\') {
// Mark a new end, but only if we didn't just advance past a `\` symbol, since we
// don't want to consume that. Exception: if this is a `\` that happens _right
// after_ another `\`, we for some reason _do_ want to consume that, because
// apparently that is parsed as a literal `\` followed by something escaped.
lexer->mark_end(lexer);
}
}
// We hit at least one hash - count them and see if they match.
uint32_t current_hash_count = 0;
while (lexer->lookahead == '#' && current_hash_count < hash_count) {
current_hash_count += 1;
advance(lexer);
}
// If we saw exactly the right number of hashes, one of three things is true:
// 1. We're trying to interpolate into this string.
// 2. The string just ended.
// 3. This was just some hash characters doing nothing important.
if (current_hash_count == hash_count) {
if (last_char == '\\' && lexer->lookahead == '(') {
// Interpolation case! Don't consume those chars; they get saved for grammar.js.
*symbol_result = RAW_STR_PART;
state->ongoing_raw_str_hash_count = hash_count;
return true;
} else if (last_char == '"') {
// The string is finished! Mark the end here, on the very last hash symbol.
lexer->mark_end(lexer);
*symbol_result = RAW_STR_END_PART;
state->ongoing_raw_str_hash_count = 0;
return true;
}
// Nothing special happened - let the string continue.
}
}
return false;
}
bool tree_sitter_swift_external_scanner_scan(
void *payload,
TSLexer *lexer,
const bool *valid_symbols
) {
// Figure out our scanner state
struct ScannerState *state = (struct ScannerState *)payload;
// Consume any whitespace at the start.
enum TokenType ws_result;
enum ParseDirective ws_directive = eat_whitespace(lexer, valid_symbols, &ws_result);
if (ws_directive == STOP_PARSING_TOKEN_FOUND) {
lexer->result_symbol = ws_result;
return true;
}
if (ws_directive == STOP_PARSING_NOTHING_FOUND || ws_directive == STOP_PARSING_END_OF_FILE) {
return false;
}
bool has_ws_result = (ws_directive == CONTINUE_PARSING_TOKEN_FOUND);
// Now consume comments (before custom operators so that those aren't treated as comments)
enum TokenType comment_result;
enum ParseDirective comment = ws_directive == CONTINUE_PARSING_SLASH_CONSUMED ? ws_directive : eat_comment(lexer, valid_symbols, /* mark_end */ true, &comment_result);
if (comment == STOP_PARSING_TOKEN_FOUND) {
lexer->mark_end(lexer);
lexer->result_symbol = comment_result;
return true;
}
if (comment == STOP_PARSING_END_OF_FILE) {
return false;
}
// Now consume any operators that might cause our whitespace to be suppressed.
enum TokenType operator_result;
bool saw_operator = eat_operators(
lexer,
valid_symbols,
/* mark_end */ !has_ws_result,
comment == CONTINUE_PARSING_SLASH_CONSUMED ? '/' : '\0',
&operator_result
);
if (saw_operator && (!has_ws_result || is_cross_semi_token(operator_result))) {
lexer->result_symbol = operator_result;
if (has_ws_result) lexer->mark_end(lexer);
return true;
}
if (has_ws_result) {
// Don't `mark_end`, since we may have advanced through some operators.
lexer->result_symbol = ws_result;
return true;
}
// NOTE: this will consume any `#` characters it sees, even if it does not find a result. Keep
// it at the end so that it doesn't interfere with special literals or selectors!
enum TokenType raw_str_result;
bool saw_raw_str_part = eat_raw_str_part(state, lexer, valid_symbols, &raw_str_result);
if (saw_raw_str_part) {
lexer->result_symbol = raw_str_result;
return true;
}
return false;
}

View file

@ -0,0 +1,54 @@
#ifndef TREE_SITTER_ALLOC_H_
#define TREE_SITTER_ALLOC_H_
#ifdef __cplusplus
extern "C" {
#endif
#include <stdbool.h>
#include <stdio.h>
#include <stdlib.h>
// Allow clients to override allocation functions
#ifdef TREE_SITTER_REUSE_ALLOCATOR
extern void *(*ts_current_malloc)(size_t);
extern void *(*ts_current_calloc)(size_t, size_t);
extern void *(*ts_current_realloc)(void *, size_t);
extern void (*ts_current_free)(void *);
#ifndef ts_malloc
#define ts_malloc ts_current_malloc
#endif
#ifndef ts_calloc
#define ts_calloc ts_current_calloc
#endif
#ifndef ts_realloc
#define ts_realloc ts_current_realloc
#endif
#ifndef ts_free
#define ts_free ts_current_free
#endif
#else
#ifndef ts_malloc
#define ts_malloc malloc
#endif
#ifndef ts_calloc
#define ts_calloc calloc
#endif
#ifndef ts_realloc
#define ts_realloc realloc
#endif
#ifndef ts_free
#define ts_free free
#endif
#endif
#ifdef __cplusplus
}
#endif
#endif // TREE_SITTER_ALLOC_H_

View file

@ -0,0 +1,290 @@
#ifndef TREE_SITTER_ARRAY_H_
#define TREE_SITTER_ARRAY_H_
#ifdef __cplusplus
extern "C" {
#endif
#include "./alloc.h"
#include <assert.h>
#include <stdbool.h>
#include <stdint.h>
#include <stdlib.h>
#include <string.h>
#ifdef _MSC_VER
#pragma warning(disable : 4101)
#elif defined(__GNUC__) || defined(__clang__)
#pragma GCC diagnostic push
#pragma GCC diagnostic ignored "-Wunused-variable"
#endif
#define Array(T) \
struct { \
T *contents; \
uint32_t size; \
uint32_t capacity; \
}
/// Initialize an array.
#define array_init(self) \
((self)->size = 0, (self)->capacity = 0, (self)->contents = NULL)
/// Create an empty array.
#define array_new() \
{ NULL, 0, 0 }
/// Get a pointer to the element at a given `index` in the array.
#define array_get(self, _index) \
(assert((uint32_t)(_index) < (self)->size), &(self)->contents[_index])
/// Get a pointer to the first element in the array.
#define array_front(self) array_get(self, 0)
/// Get a pointer to the last element in the array.
#define array_back(self) array_get(self, (self)->size - 1)
/// Clear the array, setting its size to zero. Note that this does not free any
/// memory allocated for the array's contents.
#define array_clear(self) ((self)->size = 0)
/// Reserve `new_capacity` elements of space in the array. If `new_capacity` is
/// less than the array's current capacity, this function has no effect.
#define array_reserve(self, new_capacity) \
_array__reserve((Array *)(self), array_elem_size(self), new_capacity)
/// Free any memory allocated for this array. Note that this does not free any
/// memory allocated for the array's contents.
#define array_delete(self) _array__delete((Array *)(self))
/// Push a new `element` onto the end of the array.
#define array_push(self, element) \
(_array__grow((Array *)(self), 1, array_elem_size(self)), \
(self)->contents[(self)->size++] = (element))
/// Increase the array's size by `count` elements.
/// New elements are zero-initialized.
#define array_grow_by(self, count) \
do { \
if ((count) == 0) break; \
_array__grow((Array *)(self), count, array_elem_size(self)); \
memset((self)->contents + (self)->size, 0, (count) * array_elem_size(self)); \
(self)->size += (count); \
} while (0)
/// Append all elements from one array to the end of another.
#define array_push_all(self, other) \
array_extend((self), (other)->size, (other)->contents)
/// Append `count` elements to the end of the array, reading their values from the
/// `contents` pointer.
#define array_extend(self, count, contents) \
_array__splice( \
(Array *)(self), array_elem_size(self), (self)->size, \
0, count, contents \
)
/// Remove `old_count` elements from the array starting at the given `index`. At
/// the same index, insert `new_count` new elements, reading their values from the
/// `new_contents` pointer.
#define array_splice(self, _index, old_count, new_count, new_contents) \
_array__splice( \
(Array *)(self), array_elem_size(self), _index, \
old_count, new_count, new_contents \
)
/// Insert one `element` into the array at the given `index`.
#define array_insert(self, _index, element) \
_array__splice((Array *)(self), array_elem_size(self), _index, 0, 1, &(element))
/// Remove one element from the array at the given `index`.
#define array_erase(self, _index) \
_array__erase((Array *)(self), array_elem_size(self), _index)
/// Pop the last element off the array, returning the element by value.
#define array_pop(self) ((self)->contents[--(self)->size])
/// Assign the contents of one array to another, reallocating if necessary.
#define array_assign(self, other) \
_array__assign((Array *)(self), (const Array *)(other), array_elem_size(self))
/// Swap one array with another
#define array_swap(self, other) \
_array__swap((Array *)(self), (Array *)(other))
/// Get the size of the array contents
#define array_elem_size(self) (sizeof *(self)->contents)
/// Search a sorted array for a given `needle` value, using the given `compare`
/// callback to determine the order.
///
/// If an existing element is found to be equal to `needle`, then the `index`
/// out-parameter is set to the existing value's index, and the `exists`
/// out-parameter is set to true. Otherwise, `index` is set to an index where
/// `needle` should be inserted in order to preserve the sorting, and `exists`
/// is set to false.
#define array_search_sorted_with(self, compare, needle, _index, _exists) \
_array__search_sorted(self, 0, compare, , needle, _index, _exists)
/// Search a sorted array for a given `needle` value, using integer comparisons
/// of a given struct field (specified with a leading dot) to determine the order.
///
/// See also `array_search_sorted_with`.
#define array_search_sorted_by(self, field, needle, _index, _exists) \
_array__search_sorted(self, 0, _compare_int, field, needle, _index, _exists)
/// Insert a given `value` into a sorted array, using the given `compare`
/// callback to determine the order.
#define array_insert_sorted_with(self, compare, value) \
do { \
unsigned _index, _exists; \
array_search_sorted_with(self, compare, &(value), &_index, &_exists); \
if (!_exists) array_insert(self, _index, value); \
} while (0)
/// Insert a given `value` into a sorted array, using integer comparisons of
/// a given struct field (specified with a leading dot) to determine the order.
///
/// See also `array_search_sorted_by`.
#define array_insert_sorted_by(self, field, value) \
do { \
unsigned _index, _exists; \
array_search_sorted_by(self, field, (value) field, &_index, &_exists); \
if (!_exists) array_insert(self, _index, value); \
} while (0)
// Private
typedef Array(void) Array;
/// This is not what you're looking for, see `array_delete`.
static inline void _array__delete(Array *self) {
if (self->contents) {
ts_free(self->contents);
self->contents = NULL;
self->size = 0;
self->capacity = 0;
}
}
/// This is not what you're looking for, see `array_erase`.
static inline void _array__erase(Array *self, size_t element_size,
uint32_t index) {
assert(index < self->size);
char *contents = (char *)self->contents;
memmove(contents + index * element_size, contents + (index + 1) * element_size,
(self->size - index - 1) * element_size);
self->size--;
}
/// This is not what you're looking for, see `array_reserve`.
static inline void _array__reserve(Array *self, size_t element_size, uint32_t new_capacity) {
if (new_capacity > self->capacity) {
if (self->contents) {
self->contents = ts_realloc(self->contents, new_capacity * element_size);
} else {
self->contents = ts_malloc(new_capacity * element_size);
}
self->capacity = new_capacity;
}
}
/// This is not what you're looking for, see `array_assign`.
static inline void _array__assign(Array *self, const Array *other, size_t element_size) {
_array__reserve(self, element_size, other->size);
self->size = other->size;
memcpy(self->contents, other->contents, self->size * element_size);
}
/// This is not what you're looking for, see `array_swap`.
static inline void _array__swap(Array *self, Array *other) {
Array swap = *other;
*other = *self;
*self = swap;
}
/// This is not what you're looking for, see `array_push` or `array_grow_by`.
static inline void _array__grow(Array *self, uint32_t count, size_t element_size) {
uint32_t new_size = self->size + count;
if (new_size > self->capacity) {
uint32_t new_capacity = self->capacity * 2;
if (new_capacity < 8) new_capacity = 8;
if (new_capacity < new_size) new_capacity = new_size;
_array__reserve(self, element_size, new_capacity);
}
}
/// This is not what you're looking for, see `array_splice`.
static inline void _array__splice(Array *self, size_t element_size,
uint32_t index, uint32_t old_count,
uint32_t new_count, const void *elements) {
uint32_t new_size = self->size + new_count - old_count;
uint32_t old_end = index + old_count;
uint32_t new_end = index + new_count;
assert(old_end <= self->size);
_array__reserve(self, element_size, new_size);
char *contents = (char *)self->contents;
if (self->size > old_end) {
memmove(
contents + new_end * element_size,
contents + old_end * element_size,
(self->size - old_end) * element_size
);
}
if (new_count > 0) {
if (elements) {
memcpy(
(contents + index * element_size),
elements,
new_count * element_size
);
} else {
memset(
(contents + index * element_size),
0,
new_count * element_size
);
}
}
self->size += new_count - old_count;
}
/// A binary search routine, based on Rust's `std::slice::binary_search_by`.
/// This is not what you're looking for, see `array_search_sorted_with` or `array_search_sorted_by`.
#define _array__search_sorted(self, start, compare, suffix, needle, _index, _exists) \
do { \
*(_index) = start; \
*(_exists) = false; \
uint32_t size = (self)->size - *(_index); \
if (size == 0) break; \
int comparison; \
while (size > 1) { \
uint32_t half_size = size / 2; \
uint32_t mid_index = *(_index) + half_size; \
comparison = compare(&((self)->contents[mid_index] suffix), (needle)); \
if (comparison <= 0) *(_index) = mid_index; \
size -= half_size; \
} \
comparison = compare(&((self)->contents[*(_index)] suffix), (needle)); \
if (comparison == 0) *(_exists) = true; \
else if (comparison < 0) *(_index) += 1; \
} while (0)
/// Helper macro for the `_sorted_by` routines below. This takes the left (existing)
/// parameter by reference in order to work with the generic sorting function above.
#define _compare_int(a, b) ((int)*(a) - (int)(b))
#ifdef _MSC_VER
#pragma warning(default : 4101)
#elif defined(__GNUC__) || defined(__clang__)
#pragma GCC diagnostic pop
#endif
#ifdef __cplusplus
}
#endif
#endif // TREE_SITTER_ARRAY_H_

View file

@ -0,0 +1,266 @@
#ifndef TREE_SITTER_PARSER_H_
#define TREE_SITTER_PARSER_H_
#ifdef __cplusplus
extern "C" {
#endif
#include <stdbool.h>
#include <stdint.h>
#include <stdlib.h>
#define ts_builtin_sym_error ((TSSymbol)-1)
#define ts_builtin_sym_end 0
#define TREE_SITTER_SERIALIZATION_BUFFER_SIZE 1024
#ifndef TREE_SITTER_API_H_
typedef uint16_t TSStateId;
typedef uint16_t TSSymbol;
typedef uint16_t TSFieldId;
typedef struct TSLanguage TSLanguage;
#endif
typedef struct {
TSFieldId field_id;
uint8_t child_index;
bool inherited;
} TSFieldMapEntry;
typedef struct {
uint16_t index;
uint16_t length;
} TSFieldMapSlice;
typedef struct {
bool visible;
bool named;
bool supertype;
} TSSymbolMetadata;
typedef struct TSLexer TSLexer;
struct TSLexer {
int32_t lookahead;
TSSymbol result_symbol;
void (*advance)(TSLexer *, bool);
void (*mark_end)(TSLexer *);
uint32_t (*get_column)(TSLexer *);
bool (*is_at_included_range_start)(const TSLexer *);
bool (*eof)(const TSLexer *);
void (*log)(const TSLexer *, const char *, ...);
};
typedef enum {
TSParseActionTypeShift,
TSParseActionTypeReduce,
TSParseActionTypeAccept,
TSParseActionTypeRecover,
} TSParseActionType;
typedef union {
struct {
uint8_t type;
TSStateId state;
bool extra;
bool repetition;
} shift;
struct {
uint8_t type;
uint8_t child_count;
TSSymbol symbol;
int16_t dynamic_precedence;
uint16_t production_id;
} reduce;
uint8_t type;
} TSParseAction;
typedef struct {
uint16_t lex_state;
uint16_t external_lex_state;
} TSLexMode;
typedef union {
TSParseAction action;
struct {
uint8_t count;
bool reusable;
} entry;
} TSParseActionEntry;
typedef struct {
int32_t start;
int32_t end;
} TSCharacterRange;
struct TSLanguage {
uint32_t version;
uint32_t symbol_count;
uint32_t alias_count;
uint32_t token_count;
uint32_t external_token_count;
uint32_t state_count;
uint32_t large_state_count;
uint32_t production_id_count;
uint32_t field_count;
uint16_t max_alias_sequence_length;
const uint16_t *parse_table;
const uint16_t *small_parse_table;
const uint32_t *small_parse_table_map;
const TSParseActionEntry *parse_actions;
const char * const *symbol_names;
const char * const *field_names;
const TSFieldMapSlice *field_map_slices;
const TSFieldMapEntry *field_map_entries;
const TSSymbolMetadata *symbol_metadata;
const TSSymbol *public_symbol_map;
const uint16_t *alias_map;
const TSSymbol *alias_sequences;
const TSLexMode *lex_modes;
bool (*lex_fn)(TSLexer *, TSStateId);
bool (*keyword_lex_fn)(TSLexer *, TSStateId);
TSSymbol keyword_capture_token;
struct {
const bool *states;
const TSSymbol *symbol_map;
void *(*create)(void);
void (*destroy)(void *);
bool (*scan)(void *, TSLexer *, const bool *symbol_whitelist);
unsigned (*serialize)(void *, char *);
void (*deserialize)(void *, const char *, unsigned);
} external_scanner;
const TSStateId *primary_state_ids;
};
static inline bool set_contains(TSCharacterRange *ranges, uint32_t len, int32_t lookahead) {
uint32_t index = 0;
uint32_t size = len - index;
while (size > 1) {
uint32_t half_size = size / 2;
uint32_t mid_index = index + half_size;
TSCharacterRange *range = &ranges[mid_index];
if (lookahead >= range->start && lookahead <= range->end) {
return true;
} else if (lookahead > range->end) {
index = mid_index;
}
size -= half_size;
}
TSCharacterRange *range = &ranges[index];
return (lookahead >= range->start && lookahead <= range->end);
}
/*
* Lexer Macros
*/
#ifdef _MSC_VER
#define UNUSED __pragma(warning(suppress : 4101))
#else
#define UNUSED __attribute__((unused))
#endif
#define START_LEXER() \
bool result = false; \
bool skip = false; \
UNUSED \
bool eof = false; \
int32_t lookahead; \
goto start; \
next_state: \
lexer->advance(lexer, skip); \
start: \
skip = false; \
lookahead = lexer->lookahead;
#define ADVANCE(state_value) \
{ \
state = state_value; \
goto next_state; \
}
#define ADVANCE_MAP(...) \
{ \
static const uint16_t map[] = { __VA_ARGS__ }; \
for (uint32_t i = 0; i < sizeof(map) / sizeof(map[0]); i += 2) { \
if (map[i] == lookahead) { \
state = map[i + 1]; \
goto next_state; \
} \
} \
}
#define SKIP(state_value) \
{ \
skip = true; \
state = state_value; \
goto next_state; \
}
#define ACCEPT_TOKEN(symbol_value) \
result = true; \
lexer->result_symbol = symbol_value; \
lexer->mark_end(lexer);
#define END_STATE() return result;
/*
* Parse Table Macros
*/
#define SMALL_STATE(id) ((id) - LARGE_STATE_COUNT)
#define STATE(id) id
#define ACTIONS(id) id
#define SHIFT(state_value) \
{{ \
.shift = { \
.type = TSParseActionTypeShift, \
.state = (state_value) \
} \
}}
#define SHIFT_REPEAT(state_value) \
{{ \
.shift = { \
.type = TSParseActionTypeShift, \
.state = (state_value), \
.repetition = true \
} \
}}
#define SHIFT_EXTRA() \
{{ \
.shift = { \
.type = TSParseActionTypeShift, \
.extra = true \
} \
}}
#define REDUCE(symbol_name, children, precedence, prod_id) \
{{ \
.reduce = { \
.type = TSParseActionTypeReduce, \
.symbol = symbol_name, \
.child_count = children, \
.dynamic_precedence = precedence, \
.production_id = prod_id \
}, \
}}
#define RECOVER() \
{{ \
.type = TSParseActionTypeRecover \
}}
#define ACCEPT_INPUT() \
{{ \
.type = TSParseActionTypeAccept \
}}
#ifdef __cplusplus
}
#endif
#endif // TREE_SITTER_PARSER_H_