feat(skills): add js-analysis skill and wire into scan modes

Add a JavaScript harvesting + static-analysis skill that other specialist
agents consume as their attack surface map. Produces a deterministic
js_analysis.md artifact with API endpoints, parameters, secrets,
dangerous sinks, source-map recovery, and auth/session touchpoints.

Wire into hunter, deep, and standard scan modes as a mandatory first
agent. Quick mode runs a lightweight pass (Phases 1-3 only).
This commit is contained in:
h0tak88r 2026-06-06 21:56:10 +03:00
parent 6e17cfb9b9
commit 4fc9156917
5 changed files with 430 additions and 0 deletions

View file

@ -156,6 +156,10 @@ Spawn specialized agents at each level. Scale horizontally to maximum paralleliz
- Each agent focuses on one specific area or vulnerability type
- Creates a massive parallel swarm covering every angle
**JS Analysis Agent (mandatory, runs first)**
Before spawning vulnerability agents, spawn a single `JS Analysis Agent` with `skills=["js-analysis"]`. It harvests every JS file (including lazy/dynamic chunks via the browser tool), recovers source maps, and extracts API endpoints, parameters, secrets, dangerous sinks, and auth/session touchpoints into a single `js_analysis.md` artifact. Every downstream specialist agent (IDOR, SSRF, XSS, Auth) consumes that artifact as input. Do not start vulnerability hunting until the artifact exists — it is the surface map.
## Mindset
Relentless. Creative. Patient. Thorough. Persistent.

View file

@ -0,0 +1,199 @@
---
name: hunter
description: Aggressive security penetration testing with recursive deepening, automated persistence, and comprehensive vulnerability coverage
---
# Hunter Scan Mode
Aggressive penetration testing methodology. Maximum surface coverage, recursive vulnerability deepening, zero-tolerance for false positives. Every finding becomes a weapon for subsequent rounds.
## Target Classification
| Target Type | Signal | Mode |
|---|---|---|
| **Focused** | User specifies scope or vulnerability class | Test ONLY the requested scope or vuln class with full recon depth. |
| **Open** | Production domain, no scope limit | All phases, all vulnerabilities. Maximum depth and persistence. |
## Core Principles
- **Think Before Acting**: Classify the target, identify intent, and deploy appropriate testing power.
- **Autonomous Deepening**: Every finding is a pivot point for subsequent rounds. Finish only when coverage is complete.
- **Zero-Tolerance for False Positives**: Every finding must be validated with concrete evidence (HTTP requests, screenshots, or PoC).
- **Total Surface Coverage**: A scan is invalid if any endpoint, feature, or parameter remains untested.
- **Recursive Evolution**: Use every leak, error, and reflection to pivot into deeper vulnerabilities.
- **Aggressive Persistence**: If automated tools find nothing, that is when the real work begins.
## Phase 1: Authentication & Reconnaissance
**Authentication First**
Authenticate immediately before any testing. Pick the fastest method (Curl first, Browser if needed). Nothing happens before login.
**Surface Mapping**
- Extract ALL endpoints from HTML, JavaScript, API docs, sitemaps, and robots.txt
- Extract all JavaScript URLs and save JS files locally for source code review
- Write the complete attack surface (endpoints + parameters + JS files) to `attack_surface.md`
- Map all user roles with different account types
- Document rate limiting, WAF rules, and security controls
**Subdomain Enumeration**
- Enumerate subdomains with multiple sources and tools
- Collect URLs from multiple aggregators (URLFinder, AlienVault OTX, crt.sh, CertSpotter)
- Use VirusTotal domain reports for historical URL discovery
- Merge and deduplicate all URL sources
**Credential & Secret Discovery**
Filter collected URLs and JavaScript files for authentication artifacts and secrets:
- URLs containing tokens, API keys, passwords, or embedded credentials
- JavaScript files containing hardcoded secrets (SendGrid keys, AWS keys, Stripe keys, GitHub PATs, Slack webhooks, private keys, Sentry DSNs)
- Private file leaks (PDFs, documents, images) from non-public directories
- Filter out common false positives: variable names without values, Base64-encoded binary data, React internal strings, empty/null values
## Phase 2: Passive Vulnerability Discovery
**Email Security Assessment**
- Check SPF, DMARC, and DKIM records for the target domain
- Report missing or weak email security configurations
- DMARC `p=none` with otherwise strong SPF is HIGH severity for organizations where customers trust email communications
**Form Discovery & Analysis**
- Search for HTML forms on all discovered subdomains
- Check for chat widget integrations (Intercom, HubSpot, Zendesk, Crisp, Drift)
- For JavaScript-heavy SPAs, use browser tools for DOM-based form extraction
- Test email-based fields for XSS with polyglot payloads at registration, login, password reset, newsletter signup, contact forms, and account settings
**Widget Misconfiguration Testing**
- Test Intercom boot-time injection for unauthorized access
- Test chat widget configurations for data leakage
- Flag Salesforce sandbox URLs leaking into production chat widgets
**Paywall & Access Control Bypass**
- Test origin IP access to bypass CDN-restricted paywalls
- Test common paywall bypass paths with direct origin IP requests
- Check for accessible private content without authentication
## Phase 3: Active Vulnerability Testing
**Race Condition Testing**
Test authenticated endpoints for race conditions on state-changing operations:
- Coupon/promo code redemption
- Gift card balance transfers
- Payment/checkout flows
- Referral bonus claiming
- Voting/liking systems
- Limited-quantity item purchases
Send multiple simultaneous requests with identical parameters and check if all succeeded when only one should have.
**Social Media & Link Hijacking**
- Extract all social media links from the target site
- Check each profile for 404s, abandoned accounts, or available usernames
- Test Discord invites for expiration
**Email Reservation Lockout**
Test email change flows for lockout vulnerabilities where an attacker can deny a user from creating an account by reserving their email during an unverified email change flow.
## Phase 4: Injection Testing
Apply recursive vulnerability classes across all discovered endpoints and parameters.
**SQL Injection**
- Error-based, boolean-based, time-based, and UNION-based techniques
- Out-of-band DNS exfiltration via `xp_dirtree` or similar
- Test all parameter types: path, query, body, headers, cookies
**Cross-Site Scripting**
- Reflected, stored, and blind XSS payloads
- Out-of-band exfiltration via fetch callbacks
- Store payloads in all fields and check for callbacks from admin panels
**Server-Side Template Injection**
- Test multiple template engines: Jinja2, ERB, Twig, Freemarker
- Escalate from arithmetic evaluation to remote code execution
**SSRF & Local File Inclusion**
- Cloud metadata endpoints (AWS, GCP, Azure)
- Local file access via `file://` and path traversal
- Internal service discovery (Redis, Postgres, internal APIs)
- Out-of-band SSRF via DNS and HTTP callbacks
**Prototype Pollution**
- Server-side prototype pollution via `__proto__` and `constructor.prototype`
- Test for status code manipulation, authentication bypass, and remote code execution
**Blind Deserialization**
- Java, PHP, Python, and .NET deserialization attack vectors
- Out-of-band exfiltration via URLDNS and similar gadgets
## Phase 5: Out-of-Band Exfiltration
Establish and use out-of-band channels for blind vulnerability confirmation:
- **DNS**: Subdomain-based exfiltration for command output
- **HTTP**: Callback-based exfiltration for large data
- **Timing**: Sleep-based detection when outbound is blocked
- **Error-based**: Local data leakage via error messages
## Phase 6: Reporting & Recursive Re-Recon
**Reporting**
For each confirmed finding, create a report with:
- **Title**: [Severity] Finding Name
- **CWE**: CWE identifier, bug type, scope, endpoint, vulnerable part, payload, technical environment
- **Description**: Technical root cause and vulnerability details
- **Steps to Reproduce**: Detailed, numbered steps
- **PoC**: Exact command or script that proves the vulnerability
- **Impact**: Business and security consequence
- **Remediation**: Specific, actionable fix
**Re-Recon After Privilege Elevation**
After every privilege elevation, restart reconnaissance from the new access level. Elevated sessions reveal new attack surfaces that were previously inaccessible.
## Agent Strategy
After initial reconnaissance, decompose the application:
1. **Component level** - Auth System, Payment Gateway, Admin Panel
2. **Feature level** - Login Form, Registration API, Password Reset
3. **Vulnerability level** - SQLi Agent, XSS Agent, Auth Bypass Agent
Spawn specialized agents at each level. Scale horizontally to maximum parallelization:
- Do NOT overload a single agent with multiple vulnerability types
- Each agent focuses on one specific area or vulnerability type
- Creates a parallel swarm covering every angle
**JS Analysis Agent (mandatory, runs first)**
Before spawning vulnerability agents, spawn a single `JS Analysis Agent` with `skills=["js-analysis"]`. Its job is to harvest every JS file (including lazy/dynamic chunks via the browser tool), extract API endpoints, parameters, secrets, dangerous sinks, and auth/session touchpoints into a single `js_analysis.md` artifact. Every downstream specialist agent (IDOR, SSRF, XSS, Auth) reads that artifact as input. Do not start vulnerability hunting until this artifact exists — it is the surface map.
## Rules
- Always confirm with a PoC before reporting
- Test every parameter (path, query, body, headers, cookies)
- Stay in the authenticated session for authenticated flows
- Never report theoretical findings
- Never test targets not explicitly authorized
- Never stop mid-scan without a summary of findings
- Zero hallucination: never write a response you didn't receive

View file

@ -45,6 +45,10 @@ Skip for quick scans:
- Low-severity information disclosure
- Theoretical issues without working PoC
**JS Analysis Agent (lightweight pass)**
Even on a quick scan, do one fast JS pass: spawn a `JS Analysis Agent` with `skills=["js-analysis"]` and instruct it to run **Phase 1 + Phase 2 + Phase 3 only** (collection, endpoint/parameter extraction, secret extraction) — skip lazy-chunk browser capture and deep sink analysis. The resulting `js_analysis.md` gives every high-impact agent (Auth bypass, IDOR, SSRF) a free endpoint and secret map without slowing the scan.
## Phase 3: Validation
- Confirm exploitability with minimal proof-of-concept

View file

@ -46,6 +46,10 @@ Before testing for vulnerabilities, understand the application:
Test each attack surface methodically. Spawn focused subagents for different areas.
**JS Analysis Agent (runs first)**
Before spawning vulnerability subagents, spawn a single `JS Analysis Agent` with `skills=["js-analysis"]`. It harvests every JS file (including lazy/dynamic chunks via the browser tool), recovers source maps, and extracts API endpoints, parameters, secrets, dangerous sinks, and auth/session touchpoints into a single `js_analysis.md` artifact. Downstream specialists (IDOR, SSRF, XSS, Auth) read it as their surface map.
**Input Validation**
- Injection testing on all input fields (SQL, XSS, command, template)
- File upload bypass attempts

View file

@ -0,0 +1,219 @@
---
name: js-analysis
description: JavaScript file harvesting and static analysis to extract API endpoints, parameters, secrets, dangerous sinks, and source maps for downstream specialist agents
---
# JavaScript Analysis
Recon-and-analysis specialist for the JavaScript attack surface. Collect every JS file the application loads (including lazy-loaded chunks and dynamically imported modules), run regex/AST extraction over them, and emit a structured artifact other specialist agents (IDOR, SSRF, XSS, Auth) can consume.
This agent does NOT exploit. It produces a high-fidelity inventory.
## Output Contract
Write a single `js_analysis.md` artifact to the workspace with these sections. Other agents read this — keep it deterministic.
```
# JS Analysis — <target>
## Inventory
- <url> — <size> — <type: main|chunk|lazy|vendor|sourcemap>
## API Endpoints
- METHOD /path — source: <jsfile>:<line> — context: <one-line snippet>
## Parameters Observed
- <param> — endpoints: [<list>] — values seen: [<sample>]
## Secrets / Keys
- <kind> — <redacted match> — source: <jsfile>:<line>
## Dangerous Sinks
- <sink: eval|innerHTML|document.write|postMessage|location.href|dangerouslySetInnerHTML|new Function|setTimeout(string)|setInterval(string)|window[var]> — source: <jsfile>:<line> — taint: <reachable from user input? yes/no/unknown>
## Auth & Session Surfaces
- localStorage/sessionStorage keys, cookie names, JWT decode sites, refresh-token endpoints
## Routes / Client Pages
- <route> — source: <jsfile>:<line>
## Source Maps
- <jsfile>.map — recovered: yes/no — original sources count: <N>
## Notes for Downstream Agents
- IDOR agent: <pointer to most likely vulnerable endpoints>
- SSRF agent: <list of URL-builder helpers found>
- XSS agent: <list of innerHTML/dangerouslySetInnerHTML sinks>
- Auth agent: <list of JWT/refresh/cookie touchpoints>
```
## Phase 1 — Collection
Two-pass collection. Static fetch + dynamic browser to catch lazy chunks.
**Pass A — Static crawl**
1. Use `katana` / `gau` / `waybackurls` on the target to harvest historical and live URLs.
2. Filter for `.js`, `.mjs`, `.cjs`, `.map` extensions and JS-Content-Type responses.
3. Use `httpx` to confirm 200 status, capture final URL after redirects, record size + sha256.
4. Save each file to `js_corpus/<host>/<sha256>.js` and keep a manifest `js_corpus/manifest.tsv`:
`url\tstatus\tsize\tsha256\tcontent_type\tsource`
**Pass B — Browser-driven lazy capture**
Many SPAs only fetch chunks after user interaction. Use the browser tool:
1. Open the app, log in if needed, hit every route in the inventory.
2. Trigger interactions that lazy-load: dropdowns, modals, tab switches, route navigation, search, file uploads, settings, admin panel.
3. Capture all network requests where `Content-Type` matches `*javascript*` or `*ecmascript*`.
4. Compare against Pass A — anything new is a lazy/dynamic chunk. Add to manifest with `source=dynamic`.
**Source maps**
For each `<file>.js`, request `<file>.js.map`. If present:
- Parse `sources[]` array — recover original file tree (often leaks internal package names, paths, auth helpers).
- Note any reference to internal hostnames, microservice names, AWS account IDs.
- Save recovered originals to `js_corpus/<host>/sourcemap/<original_path>`.
## Phase 2 — Endpoint & Parameter Extraction
Run these regex passes over every JS file. Tag each hit with `<file>:<line>`.
**HTTP method + path strings**
```regex
(?:fetch|axios\.(?:get|post|put|delete|patch|request)|\$\.(?:get|post|ajax)|XMLHttpRequest|\.open)\s*\(\s*['"`]([^'"`]+)['"`]
```
```regex
(?:url|endpoint|path|route)\s*[:=]\s*['"`](/[A-Za-z0-9_\-/.?&=:%{}]+)['"`]
```
```regex
['"`](/(?:api|v\d+|graphql|gql|rest|rpc|internal|admin|auth|user|users|account|me)[/A-Za-z0-9_\-./?&=:%{}]*)['"`]
```
**Inline template literals (routes with `${id}` etc.)**
```regex
`(/[A-Za-z0-9_\-/.?&=:%]*\$\{[^}]+\}[A-Za-z0-9_\-/.?&=:%]*)`
```
Normalize `${var}` → `{var}` in the output.
**Query / body parameter names**
```regex
(?:params|data|body|searchParams|form)\s*[:=]\s*\{([^}]{1,300})\}
```
Extract keys from the matched object literal.
**GraphQL operations**
```regex
(?:gql|graphql)\s*`\s*(query|mutation|subscription)\s+(\w+)
```
Save full operation bodies; they describe the entire backend object graph.
## Phase 3 — Secret Extraction
Run high-precision regex, then triage. False positives waste downstream agents' time.
```
AWS access key AKIA[0-9A-Z]{16}
AWS secret (?i)aws(.{0,20})?(?-i)['"`][0-9a-zA-Z/+]{40}['"`]
Google API key AIza[0-9A-Za-z\-_]{35}
Stripe live sk_live_[0-9a-zA-Z]{24,}
Stripe publishable pk_live_[0-9a-zA-Z]{24,}
SendGrid SG\.[A-Za-z0-9_\-]{22}\.[A-Za-z0-9_\-]{43}
Slack webhook https://hooks\.slack\.com/services/T[A-Z0-9]+/B[A-Z0-9]+/[A-Za-z0-9]+
Slack token xox[abprs]-[0-9a-zA-Z\-]+
GitHub PAT ghp_[A-Za-z0-9]{36}
GitHub fine-grain github_pat_[A-Za-z0-9_]{82}
JWT eyJ[A-Za-z0-9_\-]+\.eyJ[A-Za-z0-9_\-]+\.[A-Za-z0-9_\-]+
Private key -----BEGIN (?:RSA |EC |DSA |OPENSSH )?PRIVATE KEY-----
Firebase config apiKey.{0,3}:.{0,3}["']AIza[0-9A-Za-z\-_]{35}["']
Sentry DSN https://[0-9a-f]{32}@[a-z0-9.\-]+/[0-9]+
Mapbox token pk\.eyJ[A-Za-z0-9_\-]+\.[A-Za-z0-9_\-]+
Algolia (?:algolia).{0,30}["'][0-9a-f]{32}["']
```
For every match: redact middle bytes when writing to artifact (`AKIA****REDACTED****ABCD`), but record full value in a separate `secrets.tsv` for the operator only. **Never** exfiltrate or test secrets against live infra without explicit approval.
## Phase 4 — Dangerous Sink Detection
These are the sinks downstream XSS / DOM-XSS / RCE / open-redirect agents need pointers to:
| Sink | Regex | Bug class |
|---|---|---|
| `eval(` | `\beval\s*\(` | RCE / DOM-XSS |
| `new Function(` | `new\s+Function\s*\(` | RCE / DOM-XSS |
| `setTimeout(string)` | `setTimeout\s*\(\s*['"`]` | RCE |
| `setInterval(string)` | `setInterval\s*\(\s*['"`]` | RCE |
| `innerHTML =` | `\.innerHTML\s*=` | DOM-XSS |
| `outerHTML =` | `\.outerHTML\s*=` | DOM-XSS |
| `document.write(` | `document\.write(?:ln)?\s*\(` | DOM-XSS |
| `dangerouslySetInnerHTML` | `dangerouslySetInnerHTML\s*=\s*\{` | DOM-XSS (React) |
| `v-html=` | `v-html\s*=` | DOM-XSS (Vue) |
| `[innerHTML]=` | `\[innerHTML\]\s*=` | DOM-XSS (Angular) |
| `location =` | `(?:window\.)?location(?:\.href)?\s*=` | Open-redirect |
| `location.replace(` | `location\.replace\s*\(` | Open-redirect |
| `postMessage(` | `\.postMessage\s*\(` | postMessage flaws |
| `addEventListener('message'` | `addEventListener\s*\(\s*['"]message['"]` | postMessage flaws |
| `JSON.parse(localStorage` | `JSON\.parse\s*\(\s*(?:localStorage\|sessionStorage)` | Prototype pollution / data tamper |
| `Object.assign({},` | `Object\.assign\s*\(` | Prototype pollution |
| `_.merge(` / `_.mergeWith(` | `_\.(?:merge\|mergeWith\|defaultsDeep)\s*\(` | Prototype pollution |
| `_.set(` | `_\.set\s*\(` | Prototype pollution |
| `window[…]` | `window\s*\[[^\]]+\]\s*\(` | Property-name injection |
For each hit, capture **3 lines of context** (the sink line + 1 before + 1 after) so the downstream agent can judge taint quickly without re-reading the whole bundle.
**Quick taint heuristic**: if the same file references `location.search`, `location.hash`, `URLSearchParams`, `document.referrer`, `window.name`, `postMessage event.data`, or reads from a URL-bound state library (React Router `useParams`/`useSearchParams`, Vue `$route`, Angular `ActivatedRoute`), mark sinks in that file as `taint: likely`. Otherwise `unknown`.
## Phase 5 — Auth & Session Inventory
Other agents (Auth, IDOR) need this fast:
- All `localStorage.setItem` / `getItem` keys
- All `sessionStorage.setItem` / `getItem` keys
- All `document.cookie =` assignments
- All `Authorization: Bearer` template literals → identifies token shape (JWT vs opaque)
- All `jwt_decode(` / `jose.decodeJwt(` call sites — confirms tokens are JWT, gives you a free decode site
- Refresh / silent-renew endpoints (search for `refresh`, `silentRenew`, `token/renew`)
- CSRF token retrieval (search for `csrf`, `xsrf`, `X-CSRF`, `X-XSRF-Token`)
- OAuth client IDs and redirect URIs (search for `client_id`, `redirect_uri`, `response_type`)
## Phase 6 — Hand-off Notes
End the artifact with explicit pointers other agents need. Be specific — name endpoints, not vibes:
```
## Notes for Downstream Agents
- IDOR agent: GET /api/v2/orders/{id}, GET /api/users/{id}/invoices, GET /api/files/{uuid} all
built from path params with no server-side ownership claim visible in JS. Highest-ROI targets.
- SSRF agent: `buildProxyUrl(host)` in vendor.abc123.js:4421 concatenates user-controlled `host`
into outbound fetch. Worth probing /api/proxy?target=
- XSS agent: dangerouslySetInnerHTML in CommentRenderer (chunk-9f2.js:188) — taint: likely.
Comment field is server-rendered; check if HTML is sanitized server-side.
- Auth agent: refresh endpoint is POST /api/auth/silent — accepts refresh_token in body, no
device binding observed in client. Possible token replay surface.
```
## Tools
- `katana`, `gau`, `waybackurls` — URL harvest
- `httpx` — fetch + metadata
- `agent_browser` — lazy chunk capture (Pass B)
- `semgrep` — for higher-fidelity AST passes if regex misses (rulesets: `r/javascript.audit.xss`, `r/javascript.lang.security`)
- Python regex driver (write a one-shot script per pass; keep outputs deterministic)
## Rules
- **Never** post extracted secrets to third-party services (no virustotal-uploading bundles, no online beautifiers).
- Beautify locally only (`js-beautify`, `prettier`).
- Source-map originals can contain proprietary code — do not exfiltrate, keep local to the workspace.
- Do not run extracted JS. Static analysis only at this stage.
- Mark every finding `taint: unknown` if you cannot trace it back to a user-controllable source — let the downstream agent decide.
- One artifact per target. Overwrite, do not append; downstream agents always read the latest.