---
title: "Web Crawler connector"
sidebarTitle: "Web Crawler"
description: "Crawl and sync websites automatically with scheduled recrawling and robots.txt compliance"
icon: "/icons/hugeicons/globe-02.svg"
---
Connect websites to automatically crawl and sync web pages into your Supermemory knowledge base. The web crawler respects robots.txt rules, includes SSRF protection, and automatically recrawls sites on a schedule.
The web crawler connector requires a **Scale Plan** or **Enterprise Plan**.
## Quick setup
### 1. Create Web Crawler Connector
```typescript
import { Supermemory } from "supermemory"
const supermemory = new Supermemory({ apiKey: process.env.SUPERMEMORY_API_KEY })
const connector = await supermemory.connectors.create("user-123", {
provider: "web-crawler",
config: {
startUrl: "https://docs.example.com",
crawlDepth: 3,
},
documentLimit: 5000,
})
// Web crawler doesn't require OAuth; crawling starts immediately
console.log("Connector ID:", connector.id)
// Note: connector.authorization is null for web-crawler
```
```python
from supermemory import Supermemory
import os
client = Supermemory(api_key=os.environ.get("SUPERMEMORY_API_KEY"))
connector = client.connectors.create(
"user-123",
request={
"provider": "web-crawler",
"config": {
"startUrl": "https://docs.example.com",
"crawlDepth": 3,
},
"documentLimit": 5000,
},
)
# Web crawler doesn't require OAuth; crawling starts immediately
print(f"Connector ID: {connector.id}")
# Note: connector.authorization.url is None for web-crawler
```
```bash
curl -X POST "https://api.supermemory.ai/ns/user-123/connectors" \
-H "Authorization: Bearer $SUPERMEMORY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"provider": "web-crawler",
"config": {
"startUrl": "https://docs.example.com",
"crawlDepth": 3
},
"documentLimit": 5000
}'
# Response: {
# "id": "PTzGiUYei7pgzg5buzZHgA",
# "authorization": null
# }
```
### Configuration options
| Parameter | Location | Required | Description |
|-----------|----------|----------|-------------|
| `startUrl` | `config.startUrl` | Yes | Public URL where the crawl starts |
| `crawlDepth` | `config.crawlDepth` | No | How many links deep to follow (1 to 5, default 3) |
| `documentLimit` | top-level | No | Maximum pages to sync (1 to 10,000) |
### 2. Connector Established
Unlike OAuth connectors, the web crawler doesn't require user authorization. `authorization` is `null`, the connector is visible immediately, and crawling begins automatically.
### 3. Monitor sync progress
```typescript
// Check connector details with recent sync runs
const connector = await supermemory.connectors.get("user-123", "PTzGiUYei7pgzg5buzZHgA", {
include: "syncs",
})
console.log("Start URL:", connector.config?.startUrl)
console.log("Last sync:", connector.latestRun?.system.status)
console.log("Pages synced:", connector.documentCount)
// List synced web pages in the namespace
const docs = await supermemory.list("user-123", "documents")
console.log(`Synced ${docs.pagination.totalItems} documents`)
```
```python
# Check connector details with recent sync runs
connector = client.connectors.get("user-123", "PTzGiUYei7pgzg5buzZHgA", include=["syncs"])
print(f"Start URL: {(connector.config or {}).get('startUrl')}")
print(f"Last sync: {connector.latest_run.system.status}")
print(f"Pages synced: {connector.document_count}")
# List synced web pages in the namespace
docs = client.list("user-123", "documents")
print(f"Synced {docs.pagination.total_items} documents")
```
```bash
# Get connector details with recent sync runs
curl "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA?include=syncs" \
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
# Response includes connector details:
# {
# "id": "PTzGiUYei7pgzg5buzZHgA",
# "provider": "web-crawler",
# "namespace": "user-123",
# "config": {"startUrl": "https://docs.example.com", "crawlDepth": 3},
# "documentLimit": 5000,
# "documentCount": 240,
# "latestRun": { "system": { "status": "completed", ... }, "error": null },
# "system": { "status": "active", "createdAt": "2024-01-15T10:00:00Z", "lastSuccessfulSyncAt": "..." }
# }
# List synced documents in the namespace
curl -X POST "https://api.supermemory.ai/ns/user-123/list/documents" \
-H "Authorization: Bearer $SUPERMEMORY_API_KEY" \
-H "Content-Type: application/json" \
-d '{}'
# Response: {"documents": [{"title": "Getting Started", "system": {"status": "done", ...}, ...}], "pagination": {...}}
```
There is no per-connector document list in v5. `supermemory.list(namespace, "documents")` returns every document in the namespace.
## Supported content types
### Web pages
- **HTML content** extracted and converted to markdown
- **Same-domain crawling** only (respects hostname boundaries)
- **Robots.txt compliance** - respects disallow rules
- **Content filtering** - only HTML pages (skips non-HTML content)
### URL requirements
The web crawler only processes valid public URLs:
- Must be a public URL (not localhost, private IPs, or internal domains)
- Must be accessible from the internet
- Must return HTML content (non-HTML files are skipped)
## Sync mechanism
The web crawler uses **scheduled recrawling** rather than real-time webhooks:
- **Initial Crawl**: Begins immediately after connector creation
- **Scheduled Recrawling**: Automatically recrawls sites that haven't been synced in 7+ days
- **No Real-time Updates**: Unlike other connectors, web crawler doesn't support webhook-based real-time sync
The recrawl schedule is automatically assigned when the connector is created. Sites are recrawled periodically to keep content up to date, but updates are not instantaneous.
### Manual Recrawl
Start a crawl now. The call returns `409` while a sync for that connector is already running.
```typescript
const run = await supermemory.connectors.sync("user-123", "PTzGiUYei7pgzg5buzZHgA")
console.log(run.status)
// Output: queued
```
```bash
curl -X POST "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA/sync" \
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
# Response: {"id": "PTzGiUYei7pgzg5buzZHgA", "status": "queued"}
```
## Connector Management
### List All Connectors
```typescript
// List all web crawler connectors in a namespace
const { connectors } = await supermemory.connectors.list("user-123")
const webCrawlerConnectors = connectors.filter(
connector => connector.provider === "web-crawler"
)
webCrawlerConnectors.forEach(connector => {
console.log(`Start URL: ${connector.config?.startUrl}`)
console.log(`Connector ID: ${connector.id}`)
console.log(`Created: ${connector.createdAt}`)
})
```
```python
# List all web crawler connectors in a namespace
connectors = client.connectors.list("user-123", provider="web-crawler").connectors
for connector in connectors:
print(f"Start URL: {(connector.config or {}).get('startUrl')}")
print(f"Connector ID: {connector.id}")
print(f"Created: {connector.created_at}")
```
```bash
# List all connectors in a namespace
curl "https://api.supermemory.ai/ns/user-123/connectors" \
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
# Response: {
# "connectors": [
# {
# "id": "PTzGiUYei7pgzg5buzZHgA",
# "provider": "web-crawler",
# "namespace": "user-123",
# "createdAt": "2024-01-15T10:30:00.000Z",
# "documentLimit": 5000,
# "config": {"startUrl": "https://docs.example.com", "crawlDepth": 3}
# }
# ],
# "pagination": { "currentPage": 1, "limit": 50, "totalItems": 1, "totalPages": 1 }
# }
```
### Delete Connector
Remove a web crawler connector when no longer needed:
```typescript
// Delete the connector and its crawled pages (default)
await supermemory.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA")
// Delete the connector but keep the crawled pages
await supermemory.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA", {
deleteDocuments: false,
})
```
```python
# Delete the connector and its crawled pages (default)
client.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA")
# Delete the connector but keep the crawled pages
client.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA", delete_documents=False)
```
```bash
# Delete the connector and its crawled pages (default)
curl -X DELETE "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA" \
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
# Delete the connector but keep the crawled pages
curl -X DELETE "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA?deleteDocuments=false" \
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
```
Deleting a connector will:
- Stop all future crawls of the website
- Remove the connector configuration
- Delete the crawled pages unless you pass `deleteDocuments: false`
## Security & compliance
### SSRF protection
Built-in protection against Server-Side Request Forgery (SSRF) attacks:
- Blocks private IP addresses (10.x.x.x, 192.168.x.x, 172.16-31.x.x)
- Blocks localhost and internal domains
- Blocks cloud metadata endpoints
- Only allows public, internet-accessible URLs
### URL validation
All URLs are validated before crawling:
- Must be valid HTTP/HTTPS URLs
- Must be publicly accessible
- Must return HTML content
- Response size limited to 10MB
**Important Limitations:**
- Requires Scale Plan or Enterprise Plan
- Only crawls same-domain URLs
- Scheduled recrawling means updates are not real-time
- Large websites may take significant time to crawl initially
- Robots.txt restrictions may prevent crawling some pages
- URLs must be publicly accessible (no authentication required)