mirror of
https://github.com/supermemoryai/supermemory.git
synced 2026-10-11 03:37:56 +00:00
Rewrites 339 TypeScript calls across 50 pages from the rc.5 `method({ namespace, body })` form to the shipped `method(namespace, { ... })` form, and aligns field names with the live v5 spec: `attach` to `include`, `authUrl` to `authorization`, `lastSync` to `latestRun`, `deletedCount` to `count`, and the paginated `namespaces.list()`.
Renames container tags to namespaces across concepts, connectors, integrations and snippets. The namespace pages keep container tag in the description, search keywords and a rename note so old searches still land, and the v3 reference page points at v5.
The migration guide's SDK table now covers both 5.0.0 SDKs, and the SDK integration page uses the real client options (`baseUrl`, `timeoutInSeconds`, `maxRetries`) and error classes.
342 lines
11 KiB
Text
342 lines
11 KiB
Text
---
|
|
title: "Web Crawler connector"
|
|
sidebarTitle: "Web Crawler"
|
|
description: "Crawl and sync websites automatically with scheduled recrawling and robots.txt compliance"
|
|
icon: "/icons/hugeicons/globe-02.svg"
|
|
---
|
|
|
|
Connect websites to automatically crawl and sync web pages into your Supermemory knowledge base. The web crawler respects robots.txt rules, includes SSRF protection, and automatically recrawls sites on a schedule.
|
|
|
|
<Warning>
|
|
The web crawler connector requires a **Scale Plan** or **Enterprise Plan**.
|
|
</Warning>
|
|
|
|
## Quick setup
|
|
|
|
### 1. Create Web Crawler Connector
|
|
|
|
<Tabs>
|
|
<Tab title="TypeScript">
|
|
```typescript
|
|
import { Supermemory } from "supermemory"
|
|
|
|
const supermemory = new Supermemory({ apiKey: process.env.SUPERMEMORY_API_KEY })
|
|
|
|
const connector = await supermemory.connectors.create("user-123", {
|
|
provider: "web-crawler",
|
|
config: {
|
|
startUrl: "https://docs.example.com",
|
|
crawlDepth: 3,
|
|
},
|
|
documentLimit: 5000,
|
|
})
|
|
|
|
// Web crawler doesn't require OAuth; crawling starts immediately
|
|
console.log("Connector ID:", connector.id)
|
|
// Note: connector.authorization is null for web-crawler
|
|
```
|
|
</Tab>
|
|
<Tab title="Python">
|
|
```python
|
|
from supermemory import Supermemory
|
|
import os
|
|
|
|
client = Supermemory(api_key=os.environ.get("SUPERMEMORY_API_KEY"))
|
|
|
|
connector = client.connectors.create(
|
|
"user-123",
|
|
request={
|
|
"provider": "web-crawler",
|
|
"config": {
|
|
"startUrl": "https://docs.example.com",
|
|
"crawlDepth": 3,
|
|
},
|
|
"documentLimit": 5000,
|
|
},
|
|
)
|
|
|
|
# Web crawler doesn't require OAuth; crawling starts immediately
|
|
print(f"Connector ID: {connector.id}")
|
|
# Note: connector.authorization.url is None for web-crawler
|
|
```
|
|
</Tab>
|
|
<Tab title="cURL">
|
|
```bash
|
|
curl -X POST "https://api.supermemory.ai/ns/user-123/connectors" \
|
|
-H "Authorization: Bearer $SUPERMEMORY_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"provider": "web-crawler",
|
|
"config": {
|
|
"startUrl": "https://docs.example.com",
|
|
"crawlDepth": 3
|
|
},
|
|
"documentLimit": 5000
|
|
}'
|
|
|
|
# Response: {
|
|
# "id": "PTzGiUYei7pgzg5buzZHgA",
|
|
# "authorization": null
|
|
# }
|
|
```
|
|
</Tab>
|
|
</Tabs>
|
|
|
|
### Configuration options
|
|
|
|
| Parameter | Location | Required | Description |
|
|
|-----------|----------|----------|-------------|
|
|
| `startUrl` | `config.startUrl` | Yes | Public URL where the crawl starts |
|
|
| `crawlDepth` | `config.crawlDepth` | No | How many links deep to follow (1 to 5, default 3) |
|
|
| `documentLimit` | top-level | No | Maximum pages to sync (1 to 10,000) |
|
|
|
|
### 2. Connector Established
|
|
|
|
Unlike OAuth connectors, the web crawler doesn't require user authorization. `authorization` is `null`, the connector is visible immediately, and crawling begins automatically.
|
|
|
|
### 3. Monitor sync progress
|
|
|
|
<Tabs>
|
|
<Tab title="TypeScript">
|
|
```typescript
|
|
// Check connector details with recent sync runs
|
|
const connector = await supermemory.connectors.get("user-123", "PTzGiUYei7pgzg5buzZHgA", {
|
|
include: "syncs",
|
|
})
|
|
|
|
console.log("Start URL:", connector.config?.startUrl)
|
|
console.log("Last sync:", connector.latestRun?.system.status)
|
|
console.log("Pages synced:", connector.documentCount)
|
|
|
|
// List synced web pages in the namespace
|
|
const docs = await supermemory.list("user-123", "documents")
|
|
|
|
console.log(`Synced ${docs.pagination.totalItems} documents`)
|
|
```
|
|
</Tab>
|
|
<Tab title="Python">
|
|
```python
|
|
# Check connector details with recent sync runs
|
|
connector = client.connectors.get("user-123", "PTzGiUYei7pgzg5buzZHgA", include=["syncs"])
|
|
|
|
print(f"Start URL: {(connector.config or {}).get('startUrl')}")
|
|
print(f"Last sync: {connector.latest_run.system.status}")
|
|
print(f"Pages synced: {connector.document_count}")
|
|
|
|
# List synced web pages in the namespace
|
|
docs = client.list("user-123", "documents")
|
|
|
|
print(f"Synced {docs.pagination.total_items} documents")
|
|
```
|
|
</Tab>
|
|
<Tab title="cURL">
|
|
```bash
|
|
# Get connector details with recent sync runs
|
|
curl "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA?include=syncs" \
|
|
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
|
|
|
|
# Response includes connector details:
|
|
# {
|
|
# "id": "PTzGiUYei7pgzg5buzZHgA",
|
|
# "provider": "web-crawler",
|
|
# "namespace": "user-123",
|
|
# "config": {"startUrl": "https://docs.example.com", "crawlDepth": 3},
|
|
# "documentLimit": 5000,
|
|
# "documentCount": 240,
|
|
# "latestRun": { "system": { "status": "completed", ... }, "error": null },
|
|
# "system": { "status": "active", "createdAt": "2024-01-15T10:00:00Z", "lastSuccessfulSyncAt": "..." }
|
|
# }
|
|
|
|
# List synced documents in the namespace
|
|
curl -X POST "https://api.supermemory.ai/ns/user-123/list/documents" \
|
|
-H "Authorization: Bearer $SUPERMEMORY_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
-d '{}'
|
|
|
|
# Response: {"documents": [{"title": "Getting Started", "system": {"status": "done", ...}, ...}], "pagination": {...}}
|
|
```
|
|
</Tab>
|
|
</Tabs>
|
|
|
|
<Note>
|
|
There is no per-connector document list in v5. `supermemory.list(namespace, "documents")` returns every document in the namespace.
|
|
</Note>
|
|
|
|
## Supported content types
|
|
|
|
### Web pages
|
|
- **HTML content** extracted and converted to markdown
|
|
- **Same-domain crawling** only (respects hostname boundaries)
|
|
- **Robots.txt compliance** - respects disallow rules
|
|
- **Content filtering** - only HTML pages (skips non-HTML content)
|
|
|
|
### URL requirements
|
|
|
|
The web crawler only processes valid public URLs:
|
|
- Must be a public URL (not localhost, private IPs, or internal domains)
|
|
- Must be accessible from the internet
|
|
- Must return HTML content (non-HTML files are skipped)
|
|
|
|
## Sync mechanism
|
|
|
|
The web crawler uses **scheduled recrawling** rather than real-time webhooks:
|
|
|
|
- **Initial Crawl**: Begins immediately after connector creation
|
|
- **Scheduled Recrawling**: Automatically recrawls sites that haven't been synced in 7+ days
|
|
- **No Real-time Updates**: Unlike other connectors, web crawler doesn't support webhook-based real-time sync
|
|
|
|
<Note>
|
|
The recrawl schedule is automatically assigned when the connector is created. Sites are recrawled periodically to keep content up to date, but updates are not instantaneous.
|
|
</Note>
|
|
|
|
### Manual Recrawl
|
|
|
|
Start a crawl now. The call returns `409` while a sync for that connector is already running.
|
|
|
|
<Tabs>
|
|
<Tab title="TypeScript">
|
|
```typescript
|
|
const run = await supermemory.connectors.sync("user-123", "PTzGiUYei7pgzg5buzZHgA")
|
|
|
|
console.log(run.status)
|
|
// Output: queued
|
|
```
|
|
</Tab>
|
|
<Tab title="cURL">
|
|
```bash
|
|
curl -X POST "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA/sync" \
|
|
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
|
|
|
|
# Response: {"id": "PTzGiUYei7pgzg5buzZHgA", "status": "queued"}
|
|
```
|
|
</Tab>
|
|
</Tabs>
|
|
|
|
## Connector Management
|
|
|
|
### List All Connectors
|
|
|
|
<Tabs>
|
|
<Tab title="TypeScript">
|
|
```typescript
|
|
// List all web crawler connectors in a namespace
|
|
const { connectors } = await supermemory.connectors.list("user-123")
|
|
|
|
const webCrawlerConnectors = connectors.filter(
|
|
connector => connector.provider === "web-crawler"
|
|
)
|
|
|
|
webCrawlerConnectors.forEach(connector => {
|
|
console.log(`Start URL: ${connector.config?.startUrl}`)
|
|
console.log(`Connector ID: ${connector.id}`)
|
|
console.log(`Created: ${connector.createdAt}`)
|
|
})
|
|
```
|
|
</Tab>
|
|
<Tab title="Python">
|
|
```python
|
|
# List all web crawler connectors in a namespace
|
|
connectors = client.connectors.list("user-123", provider="web-crawler").connectors
|
|
|
|
for connector in connectors:
|
|
print(f"Start URL: {(connector.config or {}).get('startUrl')}")
|
|
print(f"Connector ID: {connector.id}")
|
|
print(f"Created: {connector.created_at}")
|
|
```
|
|
</Tab>
|
|
<Tab title="cURL">
|
|
```bash
|
|
# List all connectors in a namespace
|
|
curl "https://api.supermemory.ai/ns/user-123/connectors" \
|
|
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
|
|
|
|
# Response: {
|
|
# "connectors": [
|
|
# {
|
|
# "id": "PTzGiUYei7pgzg5buzZHgA",
|
|
# "provider": "web-crawler",
|
|
# "namespace": "user-123",
|
|
# "createdAt": "2024-01-15T10:30:00.000Z",
|
|
# "documentLimit": 5000,
|
|
# "config": {"startUrl": "https://docs.example.com", "crawlDepth": 3}
|
|
# }
|
|
# ],
|
|
# "pagination": { "currentPage": 1, "limit": 50, "totalItems": 1, "totalPages": 1 }
|
|
# }
|
|
```
|
|
</Tab>
|
|
</Tabs>
|
|
|
|
### Delete Connector
|
|
|
|
Remove a web crawler connector when no longer needed:
|
|
|
|
<Tabs>
|
|
<Tab title="TypeScript">
|
|
```typescript
|
|
// Delete the connector and its crawled pages (default)
|
|
await supermemory.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA")
|
|
|
|
// Delete the connector but keep the crawled pages
|
|
await supermemory.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA", {
|
|
deleteDocuments: false,
|
|
})
|
|
```
|
|
</Tab>
|
|
<Tab title="Python">
|
|
```python
|
|
# Delete the connector and its crawled pages (default)
|
|
client.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA")
|
|
|
|
# Delete the connector but keep the crawled pages
|
|
client.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA", delete_documents=False)
|
|
```
|
|
</Tab>
|
|
<Tab title="cURL">
|
|
```bash
|
|
# Delete the connector and its crawled pages (default)
|
|
curl -X DELETE "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA" \
|
|
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
|
|
|
|
# Delete the connector but keep the crawled pages
|
|
curl -X DELETE "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA?deleteDocuments=false" \
|
|
-H "Authorization: Bearer $SUPERMEMORY_API_KEY"
|
|
```
|
|
</Tab>
|
|
</Tabs>
|
|
|
|
<Note>
|
|
Deleting a connector will:
|
|
- Stop all future crawls of the website
|
|
- Remove the connector configuration
|
|
- Delete the crawled pages unless you pass `deleteDocuments: false`
|
|
</Note>
|
|
|
|
## Security & compliance
|
|
|
|
### SSRF protection
|
|
|
|
Built-in protection against Server-Side Request Forgery (SSRF) attacks:
|
|
- Blocks private IP addresses (10.x.x.x, 192.168.x.x, 172.16-31.x.x)
|
|
- Blocks localhost and internal domains
|
|
- Blocks cloud metadata endpoints
|
|
- Only allows public, internet-accessible URLs
|
|
|
|
### URL validation
|
|
|
|
All URLs are validated before crawling:
|
|
- Must be valid HTTP/HTTPS URLs
|
|
- Must be publicly accessible
|
|
- Must return HTML content
|
|
- Response size limited to 10MB
|
|
|
|
|
|
<Warning>
|
|
**Important Limitations:**
|
|
- Requires Scale Plan or Enterprise Plan
|
|
- Only crawls same-domain URLs
|
|
- Scheduled recrawling means updates are not real-time
|
|
- Large websites may take significant time to crawl initially
|
|
- Robots.txt restrictions may prevent crawling some pages
|
|
- URLs must be publicly accessible (no authentication required)
|
|
</Warning>
|