--- title: "Web Crawler connector" sidebarTitle: "Web Crawler" description: "Crawl and sync websites automatically with scheduled recrawling and robots.txt compliance" icon: "/icons/hugeicons/globe-02.svg" --- Connect websites to automatically crawl and sync web pages into your Supermemory knowledge base. The web crawler respects robots.txt rules, includes SSRF protection, and automatically recrawls sites on a schedule. The web crawler connector requires a **Scale Plan** or **Enterprise Plan**. ## Quick setup ### 1. Create Web Crawler Connector ```typescript import { Supermemory } from "supermemory" const supermemory = new Supermemory({ apiKey: process.env.SUPERMEMORY_API_KEY }) const connector = await supermemory.connectors.create("user-123", { provider: "web-crawler", config: { startUrl: "https://docs.example.com", crawlDepth: 3, }, documentLimit: 5000, }) // Web crawler doesn't require OAuth; crawling starts immediately console.log("Connector ID:", connector.id) // Note: connector.authorization is null for web-crawler ``` ```python from supermemory import Supermemory import os client = Supermemory(api_key=os.environ.get("SUPERMEMORY_API_KEY")) connector = client.connectors.create( "user-123", request={ "provider": "web-crawler", "config": { "startUrl": "https://docs.example.com", "crawlDepth": 3, }, "documentLimit": 5000, }, ) # Web crawler doesn't require OAuth; crawling starts immediately print(f"Connector ID: {connector.id}") # Note: connector.authorization.url is None for web-crawler ``` ```bash curl -X POST "https://api.supermemory.ai/ns/user-123/connectors" \ -H "Authorization: Bearer $SUPERMEMORY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "provider": "web-crawler", "config": { "startUrl": "https://docs.example.com", "crawlDepth": 3 }, "documentLimit": 5000 }' # Response: { # "id": "PTzGiUYei7pgzg5buzZHgA", # "authorization": null # } ``` ### Configuration options | Parameter | Location | Required | Description | |-----------|----------|----------|-------------| | `startUrl` | `config.startUrl` | Yes | Public URL where the crawl starts | | `crawlDepth` | `config.crawlDepth` | No | How many links deep to follow (1 to 5, default 3) | | `documentLimit` | top-level | No | Maximum pages to sync (1 to 10,000) | ### 2. Connector Established Unlike OAuth connectors, the web crawler doesn't require user authorization. `authorization` is `null`, the connector is visible immediately, and crawling begins automatically. ### 3. Monitor sync progress ```typescript // Check connector details with recent sync runs const connector = await supermemory.connectors.get("user-123", "PTzGiUYei7pgzg5buzZHgA", { include: "syncs", }) console.log("Start URL:", connector.config?.startUrl) console.log("Last sync:", connector.latestRun?.system.status) console.log("Pages synced:", connector.documentCount) // List synced web pages in the namespace const docs = await supermemory.list("user-123", "documents") console.log(`Synced ${docs.pagination.totalItems} documents`) ``` ```python # Check connector details with recent sync runs connector = client.connectors.get("user-123", "PTzGiUYei7pgzg5buzZHgA", include=["syncs"]) print(f"Start URL: {(connector.config or {}).get('startUrl')}") print(f"Last sync: {connector.latest_run.system.status}") print(f"Pages synced: {connector.document_count}") # List synced web pages in the namespace docs = client.list("user-123", "documents") print(f"Synced {docs.pagination.total_items} documents") ``` ```bash # Get connector details with recent sync runs curl "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA?include=syncs" \ -H "Authorization: Bearer $SUPERMEMORY_API_KEY" # Response includes connector details: # { # "id": "PTzGiUYei7pgzg5buzZHgA", # "provider": "web-crawler", # "namespace": "user-123", # "config": {"startUrl": "https://docs.example.com", "crawlDepth": 3}, # "documentLimit": 5000, # "documentCount": 240, # "latestRun": { "system": { "status": "completed", ... }, "error": null }, # "system": { "status": "active", "createdAt": "2024-01-15T10:00:00Z", "lastSuccessfulSyncAt": "..." } # } # List synced documents in the namespace curl -X POST "https://api.supermemory.ai/ns/user-123/list/documents" \ -H "Authorization: Bearer $SUPERMEMORY_API_KEY" \ -H "Content-Type: application/json" \ -d '{}' # Response: {"documents": [{"title": "Getting Started", "system": {"status": "done", ...}, ...}], "pagination": {...}} ``` There is no per-connector document list in v5. `supermemory.list(namespace, "documents")` returns every document in the namespace. ## Supported content types ### Web pages - **HTML content** extracted and converted to markdown - **Same-domain crawling** only (respects hostname boundaries) - **Robots.txt compliance** - respects disallow rules - **Content filtering** - only HTML pages (skips non-HTML content) ### URL requirements The web crawler only processes valid public URLs: - Must be a public URL (not localhost, private IPs, or internal domains) - Must be accessible from the internet - Must return HTML content (non-HTML files are skipped) ## Sync mechanism The web crawler uses **scheduled recrawling** rather than real-time webhooks: - **Initial Crawl**: Begins immediately after connector creation - **Scheduled Recrawling**: Automatically recrawls sites that haven't been synced in 7+ days - **No Real-time Updates**: Unlike other connectors, web crawler doesn't support webhook-based real-time sync The recrawl schedule is automatically assigned when the connector is created. Sites are recrawled periodically to keep content up to date, but updates are not instantaneous. ### Manual Recrawl Start a crawl now. The call returns `409` while a sync for that connector is already running. ```typescript const run = await supermemory.connectors.sync("user-123", "PTzGiUYei7pgzg5buzZHgA") console.log(run.status) // Output: queued ``` ```bash curl -X POST "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA/sync" \ -H "Authorization: Bearer $SUPERMEMORY_API_KEY" # Response: {"id": "PTzGiUYei7pgzg5buzZHgA", "status": "queued"} ``` ## Connector Management ### List All Connectors ```typescript // List all web crawler connectors in a namespace const { connectors } = await supermemory.connectors.list("user-123") const webCrawlerConnectors = connectors.filter( connector => connector.provider === "web-crawler" ) webCrawlerConnectors.forEach(connector => { console.log(`Start URL: ${connector.config?.startUrl}`) console.log(`Connector ID: ${connector.id}`) console.log(`Created: ${connector.createdAt}`) }) ``` ```python # List all web crawler connectors in a namespace connectors = client.connectors.list("user-123", provider="web-crawler").connectors for connector in connectors: print(f"Start URL: {(connector.config or {}).get('startUrl')}") print(f"Connector ID: {connector.id}") print(f"Created: {connector.created_at}") ``` ```bash # List all connectors in a namespace curl "https://api.supermemory.ai/ns/user-123/connectors" \ -H "Authorization: Bearer $SUPERMEMORY_API_KEY" # Response: { # "connectors": [ # { # "id": "PTzGiUYei7pgzg5buzZHgA", # "provider": "web-crawler", # "namespace": "user-123", # "createdAt": "2024-01-15T10:30:00.000Z", # "documentLimit": 5000, # "config": {"startUrl": "https://docs.example.com", "crawlDepth": 3} # } # ], # "pagination": { "currentPage": 1, "limit": 50, "totalItems": 1, "totalPages": 1 } # } ``` ### Delete Connector Remove a web crawler connector when no longer needed: ```typescript // Delete the connector and its crawled pages (default) await supermemory.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA") // Delete the connector but keep the crawled pages await supermemory.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA", { deleteDocuments: false, }) ``` ```python # Delete the connector and its crawled pages (default) client.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA") # Delete the connector but keep the crawled pages client.connectors.delete("user-123", "PTzGiUYei7pgzg5buzZHgA", delete_documents=False) ``` ```bash # Delete the connector and its crawled pages (default) curl -X DELETE "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA" \ -H "Authorization: Bearer $SUPERMEMORY_API_KEY" # Delete the connector but keep the crawled pages curl -X DELETE "https://api.supermemory.ai/ns/user-123/connectors/PTzGiUYei7pgzg5buzZHgA?deleteDocuments=false" \ -H "Authorization: Bearer $SUPERMEMORY_API_KEY" ``` Deleting a connector will: - Stop all future crawls of the website - Remove the connector configuration - Delete the crawled pages unless you pass `deleteDocuments: false` ## Security & compliance ### SSRF protection Built-in protection against Server-Side Request Forgery (SSRF) attacks: - Blocks private IP addresses (10.x.x.x, 192.168.x.x, 172.16-31.x.x) - Blocks localhost and internal domains - Blocks cloud metadata endpoints - Only allows public, internet-accessible URLs ### URL validation All URLs are validated before crawling: - Must be valid HTTP/HTTPS URLs - Must be publicly accessible - Must return HTML content - Response size limited to 10MB **Important Limitations:** - Requires Scale Plan or Enterprise Plan - Only crawls same-domain URLs - Scheduled recrawling means updates are not real-time - Large websites may take significant time to crawl initially - Robots.txt restrictions may prevent crawling some pages - URLs must be publicly accessible (no authentication required)