API reference
API documentation for the web scraping API
A single REST API to crawl a page, render its JavaScript and extract clean markdown or typed JSON. Authenticate with a bearer key, call one of three endpoints, and get LLM-ready data back. Here is everything you need to make your first request.
Introduction
The ClawEngine API is organized around predictable, resource-oriented REST. Every request is sent over HTTPS, accepts and returns JSON, and authenticates with a bearer key. The base URL for all endpoints is:
https://clawengine.ai/v1
Authentication
Authenticate every request by passing your secret key in the Authorization header as a bearer token. Keep your key secret, on a server and out of client-side code. You get a key when you create an account.
Authorization: Bearer $CLAWENGINE_API_KEY
POST /v1/extract
Extract a single URL. ClawEngine fetches the page, renders its JavaScript, strips boilerplate, and returns clean markdown, structured JSON, or fields typed to a schema you provide.
curl https://clawengine.ai/v1/extract \
-H "Authorization: Bearer $CLAWENGINE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products/atlas",
"format": "json"
}'
import requests
r = requests.post(
"https://clawengine.ai/v1/extract",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"url": "https://example.com/products/atlas",
"format": "json"},
)
data = r.json()
const res = await fetch("https://clawengine.ai/v1/extract", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.CLAWENGINE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ url: "https://example.com/products/atlas", format: "json" }),
});
const data = await res.json();
Response · 200 OK
{
"url": "https://example.com/products/atlas",
"title": "Atlas Field Notebook",
"markdown": "# Atlas Field Notebook\n\nDurable...",
"data": { "name": "Atlas Field Notebook", "price": 24.00 },
"links": [ "/products", "/cart" ],
"metadata": { "rendered": true, "robotsTxt": "respected" }
}
Body parameters
| Field | Type | Description |
|---|---|---|
| url | string | The public page to extract. Required. |
| format | string | One of markdown, json or structured. Defaults to markdown. |
| schema | object | Optional. A field-to-type map for structured extraction. |
| render | boolean | Render JavaScript before extraction. Defaults to true. |
POST /v1/crawl
Crawl a whole site. ClawEngine starts at the URL you give, follows the links you allow, and extracts every page. Crawls run asynchronously: this endpoint returns a crawl id you poll for results, or you can register a webhook to receive pages as they complete.
curl https://clawengine.ai/v1/crawl \
-H "Authorization: Bearer $CLAWENGINE_API_KEY" \
-d '{
"url": "https://example.com/docs",
"limit": 500,
"format": "markdown",
"webhook": "https://yourapp.com/hooks/clawengine"
}'
Response · 202 Accepted
{
"id": "crawl_8f2a1c9e",
"status": "queued",
"url": "https://example.com/docs"
}
Body parameters
| Field | Type | Description |
|---|---|---|
| url | string | The starting page to crawl. Required. |
| limit | integer | Maximum number of pages to crawl. Optional. |
| format | string | One of markdown, json or structured. Defaults to markdown. |
| webhook | string | Optional public URL. ClawEngine posts the finished crawl there, signed. Startup and Scale. |
| schema | object | Optional field-to-type map, used with format: "structured". Startup and Scale. |
| render | boolean | Render JavaScript on every page before extraction. Defaults to true. |
GET /v1/crawl/{id}
Check the status of a crawl and collect its pages. Poll this endpoint with the id from the crawl request until status is completed. Each page comes back in the same shape as a single extract.
curl https://clawengine.ai/v1/crawl/crawl_8f2a1c9e \
-H "Authorization: Bearer $CLAWENGINE_API_KEY"
Response · 200 OK
{
"id": "crawl_8f2a1c9e",
"status": "completed",
"total": 128,
"pages": [
{ "url": "https://example.com/docs/quickstart",
"title": "Quickstart",
"markdown": "# Quickstart..." }
]
}
Structured extraction
Send format: "structured" with a schema and ClawEngine returns a fields object typed to your declaration. The schema is a field-to-type map: string, number, integer, boolean, date and url, plus nested objects and one-entry lists for arrays.
A field the page does not state comes back null, and a value that does not match the type you declared comes back null too. Nothing is guessed into shape, so a value you receive is a value the page carried. Structured extraction is included in the Startup and Scale plans.
{
"url": "https://example.com/products/atlas",
"format": "structured",
"schema": {
"name": "string",
"price": "number",
"inStock": "boolean",
"tags": ["string"],
"manufacturer": { "name": "string", "country": "string" }
}
}
Verifying a webhook
Every delivery carries two headers: ClawEngine-Event, which is crawl.completed, and ClawEngine-Signature, in the form t=<timestamp>,v1=<hmac>. The HMAC is SHA-256 over <timestamp>.<raw body>, keyed with your API key. Compare it in constant time and reject a timestamp that is not recent.
import hashlib, hmac, time
def verify(raw_body: bytes, header: str, api_key: str) -> bool:
parts = dict(p.split("=", 1) for p in header.split(","))
if abs(time.time() - int(parts["t"])) > 300:
return False
expected = hmac.new(
api_key.encode(),
f'{parts["t"]}.'.encode() + raw_body,
hashlib.sha256,
).hexdigest()
return hmac.compare_digest(expected, parts["v1"])
REST API, any language
Call the REST API with curl, Python's requests or Node's fetch, no client library to install. Typed JSON responses in, typed JSON responses out, so it drops straight into whatever HTTP client your stack already uses.
curl https://clawengine.ai/v1/extract \
-H "Authorization: Bearer $CLAWENGINE_API_KEY"
Webhooks
Register a webhook URL on a crawl and ClawEngine posts the finished crawl to your endpoint, so you never hold a long connection open or poll. Each delivery is signed, so you can verify it came from us before you ingest the data.
POST /hooks/clawengine
X-Clawengine-Signature: t=...,v1=...
Compliance by default. Every request reads robots.txt first, honors crawl-delay, and respects site Terms of Service. ClawEngine is for public, permitted data only. You are responsible for the URLs you crawl. Read the compliance policy.
Make your first API call
Get a key, send a URL, and get LLM-ready markdown or JSON back from one request. Public, permitted data only.