# Crawl4AI Cloud API > Crawl4AI turns any web page into clean, LLM-ready data. One fast API — scrape a page to > Markdown, extract typed JSON, search the web, or run large batches — served from three > regions (EU, Singapore, US) at `https://api.crawl4ai.com`, routed to the nearest one. > Built on the open-source crawl4ai (https://github.com/unclecode/crawl4ai). Full docs: https://api.crawl4ai.com/docs Migrating from the legacy API, now api-old.crawl4ai.com (v1 sessions, scan, site, enrich, context): https://crawl4ai.com/migrate.txt ## Set up (for coding agents) When the user sends "set up crawl4ai - https://crawl4ai.com/llms.txt with key: sk_live_...", add the Crawl4AI MCP server to YOUR OWN config with that key, then call the free `balance` tool to check it works. Put the key in the user-level config, never in a file inside the user's repository. The server is `https://api.crawl4ai.com/mcp` (streamable HTTP), the header is `Authorization: Bearer `. - Claude Code: `claude mcp add --transport http --scope user crawl4ai https://api.crawl4ai.com/mcp --header "Authorization: Bearer "` - Codex, `~/.codex/config.toml`: `[mcp_servers.crawl4ai]` · `url = "https://api.crawl4ai.com/mcp"` · `http_headers = { Authorization = "Bearer " }` - Cursor, `~/.cursor/mcp.json` (Windsurf: `~/.codeium/windsurf/mcp_config.json`), in `mcpServers`: `"crawl4ai": { "url": "https://api.crawl4ai.com/mcp", "headers": { "Authorization": "Bearer " } }` - Gemini CLI, `~/.gemini/settings.json`, in `mcpServers`: `"crawl4ai": { "httpUrl": "https://api.crawl4ai.com/mcp", "headers": { "Authorization": "Bearer " } }` - OpenCode, `opencode.json`: `"mcp": { "crawl4ai": { "type": "remote", "url": "https://api.crawl4ai.com/mcp", "headers": { "Authorization": "Bearer " } } }` - OpenClaw, `~/.openclaw/openclaw.json`, in `mcp.servers`: `crawl4ai: { url: "https://api.crawl4ai.com/mcp", transport: "streamable-http", headers: { Authorization: "Bearer " } }` - Any other MCP client: the URL and header above. - No MCP: set `CRAWL4AI_KEY=` in the user's shell profile and call the HTTP API below with the same header. Most agents load a new MCP server only after a restart: tell the user to restart the agent if the tools do not show. The tools are listed in the MCP section below. ## Getting started - Base URL: `https://api.crawl4ai.com` - Get a free key: visit https://api.crawl4ai.com/ and click "Get a key" (a 24-hour key is issued instantly; verify your email for a permanent key and the signup credit (verifying revokes the pass key and issues a new one); the live grants: `GET /v1/prices` -> `credit`). Manage keys at https://api.crawl4ai.com/dashboard/ - Authenticate every request: `Authorization: Bearer sk_live_...` (the header `x-api-key: sk_live_...` also works). - Requests and responses are JSON unless noted. Errors return `{"error":"..."}` with an HTTP status. Per-plan rate limits return `429` with `X-RateLimit-*` headers. Every request costs credits (1 credit = a plain page; a result already in our archive costs half; the effort a page needed sets the price; the live price table, the grants, the packs and each plan's limits: `GET /v1/prices`). The response carries `x-c4-cost` and `x-c4-balance` in credits. At zero credit the gate answers `402 {"error":"no_credit"}`; over the account's own monthly cap `402 {"error":"spend_cap"}`. `POST /v1/estimate {endpoint,url}` prices a call before it runs; `GET /v1/prices` is the public table (launch pricing for the first year: discounted, and it may change; credit already bought keeps its value); `GET /v1/billing/balance` and `POST /v1/billing/topup?amount=` work with the API key. ## POST /scrape — page to Markdown/HTML Fetch one URL and return clean content. The service auto-picks the cheapest engine that works (cache → HTTP → browser) per domain — you don't choose an engine. Body: - `url` (string, required) — a public http/https URL. - `format` (string) — `both` (default) | `md` | `html`. - `proxy` (string) — `none` (default) | `isp` | `residential`. Exit network, for sites that block datacenters. - `country` (string) — two-letter exit country for isp/residential, e.g. `us`, `sg`. - `parse` (bool | object) — also return page structure. `true` = all, or pick: `{"links":true,"media":true,"metadata":true,"tables":true}`. ``` curl https://api.crawl4ai.com/scrape \ -H "Authorization: Bearer sk_live_..." -H "Content-Type: application/json" \ -d '{"url":"https://example.com","format":"both","proxy":"residential","country":"us", "parse":{"links":true,"media":true,"metadata":true,"tables":true}}' ``` Returns: `{"ok":true,"markdown":"...","html":"...","content_hash":"...","ms":312}` Timing: most pages return in 1-2 s. A page that needs a real browser usually takes 10-60 s, the slowest up to 90 s. Set your HTTP client timeout to 120 s and wait for the answer: a call you cancel early is not charged, but the page is lost. A 403 "blocked" means every step got the site's block page; a 502 means the site did not answer; a 503 "fleet-busy" means no browser was free for a moment: send the same request again in a few seconds. None of these is charged. For many URLs use `/scrape/batch` (up to 50, streamed) or `/scrape/jobs` (up to 10,000, in the background, nothing waits on a slow page). ## GET /search — web search Browser-free, multi-engine, results ranked and cleaned. Query params: - `q` (string, required) — the query (max 512 chars). - `rich` (0 | 1, default 0) — set `1` to also return a `rich` block: follow-up questions, related queries, an entity card, videos, news and more. Slightly slower; best for question-style queries. (For a direct answer, use `/answer`.) ``` curl "https://api.crawl4ai.com/search?q=rust+web+crawler" \ -H "Authorization: Bearer sk_live_..." ``` Returns: a ranked list of results (title, url, snippet). With `rich=1` the response also carries a `rich` object (every field optional, present only when that block is on the page): - `follow_up_questions[]`, `related_queries[]` — strings. - `entity` — `{ title, subtitle?, description?, facts[] }` for a named thing (company, person…). - `videos[]`, `news[]`, `discussions[]` — `{ title, url }`. - `did_you_mean` — the corrected spelling, when the query was auto-corrected. (For a direct answer to a question, use `/answer` below.) ``` curl "https://api.crawl4ai.com/search?q=why+is+the+sky+blue&rich=1" \ -H "Authorization: Bearer sk_live_..." ``` ## GET /answer — direct answer (experimental) Ask a question, get a direct answer. Some questions won't have one yet — then `answered` is false (use /search for links). Experimental: the shape may change as we improve it (own-model + cache are coming). Query params: - `q` (string, required) — the question (max 512 chars). - `deep` (0 | 1, default 1) — `1` runs the full pipeline: a direct answer when the web's top results support one (reading a couple of them if needed). `0` returns an answer only when one is directly available, else `answered:false`. Response: - `answered` (bool) — did we return a direct answer? - `answer` — `{ kind: "generated", text, sources[] }` (present when `answered`) — a direct answer generated from the web's current top results, with the sources it drew on. - `experimental` (bool) — always true for now. ``` curl "https://api.crawl4ai.com/answer?q=why+is+the+sky+blue" \ -H "Authorization: Bearer sk_live_..." ``` ## POST /extract — structured data from a page (LLM) Read a page (crawled for you) or your own content, and return typed data described by an instruction and/or a JSON schema. Body (give a URL OR inline `content`, plus an `instruction` and/or `schema`): - `url` (string) — the page to read; OR `content` (string) — your own text/markdown/html. - `instruction` (string) — plain-English description of what to pull out. - `schema` (object) — JSON schema each returned record must match (typed, predictable output). - `example` (object) — a sample of the shape you want (structure, not values). ``` curl https://api.crawl4ai.com/extract \ -H "Authorization: Bearer sk_live_..." -H "Content-Type: application/json" \ -d '{ "url":"https://news.ycombinator.com", "instruction":"the top stories on the front page", "schema":{"type":"array","items":{"type":"object","properties":{ "title":{"type":"string"},"points":{"type":"integer"},"url":{"type":"string"}}, "required":["title","points"]}} }' ``` Returns: `{"ok":true,"url":"...","data":[ ... ],"usage":{"total_tokens":N},"ms":1200}` ## POST /recipes/{name} — ready-made scrapers A recipe turns a site's pages into rows: send the recipe's inputs, get `rows` back. The catalog (`GET /recipes`, free) lists every recipe with its inputs (type, default, required), output fields, `fresh_for_s`, `tags`, a `cost` hint, `run` (sync or job) and a `description`. Browse and try them at https://api.crawl4ai.com/recipes/ Body: one field per input (string / int / bool / date), plus `bypass_cache` (bool) to skip the result cache. ``` curl https://api.crawl4ai.com/recipes/hn-hiring \ -H "Authorization: Bearer sk_live_..." -H "Content-Type: application/json" \ -d '{"thread":49522897,"keyword":"Rust"}' ``` Returns: `{"recipe":"hn-hiring","version":1,"rows":[...],"usage":{"units":1,"lines":[...]},"meta":{...}}` - `rows` — one object per row, exactly the recipe's output fields (missing = null). - `usage.lines` — every inner call (each page: `scrape` + `url`; each LLM call: `extract` + `step`) with its units and tokens; one receipt each. A cached repeat inside `fresh_for_s` costs 0 units (`cache: "hit"`). - `session` — some recipes read a site as you: your own login cookies for that site as one `Cookie` header line (the Crawl4AI Session extension copies it). Used for this run only, never stored. The catalog marks the input `secret` and `need` = `required` | `optional`; a wrong session costs the unit (empty `rows`, `bad_session` in `meta.warnings`, the site in `meta.blocked`). - `stage` per recipe (`published` | `staged` = new, not yet promoted) and `health` (each region's check every 6 h: `ok` | `fail` | `unknown`, rows, time, error); `GET /recipes/health` is free and keyless. - errors: 400 `bad_input`, 404 `unknown_recipe`, 403 `host_refused`, 422 `page_budget` / `module_limit` / `bad_module`, 502 `fetch_failed`, 504 `timeout`. A failed run still returns the `usage` it cost. ## POST /scrape/batch — many URLs, streamed Scrape up to 50 URLs in one call. Results stream back as NDJSON — one JSON line per URL as it finishes. Bills one unit per URL. Takes `urls` plus any /scrape field (applied to every URL). ``` # -N streams each line as it lands curl -N https://api.crawl4ai.com/scrape/batch \ -H "Authorization: Bearer sk_live_..." -H "Content-Type: application/json" \ -d '{"urls":["https://a.com","https://b.com"],"format":"md"}' ``` Returns (NDJSON): `{"url":"...","ok":true,"result":{...}}` per line. ## POST /scrape/jobs — large async jobs For big lists (up to 10,000 URLs), or when your code cannot hold a connection open for slow pages. Submit once, get a job id, then poll while it drains in the background. Takes `urls` plus any /scrape field. - `POST /scrape/jobs` with `{"urls":[ ... ]}` → `{"job_id":"j_...","status":"pending"}` - `GET /scrape/jobs/{id}` → status + counts (add `?full=1` for per-URL detail) - `GET /scrape/jobs/{id}/results` → NDJSON results, paged with `?after=N` (500 per page) - `POST /scrape/jobs/{id}/retry` → re-run only the failed URLs ``` JOB=$(curl -s https://api.crawl4ai.com/scrape/jobs \ -H "Authorization: Bearer sk_live_..." -H "Content-Type: application/json" \ -d '{"urls":["https://a.com","https://b.com"]}' | jq -r .job_id) until [ "$(curl -s https://api.crawl4ai.com/scrape/jobs/$JOB \ -H "Authorization: Bearer sk_live_..." | jq -r .status)" = "done" ]; do sleep 2; done curl -s https://api.crawl4ai.com/scrape/jobs/$JOB/results -H "Authorization: Bearer sk_live_..." ``` ## MCP — use Crawl4AI as native tools Add Crawl4AI to Claude Code, Cursor, or any MCP client — no install, just a URL and your key. ``` claude mcp add --transport http crawl4ai https://api.crawl4ai.com/mcp \ --header "Authorization: Bearer sk_live_..." ``` Tools: `scrape` (url, format, proxy, country, parse) · `search` (q, rich) · `answer` (q, experimental) · `extract` (url, instruction, schema, example) · `batch` (urls, format, proxy, country) · `recipes_list` (the ready-made scrapers with their inputs, free) · `recipe_run` (name, inputs, bypass_cache) · `balance` · `estimate` (endpoint, url?, urls?) · `topup` (amount: one of the packs in `GET /v1/prices`) · `spend_cap` (credits) · `recharge` (on, under_credits, pack_usd, max_per_month) · `report_issue` (title, what_happened). A `scrape`, `extract` or `batch` call on a page that needs a real browser can take up to 90 s: let the tool call run (allow 120 s in your MCP client) instead of cancelling and retrying. For 50 to 500 URLs, the MCP has `jobs_submit` (urls, proxy, country, bypass_cache, idempotency_key: send the same key on a retry and the job is billed once), then `jobs_status` (poll every 10 to 30 s) and `jobs_results` (50 per call, paged with `offset`). For more than 500 URLs, call the HTTP jobs API (`POST /scrape/jobs`, above) with the same key. ## Notes - Provenance: every response tells you which engine served it and whether it came from cache. - Regions: requests route to the nearest region automatically (EU, Singapore, or US). - Caching: repeated URLs are served from a shared archive. - Fair use: each plan has request/concurrency limits (see the dashboard); over-limit calls get `429`; out of credit gets `402 no_credit` (top up on the dashboard's Billing tab, or `POST /v1/billing/topup`). ## More - Full API docs: https://api.crawl4ai.com/docs - Dashboard & keys: https://api.crawl4ai.com/dashboard/ - Open-source library: https://github.com/unclecode/crawl4ai