Crawl4AI / docs

API reference

One key, plain JSON, one fast endpoint. Turn any URL into clean Markdown, search the web, or pull typed JSON out of a page. Every call is authenticated with your API key and returns in one round trip.

Base URL https://api.crawl4ai.com  ·  Auth header Authorization: Bearer sk_live_...  ·  get a key

From the libraryScrapeSearchAnswerExtract BatchBulk jobsRecipesWatchMCPPricingRegions & data

OSSfrom the library to the cloud

You already know crawl4ai. The cloud runs the same engine behind one key: no browser to install, no proxies to buy, no cookies to babysit. Each library call has one API call.

# one page → Markdown
async with AsyncWebCrawler() as crawler:
    result = await crawler.arun("https://example.com")
    print(result.markdown)

# many pages
results = await crawler.arun_many(urls, config=run_conf)

# typed JSON out of a page
config = CrawlerRunConfig(extraction_strategy=LLMExtractionStrategy(schema=..., instruction=...))
result = await crawler.arun(url, config=config)
in the libraryin the cloudwhat changes
crawler.arun(url)POST /scrapethe engine is picked per domain (cache → HTTP → browser); bot walls and proxies are handled
crawler.arun_many(urls)POST /scrape/batch · /scrape/jobsresults stream as they land; jobs run in the background for thousands of URLs
LLMExtractionStrategy / JsonCssExtractionStrategyPOST /extracta plain-English instruction or a JSON schema; no selectors to write
PruningContentFilter → fit_markdownon by defaultevery scrape returns clean Markdown with the boilerplate stripped
(no web search)GET /searchranked results as JSON, from our own search stack
your machine's browser, proxies, cookiesoursthree regions, routed to the nearest; the library stays open source, forever

POST/scrape

Fetch a page and return clean Markdown (and/or HTML). Handles JS-heavy pages and bot walls automatically — you don't pick an engine.

curl
curl https://api.crawl4ai.com/scrape \
  -H "Authorization: Bearer sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "format": "both",
    "proxy": "residential",
    "country": "us",
    "parse": { "links": true, "media": true, "metadata": true, "tables": true }
  }'

Body

FieldTypeDescription
urlreqstringThe page to scrape. Must be a public http/https URL.
formatstringboth (default) · md · html — what content comes back.
proxystringnone (default) · isp · residential — which exit network to render through, for sites that block datacenters.
countrystringTwo-letter exit country (e.g. us, sg). Applies with isp/residential.
parsebool | objectAlso return structured page data. true for all, or pick: {"links":true,"media":true,"metadata":true,"tables":true}.

How long it takes

Most pages come back in 1 to 2 seconds. A page that needs a real browser usually takes 10 to 60 seconds, and the slowest up to 90. Set your client's timeout to 120 seconds and wait for the answer. A call you cancel early is not charged, but you lose the page. A 403 with "blocked" means every step got the site's block page; a 502 means the site did not answer; a 503 with "fleet-busy" means no browser was free for a moment, so send the same request again in a few seconds. None of these is charged. For many URLs, use /scrape/batch (up to 50, each result streams back when its page is ready) or /scrape/jobs (up to 10,000, runs in the background while you poll, so no connection waits on a slow page).

Web search, browser-free, results ranked and cleaned. Add rich=1 to also get the on-page extras — follow-up questions, related queries, an entity card and more (see below). For a direct answer, use /answer.

curl
curl "https://api.crawl4ai.com/search?q=rust+web+crawler" \
  -H "Authorization: Bearer sk_live_..."

Query

ParamTypeDescription
qreqstringThe search query (max 512 chars).
rich0 | 1Default 0 (ranked links only). 1 adds a rich block: a direct answer when one exists, follow-up questions, related queries, an entity card, videos, news and more. Slightly slower; best for questions.

Rich response rich=1

Every key below is optional — each appears only when that block is on the page. For a direct answer to a question, use /answer.

rich block
{
  "results": [ … ranked links, same as always … ],
  "rich": {
    "follow_up_questions": [ "Why is the sky blue at sunset?" ],
    "related_queries":    [ "why is the ocean blue" ],
    "entity":  { "title": "Apple Inc", "subtitle": "NASDAQ: AAPL" },
    "videos": [ { "title": "…", "url": "https://…" } ],
    "news":   [ { "title": "…", "url": "https://…" } ],
    "discussions":  [ … ],
    "did_you_mean": "corrected spelling"
  }
}

GET/answer experimental

Ask a question, get a direct answer. Some questions won't have one yet — then answered is false (use /search for links). Add deep=0 for a direct answer only when one is readily available (no page reading). Experimental: the shape may change as we improve it.

curl
curl "https://api.crawl4ai.com/answer?q=why+is+the+sky+blue" \
  -H "Authorization: Bearer sk_live_..."
response
{
  "answered": true,
  "answer": {
    "kind": "generated",
    "text": "The sky is blue because Earth's atmosphere scatters sunlight…",
    "sources": [ { "title": "NASA", "url": "https://…" } ]
  },
  "experimental": true
}

POST/extract

Pull structured, typed data out of a page with an instruction and/or a JSON schema. Give a URL (we fetch it) or your own content.

curl
curl https://api.crawl4ai.com/extract \
  -H "Authorization: Bearer sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://news.ycombinator.com",
    "instruction": "the top stories on the front page",
    "schema": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "title":  { "type": "string"  },
          "points": { "type": "integer" },
          "url":    { "type": "string"  }
        },
        "required": ["title", "points"]
      }
    },
    "example": [
      { "title": "Show HN: My project", "points": 128, "url": "https://..." }
    ]
  }'

Body

FieldTypeDescription
urlstringPage to read. We fetch it for you.
contentstringYour own text/markdown/html to extract from, instead of a URL.
instructionstringPlain-English description of what to pull out.
schemaobjectJSON schema each returned record must match — gives you typed, predictable output.
exampleobjectA sample of the shape you want (structure, not values).
Give a URL or content (plus an instruction and/or schema). Long pages are chunked and re-assembled automatically.

POST/scrape/batch

Scrape many URLs in one call — up to 50 — and stream a result per line as each finishes.

# -N streams each result line as it lands
curl -N https://api.crawl4ai.com/scrape/batch \
  -H "Authorization: Bearer sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://a.com", "https://b.com"], "format": "md"}'

Body

FieldTypeDescription
urlsreqstring[]The URLs to scrape (max 50 per call).
…—Any /scrape field (format, proxy, country, parse) applies to every URL.
Response is application/x-ndjson — one JSON line per URL as it completes.

POST/scrape/jobs

For big lists — up to 10,000 URLs — or when your code cannot hold a connection open for slow pages. Submit once, get a job id, then poll for results while it drains in the background.

# 1. submit -> job id
JOB=$(curl -s https://api.crawl4ai.com/scrape/jobs \
  -H "Authorization: Bearer sk_live_..." -H "Content-Type: application/json" \
  -d '{"urls": ["https://a.com", "https://b.com"]}' | jq -r .job_id)

# 2. poll until done
until [ "$(curl -s https://api.crawl4ai.com/scrape/jobs/$JOB \
       -H "Authorization: Bearer sk_live_..." | jq -r .status)" = "done" ]; \
  do sleep 2; done

# 3. fetch results (NDJSON, one line per URL)
curl -s https://api.crawl4ai.com/scrape/jobs/$JOB/results \
  -H "Authorization: Bearer sk_live_..."

Endpoints

RouteDoes
POST /scrape/jobsSubmit urls (max 10,000) + any /scrape field. Returns a job_id.
GET /scrape/jobs/{id}Status + counts (pending / done / error). Add ?full=1 for per-URL detail.
GET /scrape/jobs/{id}/resultsStreamed results, paged with ?after=N (500 per page).
POST /scrape/jobs/{id}/retryRe-run just the failed URLs.

POST/recipes/{name}

Ready-made scrapers. Pick a recipe from the catalog, send its inputs, get rows back. Browse and try them at /recipes/.

curl
curl https://api.crawl4ai.com/recipes/hn-hiring \
  -H "Authorization: Bearer sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{ "thread": 49522897, "keyword": "Rust" }'

Body

FieldTypeDescription
<input>string · int · bool · dateOne field per recipe input, as the catalog lists them. A missing input takes the recipe's default; a required one without a default fails with bad_input.
bypass_cacheboolRun again even when a fresh result is cached (see below).

Response

json
{
  "recipe": "hn-hiring", "version": 1,
  "rows": [ { "company": "Senzing", "author": "samlk", "text": "Senzing | Platform Engineer | Remote (USA) ..." } ],
  "usage": {
    "units": 1, "boxes": ["crawl"], "cache": "miss", "engine": "recipe",
    "lines": [ { "call": "scrape", "url": "https://news.ycombinator.com/item?id=49522897", "units": 1, "cache": "archive", "engine": "archive" } ]
  },
  "meta": { "pages_fetched": 1, "from_cache": false, "elapsed_ms": 812, "llm_calls": 0, "warnings": [] }
}
FieldDescription
rowsOne object per row, exactly the recipe's output fields, in order. A field the page did not have is null.
usageWhat this run deducted, the same object every endpoint returns: units, llm_tokens (only when an LLM ran), boxes, remaining_before on capped tiers. lines lists every inner call in order: each page (scrape, its url) and each LLM call (extract, its step) with its own units and tokens. Each line is one receipt in your usage history.
metapages_fetched, llm_calls, elapsed_ms, from_cache, and warnings (fields that came back empty, unmatched joins, a module's log lines).

The result cache

Every recipe declares fresh_for_s. A second call with the same inputs inside that window is served from the cache: usage.cache is "hit", usage.units is 0, meta.from_cache is true. Send "bypass_cache": true to run again. Recipes with a secret input are never cached.

Your session

Some recipes read a site as you: they take a session input, your own login cookies for that site as one Cookie header line. The Crawl4AI Session extension copies it in one click. Crawl4AI uses it for this run only and does not store it: no archive, no cache, no log. Your account and the site's terms are your responsibility. The catalog says per input whether it is a secret and whether the session is required or optional; without an optional session the recipe reads what a logged-out visitor sees. A wrong or expired session costs the run's unit: rows is empty, meta.warnings carries bad_session, and meta.blocked names the site.

Staged recipes and health

Every recipe in the catalog carries a stage: published, or staged for a new recipe that runs but is not yet promoted. Each region checks every recipe every six hours and the catalog carries the result as health (ok, fail or unknown, with the last check's rows, time and error). GET /recipes/health returns the same, without a key.

Errors

StatuserrorMeaning
400bad_inputAn unknown input, a required one missing, or a wrong type.
404unknown_recipeNo recipe with that name.
403host_refusedThe recipe tried a host outside its allowed list.
422page_budget · module_limit · bad_moduleThe run needed more pages than the recipe allows, or its module hit a limit.
502 · 504fetch_failed · timeoutA page could not be fetched, or the run passed 300 s.
A failed run still returns usage: the pages fetched before the failure were billed, and it says which. A partial result is never returned as success.

The catalog GET /recipes

curl
curl https://api.crawl4ai.com/recipes -H "Authorization: Bearer sk_live_..."

Free (no unit). Returns {"recipes": [...]}: for each recipe its name, version, title, stage, health, kind (json, module or script), inputs (type, default, required), output fields, fresh_for_s, tags, a cost hint, run (sync or job) and a description.

POST/monitors

Watch a page and get told when it changes in the way you care about. Give the URL and a plain sentence; every real change that matches comes back as typed facts (what changed, from what, to what) with links to both page versions. Try it in the dashboard's Playground, mode watch.

curl
curl https://api.crawl4ai.com/monitors \
  -H "Authorization: Bearer sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://competitor.com/pricing",
        "watch": "a plan gets more expensive, or a plan is removed",
        "every": "1h",
        "notify": { "webhook": "https://you.example/hook", "secret": "s3cr3t" } }'

Body

FieldTypeDescription
urlstringThe page to watch. http or https, a public host.
watchstringWhat to tell you about, in your words. Several sentences are fine: "a new article about OpenAI", "the product is back in stock", "a plan gets more expensive, or a plan is removed". We split it into conditions and show them back in the response. Empty = any real change on the page.
everystringHow often to check: 4m, 15m, 1h, 6h, 1d (default), 1w. At least 4 minutes. A page that changes often is checked more often on its own; every is the longest you wait.
notifyobjectwebhook (a URL we POST each confirmed change to), secret (signs the POST, see below), email. All optional; without them, read the changes with the events call.
schemaobjectYour own JSON shape, filled on every change next to the standard facts, as custom. A field the page does not show is null.
namestringA label for you.
regionstringeu, sin or us: where the page is fetched from. Default: the region of the gate you called.

Response

json
{
  "id": "mon_7f3a20c1e9b04d55", "created": true,
  "url": "https://competitor.com/pricing",
  "watch": "a plan gets more expensive, or a plan is removed",
  "conditions": { "any": [ { "id": "c1", "text": "the price of a plan goes up" },
                           { "id": "c2", "text": "a plan is removed" } ] },
  "every": "1h", "region": "eu", "health": "ok",
  "baseline": { "status": "pending" },
  "cost": { "per_change_credits": 3 }
}
The same url and watch return the existing monitor with created: false. baseline becomes { "status": "ok", "version": "…" } after the first check, within a minute. Read conditions: that is how we understood your sentence; call again with a clearer one if it is not what you meant.

The other calls

CallDescription
GET /monitorsYour monitors, newest first.
GET /monitors/{id}One monitor.
PATCH /monitors/{id}Change watch, every, notify, schema, name, or paused: true|false. A paused monitor keeps its history and costs nothing.
DELETE /monitors/{id}Remove it with its history.
GET /monitors/{id}/events?since=&limit=The timeline, oldest first: every change with its facts, and health events. since is ISO 8601 or unix seconds; limit 1 to 500, default 100.
GET /monitors/{id}/versions/{version}The page as we saw it at that version, as Markdown: the two evidence links of a change.

A change

json · one event, the same shape on the timeline and in the webhook
{
  "id": 4181, "event": "change", "monitor": "mon_7f3a20c1e9b04d55",
  "url": "https://competitor.com/pricing",
  "matched": ["c1"],
  "summary": "The Pro plan's price went from $49 / month to $59 / month.",
  "facts": [ { "kind": "changed", "entity": "Pro", "field": "price",
               "before": "$49 / month", "after": "$59 / month", "section": "Pricing > Pro" } ],
  "custom": null,
  "first_seen": "2026-09-26T12:05:32Z", "confirmed_at": "2026-09-26T12:35:41Z",
  "evidence": { "before": "https://api.crawl4ai.com/monitors/mon_7f3a…/versions/75af60efc867a0cb",
                "after":  "https://api.crawl4ai.com/monitors/mon_7f3a…/versions/e70482929f228688" },
  "cost_credits": 3,
  "delivery": { "webhook": { "status": 200, "attempts": 1, "at": "2026-09-26T12:35:42Z" }, "done": true }
}
FieldDescription
matchedThe ids of the conditions that held. A change that matched nothing is on the timeline too, with [] and no cost, so you can widen the sentence.
factsOne per thing that changed: kind is added, removed or changed; entity and field name it; before and after are the values as the page shows them, or "not shown" when the entity stays but the value is not visible.
first_seen · confirmed_atA change is delivered 30 minutes after it was first seen, when the page has not gone back. The timeline shows it at once.
event: "health"Instead of a change: health is blocked or error with a plain message, when the site refused three checks in a row; and ok when it answers again.

The webhook

One POST per confirmed change, the event above as the body, with the headers x-c4-event: change, x-c4-monitor, x-c4-timestamp (unix seconds) and, when you gave a secret, x-c4-signature: sha256=<hex HMAC-SHA256(secret, timestamp + "." + body)>. Answer 2xx. Otherwise we try again after 1, 5, 30 minutes, 2 and 12 hours, then stop; the timeline keeps the last status.

Credits

1 credit per check, once per every (a page checked more often than that inside the window costs no more). 3 credits per confirmed change that matched. A blocked or failed check, a change that matched nothing, and a paused monitor cost nothing. Up to 100 monitors per account.

Errors

StatuserrorMeaning
400url is not valid · url must start with http:// or https:// · url host is not public · every must be at least 4mThe URL, a private or local host, or an interval under 4 minutes. The error text says which.
404no such monitor · no such versionNot on your account.
429at most 100 monitors per account · at most 20 running monitors checked more often than every 1h per account · this site has too many monitors alreadyYour account's limits (100 monitors, of which 20 may run faster than hourly), or 500 on one host across all accounts.
503could not read the sentence, try againThe split of your sentence failed; call again.

MCP/mcp

Use Crawl4AI as native tools inside Claude, Cursor, or any MCP client — no install, just a URL and your key.

# one line in your terminal
claude mcp add --transport http crawl4ai \
  https://api.crawl4ai.com/mcp \
  --header "Authorization: Bearer sk_live_..."
Tools exposed: scrape · search · answer · extract · batch · jobs_submit · jobs_status · jobs_results (the bulk jobs, 50 to 500 URLs) · recipes_list · recipe_run (the recipes) · the billing tools balance · estimate · topup · spend_cap · recharge · and report_issue (sends a bug report to the Crawl4AI team). Your dashboard's MCP tab pre-fills this with your key.

OPSregions & data

Three regions, one address. Your request goes to the nearest region by DNS; every region runs the whole stack.

regionaddresswhen to pin it
Europeeu-gate.crawl4ai.comapi.crawl4ai.com picks the nearest region for you. Pin a region when the page you fetch answers differently per country, or when your data must stay in one region.
Singaporesin-gate.crawl4ai.com
United Statesus-gate.crawl4ai.com
What we keep: every request's parameters and result summary for 90 days, so your usage history and our reliability work have the facts; the pages we fetch for you sit in our cache so a repeat is instant. We never sell or share what you fetch. Your API key is shown once in the dashboard and can be regenerated at any time.

CREDITSpricing

One unit, the credit. Every request costs credits, and its receipt says why. You pay for the effort a page needed, never more: a result already in our archive costs half (we keep it fresh), a plain page is the base price, a page that fought back costs more. Money enters once, at a top-up. The numbers below are read live from GET /v1/prices.

Note
Launch pricing for the first year: discounted, and it may change. Credit you already hold keeps its value.

The table

whatcreditsnote
1 credit…the dollar value of one credit; a top-up buys credits at this rate
scrape…one page at the light level; × the effort multiplier below
search…one live search; a search already in our archive costs half
extract…the fetch, plus the tokens
answer…the search, plus the tokens; an answer already in our archive costs half
tokens…per 5,000 tokens, input and output each have their own price (llm.models in GET /v1/prices). At launch every extract runs on Gemini 3.1 Flash-Lite; model, provider and api_key are ignored for now. The response and the receipt show the count
jobs…per URL, at the light level
recipes…per unit of a run, at the light level
watch1 · 31 credit per check, once per interval; 3 per confirmed change that matched

The effort levels

The effort a page needed sets its multiplier. The receipt names the level; the engine that did the work is ours.

level×what it means

Free credit and the packs

whatcredits
at signup, once…
a 7-day pass…
top-up packs…

The signup credit and bought credit never expire; the 7-day pass expires with its key. A paid top-up turns on the Supporter limits (10 at a time, 120 a minute, 10 keys, no monthly cap). Your own monthly spend cap and auto-recharge are on the dashboard's Billing tab.

On every response

headers
x-c4-cost: 1.500        # what this call cost, in credits
x-c4-balance: 14312.000  # your credit after it

Out of credit, the gate answers 402 with the reason, a message you can show as it is, and every way on: {"error":"no_credit","balance":-0.50,"message":"…","actions":{"topup":"","topup_api":"POST …/v1/billing/topup?amount=<a pack> … with a saved card the pack is charged at once, otherwise the answer carries checkout_url to open","mcp":"call the topup tool"}}. Over your own monthly cap: {"error":"spend_cap","cap_left":…,"actions":{"raise":…,"raise_api":"POST /v1/billing/settings {\"spend_cap_mc\":…}","mcp":"call the spend_cap tool"}}. Your key can call both billing routes, so an agent can act on a 402 by itself. A request the gate refused itself (a bad URL, a rate limit) costs nothing; a 5xx costs nothing.

Estimate before a call POST /v1/estimate

The same body as the call, plus the endpoint. The answer is a range: the fixed part is exact, the level and the tokens are a band from real traffic.

curl
curl -X POST "https://api.crawl4ai.com/v1/estimate" \
  -H "Authorization: Bearer sk_live_..." -H "Content-Type: application/json" \
  -d '{"endpoint":"scrape","url":"https://news.ycombinator.com"}'
# → {"min":"1.00","max":"3.00","text":"about 1.00 to 3.00 credits","parts":[…]}

Balance and top-up from a key

curl
curl "https://api.crawl4ai.com/v1/billing/balance" -H "Authorization: Bearer sk_live_..."
curl -X POST "https://api.crawl4ai.com/v1/billing/topup?amount=25&checkout=1" -H "Authorization: Bearer sk_live_..."
# → {"checkout_url": "https://checkout.stripe.com/…", "credits": …, "usd": 25}  open it, pay, the credits land in seconds

Without checkout=1, an account with a card saved from an earlier top-up is charged at once and there is no page to open: the answer is {"paid": true, "usd": 25, "credits": …, "invoice_url": "…", "balance": {…}}. If that card does not pay, the answer is 402: send the request again with checkout=1, or update the card on the dashboard's Billing tab. An account with no saved card always gets the checkout_url.

Get a key → Dashboard llms.txt Open source ↗