POST /v1/web/extract turns pages into structured JSON. You describe the shape you want — a JSON Schema, a natural-language prompt, or both — and an LLM maps each page onto it. Point it at one URL, a list, or a whole crawl scope, and get back typed data instead of Markdown.
Extract is the premium tier. It’s what you reach for when you don’t just want the page cleaned, you want the specific facts out of it: a product’s price and stock, a company’s funding round, a directory’s every listing.
When to use it
- You need typed fields, not prose — feed a schema and get exactly those keys back.
- You want to extract the same shape from many pages, or from every page under a path.
- You need to trust the result: return per-field confidence and the exact passage each value came from.
- You want many pages collapsed into one deduplicated collection, one row per real-world entity.
Example request
Extract a product’s details from a single page, with confidence and sources turned on.Example response
One entry inresults per extracted page. With showConfidence, each result also carries a fields map: for every field, how certain the value is and the passage it was drawn from. With showSources, the concrete URLs that were extracted are listed at the top level.
What makes it different
Per-field confidence + evidence
With
showConfidence, every field comes back with a score from 0 to 1 and the exact source passage it was drawn from — so you can gate on trust instead of guessing.Merge into one collection
With
mergeEntities, results across many pages collapse into one deduplicated collection — one row per entity, each carrying the source URLs that contributed to it.Crawl scope with /*
Any URL ending in /* is a crawl scope, not a literal address. Every page discovered under that path is extracted and merged into the result. Mix literal URLs and scopes freely in the same urls list.
Key options
string[]
required
The pages to extract from — up to 10 entries. Each must be an
http(s) URL. A trailing /* marks a crawl scope: every page discovered under that path is extracted and merged in.object
A JSON Schema describing the shape you want back. Optional if
prompt is given. When both are present, the schema fixes the field names and types while the prompt guides what to pull.string
A natural-language instruction for what to extract, up to 2000 characters. Use with or instead of a schema.
boolean
default:"false"
Pull in extra source pages by web-searching your prompt, to fill fields your URLs don’t cover. Requires a
prompt.boolean
default:"false"
Return the concrete list of URLs that were actually extracted, after any wildcard and web-search expansion.
boolean
default:"false"
For each field, return a confidence score (0 to 1) and the exact source passage the value was drawn from.
boolean
default:"false"
Merge the per-page results into one deduplicated collection — one row per entity, with its contributing source URLs — instead of a separate result per page.
boolean
default:"false"
Preserve document structure (headings, lists, tables) over prose density when reading the page. Good for listing and catalog pages.
number
Reuse a recent capture of each page if it’s younger than this many milliseconds. Omit or set
0 to always fetch fresh. Capped at 7 days (604800000 ms).Response fields
object[]
One result per extracted page. Each has
url, data (the extracted object, or null when nothing matched), and error (null on success). Includes fields when showConfidence is set.object
Per-field map keyed by field name, each with
confidence (0–1) and evidence (the source passage). Present only when showConfidence is set.string[]
The concrete URLs actually extracted, after wildcard and web-search expansion. Present only when
showSources is set.object[]
The deduplicated collection, one row per entity. Each row has
data (the unioned fields for that entity) and sources (every URL that contributed). Present only when mergeEntities is set.Billing
Extract costs 5 credits per page that returns data. Pages that yield nothing aren’t charged. A crawl scope orenableWebSearch can expand the page count, so the total scales with how many pages actually produce results — the response tells you which ones. See Credits.
Related
Search
Find the pages to extract from when you don’t have their URLs.
Crawl
The site-discovery behind a
/* extract scope, as a standalone job.Formats
How extracted JSON compares to Markdown and structured formats.
API Reference
Full
POST /v1/web/extract schema and a live playground.