Scrape
Any single page as markdown, JSON, HTML, links or a screenshot. Renders JavaScript, waits for the content that matters, strips the chrome.
The context API for AI agents. Search, scrape, crawl and interact with any site at scale — proxies, JavaScript rendering, anti-bot and rate limits already handled. Open source, MIT.
Retrieval infrastructure for teams shipping AI in production
One API · six endpoints
The parts of the web that break pipelines — rendering, pagination, blocks, PDFs, stale caches — resolved before the payload reaches you. Same auth, same schema, same billing unit.
Any single page as markdown, JSON, HTML, links or a screenshot. Renders JavaScript, waits for the content that matters, strips the chrome.
Point at a domain, get every reachable page. Depth limits, include and exclude globs, sitemap-aware, resumable, streamed as it completes.
Every URL on a site in under a second, before you spend a single crawl credit. Search the map by keyword and crawl only what you actually need.
Web search that returns full page content instead of ten blue links. News, images and scholarly sources included, results already scraped.
Describe the shape you want with a JSON schema or a prompt. Get typed, validated objects back — across one page or ten thousand.
Log in, click, scroll, fill forms, paginate — in plain language or with explicit actions. The session persists and the data still comes back clean.
Output formats
Ask for as many as you like in a single call — you are billed once for the page, never once per format.
Pages delivered to production agents
Success rate on sites that block naive scrapers
p95 latency with JavaScript rendering on
Developers building on the API and SDKs
How teams use it
The order is the point: map before you crawl and you only pay for pages you actually want. Most production pipelines are these four steps on a schedule.
Pull every URL in under a second and filter down to the sections worth indexing.
map → 4,812 urlsQueue the filtered set. Results stream back as markdown while the job is still running.
crawl → 1,203 pagesHand over a schema, get typed objects — no regex and no brittle selectors to maintain.
extract → typed JSONRe-crawl nightly. A webhook fires only on real changes, with a diff attached.
monitor → webhookBuilt for production
Everything that turns a weekend scraper into a system your team is willing to be paged about at 3am.
Residential and datacenter pools, automatic rotation, fingerprint randomization and CAPTCHA handling. No proxy vendor, no per-gigabyte surprise.
Full JavaScript rendering on every request, with wait conditions, custom headers and session reuse. Thousands of concurrent browsers, kept warm.
Re-crawl on a schedule and get a webhook only when content actually changed — with a diff, so your index stays fresh without a full re-run.
Submit ten thousand URLs in one call. Jobs are idempotent, resumable and streamed back as each page lands — no polling loop to babysit.
SOC 2 Type II, GDPR-aligned processing, EU and US data residency, robots.txt respected by default and zero retention on request.
MIT-licensed core with a Docker image and a Helm chart. Run it inside your VPC, keep the SDKs, move to the cloud whenever you would rather not.
Pricing
Credits never expire on paid plans and annual billing takes two months off. Start on the free tier — no card, every endpoint unlocked.
1,000 credits, one time
3,000 credits per month
100,000 credits per month
500,000 credits per month
Scrape · Crawl · Map · Monitor — 1 credit / page Search — 2 credits / 10 results Interact — 2 credits / browser-minute
Volume pricing, dedicated proxy pools, single-tenant regions, custom data residency, SSO and SCIM — and an engineer who knows your workload by name.
Customers
Fewer moving parts, fewer 3am pages, and a retrieval layer that stops being the reason the demo fell over.
“We ran our own browser fleet for two years. Moving retrieval to DoubleCrawl deleted a service, a proxy contract and about four thousand lines of code — and our success rate went up.”
MRMira ReinholtStaff Engineer, Northwind AI
“Map first, then crawl only what changed. That single pattern took our nightly index refresh from six hours to eleven minutes.”
“Extract returns typed objects straight from a schema. Our agents stopped inventing fields because there are no loose strings left to invent from.”
“We evaluated four vendors in a week. DoubleCrawl was the only one that handled logged-in pages without us writing a single selector.”
1,000 credits free, every endpoint unlocked, no card. Most teams are in production within the same week.