New DoubleCrawl raises a $30M Series A to build the retrieval layer for AI agents
Interact v2 — drive a live browser with a prompt

Turn the web into clean, LLM-ready data

The context API for AI agents. Search, scrape, crawl and interact with any site at scale — proxies, JavaScript rendering, anti-bot and rate limits already handled. Open source, MIT.

1,000 credits free No credit card SOC 2 Type II
api.doublecrawl.com 200 · 1.9s
Request

          
Response

          

Retrieval infrastructure for teams shipping AI in production

One API · six endpoints

Everything between a URL and your model's context window

The parts of the web that break pipelines — rendering, pagination, blocks, PDFs, stale caches — resolved before the payload reaches you. Same auth, same schema, same billing unit.

Scrape

Any single page as markdown, JSON, HTML, links or a screenshot. Renders JavaScript, waits for the content that matters, strips the chrome.

POST /v2/scrape 1 credit

Crawl

Point at a domain, get every reachable page. Depth limits, include and exclude globs, sitemap-aware, resumable, streamed as it completes.

POST /v2/crawl 1 credit / page

Map

Every URL on a site in under a second, before you spend a single crawl credit. Search the map by keyword and crawl only what you actually need.

POST /v2/map 1 credit

Search

Web search that returns full page content instead of ten blue links. News, images and scholarly sources included, results already scraped.

POST /v2/search 2 credits / 10 results

Extract

Describe the shape you want with a JSON schema or a prompt. Get typed, validated objects back — across one page or ten thousand.

POST /v2/extract 1 credit / page

Interact

New

Log in, click, scroll, fill forms, paginate — in plain language or with explicit actions. The session persists and the data still comes back clean.

POST /v2/interact 2 credits / minute

Output formats

One request. Every format your pipeline expects.

Ask for as many as you like in a single call — you are billed once for the page, never once per format.

  • Navigation, boilerplate and cookie walls removed before tokenization — typically 93% fewer input tokens than raw HTML.
  • PDFs, DOCX and slide decks parsed on the very same endpoint.
  • SDKs for Python, Node, Go and Rust, plus first-party LangChain, LlamaIndex and MCP integrations.
0B

Pages delivered to production agents

0%

Success rate on sites that block naive scrapers

0s

p95 latency with JavaScript rendering on

0M

Developers building on the API and SDKs

How teams use it

Four calls from a domain name to a live index

The order is the point: map before you crawl and you only pay for pages you actually want. Most production pipelines are these four steps on a schedule.

01

Map the domain

Pull every URL in under a second and filter down to the sections worth indexing.

map → 4,812 urls
02

Crawl what matters

Queue the filtered set. Results stream back as markdown while the job is still running.

crawl → 1,203 pages
03

Extract structure

Hand over a schema, get typed objects — no regex and no brittle selectors to maintain.

extract → typed JSON
04

Keep it fresh

Re-crawl nightly. A webhook fires only on real changes, with a diff attached.

monitor → webhook

Built for production

The infrastructure you would otherwise build twice

Everything that turns a weekend scraper into a system your team is willing to be paged about at 3am.

Proxies and stealth, included

Residential and datacenter pools, automatic rotation, fingerprint randomization and CAPTCHA handling. No proxy vendor, no per-gigabyte surprise.

Real browsers at scale

Full JavaScript rendering on every request, with wait conditions, custom headers and session reuse. Thousands of concurrent browsers, kept warm.

Monitor for changes

Re-crawl on a schedule and get a webhook only when content actually changed — with a diff, so your index stays fresh without a full re-run.

Batching and queues

Submit ten thousand URLs in one call. Jobs are idempotent, resumable and streamed back as each page lands — no polling loop to babysit.

Compliance you can show

SOC 2 Type II, GDPR-aligned processing, EU and US data residency, robots.txt respected by default and zero retention on request.

Open source, self-hostable

MIT-licensed core with a Docker image and a Helm chart. Run it inside your VPC, keep the SDKs, move to the cloud whenever you would rather not.

Pricing

One credit is one page. That is the whole model.

Credits never expire on paid plans and annual billing takes two months off. Start on the free tier — no card, every endpoint unlocked.

Free
$0

1,000 credits, one time

  • 2 concurrent browsers
  • All six endpoints
  • Community support
Get an API key
Hobby
$19/mo

3,000 credits per month

  • 5 concurrent browsers
  • Scheduled monitoring
  • Email support
Choose Hobby
Most popular
Standard
$99/mo

100,000 credits per month

  • 50 concurrent browsers
  • Residential proxies
  • Webhooks and diffs
  • 99.99% uptime SLA
Choose Standard
Growth
$399/mo

500,000 credits per month

  • 100 concurrent browsers
  • Priority queue and routing
  • Shared Slack channel
Choose Growth

Scrape · Crawl · Map · Monitor — 1 credit / page Search — 2 credits / 10 results Interact — 2 credits / browser-minute

Enterprise

Volume pricing, dedicated proxy pools, single-tenant regions, custom data residency, SSO and SCIM — and an engineer who knows your workload by name.

Talk to sales

Customers

Teams that deleted their scraping stack

Fewer moving parts, fewer 3am pages, and a retrieval layer that stops being the reason the demo fell over.

“We ran our own browser fleet for two years. Moving retrieval to DoubleCrawl deleted a service, a proxy contract and about four thousand lines of code — and our success rate went up.”

MR
Mira Reinholt
Staff Engineer, Northwind AI

“Map first, then crawl only what changed. That single pattern took our nightly index refresh from six hours to eleven minutes.”

DA
Devon Achebe
Head of Data, Corvus Labs

“Extract returns typed objects straight from a schema. Our agents stopped inventing fields because there are no loose strings left to invent from.”

SK
Sanne Køhler
Founder, Fathom Retrieval

“We evaluated four vendors in a week. DoubleCrawl was the only one that handled logged-in pages without us writing a single selector.”

TO
Tomás Oyelaran
CTO, Meridian Intelligence

Ship the agent.
We'll handle the web.

1,000 credits free, every endpoint unlocked, no card. Most teams are in production within the same week.

$ pip install doublecrawl