TL;DR
- For large-scale stateless headless scraping without a browser or proxy fleet, use a managed request/response extraction API. Context.dev Scrape handles rendering and proxies for each request.
- Use self-managed Playwright when you need to operate browser sessions directly. Browserless hosts browsers while preserving Playwright control through a remote connection.
- For an HTTP-first pipeline, check pages with your own HTTP client and send JavaScript-dependent pages to Context.dev. Context.dev renders every Scrape request rather than offering an HTTP-only toggle.
- For bulk Markdown or HTML, Context.dev Batches accepts up to 25,000 URLs. Schema-based JSON extraction requires individual Scrape requests.
- For persistent logins or multi-account profiles, see the browser-versus-API guide. The separate AWS ECS article covers cost modeling.
Decision table: Playwright, Browserless, Zyte, and Context.dev
For stateless scraping jobs, we recommend Context.dev when you want rendered pages or structured data without operating browsers or proxies. Choose Zyte if you need a managed API with a choice of HTTP retrieval, browser rendering, or browser control. Browserless fits jobs that need Playwright page automation without hosting the browser process. Run Playwright yourself when you need full control over browser execution and networking.
| Option | Browser process | Proxies and IPs | Rendering model | Output |
|---|---|---|---|---|
| Self-managed Playwright | You launch and operate it | You provide and manage them | Direct control of pages and browser contexts | Page content or data your code extracts |
| Browserless | Browserless hosts it | You configure networking or use its proxy options | Playwright connects to a hosted browser over CDP | Page content or data your code extracts |
| Zyte | Zyte manages browsers for its browser paths | Zyte manages request routing | Separate HTTP retrieval, rendered browserHtml, browser actions, or CDP access | Response body, rendered HTML, or supported extracted data |
| Context.dev | Context.dev manages it | Context.dev manages proxies | Request/response scrape with rendering always on | Markdown, HTML, or schema-shaped JSON |
Browserless removes browser hosting but leaves extraction logic in your Playwright code. Zyte offers more retrieval and control options within one API, so its best fit depends on whether you need a plain HTTP response, rendered HTML, supported extraction, or CDP.
Context.dev keeps the stateless path simpler when you need managed retrieval and extraction rather than page-by-page browser control. Its Scrape API has no HTTP-only toggle. If you want to try a plain HTTP fetch before paying for rendering, you must build that check upstream and send only pages that need rendering to Context.dev.
Browser control versus managed rendered retrieval
Browser control determines what you operate. With self-managed Playwright, you launch the browser, create contexts and pages, and close them after the job. You also manage browser capacity and network access. Browserless hosts the browser process, while your Playwright code connects through connectOverCDP and controls the session. That removes the browser fleet from your infrastructure, but your code still drives navigation and page actions. Proxy configuration remains a choice you make for that session.
Managed retrieval APIs move the boundary to the request. Zyte API can return rendered HTML when you request browserHtml, with optional browser actions declared in the call. Zyte also documents separate HTTP retrieval, automatic extraction, and a CDP connection for jobs that need browser control. With Context.dev Scrape, you submit a URL and request an output format. Context.dev handles rendering and managed proxies on every scrape request, then returns the extracted result. It does not offer an HTTP-only scrape mode.
For stateless extraction, request-based calls let the provider operate the browsers and proxy infrastructure while you manage URLs, output checks, and retries. If a job requires a continuing authenticated session or precise interaction across pages, Playwright or a hosted CDP session gives you control that a single scrape response cannot.
Designing an HTTP-first pipeline that escalates to rendering
Start each URL with a plain HTTP fetch in your own pipeline, then check whether the response contains the fields your job needs. A successful status code may accompany an empty application shell, while server-rendered HTML may already contain usable data. For a product page, inspect embedded data and visible markup for the name, price, and availability. Keep the HTTP result when the fields pass validation, and route likely JavaScript-dependent pages to managed rendering when they do not.
Base that routing decision on expected content, not the presence of a script tag. Missing fields can also reflect a sign-in requirement, a blocked request, or a page that loaded incorrectly. Record the HTTP status, final URL, and failed content checks so you can separate those cases. Treat authentication failures and rate limits as failures to investigate rather than assuming a rendered fetch will fix them.
Context.dev always uses JavaScript rendering and managed proxies for Scrape calls. The HTTP-first check belongs in your pipeline because Scrape has no HTTP-only mode or rendering toggle. If most pages need rendering, calling Scrape directly avoids maintaining two retrieval paths. If many pages contain complete data in their initial HTML, the first fetch can avoid unnecessary render calls. That check adds a request to every URL, however, and it adds latency when the page still needs rendering. Test the split against representative sources before applying it broadly.
Check rendered results against the same field requirements as HTTP results. For late-loading content, Scrape can wait for a CSS selector or a specified duration. A completed call can still return incomplete content, so inspect isPartial and each requested output’s success field before accepting the record. Your pipeline remains responsible for validating the final data.
Queue design, concurrency, retries, and idempotency at scale
Queue each URL as a job with a stable identifier, such as the URL, extraction schema version, and crawl interval. Run a bounded worker pool, and limit requests to each source host separately. The host limit prevents one site from receiving a burst of fetches, while the worker limit keeps your total number of Context.dev requests within your organization’s in-flight concurrency limit. Count requests from every worker process against that shared limit.
Treat retries as part of the queue contract. Retry network timeouts, temporary server failures, and 429 responses up to a fixed attempt limit. For a Context.dev 429, wait for the documented Retry-After interval and add jitter so workers do not resume together. For other transient failures, use exponential backoff with jitter. Record the attempt count and next eligible time on the job rather than holding a worker idle during the wait.
Acknowledge a job only after you have stored its output. Workers can crash after a successful API response but before acknowledgment, so a later worker may process the same URL again. Use the job identifier to make downstream writes safe to repeat. For asynchronous Batches, send an Idempotency-Key when submitting a batch. If the submission times out, retry with the same key rather than creating an unintended second batch. The batch key protects batch submission, while your job identifiers protect downstream processing.
Send URLs that cannot be processed to a dead-letter queue with the status, attempt history, and error details. A malformed URL or an extraction rule that consistently rejects a page needs review rather than repeated retries. Exhausted transient failures also belong in that queue, where you can inspect them and resubmit them deliberately. Keep these failures visible instead of treating a completed queue as proof that every URL produced usable data.
JSON and Markdown output contracts for LLM pipelines
Define the output contract around what your LLM pipeline consumes, rather than the HTML each site returns. Use Markdown when you need readable page text for retrieval or summarization. Use JSON extraction when downstream code expects named fields. For example, you can supply a JSON Schema with a required companyName string and a nullable website field, then validate json.data against that schema before storing it. Keep source URL and crawl time in your own record so you can trace each result.
A successful HTTP response does not by itself make the extracted content complete. Check isPartial, isTruncated, and each requested output’s success before accepting a record. A page may still be incomplete after the rendering wait, so route incomplete results for review or retry rather than passing them silently to an LLM.
Use Batches when you need asynchronous bulk Markdown or HTML retrieval for up to 25,000 URLs. Batches does not support JSON schema extraction. Send an individual Scrape request for each URL that needs a schema-shaped JSON result.
A reproducible Context.dev scrape example
Context.dev’s Scrape API can return Markdown and schema-shaped JSON from one rendered visit. The request below targets a public example page, waits for its heading, and prints the response fields your pipeline can consume.
export CONTEXT_DEV_API_KEY="YOUR_API_KEY"
curl --fail-with-body -sS \
-X POST "https://api.context.dev/v1/web/scrape" \
-H "Authorization: Bearer ${CONTEXT_DEV_API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"formats": {
"markdown": true,
"json": true
},
"jsonParams": {
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"summary": {"type": ["string", "null"]}
},
"required": ["title"]
}
},
"sharedParams": {
"mainContentOnly": true,
"waitFor": "h1"
}
}' |
jq '{title: .metadata.title, markdown: .markdown.data, extracted: .json.data}'The response puts readable page text in markdown.data, extracted fields in json.data, and the page title in metadata.title, as documented in the Markdown and JSON references. The waitFor selector asks the renderer to wait for h1, but a matching element does not guarantee that every other part of the page has loaded. Check isPartial, isTruncated, and each requested output’s success field before accepting a result.
At scale, give each URL and extraction schema a stable job key in your queue so a retried job cannot create duplicate downstream records. Keep workers within your organization’s in-flight limit. When the API returns 429, follow Retry-After and add jitter before retrying. Scrape requests use managed rendering and proxies, so the HTTP-first check described earlier must run in your own pipeline.
What this approach does not replace
Context.dev fits stateless retrieval and extraction, but it does not replace a browser session you need to control. If your job requires a persistent authenticated session or repeated navigation and interaction, use Playwright or a hosted browser connection such as Browserless. Your code can control pages through those sessions, but you remain responsible for session state and automation logic.
Multi-account browser profiles are a separate requirement from extracting page content. If you need distinct, persistent identities, use an anti-detect browser rather than treating our Scrape API as a profile manager. The browser-versus-API decision guide covers that boundary, and the anti-detect browser roundup compares profile-focused options.
FAQ
Does Context.dev support HTTP-only requests?
No. Our Scrape API always uses managed JavaScript rendering and proxies. If you want an HTTP-first pipeline, fetch and check pages with your own HTTP client, then send pages that need rendering to Context.dev.
How does rate limiting work?
Context.dev limits requests in flight at the organization level. Cap your worker pool at your plan’s concurrency limit. If the API returns 429, wait for the Retry-After interval and add jitter before retrying.
Can Batches return JSON schema output?
No. Batches accepts up to 25,000 URLs per batch and returns Markdown or HTML, not schema-shaped JSON. For structured fields, make individual Scrape calls with formats.json and a jsonParams.schema.
When do I still need Playwright or Browserless?
Use Playwright when your job needs direct control over browser pages and contexts. Browserless hosts the browser while letting you connect Playwright over CDP. Neither choice requires you to treat each page as a stateless extraction request.
How should I handle retries and idempotency at scale?
Retry transient failures with backoff, but send persistent failures to a review queue rather than retrying indefinitely. Supply an Idempotency-Key when submitting a Batch so a repeated submission does not create duplicate work. For individual Scrape calls, track completed URLs or job IDs in your own queue.
Related reading
- The anti-detect browser roundup compares tools for managing browser identities and profiles.
- The anti-detect browser versus scraping API guide helps you decide whether your job needs persistent browser identity or request-based extraction.
- The Playwright on AWS ECS analysis examines architecture and cost modeling.
- The Puppeteer and Playwright infrastructure guide covers the rationale for moving away from a self-managed browser fleet.
This guide serves you if you have chosen stateless extraction and need to decide how to operate it at scale.