TL;DR
- Crawlee, a crawling framework, fits large self-managed crawls that need request queues, concurrency controls, retries, and browser or HTTP-based workers.
- Playwright, a browser automation library, fits JavaScript-rendered sites, interactive workflows, persistent sessions, and cross-browser control.
- Puppeteer, a browser automation library, fits Chrome and Chromium workflows that benefit from direct DevTools Protocol access.
- Cheerio with Axios, a parser and HTTP client, fits fast, lightweight extraction from static HTML that does not require JavaScript execution.
- Context.dev, a managed API, fits teams that need rendered retrieval and clean Markdown or JSON without maintaining browsers, proxies, retries, or anti-bot infrastructure.
Choosing between a scraper library and a managed API
The five tool categories solve different parts of a Node.js scraping stack. Choose the category based on how the target site delivers content and how much infrastructure you want to operate.
HTTP clients such as Axios and undici download pages without executing JavaScript. They work well when the server returns the required content in its initial HTML response.
Parsers such as Cheerio turn downloaded HTML into a structure you can query with CSS selectors. A static scraping script commonly pairs Axios with Cheerio because neither tool renders pages or controls a browser.
Browser automation libraries such as Playwright and Puppeteer execute client-side JavaScript and support page interactions. They suit dynamic sites, persistent sessions, and workflows that require actions such as clicking buttons or completing forms. You must operate the browsers and handle proxy rotation yourself.
Crawling frameworks such as Crawlee coordinate work across many pages. Crawlee manages request queues and concurrency while using an HTTP parser or browser automation library underneath. You still own the runtime infrastructure unless you deploy it through a hosted service.
Managed APIs operate browser, proxy, retry, and anti-bot infrastructure for you. Context.dev combines retrieval and rendering with structured extraction through one API. Managed retrieval fits teams that want clean Markdown or JSON for AI pipelines without maintaining a scraping stack. Playwright or Puppeteer remains a better choice when you need full browser control, persistent logged-in sessions, or custom interactions.
Six criteria help compare these categories fairly. JavaScript rendering determines whether a tool can access client-rendered content. Concurrency and crawl orchestration govern multi-page workloads. Anti-bot and proxy handling determine which infrastructure you must supply. Extraction formats range across raw HTML, Markdown, JSON, and schema-shaped records. Maintenance burden covers browsers, retries, selectors, and deployment. AI and LLM pipeline fit depends on how much cleanup and normalization the output requires.
Decision table: which Node.js scraping tool fits your job
Choose the category that matches how the target site delivers content and how much infrastructure you want to maintain.
| Tool | Category | Best-fit use case | JS rendering | Anti-bot and proxy handling | Output format | Maintenance |
|---|---|---|---|---|---|---|
| Cheerio + Axios | HTTP client and parser | Small, static HTML jobs | No | None | Custom JSON or text | Low |
| Playwright | Browser automation | Interactive sites and stateful workflows | Yes | Self-managed | DOM data, files, or custom JSON | High |
| Puppeteer | Browser automation | Chrome-focused automation | Yes | Self-managed | DOM data, files, or custom JSON | High |
| Crawlee | Crawling framework | Large self-hosted crawls with queues and concurrency | Optional | Self-managed | Custom datasets | Medium to high |
| Context.dev | Managed API | Rendered retrieval and structured AI pipeline data | Managed | Managed | Markdown or JSON | Low |
Playwright and Puppeteer provide direct browser control. Crawlee adds crawl orchestration around HTTP or browser workers. Context.dev handles retrieval infrastructure when persistent sessions and custom browser interactions are not required.
Crawlee
Best for: Crawlee suits developers building large, resilient Node.js crawls who want to control the crawling logic and run their own infrastructure.
What it is: Crawlee is an open-source crawling framework that adds orchestration around HTTP clients and browser automation libraries. Its request queue stores discovered URLs, avoids duplicate work, and preserves crawl state for retries or restarts. Its autoscaled pool adjusts concurrency according to available CPU and memory, which helps prevent a large crawl from overwhelming its host.
Crawlee provides several crawler implementations for different page types. CheerioCrawler fetches and parses static HTML without launching a browser. PlaywrightCrawler and PuppeteerCrawler render JavaScript and support browser interactions. You can use different implementations while keeping Crawlee’s routing, storage, retry, and request-management patterns.
Pros: Crawlee handles operational concerns that raw Playwright or Puppeteer scripts leave to you. Built-in request queues support multi-page discovery, while retry controls help recover from temporary failures. Crawlee also supports proxy configuration and session management, so you can apply consistent policies across a crawl.
Cons: Crawlee remains a self-managed framework. You must supply compute capacity and maintain any browsers or proxies required by the target sites. CAPTCHA handling and browser fingerprint management also remain your responsibility. Running Crawlee on Apify can move hosting, scheduling, and storage into Apify’s cloud, but you still own the crawler’s behavior and extraction code.
Crawlee can also introduce unnecessary complexity for a small static job. Axios with Cheerio usually requires less code when you only need to fetch and parse a few pages.
Pricing: The Crawlee package is free and open source. You pay for the servers, browser resources, proxies, and related services used by your deployment. Apify Cloud offers separate usage-based hosting for Crawlee projects.
Playwright
Best for: Playwright fits Node.js scraping jobs that require JavaScript rendering, browser interactions, or testing across browser engines. Choose it when your scraper must click controls, wait for client-side content, intercept requests, or preserve session state.
What it is: Playwright is an open-source browser automation library for Node.js and other languages. It exposes one API for Chromium and Firefox, and it also supports WebKit. Browser contexts let one process create isolated sessions without launching a separate browser for every session.
Pros: Playwright waits for elements to become actionable before clicking or reading them, which reduces brittle sleep calls in dynamic-page scripts. Its network controls can inspect API responses, modify requests, or block unnecessary assets. Playwright Test adds tracing and debugging tools even when you use the library primarily for scraping. Trace Viewer records browser actions and network activity for later inspection.
Playwright generally offers the broader option in a Puppeteer vs Playwright decision. Its consistent multi-browser API suits projects that need WebKit or Firefox coverage. Puppeteer remains a simpler choice for Chrome-focused scripts and workflows built closely around Chrome DevTools.
Cons: Playwright gives you browser control, but you must operate the scraping infrastructure. You remain responsible for browser deployment, proxy rotation, retries, concurrency limits, and CAPTCHA handling. Browser processes also consume substantially more memory and CPU than Axios paired with Cheerio.
Auto-waiting cannot fix unstable selectors or site-specific interaction logic. Long-running crawlers still need request queues, persistence, and recovery mechanisms. Crawlee can provide that orchestration while using Playwright as its browser layer. A managed retrieval API may require less maintenance when you only need rendered content or structured output.
Pricing: Playwright is free under the Apache 2.0 license. Your costs come from compute, storage, proxies, monitoring, and the engineering work required to keep browser automation running.
Puppeteer
Best for: Chrome-focused scraping that needs direct browser control, custom interactions, persistent sessions, or access to Chrome DevTools Protocol features.
What it is: Puppeteer is an open-source Node.js library for controlling Chrome and Chromium through a high-level API. You can navigate pages, execute JavaScript, click elements, intercept network requests, create PDFs, and extract content after client-side rendering. Its close connection to the Chrome DevTools Protocol gives you detailed control over browser behavior and network activity.
Pros: Puppeteer offers a focused API for Chrome-based workflows and has a mature package ecosystem, extensive documentation, and broad community use. Developers who already depend on Chrome DevTools Protocol commands may find Puppeteer simpler than adding Playwright's cross-browser abstractions. Puppeteer also works well for persistent login sessions, custom page interactions, and Chrome-specific debugging.
Cons: Puppeteer remains primarily Chrome and Chromium focused. Firefox support exists, but Puppeteer does not provide Playwright's coverage across Chromium, Firefox, and WebKit. Its waiting behavior also requires more deliberate handling in some workflows, while Playwright applies broader actionability checks and automatic waiting before interactions.
Puppeteer handles browser automation rather than crawl orchestration. Large crawls still need request queues, concurrency controls, retries, and URL deduplication. Crawlee can add those capabilities around Puppeteer. You must also operate browsers, rotate proxies, manage resource consumption, and handle blocking or CAPTCHA challenges yourself.
Pricing: Puppeteer is free under the Apache 2.0 license. Production costs come from compute, browser infrastructure, proxies, storage, and the engineering work required to keep scraping jobs reliable.
Cheerio (with Axios)
Best for: Cheerio with Axios fits small, high-volume scraping jobs where the server returns complete HTML. Common examples include product listings, article archives, and documentation pages that do not depend on client-side rendering.
What it is: Axios retrieves the page, and Cheerio parses the returned HTML. Cheerio provides a jQuery-like selector API, so you can locate elements with familiar CSS selectors such as $('.product-title'). Node.js developers can replace Axios with the built-in fetch API or Undici without changing Cheerio’s role.
Pros: A Cheerio web scraping stack starts quickly and consumes far less memory than a headless browser. Each request downloads HTML without launching Chromium, which makes higher concurrency practical on modest infrastructure. Cheerio also gives you direct control over selectors, data cleanup, and output formats.
Cons: Cheerio cannot execute JavaScript, render a page, click controls, or wait for client-side requests. If the data appears in the browser’s live DOM but not in the original HTML response, Cheerio alone cannot retrieve it. Axios and Cheerio also leave proxy rotation, retries, rate limiting, anti-bot handling, and crawl orchestration to your code.
Pricing: Cheerio and Axios are free, open-source packages. Your costs come from compute, proxy services, and the engineering work required to operate the scraper. For static pages, those costs can remain low. Dynamic or heavily protected sites usually call for Playwright, Puppeteer, Crawlee, or a managed retrieval API instead.
Context.dev
Best for: Context.dev fits AI engineering teams that need rendered pages or structured web data without maintaining scraping infrastructure. It works well when fast deployment and consistent output matter more than direct browser control.
What it is: Context.dev provides managed retrieval, rendering, and extraction through one API. You can retrieve an individual page or crawl groups of URLs, then receive clean Markdown or schema-shaped JSON instead of raw browser output.
For larger datasets, Map discovers pages through sitemap-driven URL mapping, and Batches processes those URLs as a managed job. MCP integration also lets compatible AI agents retrieve current web content without a separate scraping service between the agent and Context.dev.
Pros:
- Context.dev runs the browser, proxy, retry, and concurrency infrastructure. We also maintain the systems used to handle CAPTCHAs and related anti-bot challenges, though success still varies by target.
- Markdown output removes navigation and other page clutter before content enters an embedding, retrieval, or summarization pipeline.
- Schema-shaped JSON can keep fields consistent when source sites use different page layouts. You still own business rules, validation, and decisions about acceptable data.
- Map and Batches support multi-page dataset jobs without requiring you to build request queues, worker pools, or browser scheduling.
- MCP gives AI agents a direct retrieval path when an agent needs current web content during a task.
Cons: Context.dev offers less browser-level control than Playwright or Puppeteer. Choose those libraries when you need persistent logged-in sessions, detailed interaction scripts, custom browser events, or deep network inspection.
Context.dev also does not replace multi-account browser management or a custom anti-detect browser environment. Crawlee remains a stronger fit when you want an extensible Node.js crawling framework and are prepared to operate its browser and proxy infrastructure.
A managed API also creates an external service dependency and usage-based costs. Self-hosted libraries may cost less for small, predictable jobs when you already have the required infrastructure and engineering capacity.
Pricing: Context.dev uses a managed API pricing model rather than an open-source package model. Review the current Context.dev plans against expected page volume, rendering needs, and extraction workload. Compare the API cost with the engineering and infrastructure expense of operating browsers, proxies, retries, and anti-bot systems yourself.
Comparison table
Pricing alone does not determine fit because each tool handles a different part of the scraping stack.
| Tool | Category | Starting price or model | Best-fit user | Notable strength |
|---|---|---|---|---|
| Context.dev | Managed API | Paid managed service | Teams minimizing infrastructure work | Rendered retrieval with Markdown or JSON output |
| Crawlee | Crawling framework | Free and open source | Developers running large self-managed crawls | Request orchestration across HTTP and browser crawlers |
| Playwright | Browser automation | Free and open source | Developers automating dynamic, interactive sites | Chromium, Firefox, and WebKit support |
| Puppeteer | Browser automation | Free and open source | Developers focused on Chrome workflows | Direct Chrome DevTools Protocol integration |
| Cheerio with Axios | Parser and HTTP client | Free and open source | Developers scraping static HTML | Fast fetching and jQuery-style parsing |
Node.js scraping examples by scenario
Each example extracts product titles and prices. Use Axios and Cheerio when the server returns complete HTML, Playwright when JavaScript creates the content, Crawlee when you need crawl orchestration, and Context.dev when you want managed retrieval.
Static HTML with Axios and Cheerio
Axios retrieves the page, and Cheerio parses its HTML without launching a browser.
npm install axios cheerio// static.mjs
import axios from 'axios'
import * as cheerio from 'cheerio'
const url = 'https://books.toscrape.com/'
const { data: html } = await axios.get(url)
const $ = cheerio.load(html)
const products = $('.product_pod').map((_, el) => ({
title: $(el).find('h3 a').attr('title'),
price: $(el).find('.price_color').text().trim()
})).get()
console.log(products)node static.mjsThis approach needs no browser infrastructure, but Cheerio cannot execute page JavaScript.
JavaScript rendering with Playwright
Playwright launches Chromium and reads the DOM after the page loads. The example uses the same page so you can compare the extraction code directly.
npm install playwright
npx playwright install chromium// rendered.mjs
import { chromium } from 'playwright'
const browser = await chromium.launch()
const page = await browser.newPage()
await page.goto('https://books.toscrape.com/', {
waitUntil: 'networkidle'
})
const products = await page.locator('.product_pod').evaluateAll(cards =>
cards.map(card => ({
title: card.querySelector('h3 a')?.getAttribute('title'),
price: card.querySelector('.price_color')?.textContent?.trim()
}))
)
console.log(products)
await browser.close()node rendered.mjsPlaywright handles rendered content and interactions, but you still own browser deployment, retries, proxies, and anti-bot handling.
Multi-page crawling with Crawlee
Crawlee adds request scheduling, concurrency, retries, and link discovery around the extraction code.
npm install crawlee// crawl.mjs
import { CheerioCrawler } from 'crawlee'
const products = []
const crawler = new CheerioCrawler({
maxRequestsPerCrawl: 3,
async requestHandler({ $, enqueueLinks }) {
$('.product_pod').each((_, el) => {
products.push({
title: $(el).find('h3 a').attr('title'),
price: $(el).find('.price_color').text().trim()
})
})
await enqueueLinks({ selector: '.next a' })
}
})
await crawler.run(['https://books.toscrape.com/'])
console.log(products)node crawl.mjsCheerioCrawler suits static pages. You can replace it with PlaywrightCrawler when each page requires JavaScript rendering, although that adds browser infrastructure.
Managed retrieval with Context.dev
Context.dev can return schema-shaped JSON while managing rendering and retrieval infrastructure. Set the endpoint and key supplied by your account, and confirm the current authentication header in the API documentation.
npm install --save-dev tsx @types/node// managed.ts
const endpoint = process.env.CONTEXT_API_ENDPOINT
const apiKey = process.env.CONTEXT_API_KEY
if (!endpoint || !apiKey) {
throw new Error('Set CONTEXT_API_ENDPOINT and CONTEXT_API_KEY')
}
const response = await fetch(endpoint, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${apiKey}`
},
body: JSON.stringify({
url: 'https://books.toscrape.com/',
output: 'json',
schema: {
type: 'object',
properties: {
products: {
type: 'array',
items: {
type: 'object',
properties: {
title: { type: 'string' },
price: { type: 'string' }
},
required: ['title', 'price']
}
}
},
required: ['products']
}
})
})
if (!response.ok) {
throw new Error(`Request failed with ${response.status}`)
}
console.log(await response.json())CONTEXT_API_ENDPOINT="YOUR_ENDPOINT" \
CONTEXT_API_KEY="YOUR_KEY" \
npx tsx managed.tsThe managed approach reduces browser, proxy, retry, and extraction maintenance. Playwright remains the better choice when you need persistent sessions or custom interactions.
Production tradeoffs: anti-bot handling, concurrency, and maintenance
A self-managed scraper requires ongoing work once target sites begin limiting automated traffic. Playwright and Puppeteer give you browser control, but you must operate proxy pools and maintain credible browser fingerprints. Proxy rotation must preserve session affinity while replacing blocked addresses. Browser updates can also change fingerprint signals and break configurations that previously worked.
CAPTCHAs and Cloudflare Turnstile add another failure path. Your code must detect a challenge before routing the request through an approved handling path. Repeated challenges should stop retries because aggressive retry loops can increase blocking and consume proxy or compute capacity without producing data.
Concurrency requires tuning for both your infrastructure and each target site. Crawlee provides request queues, retry controls, and autoscaling, but you still choose rate limits and concurrency policies. A setting that works for a static catalog may overwhelm a slower site or trigger defenses on a protected one. Browser crawls also consume far more memory than HTTP requests, so worker counts must account for page crashes and long-running sessions.
Maintenance continues after a crawler reaches production. You must monitor proxy health, browser crashes, blocked responses, and extraction failures. Site changes can require new selectors or interaction logic, while browser and dependency upgrades can introduce separate regressions.
Managed APIs transfer browser operations, proxy routing, challenge handling, and retry infrastructure to the provider. You give up some control over browser configuration and request execution in exchange for less infrastructure ownership. Playwright or Puppeteer remains the better fit for persistent sessions and custom interactions. Crawlee remains useful when you want to control crawl orchestration yourself. A managed API becomes attractive when infrastructure maintenance consumes more engineering time than the custom browser behavior delivers.
Which Node.js scraper fits AI and LLM data pipelines
AI-ready scraping produces clean Markdown or JSON that downstream models can consume without page-specific cleanup. For structured extraction, each record should follow a defined schema even when source sites use different HTML layouts. Consistent fields help embedding pipelines, retrieval systems, and agents process records without guessing which selector or property contains the needed value.
Playwright and Puppeteer give you rendered pages, DOM access, and full interaction control, but you must write the extraction and normalization logic. Cheerio offers efficient parsing for static HTML and requires the same page-specific selectors. Crawlee manages queues, retries, and concurrency across larger crawls, but its crawlers still need custom code to turn inconsistent pages into a shared Markdown or JSON structure.
Context.dev fits pipelines that need managed retrieval, rendering, and schema-shaped extraction. Its API can return clean Markdown or JSON without requiring you to operate browsers, proxies, CAPTCHA handling, or selector registries. MCP integration also lets compatible agents retrieve web content through a standard tool interface.
Managed extraction reduces glue code, but it does not replace customer-specific validation. You still need to check required fields, enforce business rules, and handle uncertain source data. Choose Playwright, Puppeteer, or Crawlee when your pipeline depends on persistent sessions or custom interactions. Choose Context.dev when consistent model-ready output and low infrastructure ownership matter more than direct browser control.
FAQs
What are the main differences between Puppeteer and Playwright?
Playwright supports Chromium, Firefox, and WebKit, and its auto-waiting features help with dynamic interfaces. Puppeteer offers a simpler, Chrome-focused API with close DevTools Protocol integration. Choose Playwright for cross-browser coverage and complex interactions. Choose Puppeteer when your workflow targets Chrome or Chromium and direct browser control matters.
When is Cheerio enough for web scraping?
Cheerio works well when the server returns the required content in the initial HTML. Pair it with Axios or Node.js fetch to retrieve pages, then extract fields with CSS selectors. Use a browser when JavaScript loads the content, user interactions reveal data, or the site requires browser state.
Does Crawlee replace Playwright or Puppeteer?
Crawlee builds on Playwright, Puppeteer, or HTTP-based crawlers rather than replacing their core capabilities. It adds request queues, concurrency controls, retries, session management, and crawl orchestration. Raw Playwright or Puppeteer scripts suit focused browser tasks, while Crawlee suits multi-page crawls that need repeatable scheduling and failure handling.
Should I self-host scraping or use a managed API?
Self-hosting gives you direct control over browser behavior, sessions, proxies, and extraction logic. It also makes you responsible for browser deployment, retries, proxy rotation, CAPTCHA handling, monitoring, and scaling. A managed API such as Context.dev fits teams that prefer clean Markdown or JSON without maintaining that infrastructure. Self-hosted tools remain better for persistent logins, custom interaction sequences, and multi-account browser workflows.
How can I scrape JavaScript-rendered pages without maintaining browsers?
Use a managed rendering API that runs browsers and retrieval infrastructure on your behalf. Context.dev accepts a page or crawl request and returns rendered content or structured output through one API. Playwright and Puppeteer provide more interaction control, but you must operate the browser, proxy, retry, and anti-bot layers yourself.
Conclusion
Choose the tool category that matches the job. Axios and Cheerio suit static HTML. Playwright or Puppeteer handle interactive browser workflows. Crawlee adds orchestration for self-managed, multi-page crawls. A managed API fits when you want rendered or structured data without operating browsers, proxies, retries, and anti-bot infrastructure.
After choosing a category, decide how much control and maintenance you want to own. Browser libraries give you deeper control over sessions and interactions, while managed retrieval reduces infrastructure work. If the managed path fits your pipeline, try Context.dev for scraping, crawling, and clean Markdown or JSON through one API.