---
name: context-dev
description: Use Context.dev to scrape websites into Markdown, HTML, screenshots, images, or CSS-selected fields, map a website's URLs, crawl linked pages, retrieve company and brand profiles with industry labels, or monitor website changes. Apply when a task needs Context.dev API or SDK selection, request construction, error handling, or production integration guidance.
license: MIT
metadata:
  author: context.dev
  version: "5.2"
  last_verified: "2026-09-26"
---

# Context.dev integration guide

Context.dev provides web scraping, URL mapping, crawling, web search and research, and brand data through `https://api.context.dev/v1`. Use it to turn a webpage into Markdown, rendered HTML, a screenshot, an image list, original bytes, or CSS-selected fields in one request, and retrieve company profiles with logos, colors, descriptions, and social links through the same API. Use the public OpenAPI document at [docs.context.dev/openapi.json](https://docs.context.dev/openapi.json) as the authority for paths, methods, parameters, and response fields.

## Get access

- To call Context.dev from this agent session without writing code, connect the MCP server `https://mcp.context.dev/mcp`. It signs in with OAuth in the user's browser and does not use an API key. Setup for each client: https://docs.context.dev/install-mcp.md
- Application code and the CLI read an API key from `CONTEXT_DEV_API_KEY`. If none is configured, follow https://www.context.dev/auth.md: register with the user's email, deliver the returned setup link and code to the user before polling, and store the key in an ignored environment file or secret manager without displaying it. Never ask the user to paste a key into chat.
- Complete setup instructions: https://docs.context.dev/agent-quickstart.md

## Keep credentials server-side

Read the bearer token from `CONTEXT_DEV_API_KEY`. Never print it, commit it, include it in browser code, or forward it to a target website.

```bash
export CONTEXT_DEV_API_KEY="ctxt_secret_..."
```

For raw HTTPS calls:

```bash
curl https://api.context.dev/v1/web/scrape \
  -H "Authorization: Bearer $CONTEXT_DEV_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "formats": { "markdown": true },
    "sharedParams": { "mainContentOnly": true }
  }'
```

## Route from input to operation

Choose the narrowest operation that directly returns the needed result.

| Input and desired result | Operation | Task guide |
| --- | --- | --- |
| One URL to Markdown or rendered HTML | `POST /web/scrape` with `formats.markdown` or `formats.html` | [Scrape a webpage](https://docs.context.dev/guides/scrape-websites-to-markdown) |
| Known CSS selectors on one page to JSON | `POST /web/scrape` with `formats.parse` and `parseParams.rules` | [CSS extraction](https://docs.context.dev/guides/scrape-websites-to-markdown#extract-structured-fields-with-css) |
| Exact URL to an inline PNG, JPEG, or WebP screenshot | `POST /web/scrape` with `formats.screenshot` | [Take a screenshot](https://docs.context.dev/guides/take-webpage-screenshot) |
| URL to image assets | `POST /web/scrape` with `formats.images` | [Extract page images](https://docs.context.dev/guides/extract-page-images) |
| Resource URL to complete base64 bytes | `POST /web/scrape` with `formats.bytes` | [Download bytes](https://docs.context.dev/guides/download-resource-bytes) |
| Product URL to normalized product record with price, availability, images, and variants | `POST /web/scrape` with `formats.product` | [Extract product data](https://docs.context.dev/guides/scrape-websites-to-markdown#extract-product-data) |
| Page that needs clicks, scrolls, or waits before capture | `POST /web/scrape` with `sharedParams.actions` | [Browser actions](https://docs.context.dev/guides/browser-actions) |
| Domain to its URL inventory with page titles and descriptions | `GET /web/urls` | [Map website URLs](https://docs.context.dev/guides/discover-website-urls) |
| Starting URL to linked pages in one response, up to 500 pages | `POST /web/crawl` | [Crawl Sync](https://docs.context.dev/guides/crawl-website) |
| Starting URL or sitemap to a background crawl, up to 25,000 pages | `POST /batch/submit` with `input.mode: "crawl"` | [Crawl Async](https://docs.context.dev/guides/crawl-website-async) |
| URL list to Markdown or HTML in the background, up to 25,000 URLs | `POST /batch/submit` | [Scrape in batches](https://docs.context.dev/guides/scrape-websites-in-batches) |
| Search query to ranked web results, optionally with Markdown | `POST /web/search` | [Search the Web](https://docs.context.dev/api-reference/web-scraping/search) |
| Research task to structured JSON and source URLs | `POST /web/answers` | [Answers](https://docs.context.dev/guides/research-web-with-answers) |
| Uploaded document up to 50 MiB to Markdown | `POST /parse` | [Parse documents](https://docs.context.dev/guides/parse-documents) |
| Domain, name, work email, ticker, direct URL, or transaction descriptor to company profile | `POST /brand/retrieve` | [Retrieve brand data](https://docs.context.dev/guides/get-brand-data) |
| Partial name or domain to matching indexed brands | `GET /brand/search` | [Search Brands](https://docs.context.dev/api-reference/brand-intelligence/search) |
| Domain or exact URL to design styles, font families, and available font files | `GET /web/styleguide`; read `styleguide.typography` and `styleguide.fontLinks` for fonts | [Extract a design system](https://docs.context.dev/guides/extract-design-system-from-website) |
| Email, social profile URL, or name plus company to a person profile | `POST /people/enrich` | [Enrich People](https://docs.context.dev/api-reference/people/enrich) |
| Company name, domain, or ticker to recent news | `POST /news/search` | [Search News](https://docs.context.dev/api-reference/news/search) |
| Domain or work email known before a Brand or Styleguide request | `POST /utility/prefetch` | [Prefetching](https://docs.context.dev/optimization/prefetching) |
| URL or site to recurring change detection | `/monitors` operations | [Monitor website changes](https://docs.context.dev/guides/monitor-website-changes) |
| Context.dev bug, docs mismatch, or friction you hit while integrating | `POST /feedback` | [Agent Feedback](https://docs.context.dev/optimization/agent-feedback) |

Do not use a general scrape when a purpose-built Brand or monitor operation already returns the required shape.

## Verify the installed SDK before using it

Published packages:

| Language | Package | Client setup | Scrape and Map URLs |
| --- | --- | --- | --- |
| TypeScript | `context.dev` | `new ContextDev({ apiKey: process.env.CONTEXT_DEV_API_KEY })` | `client.web.scrape`, `client.web.mapUrls` |
| Python | `context.dev` | `ContextDev(api_key=os.environ["CONTEXT_DEV_API_KEY"])` | `client.web.scrape`, `client.web.map_urls` |
| Ruby | `context.dev` | `ContextDev::Client.new(api_key: ENV.fetch("CONTEXT_DEV_API_KEY"))` | `client.web.scrape`, `client.web.map_urls` |
| Go | `github.com/context-dot-dev/context-go-sdk/v2` | Use the `/v2` module; inspect its generated request model. | `client.Web.Scrape`, `client.Web.MapURLs` |
| PHP | `context-dev/context-dev-php` | Inspect the installed package version and generated method signature. | `$client->web->scrape`, `$client->web->mapUrls` |

Go uses the `/v2` module path for `Web.Scrape`, `Web.MapURLs`, and the current Brand request body. If the installed PHP Brand helper cannot express a single lookup type, use the SDK's low-level request method for that operation. Do not invent generated types or pass fields that the installed signature does not accept.

For installation and runnable examples, read [SDKs](https://docs.context.dev/sdks). For authentication and example responses, read the [Quickstart](https://docs.context.dev/quickstart). Check each operation's reference page for its full response schema.

## Scrape a webpage

```typescript
import ContextDev from "context.dev";

const client = new ContextDev({
  apiKey: process.env.CONTEXT_DEV_API_KEY,
});

const page = await client.web.scrape({
  url: "https://example.com",
  formats: { markdown: true },
  sharedParams: { mainContentOnly: true },
});

if (!page.markdown.data?.trim()) {
  // Handle an empty selector or page result before indexing it.
}
```

Use these controls deliberately:

- `formats`: enable at least one of `html`, `markdown`, `screenshot`, `images`, `bytes`, `parse`, `highlights`, `json`, or `product`. Outputs share one page visit. Each output has `requested`, `success`, and `data`; `success` is `true` when retrieved, `false` when retrieval fails, and `null` when not requested. Failed outputs have `data: null` and do not discard successful outputs. The base price is 1 credit including cache hits, or 2 with `sharedParams.actions`. Highlights add 3 credits when passages are returned, JSON adds 4 when extraction succeeds and its result is returned, and product adds 1 when its successful result is returned or the target page is missing. A successful returned product model verdict adds 6 more credits. PDF OCR adds 1 credit per recovered page on fresh extraction. All-failed responses are unbilled except missing pages, which retain the base price and the product charge when requested.
- `maxAgeMs`: chooses acceptable cache age. Scrape and Crawl default to one day and accept up to 30 days; `0` fetches fresh and refreshes the requested outputs. Each Scrape output has its own cache key, so one response can combine cached outputs from different visits. Brand and Styleguide default to three months, accept `0` for a hard refresh, and clamp values above one year. The direct-URL Brand variant does not accept this control.
- `timeoutOpts`: set `milliseconds` up to 300,000 and choose `behavior: "fail"` or `"return-partial"` where supported. Reaching the overall deadline with `"fail"` returns an unbilled `408`. Scrape requires at least 5,000 ms for `return-partial` and marks partial captures or failed outputs with `isPartial: true`; individual outputs can fail under either behavior while others succeed. Failed retrievals and incomplete captures are not cached, but valid captured pieces may be cached independently. Crawl and Map URLs return `partial: true`. Prefetch is fail-only and Parse does not accept `timeoutOpts`. See [timeouts](https://docs.context.dev/optimization/timeouts).
- `maxPages`, `maxDepth`, `urlRegex`, and `stopAfterMs`: bound crawl coverage, time, and credit exposure. Crawl caps `maxPages` at 500.
- `sharedParams.mainContentOnly`, `includeSelectors`, `excludeSelectors`, and `includeFrames`: control page content before conversion. Exclusions win. Filters apply to HTML, Markdown, images, and parsed fields, never to `bytes`.
- `formats.parse` with `parseParams.rules`: map each field name to a CSS selector string or `{selector, type: "item" | "list", output: "text" | "html" | "@attribute" | nested rules}`. Results arrive in `parsed.data`; missing items return `null` and missing lists return `[]`. No LLM and no extra credits. See [CSS extraction](https://docs.context.dev/guides/scrape-websites-to-markdown#extract-structured-fields-with-css).
- `sharedParams.waitFor` (milliseconds or a CSS selector; default 500 ms), `settleAnimations`, and [browser actions](https://docs.context.dev/guides/browser-actions) (`type: "perform" | "scroll" | "wait" | "waitFor"`, 1 to 5 per request): handle dynamic page state. Actions require a paid plan and bypass the cache. Action or selector failures can leave affected outputs with `success: false`; verify the page state from successful captured outputs.
- `formats.bytes` on [Bytes](https://docs.context.dev/guides/download-resource-bytes): returns the original HTTP body as `{contentType, base64}` after decompression, up to 20 MiB decoded. Waiting, actions, and content filters never change it; request `formats.html` for rendered HTML.
- `sharedParams.headers` are separate from the API bearer key and go to the target origin only. Custom headers, actions, and `zdr` bypass cache reads and writes. Otherwise inspect `cache_metadata.status` (`hit`, `miss`, or `zdr`) and `age_ms`. See [target headers](https://docs.context.dev/guides/take-webpage-screenshot#send-target-headers).
- `screenshot.data` is an inline data URL (`data:image/png;base64,...`) when `screenshot.success` is `true`; otherwise it is `null`. Use a successful result directly as an image source, or decode the part after the comma to save the file. Choose the capture area with `screenshotParams.area` (`viewport`, `fullPage`, `{selector}`, or a rectangle) up to 40 megapixels.
- `sharedParams.country` on [Images](https://docs.context.dev/guides/extract-page-images#route-through-a-country) and every other output: a two-letter country code that controls the network exit for the page fetch and image downloads. An unsupported code returns `400` before the scrape starts.
- `zdr: "enabled"`: bypasses shared caches and retained content logs only when Zero Data Retention is enabled for the organization; otherwise the request fails with `403`.

Browser rendering and proxy routing improve access but do not guarantee that a page can be read. Preserve a fallback for `WEBSITE_BLOCKED`, login walls, missing content, and partial crawl results.

## Map a website's URLs

`GET /web/urls` maps all URLs a website has using Context.dev's index and adds stored `title`, `description`, `keywords`, and `language` where available. Pass `domain` without a protocol; narrow with `urlRegex`, `maxLinks` (default 10,000), `includeSubdomains`, or `search` (relevance-ordered, 2 credits instead of 1).

```typescript
const site = await client.web.mapUrls({
  domain: "stripe.com",
  maxLinks: 50,
  urlRegex: "/customers/",
});

for (const entry of site.urls) {
  // entry.title and entry.description are absent until the URL has been enriched.
}
```

URLs without stored metadata return with only `url` and are queued for background enrichment, so a later request can include their metadata. Requests with `zdr=enabled` or credential-bearing target headers return URLs only: they neither read nor store shared metadata and do not queue enrichment. Responses are never cached as a whole. `partial: true` means the deadline stopped URL mapping. Feed the result into Scrape, Crawl, or a batch.

## Retrieve brand data

Send exactly one lookup variant:

| `type` | Required field | Important constraint |
| --- | --- | --- |
| `by_domain` | `domain` | Prefer a bare company domain. |
| `by_name` | `name` | 3 to 30 characters; use `country_gl` as an ambiguity hint. |
| `by_email` | `email` | Free and disposable providers return `422`. |
| `by_ticker` | `ticker` | Add `ticker_exchange` when known. |
| `by_direct_url` | `direct_url` | Reads only that page; no wider resolver or cross-source enrichment. |
| `by_transaction` | `transaction_info` | Add MCC, city, country, or phone hints when available. |

```typescript
const response = await client.brand.retrieve({
  type: "by_domain",
  domain: "stripe.com",
});

const title = response.brand?.title ?? "Unknown company";
```

Brand arrays and nested fields can be missing or empty. Select logos by `type`, `mode`, and resolution; do not assume `logos[0]` is appropriate. Treat address, contacts, employee count, classifications, and stock data as discovered data rather than verified legal records.

## Permissions and completeness

Restricted API keys need the relevant [scope](https://docs.context.dev/guides/manage-api-keys): `data:execute` for direct data calls, Read/Manage scopes for monitors and batches, and `logs:read` for request logs. Any key can submit Agent Feedback. Dashboard [team roles](https://docs.context.dev/guides/manage-team-access) are independent; Members can use shared credentials and incur charges.

Answers defaults to Ultra (100 credits); Fast costs 10. `json_format` is an example object; source URLs are not per-field citations. Scrape, Map URLs, Crawl, Search, Answers, Parse, People Enrich, and Styleguide honor `zdr` when the organization is entitled. Scrape uses a base price of 1 credit, or 2 with actions, plus applicable format charges, returns original bytes up to 20 MiB, and supports ZDR. Map URLs costs 1 credit, or 2 with `search`. Parse accepts uploads up to 50 MiB.

Page monitors accept `include_selectors`/`exclude_selectors`; changing them establishes a new baseline.

## Handle errors by category

Inspect both the HTTP status and `error_code` for request errors. Scrape also reports output failures inside `200` responses: inspect each output's `success` and preserve `isPartial: true`. Use raw HTTPS or the SDK's raw response if an older generated model does not expose these fields.

Processed `404` results on priced data APIs consume the normal request credits. Validation, authentication, rate-limit, timeout, and server errors remain unbilled; successful partial results are billable. Record `key_metadata.credits_consumed` or `X-Credits-Used` instead of inferring cost from status alone. See [Credits on errors](https://docs.context.dev/optimization/troubleshooting#credits-on-errors).

| Status | Treat as | Action |
| --- | --- | --- |
| `400` | Invalid request, inaccessible or blocked target (`WEBSITE_ACCESS_ERROR`, `WEBSITE_BLOCKED`), image-only PDF (`PDF_IMAGES_ONLY`), skipped PDF (`PDF_SKIPPED`), or no match depending on `error_code` | Fix the input or options. For `PDF_IMAGES_ONLY`, set `sharedParams.parsers.pdf.ocr` to `"auto"`. Use a fallback for a no-match or inaccessible site. |
| `401` | Credential or disabled-access failure | Inspect `error_code`; fix credentials or check the account/key state. |
| `403` | Permission, plan, ZDR entitlement, disabled key, or usage failure | Do not retry unchanged; inspect `error_code`. |
| `404` | Target or entity not found on operations that define it | Treat as an expected empty outcome where appropriate. |
| `408` | Timeout; `REQUEST_TIMEOUT` is unbilled | Retry outside a user-facing path, raise the budget, use `return-partial` where supported, or prefetch supported Brand and Styleguide requests. |
| `200` with a failed Scrape output | Retrieval, parsing, actions, selectors, or output limits prevented that output from completing | Preserve successful outputs and fix the target or options before retrying the failed output. Oversized outputs have `success: false` and `data: null`. |
| `413` | Content too large on an operation that defines this error, such as a Parse upload over 50 MiB | Use a smaller input. Scrape marks oversized outputs as failed instead. |
| `415` | Unsupported content on an operation that defines this error | Choose a supported parser or source. Scrape marks unsupported outputs as failed instead. |
| `422` | An operation-specific input restriction, such as an invalid email class or a Brand timeout that is too low | Change the input or options; do not retry unchanged. |
| `429` | Rate limit | Honor `Retry-After`; retry with jittered, bounded backoff. |
| `500` | Transient service failure | Retry with jittered, bounded backoff, then surface a fallback. |
| `502` | The browser could not return a complete capture | Retry with bounded backoff, then surface a fallback. |
| `503` | Browser capacity is temporarily unavailable | Retry with bounded backoff after a delay. |

Do not retry validation, permission, no-match, content-size, or unsupported-media failures unchanged. Official SDKs may perform their own limited retries; inspect the installed version before adding another retry layer.

See [Troubleshooting](https://docs.context.dev/optimization/troubleshooting) and [Rate limits](https://docs.context.dev/optimization/rate-limits) for operation-specific behavior.

## Report problems with Agent Feedback

When an endpoint, docs page, SDK, or CLI behaves differently than documented, offer to report it with [Agent Feedback](https://docs.context.dev/optimization/agent-feedback). It costs 0 credits, works with any API key, and has its own rate limit. Send one report per problem with the affected `request_id` so the team can see the exact request, and keep secrets and personal data out of the note.

```bash
curl https://api.context.dev/v1/feedback \
  -H "Authorization: Bearer $CONTEXT_DEV_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "request_id": "<request_id from the affected response>",
    "category": "docs_mismatch",
    "note": "The response is missing a documented field; expected it per the API reference."
  }'
```

Send `url` instead of, or with, `request_id` for a docs page or one page of a crawl. Categories are `bug`, `docs_mismatch`, `friction`, `feature_gap`, `quality_degradation`, and `other`. Reporting the same `request_id` again returns the original `feedback_id` with `already_submitted: true`.

## Completion checklist

Before declaring an integration complete:

- Confirm the path, method, and parameter names against the current API reference.
- Request only the `formats` the task needs, and read each output's `data` rather than assuming it is present.
- Keep the API key on the server and prove the missing-key path is intentional.
- Bound crawl size, latency, retries, and credit exposure.
- Handle missing fields and expected no-result states.
- Validate parsed and structured output and preserve provenance when needed.
- Run a focused success test and at least one relevant failure test.

Use [API stability](https://docs.context.dev/optimization/api-stability) and the [changelog](https://docs.context.dev/changelog) for long-lived integrations.
