Agents are making more web searches and fetching more pages per task. As that grows, ensuring the web data entering their context is relevant becomes essential for long-running performance, cost, and task latency.
Today, we’re introducing Context Highlights. Given a webpage and a query, it returns relevant passages, substantially reducing the number of tokens an agent needs to process.
Keeping the right context
Extracting relevant text is harder than finding sentences that match a query. A code example might have an exception explained further down the page, and a sentence may depend on the paragraphs around it.
The query is also an imperfect description of the agent’s intent. Since we don’t receive the full conversation that led to the query, we aim to preserve enough surrounding context for the agent to interpret each passage and decide whether it needs more of the page.
Evaluating Highlights
We adapted Exa’s WebCode highlights benchmark to compare six tools on the same 250 URL-and-question pairs.
We measured groundedness: the percentage of cases where the returned text supports every substantive fact in the reference answer. We also measured the number of tokens returned.
View WebCode results as a table
| Tool | Groundedness | Average output tokens |
|---|---|---|
| Context Highlights | 76.0% | 229 |
| Exa Highlights | 75.6% | 286 |
| Parallel Extract | 70.0% | 482 |
| Codex native open | 57.6% | 6,889 |
| Claude Code WebFetch | 50.4% | 231 |
| Firecrawl Highlights | 48.8% | 52 |
Shorter output isn’t automatically better. The useful comparison is how much evidence each tool preserved and how many tokens it returned.
Among the failed cases, Firecrawl returned too little detail to fully support the reference answer, while Claude Code’s WebFetch summaries left out relevant information. Our Codex baseline made a single page-open request, which could miss relevant passages further down the page; an agent could retrieve those with additional calls.
Context and Exa achieved similar groundedness, while Context returned approximately 20% fewer tokens on average. Retrieval failures count against groundedness but are excluded from token averages. See the appendix for methodology and shared-subset results.
SimpleQA Verified
We also compared Context with zilliz/semantic-highlight-bilingual-v1 on all 1,000 questions in SimpleQA Verified. Both extractors received the same source text, then GPT-4.1 mini answered each question using their highlights, with GPT-4.1 judging the answers. Here, we measured answer accuracy: the percentage of questions answered correctly.
View SimpleQA Verified results as a table
| Extractor | Answer accuracy | Average output tokens |
|---|---|---|
| Context Highlights | 76.7% | 183 |
| Zilliz, threshold 0.1 | 74.0% | 577 |
| Zilliz, threshold 0.5 | 60.0% | 133 |
Context achieved 2.7 percentage points higher accuracy than Zilliz at threshold 0.1 while returning 68% fewer tokens. At threshold 0.5, Zilliz returned fewer tokens than Context, but accuracy fell to 60.0%.
Try Highlights
Use Highlights in Claude Code or Codex, or build it into your own agent with the Context API.
What’s next
You can already combine Highlights with Context Search: search for relevant pages, then request Highlights from the URLs you want to read.
Next, we’re bringing Highlights directly into search and research workflows. Longer term, we’re interested in extraction that understands more of the agent’s actual context: what it’s working on, what it already knows, and what it still needs to find.
Appendix
WebCode
We ran all 250 pairs from Exa’s original WebCode benchmark on September 25, 2026. A provider-blind gpt-6-luna judge with medium reasoning checked whether each response supported every reference fact, and we counted tokens with cl100k_base. Failed requests and error pages count against groundedness but are excluded from token averages; valid responses are included even when they don’t support the answer.
Context used live Markdown and our local extractor at commit d956454c1a62, with Jev 1.13.0 and a 4,000-character output setting, rather than the deployed endpoint. We passed the question to Exa Highlights, Parallel Extract, Firecrawl Highlights, and Claude Code 2.1.280 WebFetch, using the hosted APIs’ unpinned versions and requesting fresh content where supported. Codex used gpt-6-sol with a single native open and a long response, without a question or follow-up calls; its token count includes line and citation markers.
In a separate source audit, we requested source text alongside highlights from the four API providers. On 173 cases, the judge found that all four returned sources supported the full reference answer. The highlights from those same responses scored as follows:
| Tool | Groundedness | Average output tokens |
|---|---|---|
| Context Highlights | 96.5% | 209 |
| Exa Highlights | 96.0% | 281 |
| Parallel Extract | 91.9% | 486 |
| Firecrawl Highlights | 65.9% | 61 |
Codex and Claude are excluded because their recorded calls didn’t expose separate source text before selection or summarization. This audit establishes what the APIs returned, not exactly what their models saw.
We think link rot and the different judge model explain much of the gap from Exa’s published results, though we haven’t measured their effects separately. We’ve used WebCode during development, and the small groundedness difference between Context and Exa doesn’t establish a quality winner.
SimpleQA Verified
We ran all 1,000 questions from SimpleQA Verified on September 25, 2026, selecting each question’s first reference URL before fetching it. Context used the production Scrape API’s 4,000-character Highlights default; Zilliz used revision 6dfd9cbee6d93 with its released chunking and thresholds of 0.1 and 0.5, fixed before the run. Both received the same Markdown and title.
We used gpt-4.1-mini-2025-04-14 to answer and gpt-4.1-2025-04-14 to judge, both at temperature 0, with the Verified paper’s grading prompt. We counted tokens with o200k_base and included all 1,000 cases in both metrics, including 92 source failures with zero output tokens. This measures answers produced with supplied highlights, not closed-book SimpleQA performance or a guarantee that every answer fact appears in the highlights.