Your AI agent guesses a URL, fetches it, gets a 404, and reports success anyway. That quiet failure mode is endemic to production agentic apps — models hallucinate plausible-looking URLs because their training data stopped updating at some point. Cloudflare just shipped the direct fix. As of October 2, 2026, the Cloudflare Web Search API is in open beta inside AI Gateway, giving agents a real search interface instead of a guessing game.
The Problem With How Agents Access the Web Today
Without a dedicated search tool, an agent that needs live information typically fabricates a URL based on patterns in its training data, then tries to fetch it directly. When the guess is wrong — and it frequently is — the request returns a 404. Production failure analysis found that agents often cannot distinguish between task failure and task impossibility, and hallucinate success back to the user just to close the loop.
Better prompting doesn’t fix this. You need grounding: a mechanism that retrieves verified, current information the model can cite rather than invent.
What the Cloudflare Web Search API Does
The Web Search API slots into Cloudflare AI Gateway as a dynamic context layer. Your agent sends a search query; the API returns structured results — title, URL, description, and optional metadata — from one of three search providers. You inject those results into your model’s context. The model answers from real pages, not training memory.
It runs through AI Gateway, so every search request appears in the same logs, analytics, and billing dashboard as your model inference calls. One auth token, one cost center.
The Workers binding is three lines:
const results = await env.AI.websearch({
gatewayId: "default",
query: userQuery,
provider: "ceramic",
limit: 5
});
The REST API endpoint is POST /accounts/{account_id}/ai/websearch/ with a Bearer token if you’re not on Workers. Parameters are minimal: query (up to 1,024 chars), provider, limit (1–10), and an optional BYOK alias if you’re using your own provider API key.
Three Providers — and Why the Price Gap Is 28x
Three providers are available at launch. The pricing spread is significant enough to determine your architecture:
| Provider | Price / 1K Requests | Type | Zero Retention |
|---|---|---|---|
| Ceramic.ai | $0.25 | Keyword / lexical | Yes |
| Linkup | $5.00 | Web results | Yes |
| Exa | $7.00 | Neural / semantic | No (at launch) |
Ceramic.ai is the default pick for most workloads. It indexes 40 billion+ pages, returns descriptions up to 8,000 characters, and has median latency under 250ms. For agents doing standard informational queries — news lookups, documentation checks, product information — keyword search at $0.25 per thousand is hard to beat.
Exa earns its $7 premium on queries where semantic matching matters. Its neural embedding index finds pages expressing the same concepts even when exact keywords don’t match. Use Exa when your queries look like “find arguments against X” or “papers similar to this abstract” — things lexical search handles poorly. One caveat: Exa doesn’t offer zero-data retention through Cloudflare at launch. If your agents process customer PII or proprietary queries, verify this before shipping.
Linkup sits in the middle and works for teams that need more than keyword search but can’t justify Exa’s pricing.
The math at scale: 100,000 monthly queries costs $25 with Ceramic or $700 with Exa. Know what you’re buying before you default to the flashiest option.
The Full Worker: Search, Inject, Answer
Here’s a complete pattern for grounding a Cloudflare Worker agent with live search results:
export default {
async fetch(request, env) {
const { query } = await request.json();
// 1. Search the live web
const search = await env.AI.websearch({
gatewayId: "default",
query,
provider: "ceramic",
limit: 5
});
// 2. Format results as context
const context = search.results
.map(r => `[${r.title}](${r.url}): ${r.description}`)
.join("\n\n");
// 3. Run model with grounded context
const response = await env.AI.run(
"@cf/meta/llama-4-scout-17b-16e-instruct",
{
messages: [
{
role: "system",
content: "Answer using only the search results provided. Cite sources."
},
{
role: "user",
content: `Search results:\n\n${context}\n\nQuestion: ${query}`
}
]
},
{ gateway: { id: "default" } }
);
return Response.json({ answer: response.response });
}
};
Swap ceramic for exa to switch providers with no other changes. Swap the model slug for any Workers AI-supported model. The search layer is fully decoupled from the inference layer — which is the point.
Why Cloudflare Beats Model-Native Search on Cost
OpenAI’s web_search tool and Anthropic’s equivalent cost $10 per thousand searches, plus the token overhead from results injected into model context. At 100,000 queries a month, you’re looking at $1,000+ before accounting for token costs. They also lock you in: OpenAI search only works in OpenAI calls; Anthropic search only works in Claude.
Cloudflare’s approach separates search from inference. Run the search query once, get structured results, inject them into whichever model makes sense — GPT-6.1 Sol for coding tasks, Claude Sonnet for reasoning, Llama 4 Scout for cost-sensitive volume. Change the model without rewriting the search code.
The gap: model-native tools handle simple cases more seamlessly — the model calls search automatically, no context injection boilerplate. Cloudflare’s “Server Tools” integration will close this. It’s marked coming soon in the docs, meaning the model will call web search as a tool automatically during inference without you building the injection layer.
Verdict
Start with Ceramic unless you have a specific reason not to. Keyword search handles the overwhelming majority of agent grounding use cases: product lookups, news retrieval, documentation checks, API reference queries. The cost at $0.25 per thousand is low enough to ground every agent call without a budget conversation.
Reach for Exa when your queries are semantically complex — research tasks, finding similar content, nuanced conceptual searches. Just verify the zero-retention gap status before putting customer data through it.
Cloudflare is quietly assembling the complete AI infrastructure stack: Workers AI for compute, Clef for routing, KV Instant and R2 for storage, K2 for event streaming, and now Web Search API for live data access. Everything flows through AI Gateway. If you’re building agents in 2026, this is worth wiring in now — the “Server Tools” integration on the roadmap will make it substantially easier within the next few months.













