Cloudflare's Web Search API: When the Edge Network Became the Search Engine for AI Agents
Cloudflare's new Web Search API gives AI agents real-time web search through AI Gateway with three providers, unified billing, and zero-config Workers integration. Here is what developers need to know.
Every AI agent eventually hits the same wall: its training data has a cutoff date, and the real world keeps moving. Models can reason, write code, and parse documents, but they cannot know what happened yesterday. Until now, developers building agents on Cloudflare had to cobble together their own web search integrations, manage API keys for multiple search providers, and handle billing separately from their AI infrastructure. Cloudflare's new Web Search API, launched in open beta on October 2, 2026, changes that equation by bringing web search directly into the AI Gateway ecosystem — and the developer community is paying attention, with the launch announcement climbing to 548 points on Hacker News within 24 hours.
The announcement matters because it solves a problem that every team building AI agents eventually confronts: how to give models access to live, current information without bolting on a fragile pipeline of third-party search APIs, custom retry logic, and separate billing arrangements. Cloudflare has wrapped web search into the same infrastructure that already handles model inference, logging, and rate limiting. The result is a single API call that returns structured search results — titles, URLs, and descriptions — ready to inject into any model's context window.
What Cloudflare Web Search API Actually Does
At its core, Web Search API is a managed search proxy that sits inside Cloudflare's AI Gateway. You send a query, choose a search provider, and get back structured results — URLs, page titles, and descriptions — in a consistent JSON format regardless of which provider you select. The API supports three providers at launch: Ceramic.ai, Exa, and Linkup. Each has distinct characteristics that make it suitable for different use cases, but the response format remains identical across all three, meaning you can swap providers by changing a single parameter without touching your response parsing code.
The API is accessible through two paths. First, a standard REST API endpoint at /ai/websearch/ that accepts POST requests with a bearer token — the same Cloudflare API token you already use for other Cloudflare services. Second, a Workers AI binding method called env.AI.websearch() that lets you call search directly from within a Cloudflare Worker, with no HTTP overhead. Both paths route through AI Gateway, which means every search request gets logged, analyzed, and billed alongside your existing model inference traffic.
The request format is deliberately minimal. You specify a query string (up to 1,024 characters), a provider name (defaults to Ceramic.ai if omitted), a limit on results (1 to 10, defaulting to 10), and optional gateway configuration. The response contains an array of items, each with a URL, title, and description, plus metadata about the request including latency. That is it. No pagination tokens, no complex filtering syntax, no async job management. For the vast majority of agent use cases — looking up current information to ground a model's response — this is exactly the right level of abstraction.
Three Search Providers, One API: Choosing the Right Engine
Cloudflare could have built a single search index and called it a day. Instead, they partnered with three specialized search providers, each targeting a different segment of the AI search market. Understanding the differences between them is critical for choosing the right one for your workload.
Ceramic.ai: The Default, Optimized for Cost and Speed
Ceramic.ai is the default provider, and for good reason. It runs its own independent web index of more than 40 billion pages, built specifically for AI agents and LLM applications rather than human search. At $0.25 per 1,000 requests, it is dramatically cheaper than the other two options — 28 times less expensive than Exa and 20 times less than Linkup. Results include long descriptions of up to 8,000 characters per page, which is substantial context for grounding model responses. Ceramic.ai also supports Zero Data Retention through Cloudflare, meaning your search queries are not stored by the provider after results are returned. For high-volume agent workloads where cost per query matters — think chatbots answering customer questions with current information, or agents running multiple searches per task — Ceramic.ai is the natural choice.
Exa: Hybrid Search for Quality and Relevance
Exa takes a different approach. Its search engine combines traditional keyword search with embeddings-based semantic search, using an 'auto' mode that balances result quality and speed. Instead of returning full page descriptions, Exa returns highlights — the specific text snippets from each page that are most relevant to your query. This makes Exa particularly well-suited for use cases where you want to pass concise, query-relevant excerpts directly into a model's context without processing full page descriptions. At $7.00 per 1,000 requests, it is significantly more expensive than Ceramic.ai, but the semantic search capabilities may produce better results for complex, nuanced queries where keyword matching alone falls short. Note that Exa does not currently support Zero Data Retention, which may be a dealbreaker for organizations with strict data handling requirements.
Linkup: Fast, Cited Results for Agent Tool Calls
Linkup occupies a middle ground at $5.00 per 1,000 requests. It uses a 'fast' search depth mode that returns raw search results without generating a synthesized answer, making it ideal for agent tool calls that need quick, cited results from trusted sources. Like Ceramic.ai, Linkup supports Zero Data Retention. The focus on speed and citation makes it a strong fit for agentic workflows where the agent needs to verify claims against primary sources, or where you need to present users with links to back up an AI-generated answer.
The price spread between these three providers is significant. Running 10,000 searches per day on Ceramic.ai costs $2.50. The same volume on Exa costs $70.00, and on Linkup it costs $50.00. For a startup building a consumer-facing AI assistant, that difference compounds quickly. The ability to switch providers by changing one parameter — and to bring your own API key for any of them — gives developers flexibility that no single-provider solution can match.
Why AI Gateway Integration Changes the Game
The real significance of Web Search API is not the search itself — web search APIs have existed for years. The significance is that search is now a first-class citizen inside Cloudflare's AI Gateway, alongside model inference. This has several practical implications that developers should understand.
First, observability. Every search request appears in your AI Gateway logs and analytics dashboard, right next to your model inference requests. You can see which agents are making searches, what queries they are running, how long searches take, and how much they cost. This is not a separate billing dashboard or a different logging system — it is the same unified view that already tracks your AI spending. For teams trying to understand and optimize their AI costs, having search and inference in one dashboard is a meaningful operational improvement.
Second, unified billing. Searches are billed to your AI Gateway credits at each provider's list price, with no markup from Cloudflare. You can also bring your own provider API key and have the provider bill you directly, which is valuable if you have negotiated custom pricing with Exa or Linkup. The BYOK system stores your API keys in Cloudflare's Secrets Store with encryption, and keys are never sent in the request itself — AI Gateway retrieves them server-side based on an alias you configure.
Third, the Workers AI binding. If you are already running AI agents on Cloudflare Workers, adding web search requires no new dependencies, no external HTTP calls, and no additional authentication management. You call env.AI.websearch() the same way you call env.AI.run() for model inference. The search runs on Cloudflare's network, close to your Worker, which minimizes latency. For agent architectures that involve multiple tool calls per user interaction, keeping everything on the same platform reduces failure modes and simplifies deployment.
Building a Grounded AI Agent: Practical Implementation
The most compelling use case for Web Search API is giving AI agents the ability to look up current information and ground their responses in real data. Cloudflare's documentation includes an example of using web search as a tool that a model can call, and the pattern is worth examining in detail because it represents a clean architecture for retrieval-augmented generation.
The flow works in three steps. First, you send the user's prompt to a model along with a tool definition for web_search, telling the model it can call this function when it needs current information. Second, when the model decides to call the tool, you execute the search using env.AI.websearch() and get structured results back. Third, you pass those results back to the model as a tool response, and the model generates its final answer grounded in the search results.
This pattern — model calls tool, tool calls search, results go back to model — is the standard agentic RAG architecture, but Web Search API makes it simpler to implement because the search step is a single function call within the same runtime. No external API client, no separate authentication, no network hops to a third-party service. The entire loop can run inside a single Cloudflare Worker.
Here is what a minimal implementation looks like using the Workers AI binding:
- Define an AI binding in your wrangler.toml or wrangler.jsonc file
- Call env.AI.run() with your model and a tools array that includes a web_search function definition
- Check if the model returned a tool_call for web_search
- If so, call env.AI.websearch() with the query from the tool call arguments
- Pass the search results back to the model as a tool response message
- Return the model's final response to the user
The supported models for tool calling include options like Google's Gemma 4 26B, and Cloudflare's documentation shows the pattern working with Workers AI models directly. For developers using external models through AI Gateway — say, calling OpenAI or Anthropic through Cloudflare's proxy — the same REST API approach works, with search results formatted as tool response messages in whatever schema the external model expects.
What This Means for Developers and IT Leaders
For developers building AI agents, Web Search API eliminates a category of infrastructure work that has been annoying everyone. Instead of evaluating search providers, signing up for API keys, building retry logic, managing rate limits, and reconciling separate bills, you get a single endpoint with three providers behind it and unified billing through credits you may already have. The ability to switch providers by changing one parameter is particularly valuable during development — you can test with the cheap Ceramic.ai provider and switch to Exa for production if result quality justifies the cost.
For IT leaders and platform teams, the AI Gateway integration provides something that has been hard to get with AI workloads: visibility. Knowing exactly how many searches each agent is making, what queries are being run, and how much they cost — all in the same dashboard as model inference — makes it possible to set budgets, identify runaway agents, and audit what your AI systems are actually doing. The Zero Data Retention support from Ceramic.ai and Linkup also addresses data governance concerns that have made legal teams nervous about third-party search APIs.
For the self-hosting and homelab community, the calculus is different. Web Search API is a managed service — you are paying Cloudflare to run search for you, and the search providers are closed-source indices. If you are already invested in self-hosted search solutions like SearXNG or running your own crawlers, Web Search API does not replace that infrastructure. But for anyone building AI agents on Cloudflare Workers or using AI Gateway for model inference, it is a natural addition that removes friction without adding vendor lock-in beyond what you already have with Cloudflare. The BYOK option also means you can use existing relationships with search providers while still getting the Gateway integration benefits.
The Bigger Picture: Search as AI Infrastructure
Cloudflare's Web Search API arrives at a moment when the AI industry is collectively realizing that retrieval is the bottleneck for useful agents. Models are getting better at reasoning, but reasoning over stale information produces confidently wrong answers. The rush to build retrieval-augmented generation systems has exposed a gap: the search infrastructure that exists for human use (Google, Bing) was not designed for programmatic, high-volume, machine-to-machine queries. A new category of search providers — Ceramic.ai, Exa, Linkup, Tavily, Brave Search API — has emerged specifically to serve AI workloads, with pricing models, response formats, and rate limits optimized for agents rather than human users.
What Cloudflare has done is not invent a new search engine. They have built the orchestration layer. By aggregating multiple AI-native search providers behind a single API with unified billing, logging, and access controls, they are positioning AI Gateway as the control plane for AI infrastructure — not just model inference, but the full stack of capabilities an agent needs. Search is the first non-inference capability added to AI Gateway, but it is unlikely to be the last. Code execution, file retrieval, and other agent primitives could follow the same pattern.
This also raises competitive questions. If Cloudflare becomes the default routing layer for AI search, what happens to direct relationships between developers and search providers? The BYOK option is a hedge against this concern — providers can still have direct billing relationships with users while Cloudflare handles the routing. But the unified billing path is frictionless, and most developers will take the path of least resistance. Search providers may find themselves in a position similar to cloud marketplace vendors: technically independent, but practically dependent on the platform that owns the customer relationship.
Looking Forward: What to Watch
Web Search API is in open beta, which means several things are likely to change. The provider list will almost certainly expand — expect to see Tavily, Brave Search API, and possibly Google's Programmable Search Engine added in the coming months. Pricing may shift as Cloudflare negotiates volume discounts with providers and passes some of those savings through. The current no-markup model is a land grab strategy to establish AI Gateway as the default routing layer; once market share is established, a small margin is likely.
Feature-wise, the most obvious gap is the absence of async or streaming search. The current API is synchronous with a maximum of 10 results per request, which is fine for most agent use cases but will not scale for bulk indexing or research workflows that need hundreds of results per query. Search filtering — by date, domain, language, or content type — is also missing at launch and will be needed for production use cases like news monitoring or competitive analysis.
For developers evaluating whether to adopt Web Search API today, the decision is straightforward. If you are already using Cloudflare AI Gateway for model inference, adding web search is a no-brainer — it is the same API, same billing, same logs. If you are building agents on Cloudflare Workers, the native binding makes integration trivial. If you are running your AI infrastructure on AWS, GCP, or Azure, the REST API is still accessible, but the value proposition is weaker since you do not get the same unified observability benefits.
The HN community reaction — 548 points and 248 comments in under 24 hours — suggests this struck a nerve. Developers have been waiting for someone to simplify the search-for-agents problem, and Cloudflare's approach of aggregating providers behind a managed gateway is a clean solution. Whether it becomes the standard depends on execution: reliability, latency, and whether the provider ecosystem stays diverse or consolidates. For now, it is one of the most practical new tools for anyone building AI agents in 2026.
Key Takeaways
- Cloudflare Web Search API is now in open beta, providing AI agents with real-time web search through AI Gateway
- Three providers available at launch: Ceramic.ai ($0.25/1K requests), Exa ($7.00/1K), and Linkup ($5.00/1K) — all behind a single API with identical response format
- Integration with AI Gateway means unified logging, billing, and observability alongside model inference requests
- Workers AI binding allows calling search directly from Cloudflare Workers with no HTTP overhead
- BYOK support lets you use existing provider relationships while still getting Gateway integration benefits
- Zero Data Retention supported by Ceramic.ai and Linkup for organizations with data governance requirements
- The API is a beta — expect more providers, pricing changes, and feature additions like filtering and async search in coming months
Web Search API is available now in open beta. If you have a Cloudflare account with AI Gateway enabled, you can start making search requests immediately using your existing API token. The documentation at developers.cloudflare.com/web-search/ includes quickstart guides for both REST API and Workers AI binding approaches, along with provider comparison details to help you choose the right search engine for your workload.
Related Posts
Strata: When a 125B AI Model Ran on a Gaming PC at 100 Tokens Per Second
A new open-source tool called Strata lets you run Qwen 3.8 Flash Next — a 125-billion-parameter model — on an ordinary gaming PC with an RTX 4090. Nothing leaves your machine, and it's faster than you can read.
When AI Agents Spend Your Money While You Sleep: Why Hard Budget Caps Are Becoming Non-Negotiable
AWS and Google Cloud finally launched hard spending limits in the same month. It's not a coincidence — it's a response to AI agents that can rack up thousands of dollars before you wake up.
When Utah Banned VPNs: How a Court Stopped a Law That Demanded the Technically Impossible
A federal judge just blocked Utah's unprecedented anti-VPN law, ruling that lawmakers cannot mandate perfect geolocation — a technical impossibility. The case reveals a deeper problem: when legislation outruns engineering.