If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.
The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:
> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.
https://www.ceramic.ai/terms-of-service
Am I alone in caring about this?
> provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...
but it continues:
> ... and is not independently accessible, extractable, or downloadable by end users or third parties
How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?
The weird attitude in the Internet Tech company scene is akin to Gold Rush scenarios.
Who are the native people?
These days, this seems to be "modus operandi". The bet is who can get closer to the administration to suddenly enforce the un-enforce-able. For your own safety. You wouldn't steal a car now, would you?
You make yourself subject to a social reality you presume outside of your control. But you enabled them, if by nothing else, by your silent acceptance.
Moral and ethical judgements cannot be left to the very same people they are supposed to restrain in the first place.
They don't make it available "easily" either, they place all kinds of hurdles on it, including you having to pay for stolen goods.
You then go on, equating wildly different situations: on one hand, a multi-billion dollar company, easily able to set up such a scheme. On the other (me representing) the public, regularly anything but destitute.
So no, you're the unreasonable one here. So are they.
If someone does not want on a search engine index and they say not to index, and then get indexed anyhow is stealing.
But a “please read my content and make available to your users” then claim doing so is stealing seems a bit out there.
What am I missing?
- Tony Soprano
Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”
It’s really really good for low cost search!
I am currently using Perplexity fast search and fetch, and I am happy with that. I would try our Ceramic.ai, but I need to be able to fetch the pages as well (I do not want summaries).
1. Google News API now returns only Google links that don't resolve to anything in code. 2. Google Search results are atrocious and only unearth non-authoritative blogspam and aggregator sites.
Don't give them (G) ideas.
Google Search only works as a business because human eyeballs see (and brains choose to click on) ads at rates that justify advertisers’ (massive in aggregate) dollars.
If Google instead spun out Google Search into it's independent company, with actual focus on search instead of whatever they're doing now, with an actual business model, it might actually work, granted they get back the old search quality.
Kagi is an example of search engine that manages to also be a business today and not theoretically.
(Was your comment a joke? Or did google announce this through different channels?)
https://docs.cloud.google.com/gemini-enterprise-agent-platfo...
- Company has great initial product
- Company gets popular
- Shareholders demand infinite growth
- Company becomes rent-seeker
- GOTO 10
I'm with OP - a company that wants to insert itself in the middle of everybody's business is not being altruistic, they're playing the long game.It means "hi spending approver, I'm going to add $100 to our CF account" instead of "hi accounting+management+security, please initiate the process of evaluating new third party vendor Foo for use in my project, I hope we can get it approved and integrated into SSO sometime next month".
Instead I put my API keys to cloudflare, set limits, and gave the agent the CloudFlare token, and in minutes it could contact tens of services.
edit: not to mention instead of loading balance to each service I could just keep balance on cloudflare that covers them all
Extremely hard to verify it's been done properly across a large organization.
Put it this way: I'd rather Cloudflare owns the Internet than Google, Meta, Amazon or Alibaba.
Cloudflare is however known to deliberately screw over some of their clients.
> I'd rather Cloudflare owns the Internet
I'd rather no one does, certainly not a firm that feeds into the NSA.
The government always wins this given enough time.
[0] https://blog.cloudflare.com/why-we-terminated-daily-stormer/
> Our terms of service reserve the right for us to terminate users of our network at our sole discretion. The tipping point for us making this decision was that the team behind Daily Stormer made the claim that we were secretly supporters of their ideology.
> Our team has been thorough and have had thoughtful discussions for years about what the right policy was on censoring. Like a lot of people, we’ve felt angry at these hateful people for a long time but we have followed the law and remained content neutral as a network. We could not remain neutral after these claims of secret support by Cloudflare.
https://blog.cloudflare.com/why-we-terminated-daily-stormer/
Per Matthew Prince:
"This was my decision. Our terms of service reserve the right for us to terminate users of our network at our sole discretion. My rationale for making this decision was simple: the people behind the Daily Stormer are assholes and I’d had enough.
Let me be clear: this was an arbitrary decision. It was different than what I’d talked talked with our senior team about yesterday. I woke up this morning in a bad mood and decided to kick them off the Internet. I called our legal team and told them what we were going to do. I called our Trust & Safety team and had them stop the service. It was a decision I could make because I’m the CEO of a major Internet infrastructure company."
This is one of the best, most honest things I've seen a tech CEO write. Contrast this with Zuckerberg's mealy-mouthed weaseling about "policies"[1]. I wish more tech overlords had the honesty to say, "No, we are not a court. We are booting you because we don't like you."
[0] https://gizmodo.com/cloudflare-ceo-on-terminating-service-to...
[1] https://www.newsweek.com/read-mark-zuckerbergs-full-statemen...
https://news.ycombinator.com/item?id=44150898 (2025)
https://news.ycombinator.com/item?id=40481808 (2024)
https://robindev.substack.com/p/cloudflare-took-down-our-web...
Reddit comments report more such incidents, with Cloudflare demanding an upgrade to Enterprise. This has also come up for other sites, such as gambling sites.
Also, Cloudflare implicitly screws over everyone by leaking data to the NSA.
I have always operated under the assumption that every cloud provider and telco does this, so this claim has always seemed very silly to me.
That said, these are US companies subject to FISA court orders and NSLs or National Security Letters. If they want your data, they can just pull it from memory in real time or pull it directly from the hypervisor and dump it wherever they're instructed to. Any idea that your data is protected because you're not even using a provider WAF or doing TLS termination for load balancing is a fantasy.
I control which DNS server I use. It is not relevant to the matter at hand.
> If they want your data, they can just pull it from memory in real time or pull it directly from the hypervisor and dump it wherever they're instructed to.
You're confusing bulk data collection with highly selective court-ordered data collection. The two are not alike. Attempting to equate them is a dumb attempt at deception on your part. There is no obligation for a firm to share bulk web data with the NSA.
Yes that's precisely how I want infrastructure to operate.
Given the number of people on HN who report massive problems from scrapers and other bots, it sounds like if Cloudflare doesn't do this, someone else will need to. I might have thought bandwidth was cheap enough now for it not to matter, but I guess the bots are costing some sites a lot of money.
As for the bots, I thought the same thing, but it is indeed a huge problem. They've brought my websites down pretty frequently recently. I tried Cloudflare but visitors complained, and I think you can't win against the bots anyway, so I've resorted to performance improvements and serving every request.
- let more traffic in, eat the compute cost
- block larger cohorts of traffic, affect many real users
- babysit the rules to get them just right, lose time doing that
In my experience the residential proxies exist but are not that common and many aren't trying as hard as they could. It's really a war of which side wants to spend more attention on the problem.
Why would I use bonsai? Why not use ElasticSearch directly?
Thanks asciimoo for https://github.com/asciimoo/hister
does it? I am running it but was under the impression that it did not cache the content I am viewing, unlike SinglePage.
The original request via the MCP is somehow blocked?
Hister tells me my index is currently 39208 pages.
And then on the providers page:
Property Value
provider exa
Zero Data Retention NoThe providers[1] behind this Web Search API have very different rates:
Ceramic.ai: $0.25 per 1,000 requests
Linkup: $5.00 per 1,000 requests
Exa: $7.00 per 1,000 requests
[0] https://serper.dev/You also don't pay for actual data transfer, so the billing is overall simpler - AWS and GCP have similar per-request compute options, but every part of the platform has extra fees (like per-gb data transfer billing, sometimes you need a VPC to interconnect services, secrets being an extra charge, etc).
Running your stuff, even private stuff, through a tunnel is great so that you don't have to expose your VPS' IPv4.
CDN, DDoS protection, excellent DNS hosting features, web monitoring, web analytics, advanced web and service filtering and blocking, zerotrust networking / vpn options, a solid API that works great with terraform, tons of other stuff.
(Actually multiple, one for each domain)
https://developers.cloudflare.com/cloudflare-for-platforms/w...
Now there is an official paid search API, and I'm guessing the certified providers will be allowed through the Cloudflare "bot protection"?
This is very worrying.
I noticed that my API quota resets every month. Have not been charged once.
What's your plan? I'm on Duo and I get charged for every MCP hit. If an API quota is included with Ultimate, that might be worth my upgrading.
This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar.
But that's meaningless because 99% of AI crawlers are "bad bots" which ignore robots.txt and use domestic IPs to circumvent blocks.
Some other reports of this: https://github.com/TecharoHQ/anubis/issues/1565
I ended up just fully closing connections with no response from these assholes on any URL.
Then shortened the query to just "qwen-3.8 flash next" ... results came.. all unrelated. In fact, these were almost all paper links .... no relation to actual search term.
And I had thought that I finally had found a cheaper search alternative.
Then searched for "Cloudflare OHTTP Gateway" .. this text is literally in the title ... but zero link for this page.. the closest it yielded was this link: "https://developers.cloudflare.com/privacy-gateway/" ... it seems cloudflare updated this 2 days back.. the original content was last updated in 2022 ... so that's what the cutoff index seems to be.
I personally used to use Firecrawl's paid credits (got a bunch of em for free at an event) before I realized that they allow you to self-host your own instance (albeit missing some features I never use anyways).
It's been working really well for my agents, I even hosted a small observability tool that proxies the requests so I can see how many are failing and the percentages are always below 2%.
There actually is such a thing as verified bots on Cloudflare that gets through most blocks (and these services are likely are part of that), but ultimately it just depends on how the website owner has things set up in Cloudflare
Verified bot is just a label. What you do with that information is entirely up to you as the operator. It doesn't say anything anything about the service and doesn't provide any guarantees about the traffic.
Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).
So what is CF providing here? Maybe some free credits to entice us to use their router? No, not that either ("billed to your AI Gateway credits"). Maybe a comparison of which agent search yields the best results? Nope.
It's a crappy proxy- probably less efficient and more volatile than hitting the agent API directly.
This is only if I understand the product correctly (which I admittedly skimmed) due to the sheer number of screeching vibey nothingburgers coming out of CF over the past month.
Hi! Exa Head of Index here. We certainly do have our own index and it's one of the biggest among the independent players (i.e. not Google and Bing, which by the way closed off their official search APIs). [1]
Regarding the quality: search is a multi-dimensional problem, you can be better on one set of queries and worse on the other. There are tons of benchmarks in the industry, all the players in the AI search market are fighting very hard to climb to the top, updates are shared every week.
We track dozens of use cases and run evals continuously, we perform well on all the verticals we optimize for. Not only we top the ranking on e.g. financial queries, but also Claude with Exa search performs better that Claude with native search -- as measured by independent observers [2]. This means that the underlying search is materially better for the outcome, it's not just how we evaluate the search itself.
[1] https://lnkd.in/p/e9u3dyEG [2] https://lnkd.in/p/enYe4h7u
I remember as I was looking at the available web tools for hermes agent not to long ago and looked through the keyless web providers privacy policies, which exa is one of them.
ceramic.ai - $0.25 per 1,000 requests
Exa - $7.00 per 1,000 requests
Linkup - $5.00 per 1,000 requests
Does anyone have insights on the quality differences? Web search API pricing for AI agent usecases has always felt so expensive for what it is, but I have no grounding on the economics of running a web index.
EDIT: formatting
They have 2,500 free which I used, it seemed good.
I'm unaffiliated--actually a clanker told me about it so I told it "go ahead"
NO SCRAPERS (except ours) -> $$$$$$$$$$$$$$$$
https://github.com/hbmartin/agent-web-search
So this gives a unified search experience without adding another cloud hop and dependency.
But does CloudFlare itself commit to zero data retention? If not, this isn’t too meaningful.
* Extraction of main content from HTMLs, the index stores markdown representations (free of headers/footers/sidebars/menus etc)
* Serving highlights picking the most relevant part for each result to reduce token usage downstream
* Also serving dynamic highlights where we summarize all the sources at once reducing the token counts even further
see here: https://exa.ai/docs/search/highlights and https://exa.ai/docs/contents/quickstart
[0] https://developers.cloudflare.com/fundamentals/reference/mar...
I use CloudFlare developer platform and quite happy with tools, but I didn’t use the gateway API and always used OpenRouter which does support web search.
I can see it useful for those who didn’t do any integrations or like to keep logs at one place, but did customers actually ask for this?
CloudFlare's entire business is scale, cost, and complexity. They are powering like half the web at this point. Wouldn't really call them a "new entrant".
MITM service, Internet gatekeeper and robber baron.
AI firms sell security services
The absence of the Perplexity Search API is to be expected though, knowing how much these two companies despise each other.
https://developers.cloudflare.com/llms.txt
As a textmode command line and text-only browser user this textfile is faster for me to use that the usual Silicon Valley style web pages
Not quite as good as sitemap-0.xml but it's nice to have this in addition
As AI systems increasingly use search as a tool, the quality and reliability of that tool become an important part of the overall agent workflow.