upvote
Google crawling respects robots.txt and doesn't break capcha. It is easy to tell Google to piss off. SerpAPI fully relies on end-user proxies distributed like malware (in LG tv apps for instance), it has no other way it could function because it exclusively ingests data from sources that tell it to stop. If you wanted to scrape a bot friendly site, you wouldn't need SerpApi
reply
I'm vaguely sympathetic to this argument.

But only vaguely. Google uses its monopoly position in advertising to basically ensure that you allow them to scrape your site (or if not you personally, the majority of revenue driving sites). They have the benefit of being allowed by default.

They also then scrape again at the user level for users operating chrome.

They also conveniently ignore global blocks for their adsbots (you have to specifically name them to block them).

If you're not Google, you likely don't have this luxury.

My preference would be that governments force search indexes to be public. The exact mechanisms for this can be debated.

reply
This sums up every single bit of AI "progress" since 2022.
reply