If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.
It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.
There could be open source tooling to create custom private "closednets", with
- trust ring mechanism to allow invitations, flagging, banning, and banning those that invite people who were banned
- the rules of the closednet
- search engine with opt-in scraping
- portal (remember the 80s?) with all the registered nodes, perhaps by service category such as public git repo hosts, web sites etc.
etc.
The first closednet could be Hacker News.
Not sure of the effectiveness but it's there.
I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less.
The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit.
I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Are you actually saying you'd be OK with that?
Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866
(I was pointing out that "they'll train on a version that strips out you as the author" seems to be is incorrect about how training works)
Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.
All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.
GENIUS
LLMs just use everything, generate similar code with no attribution and keep users from visiting, so no bragging rights or attention.
Worse, there are some PRs that seem fully generated ...
So i mostly stopped sharing and started pulling my old repos offline.
At this pace, i don't want to compete with a clone of myself in the future that will do my work for much cheaper.
Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?
That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.
In that case the creator should welcome AIs with open arms; a human reader will forget eventually, but the AI will preserve the knowledge forever.
no, only some mangled form of it
Should be "You're thinking".
This should be "This should be 'You're thinking'." don't you think? Why bother correcting someone's grammar with a sentence fragment? You're just trading one mistake for another. I'm hoping someone finds a grammar error in my post, because continuing this would be hilarious.
Reflexively, I think it should be more like ...
javascript: `This should be "You're thinking".` ;
// to preserve the original character use and to avoid '...'...' parse foos
// however `"...".` also possibly deserves a [sic] to critique the original
// i.e. ~grammar police say the period belongs within the quote marks, no?
... but then that's just me, in [my] quirks mode.But we already have the latter case that exists - ad blockers. Ad blockers literally serve up the word-for-word original content minus the ads.
I might never blog/publish code again amidst all this. I never had ads on my sites. I am not alone in this.
It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!
obviously new content still has value because it remains the source layer for LLM agents. it just wont be ads giving you revenues thats all.
But maybe that is the future.
Every country starts erecting their own towers of babel that we talk at, and it constantly compresses our conversations down to the most effective distribution of weights.
At some point talking at the machine becomes a high status job, and we give respect to the people who whisper to it the most.
When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else.
I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way).
Cheap access came from the invention of cheap printing . The laws were passed to restrict it.
Sure, humans would benefit.
It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T.
Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype.
But the catch is that the squirrel cage must run non-stop still.
I had the similar experience to yours yesterday and it lead nowhere. Funnily enough I was also trying to configure a vpn on a router, google didn't return anything useful (besides a blog post clearly written by AI and with absolutely no information in it). Claude managed to give some interesting pointers, but its suggestions were not working and I also noticed that it started to hallucinate badly about ipv6 and gave me some suggestions that were just plain untrue. Claude Opus is smart, usually when it gets so convinced about something is after researching the internet and not just based on its training data. I wonder where it got so convinced about it. Maybe reading some other hallucinated blog post like the one I stumbled upon?
To be more precise, I hate the SEO shithole the internet has become, that Google serves up, that Google facilitated, indirectly created.
(I really don't have any tears to shed if there is a death of the Corporate Internet™.)
We're likely at the "golden age" of LLM-assisted web searching and summarization.
Hopefully open models keep it cracked open, but expecting enshittification is always the safe bet these days.
Very happy with Kagi personally
> One you pay for yourself !
SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
It's nice to be the customer instead of being the product, for once.
They do this because they benefit from their site being visited or the information they are providing being noticed.
> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I see no reason it would have that effect. It does, however, create different incentives for the search provider to improve the signals indicating page relevance since the user is the priority instead of advertisers.
>>>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
>>> One you pay for yourself !
>> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
> They do this because they benefit from their site being visited or the information they are providing being noticed.
>> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
> I see no reason it would have that effect.
Agreed.
We used to know better, Standard Oil vertical integration was dismantled.
> Oh, I should mention though. There was no advertising at all. They didn't make any money off me
Are you sure about that? Even if you didn't see ads ( remember people pay even if you don't click - just like a billboard ) - they are still profiling you to better sell you ads in the future, and using your interaction as free training data.
They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.
He said, proud of his own ignorance.
I would be the first to admit my ignorance on the absolute majority of topics. There is a limited number of things I can learn in life, and kubernetes won't be one of them - I'm just not interested in it (and all the other infra stuff, to be honest), as long as it works.
It seems to get shell scripts right most of the time.
YMMV.
But yeah, text remix machines are not a long-term solution to that problem.
It's the same reason why we don't give students the answers to things, we teach them to find the answers.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.