upvote
The crawlers are not AI. The crawlers are deterministic. They are collecting data to train AIs.
reply
But gitweb is probably the second most used method of hosting a git repo and easily recognizable through heuristics. If it’s gitweb, fallback to git access and save everyone, including the crawler, time and resources.
reply
Which, as the post notes, it's incredibly stupid. So much for artificial "intelligence"
reply
Makes me wonder how much garbage they actually collect across the web. That can't be good for the quality of the LLM.
reply