Hacker News
new
past
comments
ask
show
jobs
points
by
TonyTrapp
10 hours ago
|
comments
by
jeremyjh
2 hours ago
|
next
[-]
The crawlers are not AI. The crawlers are deterministic. They are collecting data to train AIs.
reply
by
jonhohle
1 hours ago
|
parent
|
[-]
But gitweb is probably the second most used method of hosting a git repo and easily recognizable through heuristics. If it’s gitweb, fallback to git access and save everyone, including the crawler, time and resources.
reply
by
diegocg
9 hours ago
|
prev
|
next
[-]
Which, as the post notes, it's incredibly stupid. So much for artificial "intelligence"
reply
by
emsign
5 hours ago
|
prev
|
[-]
Makes me wonder how much garbage they actually collect across the web. That can't be good for the quality of the LLM.
reply