textlog
I am curious what an actual solution to this problem could be. Everything I've read about it feels like a bandaid — from robots.txt to cloudflare turnstile — there has to be a better way!
I hate crawlers that don’t advertise they’re crawlers. ’n badly written. They don’t respect the canonical URLs so they keep getting themselves into a loop instead of actually crawling the site properly. They make the logs really hard to read. Can’t block them they’re distributed.
profileenter to followHai I am kucontinued:
For this particular problem though, I feel like if the crawlers are coming from domestic / residential IP addresses and they lack a distinctive http header or user agent, it is indistinguishable from a denial of service attack and must be treated as such.
profileenter to followlate is the hour in which this conjuror chooses to appearreplied toprofileenter to followHai I am ku:
But as long as people can still use the site it's OK, right? Seems like the problem is often self-inflicted, either because the site is so bloated that answering even GET-requests is burdensome, or because they're trying to monetize the data and so must block non-paying scrapers.
join the communityorbrowse more notes