[HN Gopher] Bots are overwhelming websites with their hunger for...
___________________________________________________________________
Bots are overwhelming websites with their hunger for AI data
Author : Bender
Score : 22 points
Date : 2025-06-17 21:26 UTC (1 hours ago)
(HTM) web link (www.theregister.com)
(TXT) w3m dump (www.theregister.com)
| tartoran wrote:
| RIP internet. It will soon make no sense to share something with
| the world unless you're in for profit. But who's gonna pay for
| it?
| superkuh wrote:
| While catchy that headline kind of misses the point. It should be
| "Corporations are overwhelming websites with their hunger for AI
| data". They're the ones doing it and corporations are by far the
| most damaging non-human persons (especially since they are formed
| nowadays to abstract away liability for the damage they cause).
|
| This is not some new enemy "bots". This is the same old non-human
| legal persons that polluted our physical world repeating things
| in the digital. Bots run by actual human persons are not the
| problem.
| Analemma_ wrote:
| I'm not sure that's true. As hardware gets cheaper, you're
| going to see more and more people wanting to build+deploy their
| own personal LLMs to avoid the guardrails/censorship (or just
| the cost) of the commercial ones, and that means scraping the
| internet themselves. I suspect the amount of scraping that's
| coming from individuals or small projects is going to increase
| dramatically in the months/years to come.
| johnea wrote:
| This is an ever growing problem.
|
| The model of the web host paying for all bandwidth was somewhat
| aligned with traditional usage models, but the wave of scrapping
| for training data is disrupting this logic.
|
| I remember reading, about 10 years ago?, of how backend website
| communications (ads and demographic data sharing) had surpassed
| the bandwidth consumed by actual users. But even in this case,
| the traffic was still primarily linked to the website hosts.
|
| Whereas with the recent scrapping frenzy the traffic is purely
| client side, and not initiated by actual website users, and not
| particularly beneficial to the website host.
|
| One has to wonder what percentage of web traffic now is generated
| by actual users, versus host backend data sharing, and the
| mammoth new wave of scrapping.
| CSMastermind wrote:
| What's the solution here? Metered usage based on network traffic
| that gets shared with the website owners?
|
| Otherwise everything moves behind a paywall?
| Analemma_ wrote:
| For now the solution is proof-of-work systems like Anubis
| combined with cookie-based rate limiting: you get throttled if
| your session cookie indicates you scraped here before, and if
| you throw the cookie out you get the POW challenge again. I
| don't know how long this will continue to work, but for my site
| at least it seems to be holding back the deluge, for the
| moment.
| rglover wrote:
| > Some of the bots identify themselves, but some don't. Either
| way, the respondents say that robots.txt directives - voluntary
| behavior guidelines that web publishers post for web crawlers -
| are not currently effective at controlling bot swarms.
|
| Is anybody tracking the IP ranges of bots or anything similar
| that's reliable?
|
| It seems like they're taking the "what are you gonna do about it"
| approach to this.
|
| Edit: Yes [1]
|
| [1] https://github.com/FabrizioCafolla/openai-crawlers-ip-ranges
| dbmikus wrote:
| Many bots use residential IP proxy networks, so they come from
| the same IPs that humans use
| josefritzishere wrote:
| I think the solution is criminal penalties.
| darekkay wrote:
| ai.robots.txt contains a big list of AI crawlers to block, either
| through robots.txt or via server rules:
|
| https://github.com/ai-robots-txt/ai.robots.tx
| Bender wrote:
| Your link is missing the t at the end of .txt. You should be
| able to edit it though.
| millipede wrote:
| Information is valuable; we just weren't charging for it. AI is
| just bringing the market for knowledge back into equilibrium.
| gnabgib wrote:
| Original source: https://www.glamelab.org/products/are-ai-bots-
| knocking-cultu... (https://news.ycombinator.com/item?id=44298771)
___________________________________________________________________
(page generated 2025-06-17 23:01 UTC)