[HN Gopher] Bots are overwhelming websites with their hunger for...
       ___________________________________________________________________
        
       Bots are overwhelming websites with their hunger for AI data
        
       Author : Bender
       Score  : 22 points
       Date   : 2025-06-17 21:26 UTC (1 hours ago)
        
 (HTM) web link (www.theregister.com)
 (TXT) w3m dump (www.theregister.com)
        
       | tartoran wrote:
       | RIP internet. It will soon make no sense to share something with
       | the world unless you're in for profit. But who's gonna pay for
       | it?
        
       | superkuh wrote:
       | While catchy that headline kind of misses the point. It should be
       | "Corporations are overwhelming websites with their hunger for AI
       | data". They're the ones doing it and corporations are by far the
       | most damaging non-human persons (especially since they are formed
       | nowadays to abstract away liability for the damage they cause).
       | 
       | This is not some new enemy "bots". This is the same old non-human
       | legal persons that polluted our physical world repeating things
       | in the digital. Bots run by actual human persons are not the
       | problem.
        
         | Analemma_ wrote:
         | I'm not sure that's true. As hardware gets cheaper, you're
         | going to see more and more people wanting to build+deploy their
         | own personal LLMs to avoid the guardrails/censorship (or just
         | the cost) of the commercial ones, and that means scraping the
         | internet themselves. I suspect the amount of scraping that's
         | coming from individuals or small projects is going to increase
         | dramatically in the months/years to come.
        
       | johnea wrote:
       | This is an ever growing problem.
       | 
       | The model of the web host paying for all bandwidth was somewhat
       | aligned with traditional usage models, but the wave of scrapping
       | for training data is disrupting this logic.
       | 
       | I remember reading, about 10 years ago?, of how backend website
       | communications (ads and demographic data sharing) had surpassed
       | the bandwidth consumed by actual users. But even in this case,
       | the traffic was still primarily linked to the website hosts.
       | 
       | Whereas with the recent scrapping frenzy the traffic is purely
       | client side, and not initiated by actual website users, and not
       | particularly beneficial to the website host.
       | 
       | One has to wonder what percentage of web traffic now is generated
       | by actual users, versus host backend data sharing, and the
       | mammoth new wave of scrapping.
        
       | CSMastermind wrote:
       | What's the solution here? Metered usage based on network traffic
       | that gets shared with the website owners?
       | 
       | Otherwise everything moves behind a paywall?
        
         | Analemma_ wrote:
         | For now the solution is proof-of-work systems like Anubis
         | combined with cookie-based rate limiting: you get throttled if
         | your session cookie indicates you scraped here before, and if
         | you throw the cookie out you get the POW challenge again. I
         | don't know how long this will continue to work, but for my site
         | at least it seems to be holding back the deluge, for the
         | moment.
        
       | rglover wrote:
       | > Some of the bots identify themselves, but some don't. Either
       | way, the respondents say that robots.txt directives - voluntary
       | behavior guidelines that web publishers post for web crawlers -
       | are not currently effective at controlling bot swarms.
       | 
       | Is anybody tracking the IP ranges of bots or anything similar
       | that's reliable?
       | 
       | It seems like they're taking the "what are you gonna do about it"
       | approach to this.
       | 
       | Edit: Yes [1]
       | 
       | [1] https://github.com/FabrizioCafolla/openai-crawlers-ip-ranges
        
         | dbmikus wrote:
         | Many bots use residential IP proxy networks, so they come from
         | the same IPs that humans use
        
       | josefritzishere wrote:
       | I think the solution is criminal penalties.
        
       | darekkay wrote:
       | ai.robots.txt contains a big list of AI crawlers to block, either
       | through robots.txt or via server rules:
       | 
       | https://github.com/ai-robots-txt/ai.robots.tx
        
         | Bender wrote:
         | Your link is missing the t at the end of .txt. You should be
         | able to edit it though.
        
       | millipede wrote:
       | Information is valuable; we just weren't charging for it. AI is
       | just bringing the market for knowledge back into equilibrium.
        
       | gnabgib wrote:
       | Original source: https://www.glamelab.org/products/are-ai-bots-
       | knocking-cultu... (https://news.ycombinator.com/item?id=44298771)
        
       ___________________________________________________________________
       (page generated 2025-06-17 23:01 UTC)