[HN Gopher] AI haters build tarpits to trap and trick AI scraper...
       ___________________________________________________________________
        
       AI haters build tarpits to trap and trick AI scrapers that ignore
       robots.txt
        
       Author : hilux
       Score  : 32 points
       Date   : 2025-01-28 22:18 UTC (41 minutes ago)
        
 (HTM) web link (arstechnica.com)
 (TXT) w3m dump (arstechnica.com)
        
       | dinobones wrote:
       | ok...
       | 
       | if depth > 5 and if sem_hash(content) in hist: return
        
         | anotherhue wrote:
         | I can generate more random content than you can store.
        
         | rossdavidh wrote:
         | It's a radar gun/radar detector kind of situation. You can
         | always change your strategies; countless top-level links that
         | are shown on the menu in a way that no human sees them (due to
         | font color or size or etc). Small number of pages with a few
         | that go on forever, in a way that any human would stop reading
         | but a bot may have trouble detecting. Real (for humans) text in
         | images, with endless invisible text for bots. Etc.
         | 
         | I think it will probably make it harder for screen readers,
         | unfortunately.
        
       | andrewfromx wrote:
       | "trap crawlers in infinite mazes of gibberish data, potentially
       | increasing AI training costs and poisoning datasets. While their
       | effectiveness is debated, creators see them as a form of
       | resistance against unchecked AI development."
        
       | BizarreByte wrote:
       | AI haters? I don't hate AI I just don't want things I've created
       | being used to enrich multi-billion dollar companies for free.
       | These companies are behaving poorly and they should expect this
       | kind of push back.
        
         | someone_eu wrote:
         | It is not even just the copyright issue.
         | 
         | The article completely misses the point that AI scrapers are
         | not a "future threat of AI domination". They already do damage
         | by DDOSing site's networking infrastructure and inflicting very
         | real costs to a site hoster.
         | 
         | Even when the data is completely free, like in case of
         | Wikipedia or OpenstreetMaps, scraping it is unethical and
         | should be illegal. Most of the open data resources have
         | procedures, which allow downloading of the data in the archived
         | form, without need for scraping. They are built with sharing in
         | mind.
         | 
         | So the arguments the article tries to use (what if it is for
         | public good?) has no sense. 1) it is not 2) there are many ways
         | to fetch the open data properly and respectfully.
        
         | senko wrote:
         | [delayed]
        
       | anotherhue wrote:
       | I wonder if we could run the cheapest smallest 1b model (or
       | smaller) llm to generate this data in a way that's just plausible
       | enough to be ingested.
       | 
       | A little note in robots.txt offering commercial terms would also
       | be available.
        
         | Dwedit wrote:
         | Usually it's just a Markov Chain rather than an LLM to generate
         | the garbage data.
        
       | Dwedit wrote:
       | There's also the other kind of AI haters who do not give any
       | anti-bot indicators about their tarpits (no robots.txt entry, no
       | "nofollow", etc), and want to intentionally feed them poisoned
       | data.
        
       | Der_Einzige wrote:
       | Gonna be funny when the AI systems autonomously engineer micro
       | kill-bots which kill those smug tarpits creators because of
       | "Roko's basilisk". Hope you laugh into the grave :p
        
       | amelius wrote:
       | What if the scrapers use breadth-first search?
        
       ___________________________________________________________________
       (page generated 2025-01-28 23:00 UTC)