[HN Gopher] AI haters build tarpits to trap and trick AI scraper...
___________________________________________________________________
AI haters build tarpits to trap and trick AI scrapers that ignore
robots.txt
Author : hilux
Score : 32 points
Date : 2025-01-28 22:18 UTC (41 minutes ago)
(HTM) web link (arstechnica.com)
(TXT) w3m dump (arstechnica.com)
| dinobones wrote:
| ok...
|
| if depth > 5 and if sem_hash(content) in hist: return
| anotherhue wrote:
| I can generate more random content than you can store.
| rossdavidh wrote:
| It's a radar gun/radar detector kind of situation. You can
| always change your strategies; countless top-level links that
| are shown on the menu in a way that no human sees them (due to
| font color or size or etc). Small number of pages with a few
| that go on forever, in a way that any human would stop reading
| but a bot may have trouble detecting. Real (for humans) text in
| images, with endless invisible text for bots. Etc.
|
| I think it will probably make it harder for screen readers,
| unfortunately.
| andrewfromx wrote:
| "trap crawlers in infinite mazes of gibberish data, potentially
| increasing AI training costs and poisoning datasets. While their
| effectiveness is debated, creators see them as a form of
| resistance against unchecked AI development."
| BizarreByte wrote:
| AI haters? I don't hate AI I just don't want things I've created
| being used to enrich multi-billion dollar companies for free.
| These companies are behaving poorly and they should expect this
| kind of push back.
| someone_eu wrote:
| It is not even just the copyright issue.
|
| The article completely misses the point that AI scrapers are
| not a "future threat of AI domination". They already do damage
| by DDOSing site's networking infrastructure and inflicting very
| real costs to a site hoster.
|
| Even when the data is completely free, like in case of
| Wikipedia or OpenstreetMaps, scraping it is unethical and
| should be illegal. Most of the open data resources have
| procedures, which allow downloading of the data in the archived
| form, without need for scraping. They are built with sharing in
| mind.
|
| So the arguments the article tries to use (what if it is for
| public good?) has no sense. 1) it is not 2) there are many ways
| to fetch the open data properly and respectfully.
| senko wrote:
| [delayed]
| anotherhue wrote:
| I wonder if we could run the cheapest smallest 1b model (or
| smaller) llm to generate this data in a way that's just plausible
| enough to be ingested.
|
| A little note in robots.txt offering commercial terms would also
| be available.
| Dwedit wrote:
| Usually it's just a Markov Chain rather than an LLM to generate
| the garbage data.
| Dwedit wrote:
| There's also the other kind of AI haters who do not give any
| anti-bot indicators about their tarpits (no robots.txt entry, no
| "nofollow", etc), and want to intentionally feed them poisoned
| data.
| Der_Einzige wrote:
| Gonna be funny when the AI systems autonomously engineer micro
| kill-bots which kill those smug tarpits creators because of
| "Roko's basilisk". Hope you laugh into the grave :p
| amelius wrote:
| What if the scrapers use breadth-first search?
___________________________________________________________________
(page generated 2025-01-28 23:00 UTC)