[HN Gopher] Cloudflare Introduces Default Blocking of A.I. Data ...
       ___________________________________________________________________
        
       Cloudflare Introduces Default Blocking of A.I. Data Scrapers
        
       Author : stephendause
       Score  : 335 points
       Date   : 2025-07-02 13:28 UTC (9 hours ago)
        
 (HTM) web link (www.nytimes.com)
 (TXT) w3m dump (www.nytimes.com)
        
       | cmg wrote:
       | Archive link: https://archive.ph/ARnyu
        
       | badlibrarian wrote:
       | Did they ever fix the auto-blocking of RSS feeds?
       | 
       | https://news.ycombinator.com/item?id=41864632
        
       | blakesterz wrote:
       | The list of bots is pretty short right now:
       | 
       | https://developers.cloudflare.com/bots/concepts/bot/#ai-bots
        
         | ZiiS wrote:
         | Enough to more than half the traffic to most sites if the
         | blocks hold.
        
         | hennell wrote:
         | Cloudflare sees a lot of the web traffic. I assume these are
         | the biggest bots they're seeing right now, and any new
         | contenders would be added as they find them. Probably
         | impossible to really block everything, but they've got the web-
         | coverage to detect more than most.
        
           | TechDebtDevin wrote:
           | They are lying. They cant detect crawlers unless we tell them
           | we are who we are.
        
         | JimDabell wrote:
         | > AI bots
         | 
         | > You can opt into a managed rule that will block bots that we
         | categorize as artificial intelligence (AI) crawlers ("AI Bots")
         | from visiting your website. Customers may choose to do this to
         | prevent AI-related usage of their content, such as training
         | large language models (LLM).
         | 
         | > CCBot (Common Crawl)
         | 
         | Common Crawl is not an AI bot:
         | 
         | https://commoncrawl.org
        
           | johneth wrote:
           | The data it collects is used by AI companies, though.
        
       | Spivak wrote:
       | Poor ChatGPT-User, nobody understands you. Blocking a real user
       | because of the, admittedly odd, browser they're using misses the
       | point.
        
       | Roark66 wrote:
       | This is a bit silly. Slowing down, yes, but blocking? People who
       | *really* want that content will find a way and this will hit
       | everyone else instead that will have to do silly riddles before
       | following every link or run crypto mining for them before being
       | shown the content .
       | 
       | I recently went to a big local auction site on which I buy
       | frequently and I got one of these "we detected unusual traffic
       | from your network" messages. And "prove you're human". Which was
       | followed by "you completed the capcha in 0.4s your IP is banned".
       | Really? Am I supposed to slow down my browsing now? I tried a
       | different browser, a different OS, logging on,clearing cookies,
       | etc. Same result when I tried a search. It took 4h after
       | contacting their customer service to unblock it. And the
       | explanation was "you're clicking too fast".
       | 
       | At some point it just becomes a farce and the hassle is not worth
       | the content. Also, while my story doesn't involve any bots
       | perhaps a time will come when local LLMs will be good enough that
       | I'll be able to tell one "reorder my cat food" and it will go and
       | do it. Why are they so determined to "stop it" (spoiler, they
       | can't).
       | 
       | For anyone who says LLMs are already capable of ordering cat food
       | I say not so fast. First the cat food has to be on sale/offer
       | (sometimes combined with extras). Second it is supposed to be
       | healthy (no grains) and third the taste needs to be to my cats
       | liking. So far I'm not going to trust a LLM with this.
        
       | Sol- wrote:
       | Do the major AI companies actually honor robots.txt? Even if some
       | of their publicly known crawlers might do it, surely they have
       | surreptitious campaigns where they do some hidden crawling, just
       | like how they illegally pirate books, images and user data to
       | train on.
        
         | px43 wrote:
         | There's a lack of clarity, but it seems likely to me that a
         | majority of this traffic is actually people asking questions to
         | the AI, and the AI going out and researching for answers. When
         | the AI tools are being used like a web browser to do research,
         | should they still be adhering to robots.txt, or is that only
         | intended for search indexing?
        
         | chasd00 wrote:
         | My thought too, honoring robots.txt is just a convention.
         | There's no requirement to follow robots.txt, or at least
         | certainly no technical requirement. I don't think there's any
         | automatic legal requirement either.
         | 
         | Maybe sites could add "you must honor policies set in
         | robots.txt" to something like a terms of service but I have no
         | idea if that would have enough teeth for a crawler to give up.
        
           | TechDebtDevin wrote:
           | Cloudflare snd their customera have been desperately for
           | years trying to kill scrapers in court. This is all.
           | Meaningless, but they are probably gearing up for another
           | legal battle to define robots.txt as a legal contract. Theyre
           | going to use this marketplace theyre scamming people with to
           | do it. They will fail.
        
         | mschuster91 wrote:
         | Cloudflare, for all I hate their role as a gatekeeper these
         | days, actually has the leverage to force the AI companies to
         | bend.
        
         | deepsun wrote:
         | Hard to tell, because minor crawlers mimic major companies to
         | not getting banned.
        
       | btown wrote:
       | The headline is somewhat misleading: sites using Cloudflare now
       | have an _opt-in_ option to quickly block all AI bots, but it won
       | 't be turned on by default for sites using Cloudflare.
       | 
       | The idea that Cloudflare _could_ do the latter at the sole
       | discretion of its leadership, though, is indicative of the level
       | of power Cloudflare holds.
        
         | bitpush wrote:
         | It is now an adversarial relationship between aibots and
         | website, and cloudflare is merely reacting to it.
         | 
         | Would you say the same for ddos protection? Isn't that the same
         | as well?
        
           | TechDebtDevin wrote:
           | They arent doing anything. They are attempting to insert
           | themselves into the middle of a marketplace (that doesnt
           | exist and never will) where scrapers pay for IP. They think
           | theyre going to profit off the bots, not protect your site.
           | Dont fall for their scam.
        
             | bitpush wrote:
             | What do you mean they are trying to insert themselves? If I
             | have a website that I host with cloudflare, I (as the
             | rightful website owner) has inserted Cloudflare in between.
             | 
             | It isnt CF going around saying, that's a nice website you
             | have there. I'm gonna put myself in between.
        
         | GrayShade wrote:
         | > sites using Cloudflare now have an opt-in option to quickly
         | block all AI bots, but it won't be turned on by default for
         | sites using Cloudflare
         | 
         | Do you have a source for that?
         | https://blog.cloudflare.com/content-independence-day-no-ai-c...
         | does say "changing the default".
        
           | mattcollins wrote:
           | "This feature is available to all customers, meaning anyone
           | can enable this today from the Cloudflare dashboard."
           | 
           | https://blog.cloudflare.com/control-content-use-for-ai-
           | train...
        
         | TechDebtDevin wrote:
         | They cant do anything other than bog down the internet. I
         | havent found a single cf provided challenge I havent been able
         | to get past in < half a day.
         | 
         | This is simply juat the first step in them implementing a
         | marketplace and trying to get into LLM SEO. They dont care
         | about your site or protecting it. They are gearing up to start
         | making a cut in the Middle between scrapers and publishers. Why
         | wouldnt I go DIRECTLY to the publisher and make a deal. So dumb
         | I hate cf so much.
         | 
         | The only thing cloudflare knows how to do is MITM attacks.
        
           | Marsymars wrote:
           | So what would you suggest as an alternative if I have a site
           | where I don't want the content used for LLM training?
        
             | fkyoureadthedoc wrote:
             | Auth? Because whatever Cloudflare is doing isn't going to
             | stop anyone serious about scraping data.
        
               | Marsymars wrote:
               | Let's say I'm talking about content that I don't want
               | behind an auth wall. Is your position simply that all
               | such sites should abandon any efforts to not have the
               | content used for LLM training?
        
               | mattl wrote:
               | If you find a solution that's not auth please let me
               | know.
        
       | postalcoder wrote:
       | > When you enable this feature via a pre-configured managed rule,
       | Cloudflare can detect and block verified AI bots that comply with
       | robots.txt and respect crawl rates, and do not hide their
       | behavior from your website. The rule has also been expanded to
       | include more signatures of AI bots that do not follow the rules.
       | 
       | We already know companies like Perplexity are masking their
       | traffic. I'm sure there's more than meets the eye, but taking
       | this at face value, doesn't punishing respectful and transparent
       | bots only incentivize obfuscation?
       | 
       | edit: This link[0], posted in a comment elsewhere, addresses this
       | question. tldr, obfuscation doesn't work.                 > We
       | leverage Cloudflare global signals to calculate our Bot Score,
       | which for AI bots like the one above, reflects that we correctly
       | identify and score them as a "likely bot."            > When bad
       | actors attempt to crawl websites at scale, they generally use
       | tools and frameworks that we are able to fingerprint. For every
       | fingerprint we see, we use Cloudflare's network, which sees over
       | 57 million requests per second on average, to understand how much
       | we should trust this fingerprint. To power our models, we compute
       | global aggregates across many signals. Based on these signals,
       | our models were able to appropriately flag traffic from evasive
       | AI bots, like the example mentioned above, as bots.
       | 
       | [0] https://blog.cloudflare.com/declaring-your-aindependence-
       | blo...
        
         | colechristensen wrote:
         | >doesn't punishing respectful and transparent bots only
         | incentivize obfuscation?
         | 
         | They're cloudflare and it's not like it's particularly easy to
         | hide a bot that is scraping large chunks of the Internet from
         | them. On top of the fact that they can fingerprint any of your
         | sneaky usage, large companies have to work with them so I can
         | only assume there are channels of communication where
         | cloudflare can have a little talk with you about your bad
         | behavior. I don't know how often lawyers are involved but I
         | would expect them to be.
        
         | jerf wrote:
         | "doesn't punishing respectful and transparent bots only
         | incentivize obfuscation?"
         | 
         | Sure, but we crossed that bridge over 20 years ago. It's not
         | creating an arms race where there wasn't already one.
         | 
         | Which is my generic response to everyone bringing similar ideas
         | up. "But the bots could just...", yeah, they've been doing it
         | for 20+ years and people have been fighting it for just as
         | long. Not a new problem, not a new set of solutions, no
         | prospect of the arms race ending any time soon, none of this is
         | new.
        
         | hombre_fatal wrote:
         | Next line:
         | 
         | > The rule has also been expanded to include more signatures of
         | AI bots that do not follow the rules.
         | 
         | The Block AI Bots rule on the Super Bot Fight Mode page does
         | filter out most bot traffic. I was getting 10x the traffic from
         | bots than I was from users.
         | 
         | It definitely doesn't rely on robots.txt or user agent. I had
         | to write a page rule bypass just to let my own tooling work on
         | my website after enabling it.
        
           | account42 wrote:
           | How many of those "bots" you are filtering are actually bots
           | and how many are regular users buttflare has misidentified as
           | bots?
        
             | hombre_fatal wrote:
             | Pretty simple to see this if you've run a website: compare
             | your analytics pre-bot to post-bot to post-bot-blocker.
             | 
             | There is a clear moment where you land on AI bot radar. For
             | my large forum, it was a month ago.
             | 
             | Overnight, "72 users are viewing General Discussion" turned
             | into "1720 users".
             | 
             | 40% requests being cached turned into 3% of requests are
             | cached.
        
         | fluidcruft wrote:
         | Cloudflare already knows how to make the web hell for people
         | they don't like.
         | 
         | I read the robots.txt entries as those AI bots that will be not
         | marked as "malicious" and that will have the opportunity to be
         | allowed by websites. The rest will be given the Cloudflare
         | special.
        
       | dougb5 wrote:
       | > Cloudflare can detect and block verified AI bots that comply
       | with robots.txt and respect crawl rates, and do not hide their
       | behavior from your website
       | 
       | It's the bots that _do_ hide their behavior -- via residential
       | proxy services -- that are causing most of the burden, for my
       | site anyway. Not these large commercial AI vendors.
        
       | alganet wrote:
       | > If A.I. companies freely use data from various websites without
       | permission or payment, people will be discouraged from creating
       | new digital content
       | 
       | I don't see a way out of this happening. AI fundamentally
       | discourages other forms of digital interaction as it grows.
       | 
       | Its mechanism of growing is killing other kinds of digital
       | content. It will eventually kill the web, which is, ironically,
       | its main source of food.
        
         | preachermon wrote:
         | just like capitalism has now turned to exploiting people as its
         | main input?
        
           | alganet wrote:
           | These kinds of comparisons rarely lead to good discussions.
           | 
           | Let's instead be focused and talk about real stuff.
           | 
           | Consider https://learnpythonthehardway.org/ for example. It
           | has influenced a generation of Python developers. Not just
           | the main website, but the tons of Python code and Python-
           | related content it inspired.
           | 
           | Why would anyone write these kinds of
           | textbooks/websites/guides if AI can replace them? AI
           | companies are effectively broadcasting you don't need the
           | hard way anymore, you can just vibe.
           | 
           | Arguibly though, without the existance of Learn Python the
           | Hard Way and similar content, AI would be worse at writing
           | Python stuff. That's what I mean by "main source of food",
           | good content that influences a lot of people. Net-positive
           | effects hard to predict or even identify except for the more
           | popular cases (such as LPTHW).
           | 
           | If my prediction is right, no one will notice that good
           | content has stopped being produced. It will appear as if
           | content is being created in generally the same way as before,
           | but in reality, these long tail initiatives like LPTHW will
           | have ceased before anyone can do anything about it.
           | 
           | Again, I don't see a way out of this scenario. Not for AI
           | companies, not for content writers. This is going to happen.
           | The world in which I'm wrong is the best one.
        
             | mfost wrote:
             | In a similar vein, I remember people advocating for
             | replacing new untrained hires with AI. After all, a
             | competent senior engineer is needed to validate the
             | contributions of the new hires anyway and they can do the
             | same checking the AI code.
             | 
             | But then, how would you even train and replace those
             | competent seniior engineers that do the filtering when they
             | retire? The whole system was predicated on having a chain
             | of new hires that gain experience in the process.
        
               | alganet wrote:
               | From what I could perceive, companies believe coding AIs
               | will eventually learn to both code and teach better than
               | seniors.
               | 
               | This is based on two assumptions:
               | 
               | - AI will get better. Developers using the system will
               | transfer their knowledge to it.
               | 
               | - Seniors in a couple of years will be different. They
               | should be those who can engage with the AI feedback loop.
               | 
               | Here's why I think it won't work:
               | 
               | - Senior developers learn more than they can produce.
               | There is knowledge they never transfer to what they work
               | on. Internalized knowledge that never materializes
               | directly into code. _But it materializes indirectly_.
               | 
               | - Senior developer knowledge come from "schools", not
               | just reading. These schools are not real physical
               | locations. They're traditions, or ideas, that form a very
               | long tail. These ideas, again, are not directly
               | transferrable to code or prose.
               | 
               | - Juniors get embarrassed. You say "stop making this
               | nonsense", and they'll stop and reflect, because they
               | respect seniors. They might disagree, but a pea was then
               | placed under their matress, and they'll think about "this
               | nonsense" you told them to stop doing and why. That is
               | how they get better. So far, AI has not demonstrated
               | being able to do that.
               | 
               | The production of quality content is an aspect of one of
               | those "schools of thought". You are supposed to bear the
               | responsibility of passing the knowledge. Keeping lean
               | codebases easy to understand is also a hallmark of many
               | schools of thought. Working from fundamentals is another
               | one of those ideas, etc.
        
           | greenchair wrote:
           | nothing's perfect but it is still better than the other
           | options
        
         | fennecfoxy wrote:
         | Additionally, ad blocker usage is apparently at 30%. So it's a
         | redundant or more nuanced argument, really.
        
           | account42 wrote:
           | Ad blockers only discourage commercialized content creation,
           | not all of it. IMO that actually improves the quality of the
           | content created.
        
         | spwa4 wrote:
         | Yes what _everyone_ wants to do with AI: generate entertainment
         | and interactions with humans, including economical ones, will
         | need to happen or AI will starve.
        
           | alganet wrote:
           | That's what is going to make it starve. Belly full, but of
           | its own shit being tossed around humans seeking cheap copouts
           | of doing actual work.
        
         | BrouteMinou wrote:
         | Just like cancer?
        
       | jasonthorsness wrote:
       | I turned this on and it adjusts the robots.txt automatically; not
       | sure what else it is doing.
       | 
       | # NOTICE: The collection of content and other data on this # site
       | through automated means, including any device, tool, # or process
       | designed to data mine or scrape content, is # prohibited except
       | (1) for the purpose of search engine indexing or # artificial
       | intelligence retrieval augmented generation or (2) with express #
       | written permission from this site's operator.
       | 
       | # To request permission to license our intellectual # property
       | and/or other materials, please contact this # site's operator
       | directly.
       | 
       | # BEGIN Cloudflare Managed content
       | 
       | User-agent: Amazonbot Disallow: /
       | 
       | User-agent: Applebot-Extended Disallow: /
       | 
       | User-agent: Bytespider Disallow: /
       | 
       | User-agent: CCBot Disallow: /
       | 
       | User-agent: ClaudeBot Disallow: /
       | 
       | User-agent: Google-Extended Disallow: /
       | 
       | User-agent: GPTBot Disallow: /
       | 
       | User-agent: meta-externalagent Disallow: /
       | 
       | # END Cloudflare Managed Content User-agent: * Disallow: /*
       | Allow: /$
        
         | xyst wrote:
         | So in addition to updating the robots.txt file, which really
         | only blocks a small number of them.
         | 
         | Seems CF has been gathering data and profiling these malicious
         | agents.
         | 
         | This post by CF elaborates a bit further:
         | https://blog.cloudflare.com/declaring-your-aindependence-blo...
         | 
         | Basically becomes a game of cat and mouse.
        
         | postalcoder wrote:
         | This is interesting. The reasoning and response don't line up.
         | > Cloudflare is making the change to protect original content
         | on the internet, Mr. Prince said. If A.I. companies freely use
         | data from various websites without permission or payment,
         | people will be discouraged from creating new digital content,
         | he said            >  prohibited except for the purpose of [..]
         | artificial intelligence retrieval augmented generation
         | 
         | This seems to be targeted at taxing training of language
         | models, but why an exclusion for the RAG stuff? That seems like
         | it has a much greater immediate impact for online content
         | creators, for whom the bots are obviating a click.
        
           | fennecfoxy wrote:
           | With that opinion, are you also suggesting that we ban ad
           | blockers? Because it's better I not click & consume resources
           | than click and not be served ads, basically just costing the
           | host money.
           | 
           | It means sense to allow for RAG in the same way that search
           | engines provide a snippet of an important chunk of the page.
           | 
           | A blog author could not complain that their blog is getting
           | ragged when they're extremely liable to be Google/whatever
           | searching all day and basically consuming others' content in
           | exactly the same way that they're trying to disparage.
        
             | postalcoder wrote:
             | I don't think we should ban ad blockers, but I also think
             | it's fair to suggest that the loss of organic traffic could
             | be affecting the incentive to create new digital content,
             | at least as much as the fear of having your content
             | absorbed into an LLM's training data.
        
               | Boldened15 wrote:
               | IMO the backlash against LLMs is more philosophical, a
               | lot of people don't like them or the idea of one learning
               | from their content. Unless your website has some unique
               | niche information unavailable anywhere else there's no
               | direct personal risk. RAG would be a more direct threat
               | if anything.
        
               | toomuchtodo wrote:
               | It's really about who is getting the value from the work
               | of the content. If content creators of all sorts have
               | their work consumed by LLMs, and LLM orgs charge for it
               | can capture all the value, why should people create to
               | have their work vacuumed up for the robot's benefit? For
               | exposure? You can't eat or pay rent with exposure. Humans
               | must get paid, and LLMs (foundational models and output
               | using RAG) cannot improve without a stream of works and
               | data humans create.
               | 
               | Whether you call it training or something else is
               | irrelevant, it's really exploitation of human work and
               | effort for AI shareholder returns and tech worker comp
               | (if those who create aren't compensated). And the
               | technocracy has not been, based on the evidence, great
               | stewards of the power they obtain through this. Pay the
               | humans for their work.
        
               | o11c wrote:
               | It's not philosophical, it's economical.
               | 
               | AI scrapers increase traffic by maybe 10x (this varies
               | per site) but provide no real value whatsoever to
               | _anyone_. If you look at various forms of  "value":
               | 
               | * Saying "this uses AI" might make numbers go up on the
               | stock market if you manage to persuade people it will
               | make numbers go up (see also: the market will remain
               | irrational longer than you can remain solvent).
               | 
               | * Saying "this uses AI" might fulfill some corporate
               | mandate.
               | 
               | * Asking AI to solve a problem (for which you would
               | actually use the solution) allows you to "launder" the
               | copyright of whatever source it is paraphrasing (it's
               | well established that LLMs fail entirely if a question
               | isn't found within their training set). Pirating it
               | directly provides the same value, with significantly less
               | errors/handholding.
               | 
               | * Asking AI to entertain you ... well, there's the
               | novelty factor I guess, but even if people refuse to
               | train themselves out of that obsession, the world is
               | still far too full of options for any individual to
               | explore them all. Even just the question of "what kind of
               | ways can I throw some kind of ball around" has more
               | answers than probably anyone here knows.
               | 
               | What am I missing?
        
             | ijk wrote:
             | What I want to know is if the flood of scraping everyone
             | has been complaining about is coming from people trying to
             | scrape for training or bots doing RAG search.
             | 
             | I get that everyone wants data, but presumably the big
             | players already scraped the web. Do they really need to do
             | it again? Or is it bit players reproducing data that's
             | likely already in the training set? Or is it really that
             | valuable to have your own scraped copy of internet scale
             | data?
             | 
             | I feel like I'm missing something here. My expectation is
             | that RAG traffic is going to be orders of magnitude higher
             | than scraping for training. Not that it would be easy to
             | measure from the outside.
        
               | mattcollins wrote:
               | I wondered about this, too.
               | 
               | Cloudflare have some recent data about traffic from bots
               | (https://blog.cloudflare.com/from-googlebot-to-gptbot-
               | whos-cr...) which indicates that, for the time being, the
               | overwhelming majority of the bot requests are for AI
               | training and not for RAG.
        
               | wiether wrote:
               | You should ask Zuck, since, for what we've seen and what
               | we were ask to act against, Meta is the main culprit in
               | scraping every single page of websites, multiple times a
               | day.
               | 
               | And I'm talking about ecommerce websites, with their bot
               | scraping every variation of each product, multiple times
               | a day.
        
           | lxgr wrote:
           | More and more people use ChatGPT for search, so blocking that
           | doesn't seem like a successful strategy long-term.
        
         | bee_rider wrote:
         | I wonder... Google scrapes for indexing and for AI, right? I
         | wonder if they will eventually say: ok, you can have me or not,
         | if you don't want to help train my AI you won't get my searches
         | either. That's a tough deal but it is sort of self-consistent.
        
           | giancarlostoro wrote:
           | "Embrace, Extend, Extinguish" Google's mantra. And yes, I
           | know about Microsoft's history with that phrase ;) But Google
           | has done this with email, browsers (Google has web apps that
           | run fine on Firefox but request you use Chrome), Linux
           | (Android), and I'm sure there's others I am forgetting about.
           | 
           | So yeah, I too could see them doing this.
        
           | mrweasel wrote:
           | Very few people seems to be complaining that Google crashes
           | their sites. Google also publish their crawlers IP ranges,
           | but you really don't need to rate-limit Google, they know how
           | to back off and not overload sites.
        
             | Symbiote wrote:
             | In theory -- in practise I've had to limit Google on two
             | large sites at work. I currently have them limited to 10/s
             | for non-cached requests.
        
         | 1vuio0pswjnm7 wrote:
         | "User-agent: CCBot disallow: /"
         | 
         | Is Common Crawl exclusively for "AI"
         | 
         | CCBot was already in so many robots.txt prior to this
         | 
         | How is CC supposed to know or control how people use the
         | archive contents
         | 
         | What if CC is relying on fair use                  # To request
         | permission to license our intellectual        # property
         | andd/or other materials, please contact this        # site's
         | operator directly
         | 
         | If the operator has no intellectual property rights in the
         | material, then do they need permission from the rights holders
         | to license such materials for use in creating LLMs and collect
         | licensing fees
         | 
         | Is it common for website terms and conditions to permit site
         | operators to sublicense other peoples' ("users") work for use
         | in creating LLMs for a fee
         | 
         | Is this fee shared with the rights holders
        
           | nemomarx wrote:
           | Read a tos and notice that you give the site operators
           | unlimited license to reproduce or spread your works, almost
           | on any site. it's required to host and show the content
           | essentially
        
           | ronsor wrote:
           | # To request permission to license our intellectual        #
           | property andd/or other materials, please contact this
           | # site's operator directly
           | 
           | Scrapers don't accept the terms of service.
           | 
           | Ironically, I've only ever scraped sites that block CCBot,
           | otherwise I'd rather go to Common Crawl for the data.
        
         | Bender wrote:
         | For my silly hobby sites I just return status 444 _close the
         | connection_ for anything that has case-insentive  "bot" in the
         | UA requesting anything other than robots.txt, humans.txt,
         | favicon.ico, etc... This would also drop search engines but I
         | blackhole route most of their CIDR blocks. I'm probably the
         | only one here that would do this.
        
           | sneak wrote:
           | How does a bot scraping your silly hobby sites for any
           | purpose harm or negatively affect you in any way?
        
         | slenk wrote:
         | I thought I saw cloudflare insert noindex links?
        
         | swyx wrote:
         | what actually are the consequences of ignoring robots.txt
         | (apart from DDOS)? have any of these cases ended up in court at
         | all?
        
         | lxgr wrote:
         | That's at least a more reasonable default than that I've seen
         | at least one newspaper do, which is to block both LLM scrapers
         | _and_ things like ChatGPT 's search feature explicitly.
        
       | cratermoon wrote:
       | I'm still not sure this is going to be very effective, as so many
       | of the worst offenders don't identify themselves as bots, and
       | often change their user agent. Has Cloudflare said anything about
       | identifying the bad actors?
        
         | chasd00 wrote:
         | i've mentioned this in a couple replies so maybe i'm wrong but
         | it's up to the client to obey robots.txt. Why would they not
         | just ignore it? Unless there's some legal consequence not
         | complying with robots.txt then why even follow it? There's no
         | technical enforcement of the policies in the file, it's up to
         | the client to honor them.
        
           | kentonv wrote:
           | > There's no technical enforcement of the policies in the
           | file, it's up to the client to honor them.
           | 
           | That's incorrect. Cloudflare does in fact enforce this at a
           | technical level. Cloudflare has been doing bot detection for
           | years and can pretty reliably detect when bots are not
           | following robots.txt and then block them.
        
         | GrayShade wrote:
         | Yes, they have over the years, for example
         | https://blog.cloudflare.com/residential-proxy-bot-
         | detection-..., https://blog.cloudflare.com/cloudflare-bot-
         | management-machin..., https://blog.cloudflare.com/introducing-
         | bot-analytics/.
        
       | lucasyvas wrote:
       | I fail to see how this won't just result in UA string or other
       | obfuscation.
        
         | chasd00 wrote:
         | a crawler doesn't have to change anything, they can just ignore
         | the robots.txt file. It's up to the client to read robots.txt
         | and follow its directives but there's no technical reason why
         | the client cannot just ignore everything in the file period.
        
         | kube-system wrote:
         | Cloudflare's filtering is already way more sophisticated than
         | just looking at UA string or other voluntary reporting. They're
         | almost certainly using fingerprinting and behavioral analytics.
        
       | gazpacho wrote:
       | From an open source projects perspective we'd want to disable
       | this on our docs sites. We actually want those to be very
       | discoverable by LLMs, during training or online usage.
        
       | rorylaitila wrote:
       | Unfortunately I think pissing into the wind. Information websites
       | are all but dead. AI contains all published human information. If
       | you have positioned your website as an answer to a question, it
       | won't survive that way.
       | 
       | "Information" is dead but content is not. Stories, empathy,
       | community, connection, products, services. Content of this
       | variety is exploding.
       | 
       | The big challenge is discoverability. Before, information
       | arbitrage was one pathway to get your content discovered, or to
       | skim a profit. This is over with AI. New means of discovery are
       | necessary, largely network and community based. AI will throw you
       | a few bones, but it will be 10% of what SEO did.
        
         | fennecfoxy wrote:
         | >AI contains all published human information
         | 
         | No, it most certainly does not. It was certainly trained on
         | large swathes of human knowledge/interactions.
         | 
         | A model that consists of a perfect representation/compression
         | of all this info is a zip file, not a model file.
        
           | rorylaitila wrote:
           | AI providers have scrapped and will continue to, all internet
           | published information or virtually so. Since "Information" is
           | infinite, AI cannot contain "all information" in a complete
           | sense. But it certainly answers almost everything that
           | matters for any existing search query that has ever been
           | targeted by a webpage that is crawlable.
           | 
           | In any case, as manifest by real world SEO, which is
           | plummeting in traffic for informational queries, the effect
           | is the same. This real world impact is what matters and will
           | not be reversed, regardless of attempts at blocking.
        
         | ozgrakkurt wrote:
         | You are assuming LLMs will replace search engines. Why is this
         | the case?
         | 
         | To me it seems like there has to be so much optimization for
         | this to happen that, it is not likely. LLM answers are slow and
         | unreliable. Even using something like perplexity doesn't give
         | much value over using a regular search engine in my experience
        
           | rorylaitila wrote:
           | LLMs will not fully replace search engines, but Google and
           | Bing are evolving to be LLM first, anyhow. So "what is a
           | search engine" today is not what it was yesterday. Let's call
           | the time before LLMs, traditional search. LLM first products
           | bundle some aspect of traditional search. And traditional
           | search is adding LLM answers.
           | 
           | Traditional search will still be highly useful for
           | transactional, product, realtime, and action oriented
           | queries. Also for discovering educational/entertainment
           | content that is valued in of itself and cannot be
           | reformulated by LLM.
        
       | Meekro wrote:
       | I've heard lots of people on HN complaining about bot traffic
       | bogging down their websites, and as a website operator myself I'm
       | honestly puzzled. If you're already using Cloudflare, some basic
       | cache configuration should guarantee that most bot traffic hits
       | the cache and doesn't bog down your servers. And even if you
       | don't want to do that, bandwidth and CPU are so cheap these days
       | that it shouldn't make a difference. Why is everyone so upset?
        
         | deepsiml wrote:
         | Not much into that kind of DevOps. What is a good basic caching
         | in this instance?
        
           | TechDebtDevin wrote:
           | Cloudflare and other CDNs will usually automatically cache
           | your static pages.
        
           | haiku2077 wrote:
           | It comes down to:
           | 
           | 1. Use the Cache-Control header to express how to cache your
           | site correctly (https://developer.mozilla.org/en-
           | US/docs/Web/HTTP/Guides/Cac...)
           | 
           | 2. Use a CDN service, or at least a caching reverse proxy, to
           | serve most of the cacheable requests to reduce load on the
           | (typically much more expensive) origin servers
        
             | mrweasel wrote:
             | Just note that many AI scrapers will go to great length to
             | do cache busting. For some reason many of them feel like
             | they need to get the absolute latest version and don't
             | trust your cache.
        
               | haiku2077 wrote:
               | You can use Cache Control headers to express that your
               | own CDN should aggressively refresh a resource but always
               | serve it to external clients from cache. It's covered in
               | the link under "Managed Caches"
        
               | cortesoft wrote:
               | A CDN can be configured to ignore cache control headers
               | in the requests and cache things anyway.
        
         | conductr wrote:
         | The presumption I'm already using cloudfare is a start. Is this
         | a requirement for maintaining a simple website now?
        
           | haiku2077 wrote:
           | Either that or Anubis (https://anubis.techaro.lol/docs), yes.
        
             | roguecoder wrote:
             | So these companies broke the internet
        
               | haiku2077 wrote:
               | Which companies?
               | 
               | OpenAI, Anthropic, Google? No, their bots are pretty well
               | behaved.
               | 
               | The smaller AI companies deploying bots that don't
               | respect any reasonable rate limits and are scraping the
               | same static pages thousands of times an hour? Yup
        
               | e3bc54b2 wrote:
               | Anecdote, but at least for tiny little server hosting
               | single public repository, _none_ of these companies had
               | 'well behaved' bots. It may be possible that they learned
               | to behave better but I wouldn't know since my only
               | possible recourse was to blacklist them all AND take the
               | repo private.
        
               | haiku2077 wrote:
               | Those are the small companies spoofing their user agent
               | as the big companies to dodge countermeasures.
        
         | noodle wrote:
         | As someone who had some outages due to AI traffic and is now
         | using CloudFlare's tools:
         | 
         | Most of my site is cached in multiple different layers. But
         | some things that I surface to unauthenticated public can't be
         | cached while still being functional. Hammering those endpoints
         | has taken my app down.
         | 
         | Additionally, even though there are multiple layers, things
         | that are expensive to generate can still slip through the
         | cracks. My site has millions of public-facing pages, and a
         | batch of misses that happen at the same time on heavier pages
         | to regenerate can back up requests, which leads to errors, and
         | errors don't result in caches successfully being filled. So the
         | AI traffic keeps hitting those endpoints, they keep not getting
         | cached and keep throwing errors. And it spirals from there.
        
         | x0x0 wrote:
         | It's not complex. I worked on a big site. We did not have the
         | compute or i/o (most particularly db iops) to live generate the
         | site. Massive crawls both generated cold pages / objects (cpu +
         | iops) and yanked them into cache, dramatically worsening cache
         | hit rates. This could easily take down the site.
         | 
         | Cache is expensive at scale. So permitting big or frequent
         | crawls by stupid crawlers either require significant
         | investments in cache or slow down and worsen the site for all
         | users. For whom we, you know, built the site, not to provide
         | training data for companies.
         | 
         | As others have mentioned, Google is significantly more
         | competent than 99.9% of the others. They are very careful to
         | not take your site down and provide, or used to provide,
         | traffic via their search. So it was a trade, not a taking.
         | 
         | Not to mention I prefer not to do business with Cloudflare
         | because I don't like companies that don't publish quota. If
         | going over X means I need an enterprise account that starts at
         | $10k/mo, I need to know the X. Cloudflare's business practice
         | appears to be letting customers exceed that quota then
         | aggressively demanding they pay or they'll be kicked off the
         | service nearly immediately.
        
         | jtolmar wrote:
         | The stories I've heard have been mostly about scraper bots
         | finding APIs like "get all posts in date range" and then
         | hammering that with every combo of start/end date.
        
         | jauntywundrkind wrote:
         | I too am a bit confused / mystified at the strong reaction. But
         | I do expect a lot of badly optimized sites that just want out.
         | 
         | I struggle to think of a web related library that has spread
         | faster than Anubis checker. It's everywhere now!
         | https://github.com/TecharoHQ/anubis
         | 
         | I'm surprised we don't see more efforts to rate limit. I assume
         | many of these are distributed crawlers, but it feels like
         | there's got to be pools of activity spinning up, on a handful
         | of IPs. And that they would be time correlated together pretty
         | clearly. Maybe that's not true. But it feels like the web, more
         | than anything else, needs some open source software to add a
         | lot more 420 Enhance Your Calm responses, as it feels like.
         | https://http.dev/420
        
           | zerocrates wrote:
           | The reaction comes from some combination of
           | 
           | - opposition to generative AI in general
           | 
           | - a view that AI, unlike search which also relies on
           | crawling, offers you no benefits in return
           | 
           | - crawlers from the AI firms being less well-behaved than the
           | legacy search crawlers, not obeying robots.txt, crawling more
           | often, more aggressively, more completely, more redundantly,
           | from more widely-distributed addresses
           | 
           | - companies sneaking in AI crawling underneath their existing
           | tolerated/whitelisted user-agents (Facebook was pretty
           | clearly doing this with "facebookexternalhit" that people
           | would have allowed to get Facebook previews; they eventually
           | made a new agent for their crawling activity)
           | 
           | - a simultaneous huge spike in obvious crawler activity with
           | spoofed user agents: e.g. a constant random cycling between
           | every version of Chrome or Firefox or any browser ever
           | released; who this is or how many different actors it is and
           | whether they're even doing crawling for AI, who knows, but
           | it's a fair bet.
           | 
           | Better optimization and caching can make this all not matter
           | so much but not everything can be cached, and plenty of small
           | operations got by just fine without all this extra traffic,
           | and would get by just fine without it, so can you really
           | blame them for turning to blocking?
        
           | jowea wrote:
           | I'm not an expert on website hosting, but after reading some
           | of the blog posts on Anubis, those people were truly at wit's
           | end trying to block AI scrappers with techniques like the
           | ones you imply.
        
         | Symbiote wrote:
         | That's a pretty big assumption.
         | 
         | The largest site I work on has 100,000s of pages, each in
         | around 10 languages -- that's already millions of pages.
         | 
         | It generally works fine. Yesterday it served just under 1000
         | RPS over the day.
         | 
         | AI crawlers have brought it down when a _single crawler_ has
         | added 100, 200 or more RPS distributed over a wide range of IPs
         | -- it 's not so much the number of extra requests, though it's
         | very disproportionate for one "user", but they can end up
         | hitting an expensive endpoint excluded by robots.txt and
         | protected by other rate-limiting measures, which didn't
         | anticipate a DDoS.
        
       | yodon wrote:
       | Discussed yesterday (270+ comments)[0]
       | 
       | [0]https://news.ycombinator.com/item?id=44432385
        
       | dawnerd wrote:
       | I've been using this for a while on my mastodon server and after
       | a few tweaks to make sure it wasn't blocking legit traffic it's
       | been really working great. Between Microsoft and Meta, they were
       | hitting my services more than any other traffic combined which
       | says a lot of you know how noisy mastodon can be. Server load
       | went down dramatically.
       | 
       | It also completely put a stop to perplexity as far as I can tell.
       | 
       | And the robots file meant nothing, they'd still request it
       | hundreds of thousands of times instead of caching it. Every
       | request they'd hit it first then hit their intended url.
        
         | TechDebtDevin wrote:
         | This does nothing dude. Literally nothing. OpenAI or whoever
         | are just going to hire people like me who dont get caught. Stop
         | ruining the experience of users and allowing cf to fill the
         | internet with more bloated javascript challenge pages and
         | privacy invading fingerprinting. Stop making cf the police of
         | the internet. We're literally handing the internet to this
         | company on a silver platter to do MITM attacks on our privacy
         | and god knows what else. Fucking wild.
        
           | fluidcruft wrote:
           | They literally said it significantly reduced their server
           | resource usage. Are you suggesting they are lying?
        
           | drowsspa wrote:
           | Why do you think you have the moral high ground here?
        
           | dawnerd wrote:
           | Well the alternative is to not have an instance at all so...
           | what do you suggest? I'm not paying for the other services,
           | it's already expensive enough to run the site.
           | 
           | The goal isn't to stop 100% of scrapers, it was to reduce
           | server load to a level that wasn't killing the site.
        
           | jowea wrote:
           | You want them to pay the server costs to serve content to AI
           | scrappers for free? The alternative is Anubis, which is maybe
           | equally annoying to users in a different way.
        
         | danielspace23 wrote:
         | Have you considered Anubis? I know it's harder to install, but
         | personally, I think the point of Mastodon is trying to avoid
         | centralization where possible, and CloudFlare is one of the
         | corporations that are keeping the internet centralized.
        
           | dawnerd wrote:
           | Haven't heard of it, will look into it. I agree, I'd rather
           | not have cloudflare but for what they provide for free it's a
           | tough offer to pass up
        
       | account42 wrote:
       | Yay, looking forward to more CAPTCHAs as a regular user.
        
       | thephotonsphere wrote:
       | account wall :-(
        
       | deadbabe wrote:
       | No one else can really do this except Cloudflare.
        
       | dirkc wrote:
       | I assume they will "protect original content online" by blocking
       | LLM clients from ingesting data as context?
       | 
       | I'm not optimistic that you can effectively block your original
       | content from ending up in training sets by simply blocking the
       | bots. For now I just assume that anything I put online will end
       | up being used to train some LLM
        
       | NullCascade wrote:
       | How would you do the opposite of this? Optimize your content to
       | be more likely crawled by AI bots? I know traditional Google-
       | focused SEO is not enough because these AI bots often use other
       | web search/indexing APIs.
        
         | TechDebtDevin wrote:
         | There are script tags you can put in your site from LLM SEO
         | companies if you want your content to be indexed by Perplexity
         | or OpenAI. Theyre kind of too new for me to reccomend.
        
       | zargath wrote:
       | Sounds very basic, sadly.
       | 
       | Anybody know why these web crawling/bot standards are not
       | evolving ? I believe robots.txt was invented in 1994(thx
       | chatgpt). People have tried with sitemaps, RSS and IndexNow, but
       | its like huge$$ organizations are depending on HelloWorld.bas
       | tech to control their entire platform.
       | 
       | I want to spin up endpoints/mcp/etc. and let intelligent bots
       | communicate with my services. Let them ask for access, ask for
       | content, pay for content, etc. I want to offer solutions for bots
       | to consume my content, instead of having to choose between full
       | or no access.
       | 
       | I am all for AI, but please try to do better. Right now the
       | internet is about to be eaten up by stupid bot farms and served
       | into chat screens. They dont want to refer back to their source
       | and when they do its with insane error rates.
        
         | TechDebtDevin wrote:
         | This comment seems like it comes from a Cloudflare employee.
         | 
         | This is clearly the first step in cf building out a marketplace
         | where they will (fail) at attempting to be the middleman in a
         | useless market between crawlers and publishers.
        
           | zargath wrote:
           | nah, disappointed cf customer
        
         | stereolambda wrote:
         | > I believe robots.txt was invented in 1994(thx chatgpt).
         | 
         | Not to pick on you, but I find it quicker to open new tab and
         | do "!w robots.txt" (for search engines supporting the bang
         | notation) or "wiki robots.txt"<click> (for Google I guess). The
         | answer is right there, no need to explain to LLM what I want or
         | verify [1].
         | 
         | [1] Ok, Wikipedia can be wrong, but at least it is a commonly
         | accessible source of wrong I can point people to if they call
         | me out. Plus my predictive model of Wikipedia wrongness gives
         | me pretty low likelihood for something like this, while for
         | ChatGPT it is more random.
        
         | reaperducer wrote:
         | _robots.txt was invented in 1994(thx chatgpt)_
         | 
         | Thought of and discussed as a possibility in 1994.
         | 
         | Proposed as a standard in 2019.
         | 
         | Adopted as a standard in 2022.
         | 
         | Thanks, IETF.
        
       | StochasticLi wrote:
       | _ehem_ https://github.com/Kaliiiiiiiiii-Vinyzu/patchright
        
       | j45 wrote:
       | This is interesting. I'm a fan of Cloudflare, and appreciate all
       | the free tiers they put out there for many.
       | 
       | Today I see this article about Cloudflare blocking scrapers.
       | There are useful and legitimate cases where I ask Claude to go
       | research something for me. I'm not sure if Cloudflare discerns
       | legitimate search/research traffic from an AI client vs scraping.
       | Of the sites that are blocked by default will include content by
       | small creators (unless on major platforms with deal?), while the
       | big guys who have something to sell like an Amazon, etc, will
       | likely be able to facilitate and afford a deal to show up more in
       | the results.
       | 
       | A few days ago, Cloudflare is also looking to charge AI companies
       | to scrape the content, which is cached copies of other people's
       | content. I'm guessing it will involve paying the owners of the
       | data at some point as well. Being able to exclude it from this
       | purpose (sell/license content, or scrape) would be a useful
       | lever.
       | 
       | Putting those two stories together:
       | 
       | - Is this a new form of showing up in the AISEO (Search
       | everywhere optimization) to show up in an AI's corpus or ability
       | to search the web, or paying licensing fees instead of
       | advertising fees.. these could be new business models which are
       | interesting, but trying to see where these steps may vector ahead
       | towards, and what to think about today.
       | 
       | - With training data being the most valuable thing for AI
       | companies, and this is another avenue for revenue for Cloudflare,
       | this can look like a solution which helps with content licensing
       | as a service.
       | 
       | I'd like to see where abstracting this out further ends up going
       | 
       | Maybe I'm missing something, is anyone else seeing it this way,
       | or another way that's illuminating to them? Is anyone thinking
       | about rolling their own service for whatever parts of Cloudflare
       | they're using?
        
         | ec109685 wrote:
         | It seems like search access is more valuable these days since
         | reasoning requires realtime access to site data.
        
       | ssijak wrote:
       | I dont want this by default. I want my website to end up in AI
       | chatbots. For SEO
        
       | abalashov wrote:
       | Few people realise that virtually everything we do online has,
       | until this point, been free training to make OpenAI, Anthropic,
       | etc. richer while cutting humans--the ones who produced the value
       | --out of the loop.
       | 
       | It might be too little, too late, at this juncture, and this
       | particular solution doesn't seem too innovative. However, it is
       | directionally 100% correct, and let's hope for massively more
       | innovation in defending against AI parasitism.
        
         | k__ wrote:
         | Is anyone suing to make the models and their weights open
         | source?
        
         | jefftk wrote:
         | I write online (comments here, open source software, blogging,
         | etc) because I have ideas I want to share. Whether it's "I did
         | a thing and here's how" or "we should change policy in this
         | specific way" or "does anyone know how to X" I'm happy for this
         | to go into training models just like I'm happy for it to go
         | into humans reading.
        
           | dolebirchwood wrote:
           | Thank you for having this attitude. I have never attempted
           | any blogging because I always figured no one is actually
           | going to read it. With LLMs, however, I know they will. I
           | actually see this as a motivation to blog, as we are in a
           | position to shape this emerging knowledge base. I don't find
           | it discouraging that others may be profiting off our freely
           | published work, just as I myself have benefited tremendously
           | from open source and the freely published works of others.
        
             | arkmm wrote:
             | This is an interesting take, thanks for sharing. I wonder
             | how someone should adjust their blogging if they believe
             | their primary audience will be LLMs.
        
               | lawlessone wrote:
               | SEO -> LLMEO
        
           | godelski wrote:
           | Tbh, that content I'm mostly fine with. My only real issue is
           | that people are making trillions off the free labor of people
           | like you and me, giving less time to create that OSS and
           | blogs. But this isn't new to AI, it is just scaled.
           | 
           | What I do care about is the theft of my identity. A person
           | may learn from the words I write but that person doesn't end
           | up mimicking the way I write. They are still uniquely
           | themselves.
           | 
           | I'm concerned that the more I write the more my text becomes
           | my identifier. I use a handle so I can talk more openly about
           | some issues.
           | 
           | We write OSS and blog because information should be free. But
           | that information is then being locked behinds paywalls and
           | becoming more difficult to be found through search. Frankly,
           | that's not okay
        
             | bob1029 wrote:
             | > OSS
             | 
             | > people are making trillions off the free labor of people
             | like you and me
             | 
             | I read "No Discrimination Against Fields of Endeavor" to
             | also include LLMs and _especially_ the cases that we most
             | deeply disagree with.
             | 
             | Either we believe in the principles of OSS or we do not. If
             | you do not like the idea of your intellectual property
             | being used for commercial purposes then this model is
             | definitely not for you.
             | 
             | There is no shame in keeping your source code and other IP
             | a secret. If you have strong expectations of being
             | compensated for your work, then perhaps a different
             | licensing and distribution model is what you are after.
             | 
             | > that information is then being locked behinds paywalls
             | and becoming more difficult to be found through search
             | 
             | Sure - If you give up and delete everything. No one is
             | forcing you to put your blog and GH repos behind a paywall.
        
               | mattl wrote:
               | Open source software typically has a license. People not
               | following the license isn't tolerated.
               | 
               | This is what AI scrapers are doing. They're taking your
               | code, your artwork and your writing without any
               | consideration for the license.
        
               | blibble wrote:
               | > Either we believe in the principles of OSS or we do
               | not. If you do not like the idea of your intellectual
               | property being used for commercial purposes then this
               | model is definitely not for you.
               | 
               | I've been writing open source for more than 20 years
               | 
               | I gave away my work for free with one condition: leave my
               | name on it (MIT license)
               | 
               | the AI parasites then strip the attribution out
               | 
               | they are the ones violating the principles of open source
               | 
               | > then perhaps a different licensing and distribution
               | model is what you are after.
               | 
               | I've now stopped producing open source entirely
               | 
               | and I suggest every developer does the same until the
               | legal position is clarified (in our favour)
        
               | godelski wrote:
               | > Either we believe in the principles of OSS or we do
               | not.
               | 
               | What about respecting licenses?
               | 
               | Seriously, don't lick the boot. We can recognize that
               | there's complexity here. Trivializing everything only
               | helps the abusers.
               | 
               | Giving credit where credit is due is not too much to ask.
               | Other people making money off my work can be good[0].
               | Taking credit for it is insulting
               | 
               | [0] If you're not making much, who cares. But if you're a
               | trillion dollar business you can afford to give a little
               | back. Here's the truth, OSS only works if we get enough
               | money and time to do the work. That's either by having a
               | good work life balance and good pay or enough donations
               | coming in. We've been mostly supported by the former, but
               | that deal seems to be going away
        
             | lxgr wrote:
             | > What I do care about is the theft of my identity. A
             | person may learn from the words I write but that person
             | doesn't end up mimicking the way I write. They are still
             | uniquely themselves.
             | 
             | Of course they do, to some extent. Just because it's been
             | infeasible to track the exact "graph of influence", that's
             | literally how humans have learned to speak and write for as
             | long as we've had language and writing.
             | 
             | > I'm concerned that the more I write the more my text
             | becomes my identifier. I use a handle so I can talk more
             | openly about some issues.
             | 
             | That's a much more serious concern, in my view. But I
             | believe that LLMs are both the problem and solution here:
             | "Remove style entropy" is just a prompt away, these days.
        
             | BeetleB wrote:
             | > A person may learn from the words I write but that person
             | doesn't end up mimicking the way I write.
             | 
             | Oh, I wish I could get AI to mimic the way I write! I'd pay
             | money for it. I often want to type up an email/doc/whatever
             | but don't because of occasional RSI issues. If I could get
             | an AI to type it up for me while still sounding like me -
             | that would be a big boon for my health.
        
         | andy99 wrote:
         | It's cloudflare and parasites like them that will make the
         | internet un-free. It's already happening, I'm either blocked or
         | back to 1998 load times be cause of "checking your browser".
         | They are destroying the internet and will make it so only
         | people who do approved things on approved browsers (meaning let
         | advertising companies monetize their online activity) will get
         | real access.
         | 
         | Cloudflare isn't solving a problem, they are just inserting
         | themselves as an intermediary to extract a profit, and making
         | everything worse.
        
           | carlhjerpe wrote:
           | I use Firefox with adblocking and some fingerprinting anti-
           | measurements and I rarely hit their challenges. Your IP
           | reputation must be bad.
           | 
           | They have an addon [1] that helps you bypass Cloudflare
           | challenges anonymously somehow, but it feels wrong to install
           | a plugin to your browser from the ones who make your web
           | experience worse
           | 
           | 1: https://developers.cloudflare.com/waf/tools/privacy-pass/
        
             | godelski wrote:
             | I'm in a pretty similar boat except I frequently hit
             | challenges. Especially if I use a VPN (which is more
             | trustworthy than my ISP). Ironically, I'm using Cloudflare
             | for DoH
        
               | lxgr wrote:
               | I'd be surprised if Cloudflare were actually correlating
               | DoH requests to HTTP requests following them, so I don't
               | think that's a signal they are likely to use.
        
           | MichaelZuo wrote:
           | If your on ipv6, I think they have to for ipv6 addresses...
           | there's just way too many bots and way too many addresses to
           | feasibly do anything more precise.
           | 
           | If your on ipv4 you should check whether your behind a NAT
           | otherwise you may have gotten an address that was previously
           | used by a bot network.
        
             | lxgr wrote:
             | > I think they have to for ipv6 addresses... there's just
             | way too many bots and way too many addresses
             | 
             | Are you really arguing that it's legitimate to consider
             | _all IPv6_ browsing traffic  "suspicious"?
             | 
             | If anything, I'd say that IPv4 is probably harder, given
             | that NATs can hide hundreds or thousands of users behind a
             | single IPv4 address, some of which might be malicious.
             | 
             | > you may have gotten an address that was previously used
             | by a bot network.
             | 
             | Great, another "credit score" to worry about...
        
           | slenk wrote:
           | How is Cloudflare a parasite? I can use Cloudflare, and get
           | their AI protection, for free. I have dozens of domains I
           | have used with Cloudflare at one point and I haven't paid
           | them a dime.
        
             | fsflover wrote:
             | They put themselves as a middle man for almost the whole
             | Internet, collect huge usage data about everyone and block
             | anybody who doesn't use mainstream tools:
             | 
             | https://news.ycombinator.com/item?id=42953508
             | 
             | https://news.ycombinator.com/item?id=13718752
             | 
             | https://news.ycombinator.com/item?id=23897705
             | 
             | https://news.ycombinator.com/item?id=41864632
             | 
             | https://news.ycombinator.com/item?id=42577076
        
               | baq wrote:
               | valid.
               | 
               | ...but OTOH it's their customers who want all of that and
               | pay to get that, because the alternative is worse.
               | 
               | rock and a hard place.
        
               | slenk wrote:
               | Right - do I want them getting some info from me, or do I
               | want my IP address exposed?
               | 
               | Besides CloudFront, which still costs money, what other
               | option is there for semi-privacy and caching for free?
        
               | mattl wrote:
               | bunny.net has some options
        
               | slenk wrote:
               | I will have to check them out I guess
        
               | aorth wrote:
               | As the old addage goes: If you're not paying for it,
               | you're the product.
               | 
               | Lots of nuance, but generally: pay for things you use.
               | Servers, engineers, and research and development are not
               | free, so someone has to pay.
        
               | qualeed wrote:
               | Lots of services don't even let me pay if I wanted to, so
               | I am forced to be the product. (Donating typically does
               | not un-productify myself).
               | 
               | Or I pay and am _still_ the product. Just with less in-
               | my-face ads.
        
               | matt-p wrote:
               | Cloud front is pretty much free for your first TB. Fastly
               | has a free plan.
               | 
               | Though why should it be for free?
        
               | slenk wrote:
               | Multiple people have brought that up. I pay for
               | everything else, why not one more.
               | 
               | Although bunny.net won't take ANY of my credit or debit
               | cards
        
               | MisterTea wrote:
               | I want to know if there is a way to design an alternative
               | that isn't controlled by a single entity which allows
               | gatekeeping.
        
               | AnthonyMouse wrote:
               | You can add another one as a result of _this_ article:
               | The data you need to train AI and the data you need to
               | build a search engine are the same data. So now they 're
               | inhibiting every new search engine that wants to compete
               | with Google.
        
             | lxgr wrote:
             | > I have dozens of domains I have used with Cloudflare at
             | one point and I haven't paid them a dime.
             | 
             | Maybe _you_ haven 't, but your users (primarily those using
             | "suspicious" operating systems and browsers) certainly have
             | - with their time spent solving captchas.
        
               | sealeck wrote:
               | But Cloudflare have removed CAPTCHAs
        
               | lxgr wrote:
               | Not sure if you're joking, but if you're not:
               | Congratulations on using a very "normal/safe"
               | OS/browser/IP.
               | 
               | I get captchas daily, without using any VPN and on
               | several different IPs (work, home, mobile). The only
               | crime I can think of is that I'm using Firefox instead of
               | Chrome.
        
               | Symbiote wrote:
               | Since a few days ago, I've been getting Captchas hourly
               | or more.
               | 
               | It's probably because I use Firefox on Linux with an ad
               | blocker.
               | 
               | For my part, I've ensured we don't use Cloudflare at
               | work.
        
               | kelvinjps10 wrote:
               | I use firefox on linux with an ad blocker and cloudfare
               | works fine
        
           | rockskon wrote:
           | LLM scrapers have dramatically been increasing the cost of
           | hosting various small websites.
           | 
           | Without something being done, the data that these scrapers
           | rely on would eventually no longer exist.
        
             | benjiro wrote:
             | I think the correct term is, that unrestricted LLM scrapers
             | have dramatically been increasing the cost of hosting
             | various small websites.
             | 
             | Its not a issue when somebody does "ethical" scraping, with
             | for instance, a 250ms delay between requests, and a active
             | cache that checks specific pages (like news article links)
             | to rescrape at 12 or 24h intervals. This type of scraping
             | results in almost no pressure on the websites.
             | 
             | The issue that i have seen, is that the more unscrupulous
             | parties, just let their scrapers go wild, constantly
             | rescraping again and again because the cost of scraping is
             | extreme low. A small VM can easily push 1000's of scraps
             | per second, let alone somebody with more dedicated
             | resources.
             | 
             | Actually building a "ethical" scraper involves more time,
             | as you need to fine tune it per website. Unfortunately,
             | this behavior is going to cost the more ethical scraper a
             | ton, as anti-scraping efforts will increase the cost on our
             | side.
        
               | Tmpod wrote:
               | The biggest issue for me is clearly masquerading their
               | User-Agent strings. Regardless of whether they are slow
               | and respectful crawlers, they should clearly identify
               | themselves, provide a documentation URL and obey
               | robots.txt. Without that, I have to play a frankly tiring
               | game of cat and mouse, wasting my time and the time of my
               | users (they have to put up with some form of captcha or
               | PoW thing).
               | 
               | I've been an active lurker in the self-hosting community
               | and I'm definitely not alone. Nearly everyone hosting
               | public facing websites, particularly those whose form is
               | rather juicy for LLMs, have been facing these issues. It
               | costs more time and money to deal with this, when
               | applying a simple User-Agent block would be much cheaper
               | and trivial to do and maintain.
               | 
               |  _sigh_
        
           | dceddia wrote:
           | Yep this terrifies me, 100%. We're slowly losing the open
           | internet and the frog is being boiled slowly enough that
           | people are very happy to defend the rising temperature.
           | 
           | If DDoS wasn't a scary enough boogeyman to get people to
           | install Cloudflare as a man-in-the-middle on all their
           | website traffic, maybe the threat of AI scrapers will do the
           | trick?
           | 
           | The thing about this slow slide is it's always defensible.
           | Someone can always say "but I don't want my site to be
           | scraped, and this service is free, or even better yet, I can
           | set up my own toll booth and collect money! They're
           | wonderful!"
           | 
           | Trouble is, one day, at this rate, almost all internet
           | traffic will be going through that same gate. And once they
           | have _literally everyone_ (and all their traffic)... well,
           | internet access is an immense amount of power to wield and I
           | can't see a world in which it remains untainted by commercial
           | and government interests forever.
           | 
           | And "forever" is what's at stake, because it'll be near
           | impossible to recover from once 99% of the population is
           | happy to use one of the 3 approved browsers on the 2 approved
           | devices (latest version only). Feels like we're already
           | accepting that future at an increasing rate.
        
             | RiverCrochet wrote:
             | The Internet is not the first global network. Before the
             | Internet, you had the global telephone network. It, too,
             | strangulated end users, but eventually became stagnant,
             | overpriced, and irrelevant. Super long-term, the current
             | Internet is not immune from this. Internet standards are
             | about getting as complicated and quirky as the old Bell
             | stuff that was trying to make miles of buried copper the
             | future, and if regulatory/commercial forces freeze this
             | stuff in place, it's going to lead to stagnation
             | eventually.
             | 
             | Something coming down the pike I think, for example, is
             | that IPv4 addresses are going to get realllly expensive
             | soon. That's going to lead to all sorts of interesting
             | things in the Internet landscape and their applications.
             | 
             | I'm sure we'll probably have to spend some decades in the
             | "approved devices and browers only" world before a next
             | wave comes.
        
             | mattl wrote:
             | We need a reasonable alternative to some of what Cloudflare
             | does that can be easily installed as a package on Linux
             | distributions without any of the following to install it.
             | 
             | * curl | bash
             | 
             | * Docker
             | 
             | * Anything that smacks of cryptocurrency or other scams
             | 
             | Just a standard repo for Debian and RHEL derived distros.
             | Fully open source so everyone can use it. (apt/dnf install
             | no-bad-actors)
             | 
             | Until that exists, using Cloudflare is inevitable.
             | 
             | It needs to be able to at least:
             | 
             | * provide some basic security (something to check for sql
             | injection, etc)
             | 
             | * rate limiting
             | 
             | * User agent blocking
             | 
             | * IP address and ASN blocking
             | 
             | Make it easy to set up with sensible defaults and a way to
             | subscribe to blocklists.
        
               | saint_yossarian wrote:
               | I remember using mod_security with Apache long ago for
               | some of this, looks like it's still around and now also
               | supports Nginx and IIS: https://modsecurity.org/
        
           | brumar wrote:
           | Correction: extract monstreous profits. When I read about the
           | revenues associated with Reddit AI deals, I can't even
           | imagine what could possibly be deals that cover half of the
           | internet. Cynically speaking, it's a genious level move.
        
           | axus wrote:
           | From the server perspective Cloudflare is solving problems
           | and not causing problems to other servers.
           | 
           | Analogy: locks for high-value items in grocery stores are
           | annoying to customers, but other stores aren't being coerced
           | by the locksmith to use them.
        
         | rramon wrote:
         | Isn't there a possibility that model makers retaliate by
         | erasing them and their frameworks from memory, hurting CF
         | adoption by devs?
        
         | cmeacham98 wrote:
         | Cutting humans out of what loop? What jobs or opportunities
         | were people posting Reddit comments or whatever getting that
         | are now going to AI?
        
           | Larrikin wrote:
           | People who used to post gained knowledge from their
           | profession or hobby. I don't bother posting any of that
           | information on large sites like Reddit anymore, for various
           | reasons but AI scraping solidified.
           | 
           | I'll still post on the increasingly fewer hobby message
           | boards that are out there.
        
           | kamarg wrote:
           | > What jobs or opportunities were people posting Reddit
           | comments or whatever getting that are now going to AI?
           | 
           | Content writing, product reviews (real & fake), creative
           | writing, customer support, photography/art to name a few off
           | the top of my head.
        
             | fkyoureadthedoc wrote:
             | Now the astroturfing is done by AI agents instead of hard
             | working serfs in a call center, you hate to see it
        
         | godelski wrote:
         | Including your comment, including this comment.
         | 
         | HN itself is routinely scraped. What makes me most
         | uncomfortable is deanonymization via speech analysis. It's
         | something we can already do but is hard to do at scale. This is
         | the ultimate tool for authoritarians. There's no hidden
         | identities because your speech is your identifier. It is
         | without borders. It doesn't matter if your government is good,
         | a bad acting government (or even large corporate entity) has
         | the power to blackmail individuals in other countries.
         | 
         | We really are quickly headed towards a dystopia. It could
         | result in the entire destruction of the internet or an
         | unprecedented level of self censorship. We already have
         | algospeak because platform censorship[0]. But this would be a
         | different type of censorship. Much more invasive, much more
         | personal. There are things worse than the dark forest
         | 
         | [0] literally yesterday YouTube gave me, a person in the 25-60
         | age bracket, a content warning because there was a video about
         | a person that got removed from a plane because they wore a
         | shirt saying "End veteran suicide".
         | 
         | [0.1] Even as I type this I'm censored! Apple will allow me to
         | swipe the word suicidal but not suicide! Jesus fuck guys! You
         | don't reduce the mental health crisis by preventing people from
         | even being able to discuss their problems, you only make it
         | worse!
        
         | Kostic wrote:
         | This would be true if not for open-weights (and even some open
         | source) LLMs that exist today. Not everything should be done
         | for profit.
        
         | giancarlostoro wrote:
         | There's a reason reddit started charging for API usage.
        
           | fkyoureadthedoc wrote:
           | It surely wasn't to force users into their shitty app where
           | they can't block ads and definitely had nothing to do with
           | their IPO. It was the AI.
        
             | giancarlostoro wrote:
             | Ah yes, its only because of ONE singular reason they
             | started charging for API usage. Are you okay? I'm listing
             | one reason out of many as to why reddit started charging
             | for API usage. After all, reddit is a for profit website.
        
               | fkyoureadthedoc wrote:
               | > There's *A* reason
               | 
               | this u chief?
        
         | dwoldrich wrote:
         | I think the parasitism goes quite a bit further than AI. We're
         | being digested not parasitized.
        
         | Dig1t wrote:
         | That was always the cost of free and open exchange of ideas
         | though. The idea of the internet in the first place was to
         | allow people to communicate in the open and publish ideas
         | freely. There was never any stipulation that using the
         | published ideas to make money was off limits.
         | 
         | Technology has advanced and now reading the sum total of the
         | freely exchanged ideas has become particularly valuable. But
         | who cares? The internet still exists and is still usable to
         | freely exchange ideas the way it's always been.
         | 
         | The value that one website provides is a minuscule amount, the
         | value of one individual poster on Reddit is minuscule. Are we
         | asking that each poster on Reddit be paid 1 penny (that's
         | probably what your posts are worth) for their individual
         | contribution? My websites were used to train these models
         | probably, but the value that each contributed is so small that
         | I wouldn't even expect a few cents for it.
         | 
         | The person who's going to profit here is Cloudflare or the
         | owners of Reddit, or any other gatekeeper site that is already
         | profiting from other people's contributions.
         | 
         | The "parasitism" here just feels like normal competition
         | between giant companies who have special access to information.
        
         | tcdent wrote:
         | Everything you did, up to this point, hopefully has made
         | _someone_ richer, otherwise you contributed literally zero
         | value to the world.
         | 
         | The internet is a free, open forum for the exchange of ideas.
         | Not a private diary for your best secrets.
        
           | friedtofu wrote:
           | What the hell...
           | 
           | Even if you're directing this at the user's blog posts
           | specifically; this is a ridiculously pessimistic, sad way to
           | view things.
           | 
           | I hope you're just having a bad day because if you sincerely
           | have this greedy, cynical mindset day to day(towards
           | blogging, software, offline/real life activities, whatever) I
           | feel sorry for you.
        
             | lxgr wrote:
             | GP's comment might be provocatively phrased, but I don't
             | think it's an invalid point to have:
             | 
             | When I publish something online for free, i.e. without
             | requiring authentication or payment, be it a Reddit
             | comment, a blog post, a Stackoverflow answer or anything
             | else, I do so hoping that it will be useful to somebody
             | somehow, without any illusions about being able to gatekeep
             | some types of current or future consumers.
        
         | risyachka wrote:
         | Maybe so, but I'll take Cloudflare over OpenAI and Meta every
         | time.
        
         | lofaszvanitt wrote:
         | Cyberpunk aged well. "You better not be on the unprotected
         | internet". Too many hazards out there. Rogue AIs and other
         | shit...
         | 
         | Cloudflare is here to protecc you from all those evils. Just
         | come under our umbrella.
        
         | bawolff wrote:
         | I think its 100% ok to freely train on public internet data.
         | 
         | What is absolutely not ok is to crawl at such an excessive
         | speed that it makes it difficult to host small scale websites.
         | 
         | Truly a tragedy of the commons.
        
           | tedd4u wrote:
           | Agree. The problem lately is that even if each single scraper
           | is doing so "reasonably," there are so many individuals and
           | groups doing this that it's still too onerous for many sites.
           | And of course many are not "reasonable."
        
         | jowea wrote:
         | Is it even possible that Cloudfare could manage to block all AI
         | data scrapping? I think this measure is just going to make it
         | harder and more expensive, which will stop AI scrappers from
         | hitting every single page every single day and creating
         | expenses for publishers, but not actually stop their data from
         | ending up in a few datasets.
        
         | mathiaspoint wrote:
         | This has been going on even since early social media. I think
         | most of the users actually prefer it.
        
       | jjangkke wrote:
       | so TLDR it adjusts your robot.txt and relies on cloudflare to
       | catch bot behavior and it doesn't actually do any sophisticated
       | residential proxy filtering or common bypass methods that works
       | on cloudflare turnstill, do I have this correct?
       | 
       | this just pushes AI agents "underground" to adopt the behavior of
       | a full blown stealth focused scraper which makes it harder to
       | detect.
        
       | nemild wrote:
       | Think this is the future, as the AI Web takes over the human web.
       | 
       | At Coinbase, we've been building tools to make the blockchain the
       | ideal payment rails for use cases like this with our x402
       | protocol:
       | 
       | https://www.x402.org/
       | 
       | Ping if you're interested in joining our open source community.
        
       | bgwalter wrote:
       | The destruction of the Web and IP theft needs to be addressed
       | legally. The opinion of a single judge notwithstanding, "AI"
       | scraping already violates copyright. This needs to be made
       | explicit in law and scrapers must get the same treatment as
       | Western governments gave to thousands of individuals who were
       | bankrupted or jailed for copyright infringement.
       | 
       | We are in the Napster phase of Web content stealing.
        
       | zackmorris wrote:
       | As usual, this is the wrong approach.
       | 
       | The open web is akin to the commons, public domain and public
       | land. So this is like putting a spy cam on a freeway billboard,
       | detecting autonomous vehicles, and shining a spotlight at their
       | camera to block them from seeing the ad. To what end?
       | 
       | Eventually these questions will need to be decided in court:
       | 
       | 1) Do netizens have the right to anonymity? If not, then we'll
       | have to disclose whether we're humans or artificial beings.
       | Spying on us and blocking us on a whim because our behavior
       | doesn't match social norms will amount to an invasion of privacy
       | (eventually devolving into papers please).
       | 
       | 2) Is blocking access to certain users discrimination? If not,
       | then a state-sanctioned market of civil rights abuse will grow
       | around toll roads (think whites-only drinking fountains).
       | 
       | 3) Is downloading copyrighted material for learning purposes by
       | AI or humans the same as pirating it and selling it for profit?
       | If so, then we will repeat the everyone-is-a-criminal torrenting
       | era of the 2000s and 2010s when "making available" was treated
       | the same as profiting from piracy, and take abuses by HBO, the
       | RIAA/MPAA and other organizations who shut off users' internet
       | connections through threat of legal actions like suing for
       | violating the DMCA (which should not have been made law in the
       | first place).
       | 
       | I'm sure there are more. If we want to live in a free society,
       | then we must be resolute in our opposition of draconian
       | censorship practices by private industry. Gatekeeping by large,
       | monopolistic companies like Cloudflare simply cannot be
       | tolerated.
       | 
       | I hope that everyone who reads this finds alternatives to
       | Cloudflare and tells their friends. If they insist on pursuing
       | this attack on our civil rights for profit, then I hope we build
       | a countermovement by organizing with the EFF and our elected
       | officials to eventually bring Cloudflare up on antitrust charges.
       | 
       | Cloudflare has shown that they lack the judgement to know better.
       | Which casts doubt on their technical merits and overall vision
       | for how the internet operates. By pursuing this course of action,
       | they have lost face like Google did when it removed its "don't be
       | evil" slogan from its code of conduct so it could implement
       | censorship and operate in China (among other ensh@ttification-
       | related goals).
       | 
       | Edit: just wanted to add that I realize this may be an opt-in
       | feature. But that's not the point - what I'm saying is that this
       | starts a bad precedent and an unnecessary arms race, when we
       | should be questioning whether spidering and training AI on
       | copyrighted materials are threats in the first place.
        
       | sct202 wrote:
       | My data served by Cloudflare has increased to 100gb /month
       | compared to <20gb like 2 years ago, and they're all fairly static
       | hobby sites. Actual people traffic is down by like half in the
       | same time frame, so I imagine a lot of this is probably cost
       | savings for Cloudflare to reduce resource usage.
        
         | Apofis wrote:
         | Makes total sense, bandwidth on this scale is expensive.
        
       | e38383 wrote:
       | Why is every second article about this claiming that it's
       | automatic? It needs to be turned on or at least there was no
       | mention of automatic in the original blog post.
       | 
       | I really hope that we can continue training AI the same way we
       | train humans - basically for free.
        
       | aunty_helen wrote:
       | I saw yesterday that they were going to allow websites to charge
       | per scrape.
       | 
       | Looks like cloudflare just invented the new App Store.
        
       | hmate9 wrote:
       | Isn't this only useful for blogs, news sites, or forums? Why
       | would I want an AI to know less about my product? I want it to
       | understand it, talk about it, and ideally recommend it. Should be
       | default off.
        
       | maximilianburke wrote:
       | Every evolution of the web, from Web 2 giving us walled gardens
       | to Web 3 giving us, well, nothing, to what we have now is taking
       | us further from a network of communities and personal
       | repositories of knowledge.
       | 
       | Sure, fidelity has gotten better but so much has been lost.
        
       | sneak wrote:
       | This idea that you can publish data for people to download and
       | read but not for people to download and store, or print, or think
       | about, or train on is a doomed one.
       | 
       | If you don't want people reading your data, don't put it on the
       | web.
       | 
       | The concept that copyright extends to "human eyeballs only" is a
       | silly one.
        
       | ChrisArchitect wrote:
       | [dupe] https://news.ycombinator.com/item?id=44432385
        
       ___________________________________________________________________
       (page generated 2025-07-02 23:01 UTC)