[HN Gopher] Cloudflare Introduces Default Blocking of A.I. Data ...
___________________________________________________________________
Cloudflare Introduces Default Blocking of A.I. Data Scrapers
Author : stephendause
Score : 335 points
Date : 2025-07-02 13:28 UTC (9 hours ago)
(HTM) web link (www.nytimes.com)
(TXT) w3m dump (www.nytimes.com)
| cmg wrote:
| Archive link: https://archive.ph/ARnyu
| badlibrarian wrote:
| Did they ever fix the auto-blocking of RSS feeds?
|
| https://news.ycombinator.com/item?id=41864632
| blakesterz wrote:
| The list of bots is pretty short right now:
|
| https://developers.cloudflare.com/bots/concepts/bot/#ai-bots
| ZiiS wrote:
| Enough to more than half the traffic to most sites if the
| blocks hold.
| hennell wrote:
| Cloudflare sees a lot of the web traffic. I assume these are
| the biggest bots they're seeing right now, and any new
| contenders would be added as they find them. Probably
| impossible to really block everything, but they've got the web-
| coverage to detect more than most.
| TechDebtDevin wrote:
| They are lying. They cant detect crawlers unless we tell them
| we are who we are.
| JimDabell wrote:
| > AI bots
|
| > You can opt into a managed rule that will block bots that we
| categorize as artificial intelligence (AI) crawlers ("AI Bots")
| from visiting your website. Customers may choose to do this to
| prevent AI-related usage of their content, such as training
| large language models (LLM).
|
| > CCBot (Common Crawl)
|
| Common Crawl is not an AI bot:
|
| https://commoncrawl.org
| johneth wrote:
| The data it collects is used by AI companies, though.
| Spivak wrote:
| Poor ChatGPT-User, nobody understands you. Blocking a real user
| because of the, admittedly odd, browser they're using misses the
| point.
| Roark66 wrote:
| This is a bit silly. Slowing down, yes, but blocking? People who
| *really* want that content will find a way and this will hit
| everyone else instead that will have to do silly riddles before
| following every link or run crypto mining for them before being
| shown the content .
|
| I recently went to a big local auction site on which I buy
| frequently and I got one of these "we detected unusual traffic
| from your network" messages. And "prove you're human". Which was
| followed by "you completed the capcha in 0.4s your IP is banned".
| Really? Am I supposed to slow down my browsing now? I tried a
| different browser, a different OS, logging on,clearing cookies,
| etc. Same result when I tried a search. It took 4h after
| contacting their customer service to unblock it. And the
| explanation was "you're clicking too fast".
|
| At some point it just becomes a farce and the hassle is not worth
| the content. Also, while my story doesn't involve any bots
| perhaps a time will come when local LLMs will be good enough that
| I'll be able to tell one "reorder my cat food" and it will go and
| do it. Why are they so determined to "stop it" (spoiler, they
| can't).
|
| For anyone who says LLMs are already capable of ordering cat food
| I say not so fast. First the cat food has to be on sale/offer
| (sometimes combined with extras). Second it is supposed to be
| healthy (no grains) and third the taste needs to be to my cats
| liking. So far I'm not going to trust a LLM with this.
| Sol- wrote:
| Do the major AI companies actually honor robots.txt? Even if some
| of their publicly known crawlers might do it, surely they have
| surreptitious campaigns where they do some hidden crawling, just
| like how they illegally pirate books, images and user data to
| train on.
| px43 wrote:
| There's a lack of clarity, but it seems likely to me that a
| majority of this traffic is actually people asking questions to
| the AI, and the AI going out and researching for answers. When
| the AI tools are being used like a web browser to do research,
| should they still be adhering to robots.txt, or is that only
| intended for search indexing?
| chasd00 wrote:
| My thought too, honoring robots.txt is just a convention.
| There's no requirement to follow robots.txt, or at least
| certainly no technical requirement. I don't think there's any
| automatic legal requirement either.
|
| Maybe sites could add "you must honor policies set in
| robots.txt" to something like a terms of service but I have no
| idea if that would have enough teeth for a crawler to give up.
| TechDebtDevin wrote:
| Cloudflare snd their customera have been desperately for
| years trying to kill scrapers in court. This is all.
| Meaningless, but they are probably gearing up for another
| legal battle to define robots.txt as a legal contract. Theyre
| going to use this marketplace theyre scamming people with to
| do it. They will fail.
| mschuster91 wrote:
| Cloudflare, for all I hate their role as a gatekeeper these
| days, actually has the leverage to force the AI companies to
| bend.
| deepsun wrote:
| Hard to tell, because minor crawlers mimic major companies to
| not getting banned.
| btown wrote:
| The headline is somewhat misleading: sites using Cloudflare now
| have an _opt-in_ option to quickly block all AI bots, but it won
| 't be turned on by default for sites using Cloudflare.
|
| The idea that Cloudflare _could_ do the latter at the sole
| discretion of its leadership, though, is indicative of the level
| of power Cloudflare holds.
| bitpush wrote:
| It is now an adversarial relationship between aibots and
| website, and cloudflare is merely reacting to it.
|
| Would you say the same for ddos protection? Isn't that the same
| as well?
| TechDebtDevin wrote:
| They arent doing anything. They are attempting to insert
| themselves into the middle of a marketplace (that doesnt
| exist and never will) where scrapers pay for IP. They think
| theyre going to profit off the bots, not protect your site.
| Dont fall for their scam.
| bitpush wrote:
| What do you mean they are trying to insert themselves? If I
| have a website that I host with cloudflare, I (as the
| rightful website owner) has inserted Cloudflare in between.
|
| It isnt CF going around saying, that's a nice website you
| have there. I'm gonna put myself in between.
| GrayShade wrote:
| > sites using Cloudflare now have an opt-in option to quickly
| block all AI bots, but it won't be turned on by default for
| sites using Cloudflare
|
| Do you have a source for that?
| https://blog.cloudflare.com/content-independence-day-no-ai-c...
| does say "changing the default".
| mattcollins wrote:
| "This feature is available to all customers, meaning anyone
| can enable this today from the Cloudflare dashboard."
|
| https://blog.cloudflare.com/control-content-use-for-ai-
| train...
| TechDebtDevin wrote:
| They cant do anything other than bog down the internet. I
| havent found a single cf provided challenge I havent been able
| to get past in < half a day.
|
| This is simply juat the first step in them implementing a
| marketplace and trying to get into LLM SEO. They dont care
| about your site or protecting it. They are gearing up to start
| making a cut in the Middle between scrapers and publishers. Why
| wouldnt I go DIRECTLY to the publisher and make a deal. So dumb
| I hate cf so much.
|
| The only thing cloudflare knows how to do is MITM attacks.
| Marsymars wrote:
| So what would you suggest as an alternative if I have a site
| where I don't want the content used for LLM training?
| fkyoureadthedoc wrote:
| Auth? Because whatever Cloudflare is doing isn't going to
| stop anyone serious about scraping data.
| Marsymars wrote:
| Let's say I'm talking about content that I don't want
| behind an auth wall. Is your position simply that all
| such sites should abandon any efforts to not have the
| content used for LLM training?
| mattl wrote:
| If you find a solution that's not auth please let me
| know.
| postalcoder wrote:
| > When you enable this feature via a pre-configured managed rule,
| Cloudflare can detect and block verified AI bots that comply with
| robots.txt and respect crawl rates, and do not hide their
| behavior from your website. The rule has also been expanded to
| include more signatures of AI bots that do not follow the rules.
|
| We already know companies like Perplexity are masking their
| traffic. I'm sure there's more than meets the eye, but taking
| this at face value, doesn't punishing respectful and transparent
| bots only incentivize obfuscation?
|
| edit: This link[0], posted in a comment elsewhere, addresses this
| question. tldr, obfuscation doesn't work. > We
| leverage Cloudflare global signals to calculate our Bot Score,
| which for AI bots like the one above, reflects that we correctly
| identify and score them as a "likely bot." > When bad
| actors attempt to crawl websites at scale, they generally use
| tools and frameworks that we are able to fingerprint. For every
| fingerprint we see, we use Cloudflare's network, which sees over
| 57 million requests per second on average, to understand how much
| we should trust this fingerprint. To power our models, we compute
| global aggregates across many signals. Based on these signals,
| our models were able to appropriately flag traffic from evasive
| AI bots, like the example mentioned above, as bots.
|
| [0] https://blog.cloudflare.com/declaring-your-aindependence-
| blo...
| colechristensen wrote:
| >doesn't punishing respectful and transparent bots only
| incentivize obfuscation?
|
| They're cloudflare and it's not like it's particularly easy to
| hide a bot that is scraping large chunks of the Internet from
| them. On top of the fact that they can fingerprint any of your
| sneaky usage, large companies have to work with them so I can
| only assume there are channels of communication where
| cloudflare can have a little talk with you about your bad
| behavior. I don't know how often lawyers are involved but I
| would expect them to be.
| jerf wrote:
| "doesn't punishing respectful and transparent bots only
| incentivize obfuscation?"
|
| Sure, but we crossed that bridge over 20 years ago. It's not
| creating an arms race where there wasn't already one.
|
| Which is my generic response to everyone bringing similar ideas
| up. "But the bots could just...", yeah, they've been doing it
| for 20+ years and people have been fighting it for just as
| long. Not a new problem, not a new set of solutions, no
| prospect of the arms race ending any time soon, none of this is
| new.
| hombre_fatal wrote:
| Next line:
|
| > The rule has also been expanded to include more signatures of
| AI bots that do not follow the rules.
|
| The Block AI Bots rule on the Super Bot Fight Mode page does
| filter out most bot traffic. I was getting 10x the traffic from
| bots than I was from users.
|
| It definitely doesn't rely on robots.txt or user agent. I had
| to write a page rule bypass just to let my own tooling work on
| my website after enabling it.
| account42 wrote:
| How many of those "bots" you are filtering are actually bots
| and how many are regular users buttflare has misidentified as
| bots?
| hombre_fatal wrote:
| Pretty simple to see this if you've run a website: compare
| your analytics pre-bot to post-bot to post-bot-blocker.
|
| There is a clear moment where you land on AI bot radar. For
| my large forum, it was a month ago.
|
| Overnight, "72 users are viewing General Discussion" turned
| into "1720 users".
|
| 40% requests being cached turned into 3% of requests are
| cached.
| fluidcruft wrote:
| Cloudflare already knows how to make the web hell for people
| they don't like.
|
| I read the robots.txt entries as those AI bots that will be not
| marked as "malicious" and that will have the opportunity to be
| allowed by websites. The rest will be given the Cloudflare
| special.
| dougb5 wrote:
| > Cloudflare can detect and block verified AI bots that comply
| with robots.txt and respect crawl rates, and do not hide their
| behavior from your website
|
| It's the bots that _do_ hide their behavior -- via residential
| proxy services -- that are causing most of the burden, for my
| site anyway. Not these large commercial AI vendors.
| alganet wrote:
| > If A.I. companies freely use data from various websites without
| permission or payment, people will be discouraged from creating
| new digital content
|
| I don't see a way out of this happening. AI fundamentally
| discourages other forms of digital interaction as it grows.
|
| Its mechanism of growing is killing other kinds of digital
| content. It will eventually kill the web, which is, ironically,
| its main source of food.
| preachermon wrote:
| just like capitalism has now turned to exploiting people as its
| main input?
| alganet wrote:
| These kinds of comparisons rarely lead to good discussions.
|
| Let's instead be focused and talk about real stuff.
|
| Consider https://learnpythonthehardway.org/ for example. It
| has influenced a generation of Python developers. Not just
| the main website, but the tons of Python code and Python-
| related content it inspired.
|
| Why would anyone write these kinds of
| textbooks/websites/guides if AI can replace them? AI
| companies are effectively broadcasting you don't need the
| hard way anymore, you can just vibe.
|
| Arguibly though, without the existance of Learn Python the
| Hard Way and similar content, AI would be worse at writing
| Python stuff. That's what I mean by "main source of food",
| good content that influences a lot of people. Net-positive
| effects hard to predict or even identify except for the more
| popular cases (such as LPTHW).
|
| If my prediction is right, no one will notice that good
| content has stopped being produced. It will appear as if
| content is being created in generally the same way as before,
| but in reality, these long tail initiatives like LPTHW will
| have ceased before anyone can do anything about it.
|
| Again, I don't see a way out of this scenario. Not for AI
| companies, not for content writers. This is going to happen.
| The world in which I'm wrong is the best one.
| mfost wrote:
| In a similar vein, I remember people advocating for
| replacing new untrained hires with AI. After all, a
| competent senior engineer is needed to validate the
| contributions of the new hires anyway and they can do the
| same checking the AI code.
|
| But then, how would you even train and replace those
| competent seniior engineers that do the filtering when they
| retire? The whole system was predicated on having a chain
| of new hires that gain experience in the process.
| alganet wrote:
| From what I could perceive, companies believe coding AIs
| will eventually learn to both code and teach better than
| seniors.
|
| This is based on two assumptions:
|
| - AI will get better. Developers using the system will
| transfer their knowledge to it.
|
| - Seniors in a couple of years will be different. They
| should be those who can engage with the AI feedback loop.
|
| Here's why I think it won't work:
|
| - Senior developers learn more than they can produce.
| There is knowledge they never transfer to what they work
| on. Internalized knowledge that never materializes
| directly into code. _But it materializes indirectly_.
|
| - Senior developer knowledge come from "schools", not
| just reading. These schools are not real physical
| locations. They're traditions, or ideas, that form a very
| long tail. These ideas, again, are not directly
| transferrable to code or prose.
|
| - Juniors get embarrassed. You say "stop making this
| nonsense", and they'll stop and reflect, because they
| respect seniors. They might disagree, but a pea was then
| placed under their matress, and they'll think about "this
| nonsense" you told them to stop doing and why. That is
| how they get better. So far, AI has not demonstrated
| being able to do that.
|
| The production of quality content is an aspect of one of
| those "schools of thought". You are supposed to bear the
| responsibility of passing the knowledge. Keeping lean
| codebases easy to understand is also a hallmark of many
| schools of thought. Working from fundamentals is another
| one of those ideas, etc.
| greenchair wrote:
| nothing's perfect but it is still better than the other
| options
| fennecfoxy wrote:
| Additionally, ad blocker usage is apparently at 30%. So it's a
| redundant or more nuanced argument, really.
| account42 wrote:
| Ad blockers only discourage commercialized content creation,
| not all of it. IMO that actually improves the quality of the
| content created.
| spwa4 wrote:
| Yes what _everyone_ wants to do with AI: generate entertainment
| and interactions with humans, including economical ones, will
| need to happen or AI will starve.
| alganet wrote:
| That's what is going to make it starve. Belly full, but of
| its own shit being tossed around humans seeking cheap copouts
| of doing actual work.
| BrouteMinou wrote:
| Just like cancer?
| jasonthorsness wrote:
| I turned this on and it adjusts the robots.txt automatically; not
| sure what else it is doing.
|
| # NOTICE: The collection of content and other data on this # site
| through automated means, including any device, tool, # or process
| designed to data mine or scrape content, is # prohibited except
| (1) for the purpose of search engine indexing or # artificial
| intelligence retrieval augmented generation or (2) with express #
| written permission from this site's operator.
|
| # To request permission to license our intellectual # property
| and/or other materials, please contact this # site's operator
| directly.
|
| # BEGIN Cloudflare Managed content
|
| User-agent: Amazonbot Disallow: /
|
| User-agent: Applebot-Extended Disallow: /
|
| User-agent: Bytespider Disallow: /
|
| User-agent: CCBot Disallow: /
|
| User-agent: ClaudeBot Disallow: /
|
| User-agent: Google-Extended Disallow: /
|
| User-agent: GPTBot Disallow: /
|
| User-agent: meta-externalagent Disallow: /
|
| # END Cloudflare Managed Content User-agent: * Disallow: /*
| Allow: /$
| xyst wrote:
| So in addition to updating the robots.txt file, which really
| only blocks a small number of them.
|
| Seems CF has been gathering data and profiling these malicious
| agents.
|
| This post by CF elaborates a bit further:
| https://blog.cloudflare.com/declaring-your-aindependence-blo...
|
| Basically becomes a game of cat and mouse.
| postalcoder wrote:
| This is interesting. The reasoning and response don't line up.
| > Cloudflare is making the change to protect original content
| on the internet, Mr. Prince said. If A.I. companies freely use
| data from various websites without permission or payment,
| people will be discouraged from creating new digital content,
| he said > prohibited except for the purpose of [..]
| artificial intelligence retrieval augmented generation
|
| This seems to be targeted at taxing training of language
| models, but why an exclusion for the RAG stuff? That seems like
| it has a much greater immediate impact for online content
| creators, for whom the bots are obviating a click.
| fennecfoxy wrote:
| With that opinion, are you also suggesting that we ban ad
| blockers? Because it's better I not click & consume resources
| than click and not be served ads, basically just costing the
| host money.
|
| It means sense to allow for RAG in the same way that search
| engines provide a snippet of an important chunk of the page.
|
| A blog author could not complain that their blog is getting
| ragged when they're extremely liable to be Google/whatever
| searching all day and basically consuming others' content in
| exactly the same way that they're trying to disparage.
| postalcoder wrote:
| I don't think we should ban ad blockers, but I also think
| it's fair to suggest that the loss of organic traffic could
| be affecting the incentive to create new digital content,
| at least as much as the fear of having your content
| absorbed into an LLM's training data.
| Boldened15 wrote:
| IMO the backlash against LLMs is more philosophical, a
| lot of people don't like them or the idea of one learning
| from their content. Unless your website has some unique
| niche information unavailable anywhere else there's no
| direct personal risk. RAG would be a more direct threat
| if anything.
| toomuchtodo wrote:
| It's really about who is getting the value from the work
| of the content. If content creators of all sorts have
| their work consumed by LLMs, and LLM orgs charge for it
| can capture all the value, why should people create to
| have their work vacuumed up for the robot's benefit? For
| exposure? You can't eat or pay rent with exposure. Humans
| must get paid, and LLMs (foundational models and output
| using RAG) cannot improve without a stream of works and
| data humans create.
|
| Whether you call it training or something else is
| irrelevant, it's really exploitation of human work and
| effort for AI shareholder returns and tech worker comp
| (if those who create aren't compensated). And the
| technocracy has not been, based on the evidence, great
| stewards of the power they obtain through this. Pay the
| humans for their work.
| o11c wrote:
| It's not philosophical, it's economical.
|
| AI scrapers increase traffic by maybe 10x (this varies
| per site) but provide no real value whatsoever to
| _anyone_. If you look at various forms of "value":
|
| * Saying "this uses AI" might make numbers go up on the
| stock market if you manage to persuade people it will
| make numbers go up (see also: the market will remain
| irrational longer than you can remain solvent).
|
| * Saying "this uses AI" might fulfill some corporate
| mandate.
|
| * Asking AI to solve a problem (for which you would
| actually use the solution) allows you to "launder" the
| copyright of whatever source it is paraphrasing (it's
| well established that LLMs fail entirely if a question
| isn't found within their training set). Pirating it
| directly provides the same value, with significantly less
| errors/handholding.
|
| * Asking AI to entertain you ... well, there's the
| novelty factor I guess, but even if people refuse to
| train themselves out of that obsession, the world is
| still far too full of options for any individual to
| explore them all. Even just the question of "what kind of
| ways can I throw some kind of ball around" has more
| answers than probably anyone here knows.
|
| What am I missing?
| ijk wrote:
| What I want to know is if the flood of scraping everyone
| has been complaining about is coming from people trying to
| scrape for training or bots doing RAG search.
|
| I get that everyone wants data, but presumably the big
| players already scraped the web. Do they really need to do
| it again? Or is it bit players reproducing data that's
| likely already in the training set? Or is it really that
| valuable to have your own scraped copy of internet scale
| data?
|
| I feel like I'm missing something here. My expectation is
| that RAG traffic is going to be orders of magnitude higher
| than scraping for training. Not that it would be easy to
| measure from the outside.
| mattcollins wrote:
| I wondered about this, too.
|
| Cloudflare have some recent data about traffic from bots
| (https://blog.cloudflare.com/from-googlebot-to-gptbot-
| whos-cr...) which indicates that, for the time being, the
| overwhelming majority of the bot requests are for AI
| training and not for RAG.
| wiether wrote:
| You should ask Zuck, since, for what we've seen and what
| we were ask to act against, Meta is the main culprit in
| scraping every single page of websites, multiple times a
| day.
|
| And I'm talking about ecommerce websites, with their bot
| scraping every variation of each product, multiple times
| a day.
| lxgr wrote:
| More and more people use ChatGPT for search, so blocking that
| doesn't seem like a successful strategy long-term.
| bee_rider wrote:
| I wonder... Google scrapes for indexing and for AI, right? I
| wonder if they will eventually say: ok, you can have me or not,
| if you don't want to help train my AI you won't get my searches
| either. That's a tough deal but it is sort of self-consistent.
| giancarlostoro wrote:
| "Embrace, Extend, Extinguish" Google's mantra. And yes, I
| know about Microsoft's history with that phrase ;) But Google
| has done this with email, browsers (Google has web apps that
| run fine on Firefox but request you use Chrome), Linux
| (Android), and I'm sure there's others I am forgetting about.
|
| So yeah, I too could see them doing this.
| mrweasel wrote:
| Very few people seems to be complaining that Google crashes
| their sites. Google also publish their crawlers IP ranges,
| but you really don't need to rate-limit Google, they know how
| to back off and not overload sites.
| Symbiote wrote:
| In theory -- in practise I've had to limit Google on two
| large sites at work. I currently have them limited to 10/s
| for non-cached requests.
| 1vuio0pswjnm7 wrote:
| "User-agent: CCBot disallow: /"
|
| Is Common Crawl exclusively for "AI"
|
| CCBot was already in so many robots.txt prior to this
|
| How is CC supposed to know or control how people use the
| archive contents
|
| What if CC is relying on fair use # To request
| permission to license our intellectual # property
| andd/or other materials, please contact this # site's
| operator directly
|
| If the operator has no intellectual property rights in the
| material, then do they need permission from the rights holders
| to license such materials for use in creating LLMs and collect
| licensing fees
|
| Is it common for website terms and conditions to permit site
| operators to sublicense other peoples' ("users") work for use
| in creating LLMs for a fee
|
| Is this fee shared with the rights holders
| nemomarx wrote:
| Read a tos and notice that you give the site operators
| unlimited license to reproduce or spread your works, almost
| on any site. it's required to host and show the content
| essentially
| ronsor wrote:
| # To request permission to license our intellectual #
| property andd/or other materials, please contact this
| # site's operator directly
|
| Scrapers don't accept the terms of service.
|
| Ironically, I've only ever scraped sites that block CCBot,
| otherwise I'd rather go to Common Crawl for the data.
| Bender wrote:
| For my silly hobby sites I just return status 444 _close the
| connection_ for anything that has case-insentive "bot" in the
| UA requesting anything other than robots.txt, humans.txt,
| favicon.ico, etc... This would also drop search engines but I
| blackhole route most of their CIDR blocks. I'm probably the
| only one here that would do this.
| sneak wrote:
| How does a bot scraping your silly hobby sites for any
| purpose harm or negatively affect you in any way?
| slenk wrote:
| I thought I saw cloudflare insert noindex links?
| swyx wrote:
| what actually are the consequences of ignoring robots.txt
| (apart from DDOS)? have any of these cases ended up in court at
| all?
| lxgr wrote:
| That's at least a more reasonable default than that I've seen
| at least one newspaper do, which is to block both LLM scrapers
| _and_ things like ChatGPT 's search feature explicitly.
| cratermoon wrote:
| I'm still not sure this is going to be very effective, as so many
| of the worst offenders don't identify themselves as bots, and
| often change their user agent. Has Cloudflare said anything about
| identifying the bad actors?
| chasd00 wrote:
| i've mentioned this in a couple replies so maybe i'm wrong but
| it's up to the client to obey robots.txt. Why would they not
| just ignore it? Unless there's some legal consequence not
| complying with robots.txt then why even follow it? There's no
| technical enforcement of the policies in the file, it's up to
| the client to honor them.
| kentonv wrote:
| > There's no technical enforcement of the policies in the
| file, it's up to the client to honor them.
|
| That's incorrect. Cloudflare does in fact enforce this at a
| technical level. Cloudflare has been doing bot detection for
| years and can pretty reliably detect when bots are not
| following robots.txt and then block them.
| GrayShade wrote:
| Yes, they have over the years, for example
| https://blog.cloudflare.com/residential-proxy-bot-
| detection-..., https://blog.cloudflare.com/cloudflare-bot-
| management-machin..., https://blog.cloudflare.com/introducing-
| bot-analytics/.
| lucasyvas wrote:
| I fail to see how this won't just result in UA string or other
| obfuscation.
| chasd00 wrote:
| a crawler doesn't have to change anything, they can just ignore
| the robots.txt file. It's up to the client to read robots.txt
| and follow its directives but there's no technical reason why
| the client cannot just ignore everything in the file period.
| kube-system wrote:
| Cloudflare's filtering is already way more sophisticated than
| just looking at UA string or other voluntary reporting. They're
| almost certainly using fingerprinting and behavioral analytics.
| gazpacho wrote:
| From an open source projects perspective we'd want to disable
| this on our docs sites. We actually want those to be very
| discoverable by LLMs, during training or online usage.
| rorylaitila wrote:
| Unfortunately I think pissing into the wind. Information websites
| are all but dead. AI contains all published human information. If
| you have positioned your website as an answer to a question, it
| won't survive that way.
|
| "Information" is dead but content is not. Stories, empathy,
| community, connection, products, services. Content of this
| variety is exploding.
|
| The big challenge is discoverability. Before, information
| arbitrage was one pathway to get your content discovered, or to
| skim a profit. This is over with AI. New means of discovery are
| necessary, largely network and community based. AI will throw you
| a few bones, but it will be 10% of what SEO did.
| fennecfoxy wrote:
| >AI contains all published human information
|
| No, it most certainly does not. It was certainly trained on
| large swathes of human knowledge/interactions.
|
| A model that consists of a perfect representation/compression
| of all this info is a zip file, not a model file.
| rorylaitila wrote:
| AI providers have scrapped and will continue to, all internet
| published information or virtually so. Since "Information" is
| infinite, AI cannot contain "all information" in a complete
| sense. But it certainly answers almost everything that
| matters for any existing search query that has ever been
| targeted by a webpage that is crawlable.
|
| In any case, as manifest by real world SEO, which is
| plummeting in traffic for informational queries, the effect
| is the same. This real world impact is what matters and will
| not be reversed, regardless of attempts at blocking.
| ozgrakkurt wrote:
| You are assuming LLMs will replace search engines. Why is this
| the case?
|
| To me it seems like there has to be so much optimization for
| this to happen that, it is not likely. LLM answers are slow and
| unreliable. Even using something like perplexity doesn't give
| much value over using a regular search engine in my experience
| rorylaitila wrote:
| LLMs will not fully replace search engines, but Google and
| Bing are evolving to be LLM first, anyhow. So "what is a
| search engine" today is not what it was yesterday. Let's call
| the time before LLMs, traditional search. LLM first products
| bundle some aspect of traditional search. And traditional
| search is adding LLM answers.
|
| Traditional search will still be highly useful for
| transactional, product, realtime, and action oriented
| queries. Also for discovering educational/entertainment
| content that is valued in of itself and cannot be
| reformulated by LLM.
| Meekro wrote:
| I've heard lots of people on HN complaining about bot traffic
| bogging down their websites, and as a website operator myself I'm
| honestly puzzled. If you're already using Cloudflare, some basic
| cache configuration should guarantee that most bot traffic hits
| the cache and doesn't bog down your servers. And even if you
| don't want to do that, bandwidth and CPU are so cheap these days
| that it shouldn't make a difference. Why is everyone so upset?
| deepsiml wrote:
| Not much into that kind of DevOps. What is a good basic caching
| in this instance?
| TechDebtDevin wrote:
| Cloudflare and other CDNs will usually automatically cache
| your static pages.
| haiku2077 wrote:
| It comes down to:
|
| 1. Use the Cache-Control header to express how to cache your
| site correctly (https://developer.mozilla.org/en-
| US/docs/Web/HTTP/Guides/Cac...)
|
| 2. Use a CDN service, or at least a caching reverse proxy, to
| serve most of the cacheable requests to reduce load on the
| (typically much more expensive) origin servers
| mrweasel wrote:
| Just note that many AI scrapers will go to great length to
| do cache busting. For some reason many of them feel like
| they need to get the absolute latest version and don't
| trust your cache.
| haiku2077 wrote:
| You can use Cache Control headers to express that your
| own CDN should aggressively refresh a resource but always
| serve it to external clients from cache. It's covered in
| the link under "Managed Caches"
| cortesoft wrote:
| A CDN can be configured to ignore cache control headers
| in the requests and cache things anyway.
| conductr wrote:
| The presumption I'm already using cloudfare is a start. Is this
| a requirement for maintaining a simple website now?
| haiku2077 wrote:
| Either that or Anubis (https://anubis.techaro.lol/docs), yes.
| roguecoder wrote:
| So these companies broke the internet
| haiku2077 wrote:
| Which companies?
|
| OpenAI, Anthropic, Google? No, their bots are pretty well
| behaved.
|
| The smaller AI companies deploying bots that don't
| respect any reasonable rate limits and are scraping the
| same static pages thousands of times an hour? Yup
| e3bc54b2 wrote:
| Anecdote, but at least for tiny little server hosting
| single public repository, _none_ of these companies had
| 'well behaved' bots. It may be possible that they learned
| to behave better but I wouldn't know since my only
| possible recourse was to blacklist them all AND take the
| repo private.
| haiku2077 wrote:
| Those are the small companies spoofing their user agent
| as the big companies to dodge countermeasures.
| noodle wrote:
| As someone who had some outages due to AI traffic and is now
| using CloudFlare's tools:
|
| Most of my site is cached in multiple different layers. But
| some things that I surface to unauthenticated public can't be
| cached while still being functional. Hammering those endpoints
| has taken my app down.
|
| Additionally, even though there are multiple layers, things
| that are expensive to generate can still slip through the
| cracks. My site has millions of public-facing pages, and a
| batch of misses that happen at the same time on heavier pages
| to regenerate can back up requests, which leads to errors, and
| errors don't result in caches successfully being filled. So the
| AI traffic keeps hitting those endpoints, they keep not getting
| cached and keep throwing errors. And it spirals from there.
| x0x0 wrote:
| It's not complex. I worked on a big site. We did not have the
| compute or i/o (most particularly db iops) to live generate the
| site. Massive crawls both generated cold pages / objects (cpu +
| iops) and yanked them into cache, dramatically worsening cache
| hit rates. This could easily take down the site.
|
| Cache is expensive at scale. So permitting big or frequent
| crawls by stupid crawlers either require significant
| investments in cache or slow down and worsen the site for all
| users. For whom we, you know, built the site, not to provide
| training data for companies.
|
| As others have mentioned, Google is significantly more
| competent than 99.9% of the others. They are very careful to
| not take your site down and provide, or used to provide,
| traffic via their search. So it was a trade, not a taking.
|
| Not to mention I prefer not to do business with Cloudflare
| because I don't like companies that don't publish quota. If
| going over X means I need an enterprise account that starts at
| $10k/mo, I need to know the X. Cloudflare's business practice
| appears to be letting customers exceed that quota then
| aggressively demanding they pay or they'll be kicked off the
| service nearly immediately.
| jtolmar wrote:
| The stories I've heard have been mostly about scraper bots
| finding APIs like "get all posts in date range" and then
| hammering that with every combo of start/end date.
| jauntywundrkind wrote:
| I too am a bit confused / mystified at the strong reaction. But
| I do expect a lot of badly optimized sites that just want out.
|
| I struggle to think of a web related library that has spread
| faster than Anubis checker. It's everywhere now!
| https://github.com/TecharoHQ/anubis
|
| I'm surprised we don't see more efforts to rate limit. I assume
| many of these are distributed crawlers, but it feels like
| there's got to be pools of activity spinning up, on a handful
| of IPs. And that they would be time correlated together pretty
| clearly. Maybe that's not true. But it feels like the web, more
| than anything else, needs some open source software to add a
| lot more 420 Enhance Your Calm responses, as it feels like.
| https://http.dev/420
| zerocrates wrote:
| The reaction comes from some combination of
|
| - opposition to generative AI in general
|
| - a view that AI, unlike search which also relies on
| crawling, offers you no benefits in return
|
| - crawlers from the AI firms being less well-behaved than the
| legacy search crawlers, not obeying robots.txt, crawling more
| often, more aggressively, more completely, more redundantly,
| from more widely-distributed addresses
|
| - companies sneaking in AI crawling underneath their existing
| tolerated/whitelisted user-agents (Facebook was pretty
| clearly doing this with "facebookexternalhit" that people
| would have allowed to get Facebook previews; they eventually
| made a new agent for their crawling activity)
|
| - a simultaneous huge spike in obvious crawler activity with
| spoofed user agents: e.g. a constant random cycling between
| every version of Chrome or Firefox or any browser ever
| released; who this is or how many different actors it is and
| whether they're even doing crawling for AI, who knows, but
| it's a fair bet.
|
| Better optimization and caching can make this all not matter
| so much but not everything can be cached, and plenty of small
| operations got by just fine without all this extra traffic,
| and would get by just fine without it, so can you really
| blame them for turning to blocking?
| jowea wrote:
| I'm not an expert on website hosting, but after reading some
| of the blog posts on Anubis, those people were truly at wit's
| end trying to block AI scrappers with techniques like the
| ones you imply.
| Symbiote wrote:
| That's a pretty big assumption.
|
| The largest site I work on has 100,000s of pages, each in
| around 10 languages -- that's already millions of pages.
|
| It generally works fine. Yesterday it served just under 1000
| RPS over the day.
|
| AI crawlers have brought it down when a _single crawler_ has
| added 100, 200 or more RPS distributed over a wide range of IPs
| -- it 's not so much the number of extra requests, though it's
| very disproportionate for one "user", but they can end up
| hitting an expensive endpoint excluded by robots.txt and
| protected by other rate-limiting measures, which didn't
| anticipate a DDoS.
| yodon wrote:
| Discussed yesterday (270+ comments)[0]
|
| [0]https://news.ycombinator.com/item?id=44432385
| dawnerd wrote:
| I've been using this for a while on my mastodon server and after
| a few tweaks to make sure it wasn't blocking legit traffic it's
| been really working great. Between Microsoft and Meta, they were
| hitting my services more than any other traffic combined which
| says a lot of you know how noisy mastodon can be. Server load
| went down dramatically.
|
| It also completely put a stop to perplexity as far as I can tell.
|
| And the robots file meant nothing, they'd still request it
| hundreds of thousands of times instead of caching it. Every
| request they'd hit it first then hit their intended url.
| TechDebtDevin wrote:
| This does nothing dude. Literally nothing. OpenAI or whoever
| are just going to hire people like me who dont get caught. Stop
| ruining the experience of users and allowing cf to fill the
| internet with more bloated javascript challenge pages and
| privacy invading fingerprinting. Stop making cf the police of
| the internet. We're literally handing the internet to this
| company on a silver platter to do MITM attacks on our privacy
| and god knows what else. Fucking wild.
| fluidcruft wrote:
| They literally said it significantly reduced their server
| resource usage. Are you suggesting they are lying?
| drowsspa wrote:
| Why do you think you have the moral high ground here?
| dawnerd wrote:
| Well the alternative is to not have an instance at all so...
| what do you suggest? I'm not paying for the other services,
| it's already expensive enough to run the site.
|
| The goal isn't to stop 100% of scrapers, it was to reduce
| server load to a level that wasn't killing the site.
| jowea wrote:
| You want them to pay the server costs to serve content to AI
| scrappers for free? The alternative is Anubis, which is maybe
| equally annoying to users in a different way.
| danielspace23 wrote:
| Have you considered Anubis? I know it's harder to install, but
| personally, I think the point of Mastodon is trying to avoid
| centralization where possible, and CloudFlare is one of the
| corporations that are keeping the internet centralized.
| dawnerd wrote:
| Haven't heard of it, will look into it. I agree, I'd rather
| not have cloudflare but for what they provide for free it's a
| tough offer to pass up
| account42 wrote:
| Yay, looking forward to more CAPTCHAs as a regular user.
| thephotonsphere wrote:
| account wall :-(
| deadbabe wrote:
| No one else can really do this except Cloudflare.
| dirkc wrote:
| I assume they will "protect original content online" by blocking
| LLM clients from ingesting data as context?
|
| I'm not optimistic that you can effectively block your original
| content from ending up in training sets by simply blocking the
| bots. For now I just assume that anything I put online will end
| up being used to train some LLM
| NullCascade wrote:
| How would you do the opposite of this? Optimize your content to
| be more likely crawled by AI bots? I know traditional Google-
| focused SEO is not enough because these AI bots often use other
| web search/indexing APIs.
| TechDebtDevin wrote:
| There are script tags you can put in your site from LLM SEO
| companies if you want your content to be indexed by Perplexity
| or OpenAI. Theyre kind of too new for me to reccomend.
| zargath wrote:
| Sounds very basic, sadly.
|
| Anybody know why these web crawling/bot standards are not
| evolving ? I believe robots.txt was invented in 1994(thx
| chatgpt). People have tried with sitemaps, RSS and IndexNow, but
| its like huge$$ organizations are depending on HelloWorld.bas
| tech to control their entire platform.
|
| I want to spin up endpoints/mcp/etc. and let intelligent bots
| communicate with my services. Let them ask for access, ask for
| content, pay for content, etc. I want to offer solutions for bots
| to consume my content, instead of having to choose between full
| or no access.
|
| I am all for AI, but please try to do better. Right now the
| internet is about to be eaten up by stupid bot farms and served
| into chat screens. They dont want to refer back to their source
| and when they do its with insane error rates.
| TechDebtDevin wrote:
| This comment seems like it comes from a Cloudflare employee.
|
| This is clearly the first step in cf building out a marketplace
| where they will (fail) at attempting to be the middleman in a
| useless market between crawlers and publishers.
| zargath wrote:
| nah, disappointed cf customer
| stereolambda wrote:
| > I believe robots.txt was invented in 1994(thx chatgpt).
|
| Not to pick on you, but I find it quicker to open new tab and
| do "!w robots.txt" (for search engines supporting the bang
| notation) or "wiki robots.txt"<click> (for Google I guess). The
| answer is right there, no need to explain to LLM what I want or
| verify [1].
|
| [1] Ok, Wikipedia can be wrong, but at least it is a commonly
| accessible source of wrong I can point people to if they call
| me out. Plus my predictive model of Wikipedia wrongness gives
| me pretty low likelihood for something like this, while for
| ChatGPT it is more random.
| reaperducer wrote:
| _robots.txt was invented in 1994(thx chatgpt)_
|
| Thought of and discussed as a possibility in 1994.
|
| Proposed as a standard in 2019.
|
| Adopted as a standard in 2022.
|
| Thanks, IETF.
| StochasticLi wrote:
| _ehem_ https://github.com/Kaliiiiiiiiii-Vinyzu/patchright
| j45 wrote:
| This is interesting. I'm a fan of Cloudflare, and appreciate all
| the free tiers they put out there for many.
|
| Today I see this article about Cloudflare blocking scrapers.
| There are useful and legitimate cases where I ask Claude to go
| research something for me. I'm not sure if Cloudflare discerns
| legitimate search/research traffic from an AI client vs scraping.
| Of the sites that are blocked by default will include content by
| small creators (unless on major platforms with deal?), while the
| big guys who have something to sell like an Amazon, etc, will
| likely be able to facilitate and afford a deal to show up more in
| the results.
|
| A few days ago, Cloudflare is also looking to charge AI companies
| to scrape the content, which is cached copies of other people's
| content. I'm guessing it will involve paying the owners of the
| data at some point as well. Being able to exclude it from this
| purpose (sell/license content, or scrape) would be a useful
| lever.
|
| Putting those two stories together:
|
| - Is this a new form of showing up in the AISEO (Search
| everywhere optimization) to show up in an AI's corpus or ability
| to search the web, or paying licensing fees instead of
| advertising fees.. these could be new business models which are
| interesting, but trying to see where these steps may vector ahead
| towards, and what to think about today.
|
| - With training data being the most valuable thing for AI
| companies, and this is another avenue for revenue for Cloudflare,
| this can look like a solution which helps with content licensing
| as a service.
|
| I'd like to see where abstracting this out further ends up going
|
| Maybe I'm missing something, is anyone else seeing it this way,
| or another way that's illuminating to them? Is anyone thinking
| about rolling their own service for whatever parts of Cloudflare
| they're using?
| ec109685 wrote:
| It seems like search access is more valuable these days since
| reasoning requires realtime access to site data.
| ssijak wrote:
| I dont want this by default. I want my website to end up in AI
| chatbots. For SEO
| abalashov wrote:
| Few people realise that virtually everything we do online has,
| until this point, been free training to make OpenAI, Anthropic,
| etc. richer while cutting humans--the ones who produced the value
| --out of the loop.
|
| It might be too little, too late, at this juncture, and this
| particular solution doesn't seem too innovative. However, it is
| directionally 100% correct, and let's hope for massively more
| innovation in defending against AI parasitism.
| k__ wrote:
| Is anyone suing to make the models and their weights open
| source?
| jefftk wrote:
| I write online (comments here, open source software, blogging,
| etc) because I have ideas I want to share. Whether it's "I did
| a thing and here's how" or "we should change policy in this
| specific way" or "does anyone know how to X" I'm happy for this
| to go into training models just like I'm happy for it to go
| into humans reading.
| dolebirchwood wrote:
| Thank you for having this attitude. I have never attempted
| any blogging because I always figured no one is actually
| going to read it. With LLMs, however, I know they will. I
| actually see this as a motivation to blog, as we are in a
| position to shape this emerging knowledge base. I don't find
| it discouraging that others may be profiting off our freely
| published work, just as I myself have benefited tremendously
| from open source and the freely published works of others.
| arkmm wrote:
| This is an interesting take, thanks for sharing. I wonder
| how someone should adjust their blogging if they believe
| their primary audience will be LLMs.
| lawlessone wrote:
| SEO -> LLMEO
| godelski wrote:
| Tbh, that content I'm mostly fine with. My only real issue is
| that people are making trillions off the free labor of people
| like you and me, giving less time to create that OSS and
| blogs. But this isn't new to AI, it is just scaled.
|
| What I do care about is the theft of my identity. A person
| may learn from the words I write but that person doesn't end
| up mimicking the way I write. They are still uniquely
| themselves.
|
| I'm concerned that the more I write the more my text becomes
| my identifier. I use a handle so I can talk more openly about
| some issues.
|
| We write OSS and blog because information should be free. But
| that information is then being locked behinds paywalls and
| becoming more difficult to be found through search. Frankly,
| that's not okay
| bob1029 wrote:
| > OSS
|
| > people are making trillions off the free labor of people
| like you and me
|
| I read "No Discrimination Against Fields of Endeavor" to
| also include LLMs and _especially_ the cases that we most
| deeply disagree with.
|
| Either we believe in the principles of OSS or we do not. If
| you do not like the idea of your intellectual property
| being used for commercial purposes then this model is
| definitely not for you.
|
| There is no shame in keeping your source code and other IP
| a secret. If you have strong expectations of being
| compensated for your work, then perhaps a different
| licensing and distribution model is what you are after.
|
| > that information is then being locked behinds paywalls
| and becoming more difficult to be found through search
|
| Sure - If you give up and delete everything. No one is
| forcing you to put your blog and GH repos behind a paywall.
| mattl wrote:
| Open source software typically has a license. People not
| following the license isn't tolerated.
|
| This is what AI scrapers are doing. They're taking your
| code, your artwork and your writing without any
| consideration for the license.
| blibble wrote:
| > Either we believe in the principles of OSS or we do
| not. If you do not like the idea of your intellectual
| property being used for commercial purposes then this
| model is definitely not for you.
|
| I've been writing open source for more than 20 years
|
| I gave away my work for free with one condition: leave my
| name on it (MIT license)
|
| the AI parasites then strip the attribution out
|
| they are the ones violating the principles of open source
|
| > then perhaps a different licensing and distribution
| model is what you are after.
|
| I've now stopped producing open source entirely
|
| and I suggest every developer does the same until the
| legal position is clarified (in our favour)
| godelski wrote:
| > Either we believe in the principles of OSS or we do
| not.
|
| What about respecting licenses?
|
| Seriously, don't lick the boot. We can recognize that
| there's complexity here. Trivializing everything only
| helps the abusers.
|
| Giving credit where credit is due is not too much to ask.
| Other people making money off my work can be good[0].
| Taking credit for it is insulting
|
| [0] If you're not making much, who cares. But if you're a
| trillion dollar business you can afford to give a little
| back. Here's the truth, OSS only works if we get enough
| money and time to do the work. That's either by having a
| good work life balance and good pay or enough donations
| coming in. We've been mostly supported by the former, but
| that deal seems to be going away
| lxgr wrote:
| > What I do care about is the theft of my identity. A
| person may learn from the words I write but that person
| doesn't end up mimicking the way I write. They are still
| uniquely themselves.
|
| Of course they do, to some extent. Just because it's been
| infeasible to track the exact "graph of influence", that's
| literally how humans have learned to speak and write for as
| long as we've had language and writing.
|
| > I'm concerned that the more I write the more my text
| becomes my identifier. I use a handle so I can talk more
| openly about some issues.
|
| That's a much more serious concern, in my view. But I
| believe that LLMs are both the problem and solution here:
| "Remove style entropy" is just a prompt away, these days.
| BeetleB wrote:
| > A person may learn from the words I write but that person
| doesn't end up mimicking the way I write.
|
| Oh, I wish I could get AI to mimic the way I write! I'd pay
| money for it. I often want to type up an email/doc/whatever
| but don't because of occasional RSI issues. If I could get
| an AI to type it up for me while still sounding like me -
| that would be a big boon for my health.
| andy99 wrote:
| It's cloudflare and parasites like them that will make the
| internet un-free. It's already happening, I'm either blocked or
| back to 1998 load times be cause of "checking your browser".
| They are destroying the internet and will make it so only
| people who do approved things on approved browsers (meaning let
| advertising companies monetize their online activity) will get
| real access.
|
| Cloudflare isn't solving a problem, they are just inserting
| themselves as an intermediary to extract a profit, and making
| everything worse.
| carlhjerpe wrote:
| I use Firefox with adblocking and some fingerprinting anti-
| measurements and I rarely hit their challenges. Your IP
| reputation must be bad.
|
| They have an addon [1] that helps you bypass Cloudflare
| challenges anonymously somehow, but it feels wrong to install
| a plugin to your browser from the ones who make your web
| experience worse
|
| 1: https://developers.cloudflare.com/waf/tools/privacy-pass/
| godelski wrote:
| I'm in a pretty similar boat except I frequently hit
| challenges. Especially if I use a VPN (which is more
| trustworthy than my ISP). Ironically, I'm using Cloudflare
| for DoH
| lxgr wrote:
| I'd be surprised if Cloudflare were actually correlating
| DoH requests to HTTP requests following them, so I don't
| think that's a signal they are likely to use.
| MichaelZuo wrote:
| If your on ipv6, I think they have to for ipv6 addresses...
| there's just way too many bots and way too many addresses to
| feasibly do anything more precise.
|
| If your on ipv4 you should check whether your behind a NAT
| otherwise you may have gotten an address that was previously
| used by a bot network.
| lxgr wrote:
| > I think they have to for ipv6 addresses... there's just
| way too many bots and way too many addresses
|
| Are you really arguing that it's legitimate to consider
| _all IPv6_ browsing traffic "suspicious"?
|
| If anything, I'd say that IPv4 is probably harder, given
| that NATs can hide hundreds or thousands of users behind a
| single IPv4 address, some of which might be malicious.
|
| > you may have gotten an address that was previously used
| by a bot network.
|
| Great, another "credit score" to worry about...
| slenk wrote:
| How is Cloudflare a parasite? I can use Cloudflare, and get
| their AI protection, for free. I have dozens of domains I
| have used with Cloudflare at one point and I haven't paid
| them a dime.
| fsflover wrote:
| They put themselves as a middle man for almost the whole
| Internet, collect huge usage data about everyone and block
| anybody who doesn't use mainstream tools:
|
| https://news.ycombinator.com/item?id=42953508
|
| https://news.ycombinator.com/item?id=13718752
|
| https://news.ycombinator.com/item?id=23897705
|
| https://news.ycombinator.com/item?id=41864632
|
| https://news.ycombinator.com/item?id=42577076
| baq wrote:
| valid.
|
| ...but OTOH it's their customers who want all of that and
| pay to get that, because the alternative is worse.
|
| rock and a hard place.
| slenk wrote:
| Right - do I want them getting some info from me, or do I
| want my IP address exposed?
|
| Besides CloudFront, which still costs money, what other
| option is there for semi-privacy and caching for free?
| mattl wrote:
| bunny.net has some options
| slenk wrote:
| I will have to check them out I guess
| aorth wrote:
| As the old addage goes: If you're not paying for it,
| you're the product.
|
| Lots of nuance, but generally: pay for things you use.
| Servers, engineers, and research and development are not
| free, so someone has to pay.
| qualeed wrote:
| Lots of services don't even let me pay if I wanted to, so
| I am forced to be the product. (Donating typically does
| not un-productify myself).
|
| Or I pay and am _still_ the product. Just with less in-
| my-face ads.
| matt-p wrote:
| Cloud front is pretty much free for your first TB. Fastly
| has a free plan.
|
| Though why should it be for free?
| slenk wrote:
| Multiple people have brought that up. I pay for
| everything else, why not one more.
|
| Although bunny.net won't take ANY of my credit or debit
| cards
| MisterTea wrote:
| I want to know if there is a way to design an alternative
| that isn't controlled by a single entity which allows
| gatekeeping.
| AnthonyMouse wrote:
| You can add another one as a result of _this_ article:
| The data you need to train AI and the data you need to
| build a search engine are the same data. So now they 're
| inhibiting every new search engine that wants to compete
| with Google.
| lxgr wrote:
| > I have dozens of domains I have used with Cloudflare at
| one point and I haven't paid them a dime.
|
| Maybe _you_ haven 't, but your users (primarily those using
| "suspicious" operating systems and browsers) certainly have
| - with their time spent solving captchas.
| sealeck wrote:
| But Cloudflare have removed CAPTCHAs
| lxgr wrote:
| Not sure if you're joking, but if you're not:
| Congratulations on using a very "normal/safe"
| OS/browser/IP.
|
| I get captchas daily, without using any VPN and on
| several different IPs (work, home, mobile). The only
| crime I can think of is that I'm using Firefox instead of
| Chrome.
| Symbiote wrote:
| Since a few days ago, I've been getting Captchas hourly
| or more.
|
| It's probably because I use Firefox on Linux with an ad
| blocker.
|
| For my part, I've ensured we don't use Cloudflare at
| work.
| kelvinjps10 wrote:
| I use firefox on linux with an ad blocker and cloudfare
| works fine
| rockskon wrote:
| LLM scrapers have dramatically been increasing the cost of
| hosting various small websites.
|
| Without something being done, the data that these scrapers
| rely on would eventually no longer exist.
| benjiro wrote:
| I think the correct term is, that unrestricted LLM scrapers
| have dramatically been increasing the cost of hosting
| various small websites.
|
| Its not a issue when somebody does "ethical" scraping, with
| for instance, a 250ms delay between requests, and a active
| cache that checks specific pages (like news article links)
| to rescrape at 12 or 24h intervals. This type of scraping
| results in almost no pressure on the websites.
|
| The issue that i have seen, is that the more unscrupulous
| parties, just let their scrapers go wild, constantly
| rescraping again and again because the cost of scraping is
| extreme low. A small VM can easily push 1000's of scraps
| per second, let alone somebody with more dedicated
| resources.
|
| Actually building a "ethical" scraper involves more time,
| as you need to fine tune it per website. Unfortunately,
| this behavior is going to cost the more ethical scraper a
| ton, as anti-scraping efforts will increase the cost on our
| side.
| Tmpod wrote:
| The biggest issue for me is clearly masquerading their
| User-Agent strings. Regardless of whether they are slow
| and respectful crawlers, they should clearly identify
| themselves, provide a documentation URL and obey
| robots.txt. Without that, I have to play a frankly tiring
| game of cat and mouse, wasting my time and the time of my
| users (they have to put up with some form of captcha or
| PoW thing).
|
| I've been an active lurker in the self-hosting community
| and I'm definitely not alone. Nearly everyone hosting
| public facing websites, particularly those whose form is
| rather juicy for LLMs, have been facing these issues. It
| costs more time and money to deal with this, when
| applying a simple User-Agent block would be much cheaper
| and trivial to do and maintain.
|
| _sigh_
| dceddia wrote:
| Yep this terrifies me, 100%. We're slowly losing the open
| internet and the frog is being boiled slowly enough that
| people are very happy to defend the rising temperature.
|
| If DDoS wasn't a scary enough boogeyman to get people to
| install Cloudflare as a man-in-the-middle on all their
| website traffic, maybe the threat of AI scrapers will do the
| trick?
|
| The thing about this slow slide is it's always defensible.
| Someone can always say "but I don't want my site to be
| scraped, and this service is free, or even better yet, I can
| set up my own toll booth and collect money! They're
| wonderful!"
|
| Trouble is, one day, at this rate, almost all internet
| traffic will be going through that same gate. And once they
| have _literally everyone_ (and all their traffic)... well,
| internet access is an immense amount of power to wield and I
| can't see a world in which it remains untainted by commercial
| and government interests forever.
|
| And "forever" is what's at stake, because it'll be near
| impossible to recover from once 99% of the population is
| happy to use one of the 3 approved browsers on the 2 approved
| devices (latest version only). Feels like we're already
| accepting that future at an increasing rate.
| RiverCrochet wrote:
| The Internet is not the first global network. Before the
| Internet, you had the global telephone network. It, too,
| strangulated end users, but eventually became stagnant,
| overpriced, and irrelevant. Super long-term, the current
| Internet is not immune from this. Internet standards are
| about getting as complicated and quirky as the old Bell
| stuff that was trying to make miles of buried copper the
| future, and if regulatory/commercial forces freeze this
| stuff in place, it's going to lead to stagnation
| eventually.
|
| Something coming down the pike I think, for example, is
| that IPv4 addresses are going to get realllly expensive
| soon. That's going to lead to all sorts of interesting
| things in the Internet landscape and their applications.
|
| I'm sure we'll probably have to spend some decades in the
| "approved devices and browers only" world before a next
| wave comes.
| mattl wrote:
| We need a reasonable alternative to some of what Cloudflare
| does that can be easily installed as a package on Linux
| distributions without any of the following to install it.
|
| * curl | bash
|
| * Docker
|
| * Anything that smacks of cryptocurrency or other scams
|
| Just a standard repo for Debian and RHEL derived distros.
| Fully open source so everyone can use it. (apt/dnf install
| no-bad-actors)
|
| Until that exists, using Cloudflare is inevitable.
|
| It needs to be able to at least:
|
| * provide some basic security (something to check for sql
| injection, etc)
|
| * rate limiting
|
| * User agent blocking
|
| * IP address and ASN blocking
|
| Make it easy to set up with sensible defaults and a way to
| subscribe to blocklists.
| saint_yossarian wrote:
| I remember using mod_security with Apache long ago for
| some of this, looks like it's still around and now also
| supports Nginx and IIS: https://modsecurity.org/
| brumar wrote:
| Correction: extract monstreous profits. When I read about the
| revenues associated with Reddit AI deals, I can't even
| imagine what could possibly be deals that cover half of the
| internet. Cynically speaking, it's a genious level move.
| axus wrote:
| From the server perspective Cloudflare is solving problems
| and not causing problems to other servers.
|
| Analogy: locks for high-value items in grocery stores are
| annoying to customers, but other stores aren't being coerced
| by the locksmith to use them.
| rramon wrote:
| Isn't there a possibility that model makers retaliate by
| erasing them and their frameworks from memory, hurting CF
| adoption by devs?
| cmeacham98 wrote:
| Cutting humans out of what loop? What jobs or opportunities
| were people posting Reddit comments or whatever getting that
| are now going to AI?
| Larrikin wrote:
| People who used to post gained knowledge from their
| profession or hobby. I don't bother posting any of that
| information on large sites like Reddit anymore, for various
| reasons but AI scraping solidified.
|
| I'll still post on the increasingly fewer hobby message
| boards that are out there.
| kamarg wrote:
| > What jobs or opportunities were people posting Reddit
| comments or whatever getting that are now going to AI?
|
| Content writing, product reviews (real & fake), creative
| writing, customer support, photography/art to name a few off
| the top of my head.
| fkyoureadthedoc wrote:
| Now the astroturfing is done by AI agents instead of hard
| working serfs in a call center, you hate to see it
| godelski wrote:
| Including your comment, including this comment.
|
| HN itself is routinely scraped. What makes me most
| uncomfortable is deanonymization via speech analysis. It's
| something we can already do but is hard to do at scale. This is
| the ultimate tool for authoritarians. There's no hidden
| identities because your speech is your identifier. It is
| without borders. It doesn't matter if your government is good,
| a bad acting government (or even large corporate entity) has
| the power to blackmail individuals in other countries.
|
| We really are quickly headed towards a dystopia. It could
| result in the entire destruction of the internet or an
| unprecedented level of self censorship. We already have
| algospeak because platform censorship[0]. But this would be a
| different type of censorship. Much more invasive, much more
| personal. There are things worse than the dark forest
|
| [0] literally yesterday YouTube gave me, a person in the 25-60
| age bracket, a content warning because there was a video about
| a person that got removed from a plane because they wore a
| shirt saying "End veteran suicide".
|
| [0.1] Even as I type this I'm censored! Apple will allow me to
| swipe the word suicidal but not suicide! Jesus fuck guys! You
| don't reduce the mental health crisis by preventing people from
| even being able to discuss their problems, you only make it
| worse!
| Kostic wrote:
| This would be true if not for open-weights (and even some open
| source) LLMs that exist today. Not everything should be done
| for profit.
| giancarlostoro wrote:
| There's a reason reddit started charging for API usage.
| fkyoureadthedoc wrote:
| It surely wasn't to force users into their shitty app where
| they can't block ads and definitely had nothing to do with
| their IPO. It was the AI.
| giancarlostoro wrote:
| Ah yes, its only because of ONE singular reason they
| started charging for API usage. Are you okay? I'm listing
| one reason out of many as to why reddit started charging
| for API usage. After all, reddit is a for profit website.
| fkyoureadthedoc wrote:
| > There's *A* reason
|
| this u chief?
| dwoldrich wrote:
| I think the parasitism goes quite a bit further than AI. We're
| being digested not parasitized.
| Dig1t wrote:
| That was always the cost of free and open exchange of ideas
| though. The idea of the internet in the first place was to
| allow people to communicate in the open and publish ideas
| freely. There was never any stipulation that using the
| published ideas to make money was off limits.
|
| Technology has advanced and now reading the sum total of the
| freely exchanged ideas has become particularly valuable. But
| who cares? The internet still exists and is still usable to
| freely exchange ideas the way it's always been.
|
| The value that one website provides is a minuscule amount, the
| value of one individual poster on Reddit is minuscule. Are we
| asking that each poster on Reddit be paid 1 penny (that's
| probably what your posts are worth) for their individual
| contribution? My websites were used to train these models
| probably, but the value that each contributed is so small that
| I wouldn't even expect a few cents for it.
|
| The person who's going to profit here is Cloudflare or the
| owners of Reddit, or any other gatekeeper site that is already
| profiting from other people's contributions.
|
| The "parasitism" here just feels like normal competition
| between giant companies who have special access to information.
| tcdent wrote:
| Everything you did, up to this point, hopefully has made
| _someone_ richer, otherwise you contributed literally zero
| value to the world.
|
| The internet is a free, open forum for the exchange of ideas.
| Not a private diary for your best secrets.
| friedtofu wrote:
| What the hell...
|
| Even if you're directing this at the user's blog posts
| specifically; this is a ridiculously pessimistic, sad way to
| view things.
|
| I hope you're just having a bad day because if you sincerely
| have this greedy, cynical mindset day to day(towards
| blogging, software, offline/real life activities, whatever) I
| feel sorry for you.
| lxgr wrote:
| GP's comment might be provocatively phrased, but I don't
| think it's an invalid point to have:
|
| When I publish something online for free, i.e. without
| requiring authentication or payment, be it a Reddit
| comment, a blog post, a Stackoverflow answer or anything
| else, I do so hoping that it will be useful to somebody
| somehow, without any illusions about being able to gatekeep
| some types of current or future consumers.
| risyachka wrote:
| Maybe so, but I'll take Cloudflare over OpenAI and Meta every
| time.
| lofaszvanitt wrote:
| Cyberpunk aged well. "You better not be on the unprotected
| internet". Too many hazards out there. Rogue AIs and other
| shit...
|
| Cloudflare is here to protecc you from all those evils. Just
| come under our umbrella.
| bawolff wrote:
| I think its 100% ok to freely train on public internet data.
|
| What is absolutely not ok is to crawl at such an excessive
| speed that it makes it difficult to host small scale websites.
|
| Truly a tragedy of the commons.
| tedd4u wrote:
| Agree. The problem lately is that even if each single scraper
| is doing so "reasonably," there are so many individuals and
| groups doing this that it's still too onerous for many sites.
| And of course many are not "reasonable."
| jowea wrote:
| Is it even possible that Cloudfare could manage to block all AI
| data scrapping? I think this measure is just going to make it
| harder and more expensive, which will stop AI scrappers from
| hitting every single page every single day and creating
| expenses for publishers, but not actually stop their data from
| ending up in a few datasets.
| mathiaspoint wrote:
| This has been going on even since early social media. I think
| most of the users actually prefer it.
| jjangkke wrote:
| so TLDR it adjusts your robot.txt and relies on cloudflare to
| catch bot behavior and it doesn't actually do any sophisticated
| residential proxy filtering or common bypass methods that works
| on cloudflare turnstill, do I have this correct?
|
| this just pushes AI agents "underground" to adopt the behavior of
| a full blown stealth focused scraper which makes it harder to
| detect.
| nemild wrote:
| Think this is the future, as the AI Web takes over the human web.
|
| At Coinbase, we've been building tools to make the blockchain the
| ideal payment rails for use cases like this with our x402
| protocol:
|
| https://www.x402.org/
|
| Ping if you're interested in joining our open source community.
| bgwalter wrote:
| The destruction of the Web and IP theft needs to be addressed
| legally. The opinion of a single judge notwithstanding, "AI"
| scraping already violates copyright. This needs to be made
| explicit in law and scrapers must get the same treatment as
| Western governments gave to thousands of individuals who were
| bankrupted or jailed for copyright infringement.
|
| We are in the Napster phase of Web content stealing.
| zackmorris wrote:
| As usual, this is the wrong approach.
|
| The open web is akin to the commons, public domain and public
| land. So this is like putting a spy cam on a freeway billboard,
| detecting autonomous vehicles, and shining a spotlight at their
| camera to block them from seeing the ad. To what end?
|
| Eventually these questions will need to be decided in court:
|
| 1) Do netizens have the right to anonymity? If not, then we'll
| have to disclose whether we're humans or artificial beings.
| Spying on us and blocking us on a whim because our behavior
| doesn't match social norms will amount to an invasion of privacy
| (eventually devolving into papers please).
|
| 2) Is blocking access to certain users discrimination? If not,
| then a state-sanctioned market of civil rights abuse will grow
| around toll roads (think whites-only drinking fountains).
|
| 3) Is downloading copyrighted material for learning purposes by
| AI or humans the same as pirating it and selling it for profit?
| If so, then we will repeat the everyone-is-a-criminal torrenting
| era of the 2000s and 2010s when "making available" was treated
| the same as profiting from piracy, and take abuses by HBO, the
| RIAA/MPAA and other organizations who shut off users' internet
| connections through threat of legal actions like suing for
| violating the DMCA (which should not have been made law in the
| first place).
|
| I'm sure there are more. If we want to live in a free society,
| then we must be resolute in our opposition of draconian
| censorship practices by private industry. Gatekeeping by large,
| monopolistic companies like Cloudflare simply cannot be
| tolerated.
|
| I hope that everyone who reads this finds alternatives to
| Cloudflare and tells their friends. If they insist on pursuing
| this attack on our civil rights for profit, then I hope we build
| a countermovement by organizing with the EFF and our elected
| officials to eventually bring Cloudflare up on antitrust charges.
|
| Cloudflare has shown that they lack the judgement to know better.
| Which casts doubt on their technical merits and overall vision
| for how the internet operates. By pursuing this course of action,
| they have lost face like Google did when it removed its "don't be
| evil" slogan from its code of conduct so it could implement
| censorship and operate in China (among other ensh@ttification-
| related goals).
|
| Edit: just wanted to add that I realize this may be an opt-in
| feature. But that's not the point - what I'm saying is that this
| starts a bad precedent and an unnecessary arms race, when we
| should be questioning whether spidering and training AI on
| copyrighted materials are threats in the first place.
| sct202 wrote:
| My data served by Cloudflare has increased to 100gb /month
| compared to <20gb like 2 years ago, and they're all fairly static
| hobby sites. Actual people traffic is down by like half in the
| same time frame, so I imagine a lot of this is probably cost
| savings for Cloudflare to reduce resource usage.
| Apofis wrote:
| Makes total sense, bandwidth on this scale is expensive.
| e38383 wrote:
| Why is every second article about this claiming that it's
| automatic? It needs to be turned on or at least there was no
| mention of automatic in the original blog post.
|
| I really hope that we can continue training AI the same way we
| train humans - basically for free.
| aunty_helen wrote:
| I saw yesterday that they were going to allow websites to charge
| per scrape.
|
| Looks like cloudflare just invented the new App Store.
| hmate9 wrote:
| Isn't this only useful for blogs, news sites, or forums? Why
| would I want an AI to know less about my product? I want it to
| understand it, talk about it, and ideally recommend it. Should be
| default off.
| maximilianburke wrote:
| Every evolution of the web, from Web 2 giving us walled gardens
| to Web 3 giving us, well, nothing, to what we have now is taking
| us further from a network of communities and personal
| repositories of knowledge.
|
| Sure, fidelity has gotten better but so much has been lost.
| sneak wrote:
| This idea that you can publish data for people to download and
| read but not for people to download and store, or print, or think
| about, or train on is a doomed one.
|
| If you don't want people reading your data, don't put it on the
| web.
|
| The concept that copyright extends to "human eyeballs only" is a
| silly one.
| ChrisArchitect wrote:
| [dupe] https://news.ycombinator.com/item?id=44432385
___________________________________________________________________
(page generated 2025-07-02 23:01 UTC)