https://go-to-hellman.blogspot.com/2025/03/ai-bots-are-destroying-open-access.html skip to main | skip to sidebar Go To Hellman This machine surrounds hate and forces it to surrender. Friday, March 21, 2025 AI bots are destroying Open Access There's a war going on on the Internet. AI companies with billions to burn are hard at work destroying the websites of libraries, archives, non-profit organizations, and scholarly publishers, anyone who is working to make quality information universally available on the internet. And the technologists defending against this broad-based attack are doing everything they can to preserve their outlets while trying to remain true to the mission of providing the digital lifeblood of science and culture to the world. Yes, many of these beloved institutions are under financial pressures in the current political environment, but politics swings back and forth. The AI armies are only growing more aggressive, more rapacious, more deceitful and ever more numerous. I'm talking about the voracious hunger of AI companies for good data to train Large Language Models (LLMs). These are the trillion-parameter sets of statistical weights that power things like Claude, ChatGPT and hundreds of systems you've never heard of. Good training data has lots of text, lots of metadata, is reliable and unbiased. It's unsullied by Search Engine Optimization (SEO) practitioners. It doesn't constantly interrupt the narrative flow to try to get you to buy stuff. It's multilingual, subject specific, and written by experts. In other words, it's like a library. At last week's Code4lib conference hosted by Princeton University Library, technologists from across the library world gathered to share information about library systems, how to make them better, how to manage them, and how to keep them running. The hot topic, the thing everyone wanted to talk about, was how to deal with bots from the dark side. robot head emoji with eyes of sauron Bots on the internet are nothing new, but a sea change has occurred over the past year. For the past 25 years, anyone running a web server knew that the bulk of traffic was one sort of bot or another. There was googlebot, which was quite polite, and everyone learned to feed it - otherwise no one would ever find the delicious treats we were trying to give away. There were lots of search engine crawlers working to develop this or that service. You'd get "script kiddies" trying thousands of prepackaged exploits. A server secured and patched by a reasonably competent technologist would have no difficulty ignoring these. The old style bots were rarely a problem. They respected robot exclusions and "nofollow" warnings. The warning helped bots avoid volatile resources and infinite parameter spaces. Even when they ignored exclusions they seemed to be careful about it. They declared their identity in "user-agent" headers. They limited the request rate and number of simultaneous requests to any particular server. Occasionally there would be a malicious bot like a card-tester or a registration spammer. You'd often have to block these based on IP address. It was part of the landscape, not the dominant feature. The current generation of bots is mindless. They use as many connections as you have room for. If you add capacity, they just ramp up their requests. They use randomly generated user-agent strings. They come from large blocks of IP addresses. They get trapped in endless hallways. I observed one bot asking for 200,000 nofollow redirect links pointing at Onedrive, Google Drive and Dropbox. (which of course didn't work, but Onedrive decided to stop serving our Canadian human users). They use up server resources - one speaker at Code4lib described a bug where software they were running was using 32 bit integers for session identifiers, and it ran out! The good guys are trying their best. They're sharing block lists and bot signatures. Many libraries are routinely blocking entire countries (nobody in china could possibly want books!) just to be able to serve a trickle of local requests. They are using commercial services such as Cloudflare to outsource their bot-blocking and captchas, without knowing for sure what these services are blocking, how they're doing it, or whether user privacy and accessibility is being flushed down the toilet. But nothing seems to offer anything but temporary relief. Not that there's anything bad about temporary relief, but we know the bots just intensify their attack on other content stores. direct.mit.edu Verifying you are human. This may take a few seconds. direct.mit.edu needs to verify the security of your connection before proceeding. Verification is taking longer than expected. Check your internet connection and refresh the page if the issue persists. The view of MIT Press's Open-Access site from the Wayback Machine. The surge of AI bots has hit Open Access sites particularly hard, as their mission conflicts with the need to block bots. Consider that Internet Archive can no longer save snapshots of one of the best open-access publishers, MIT Press because of cloudflare blocking. (see above) Who know how many books will be lost this way? Or consider that the bots took down OAPEN, the worlds most important repository of Scholarly OA books, for a day or two. That's 34,000 books that AI "checked out" for two days. Or recent outages at Project Gutenberg, which serves 2 million dynamic pages and a half million downloads per day. That's hundreds of thousands of downloads blocked! The link checker at doab-check.ebookfoundation.org (a project I worked on for OAPEN) is now showing 1,534 books that are unreachable due to "too many requests". That's 1,534 books that AI has stolen from us! And it's getting worse. Thousands of developer hours are being spent on defense against the dark bots and those hours are lost to us forever. We'll never see the wonderful projects and features they would have come up with in that time. The thing that gets me REALLY mad is how unnecessary this carnage is. Project Gutenberg makes all its content available with one click on a file in its feeds directory. OAPEN makes all its books available via an API. There's no need to make a million requests to get this stuff!! Who (or what) is programming these idiot scraping bots? Have they never heard of a sitemap??? Are they summer interns using ChatGPT to write all their code? Who gave them infinite memory, CPUs and bandwidth to run these monstrosities? (Don't answer.) We are headed for a world in which all good information is locked up behind secure registration barriers and paywalls, and it won't be to make money, it will be for survival. Captchas will only be solvable by advanced AIs and only the wealthy will be able to use internet libraries. Or maybe we can find ways to destroy the bad bots from within. I'm thinking a billion rickrolls? Notes: 1. I've found that I can no longer offer more than 2 facets of faceted search. Another problematic feature is "did you mean" links. AI bots try to follow every link you offer even if there are a billion different ones. 2. Two projects, iocaine and nepenthes are enabling the construction of "tarpits" for bots. These are automated infinite mazes that bots get stuck in, perhaps keeping the bots occupied and not bothering anyone else. I'm skeptical. 3. Here is an implementation of the Cloudflare Turnstyle service (supposedly free) that was mentioned favorably at the conference. 4. It's not just open access, it's also Open Source. 5. Cloudflare has announced an "AI honeypot". Should be interesting. Posted by Eric at 6:19 PM Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest # Labels: Bots, Code4Lib, Open Access 0 comments: Contribute a Comment Note: Only a member of this blog may post a comment. Older Post Home Subscribe to: Post Comments (Atom) Free Ebook Foundation Making the world of ebooks safe for the free. Unglue.it 150,000 Free ebooks. Bluesky Eric's Bluesky. Mastodon Eric's feed. Blog Archive + V 2025 (2) o V March (1) # AI bots are destroying Open Access o > February (1) + > 2024 (7) o > November (1) o > October (1) o > August (1) o > June (2) o > May (1) o > April (1) + > 2023 (2) o > December (1) o > August (1) + > 2022 (1) o > February (1) + > 2021 (5) o > December (1) o > July (1) o > February (3) + > 2020 (3) o > December (1) o > October (1) o > September (1) + > 2019 (9) o > December (1) o > July (1) o > May (6) o > April (1) + > 2018 (11) o > December (2) o > October (1) o > September (1) o > August (1) o > June (1) o > May (2) o > April (1) o > March (1) o > January (1) + > 2017 (13) o > December (1) o > November (1) o > October (1) o > September (1) o > August (1) o > July (1) o > June (1) o > May (1) o > April (1) o > March (1) o > February (1) o > January (2) + > 2016 (12) o > December (1) o > October (1) o > September (1) o > July (1) o > June (1) o > May (1) o > April (1) o > March (2) o > February (1) o > January (2) + > 2015 (18) o > December (2) o > November (1) o > October (1) o > September (2) o > August (1) o > July (2) o > June (2) o > May (2) o > April (1) o > March (1) o > February (2) o > January (1) + > 2014 (30) o > December (2) o > November (3) o > October (4) o > September (4) o > August (1) o > July (2) o > June (3) o > May (2) o > April (2) o > March (3) o > February (2) o > January (2) + > 2013 (49) o > December (3) o > November (5) o > October (3) o > September (2) o > August (3) o > July (4) o > June (6) o > May (12) o > April (3) o > March (2) o > February (4) o > January (2) + > 2012 (31) o > December (3) o > November (2) o > October (3) o > September (2) o > August (2) o > July (2) o > June (2) o > May (2) o > April (3) o > March (3) o > February (3) o > January (4) + > 2011 (69) o > December (3) o > November (3) o > October (7) o > September (5) o > August (4) o > July (4) o > June (6) o > May (8) o > April (5) o > March (7) o > February (9) o > January (8) + > 2010 (87) o > December (9) o > November (4) o > October (7) o > September (5) o > August (5) o > July (5) o > June (7) o > May (6) o > April (8) o > March (10) o > February (7) o > January (14) + > 2009 (82) o > December (10) o > November (7) o > October (11) o > September (8) o > August (5) o > July (14) o > June (13) o > May (10) o > April (4) Popular Posts + [eyesofbot] AI bots are destroying Open Access There's a war going on on the Internet. AI companies with billions to burn are hard at work destroying the websites of libraries, archiv... + [IMG_2057] Protect Reader Privacy with Referrer Meta Tags Back when the web was new, it was fun to watch a website monitor and see the hits come in. The IP address told you the location of the user... + [catastroph] 2010 Summary: Libraries are Still Screwed In mathematics, catastrophe theory is the study of nonlinear dynamical systems which exhibit points or curves of singularity. The behavior ... + [Scihub_rav] Sci-Hub, LibGen, and Total Information Awareness "Good thing downloads NOT trackable!" was one twitter response to my post imagining a skirmish in the imminent scholarly publi... + [contentset] How to check if your library is leaking catalog searches to Amazon I've been writing about privacy in libraries for a while now, and I get a bit down sometimes because progress is so slow. I've come ... Subscribe To [arrow_drop] [icon_feed1] Posts [subscribe-] [subscribe-] [icon_feed1] Atom [arrow_drop] [icon_feed1] Posts [arrow_drop] [icon_feed1] Comments [subscribe-] [subscribe-] [icon_feed1] Atom [arrow_drop] [icon_feed1] Comments Me + Eric + Eric Go To Hellman Fan Page Go To Hellman on Facebook Labels + ebooks (94) + Libraries (72) + book industry (52) + E-book (49) + privacy (49) + Copyright (48) + business models (45) + linked data (33) + Semantic web (28) + Open Access (26) + Ungluing Ebooks (26) + Creative Commons (23) + physics (23) + Google Book Search (21) + Publishing (21) + Twitter (21) + Web Design and Development (21) + library automation (21) + Google (20) + Unglue.it (20) + Piracy (19) + Gluejar (18) + magic (18) + social practice (18) + ALA Midwinter (17) + RDF (17) + Overdrive (16) + linking technology (16) + metadata (16) + scholarly publishing (16) + Amazon (15) + Amazon Kindle (15) + identifiers (15) + Book Use (14) + Digital rights management (14) + ALA Annual (13) + Conferences (13) + Google Book Search Settlement (13) + HarperCollins (13) + Crossref (12) + EPUB (12) + OpenURL (12) + facebook (12) + Just Kidding (11) + New York Times (11) + RDFa (11) + Truth (11) + Big Library Read (10) + Book Digitization (10) + HTTP Secure (10) + The Four Corners of the Sky: A Novel (10) + isbn (10) + Blogging (9) + Public library (9) + Bugs (8) + Denny Chin (8) + IDPF (8) + Project Gutenberg (8) + URL redirection (8) + knowledgebases (8) + languages (8) + semtech2009 (8) + social networks (8) + wikipedia (8) + Attributor (7) + Book Rights Registry (7) + Hackathon (7) + Kickstarter (7) + Library (7) + New Jersey (7) + RA21 (7) + bit.ly (7) + Apple (6) + DOI (6) + Digital library (6) + Google Books (6) + IPad (6) + India (6) + Newspaper industry (6) + Open Source (6) + Public Domain (6) + semantic technology (6) + Digital Object Identifier (5) + Entrepreneurship (5) + Intel (5) + Interlibrary loan (5) + Library journal (5) + Microdata (5) + OCLC (5) + Star Trek (5) + authentication (5) + crowdfunding (5) + public identity (5) + running (5) + Aaron Swartz (4) + Amazon Web Services (4) + American Library Association (4) + Bell Labs (4) + Bitcoin (4) + Brian O'Leary (4) + Code4Lib (4) + DPLA (4) + Electronic Journals (4) + Google Analytics (4) + J. K. Rowling (4) + Koha (4) + Liblime (4) + LibraryThing (4) + Neal Stephenson (4) + Publishing Point (4) + SOPA (4) + Sweden (4) + my attic (4) + my dad (4) + Accessibility (3) + AdWords (3) + Adobe Digital Editions (3) + Baseball (3) + Bruce Springsteen (3) + Cryptography (3) + Forms of government (3) + Geolocation (3) + GitHub (3) + Google Wave (3) + JSTOR (3) + Macmillan (3) + Network Effect (3) + New York Public Library (3) + OWL (3) + PTFS (3) + Search Engine Optimization (3) + blockchain (3) + death (3) + genealogy (3) + hashtags (3) + http-range (3) + iPhone (3) + politics (3) + security (3) + unicode (3) + Advertising (2) + Americans with Disabilities Act of 1990 (2) + Book Design (2) + Book Industry Study Group (2) + Bots (2) + Database Licensing (2) + Disruptive technology (2) + Electronic Frontier Foundation (2) + FRBR (2) + Fair use (2) + Fan Fiction (2) + File sharing (2) + Fusion Tables (2) + Gitenberg (2) + Google Book (2) + Great Gatsby (2) + Hachette Book Group (2) + Hal Varian (2) + Hurricane Sandy (2) + Internet Archive (2) + John Sundman (2) + Nook (2) + OpenID (2) + OpenSource (2) + Payments (2) + Philadelphia Phillies (2) + Proxy server (2) + Radiolab (2) + Random House (2) + Rush Holt (2) + School library (2) + Social network (2) + Spam (2) + Star trek TNG (2) + Vegetables (2) + Wolfram Alpha (2) + ebrary (2) + linkedin (2) + poetry (2) + technology (2) + tr.im (2) + AdaptiveBlue (1) + Assistive Technology (1) + Beer (1) + Bibliocommons (1) + Brewster Kahle (1) + Clay Johnson (1) + Clayton M. Christensen (1) + Comic Con (1) + DBpedia (1) + DCWG (1) + Dave Winer (1) + Digital watermarking (1) + EBL (1) + Evan Ratliff (1) + Evert Taube (1) + Firefox (1) + GNU Affero General Public License (1) + Garage sale (1) + Hugh Howie (1) + Ian Davis (1) + Infochimps (1) + Infrastructure (1) + Instant Messaging (1) + Jon Stewart (1) + Knowledge representation (1) + Kobo (1) + Lawrence Lessig (1) + Mac OS X (1) + Metcalfe's Law (1) + Neil Gaiman (1) + Neurobiology (1) + ORCID (1) + Open Database License (1) + Open Knowledge Foundation (1) + Open Library (1) + PDDL (1) + Paypal (1) + ProQuest (1) + PubMed (1) + Qin Dynasty (1) + Qin Shi Huangdi (1) + RV Guha (1) + Ralph Waldo Emerson (1) + SPARQL (1) + Simon and Schuster (1) + Single sign-on (1) + Siri (1) + Star Wars (1) + Text-To-Speech (1) + Textbooks (1) + The Hitchhiker's Guide to the Galaxy (1) + Tim O'Reilly (1) + Tor (anonymity network) (1) + Warner Oland (1) + Weeds (1) + YouTube (1) + Zemanta (1) + Zola Books (1) + dead serious (1) + design patterns (1) + family (1) + gmail (1) + h1n1 (1) + life (1) + music (1) + patents (1) + shibboleth (1) + swedish music (1) + twitterdata (1) If you are a Comment Spammer, comments are closed to you. Your use of this material is subject to the Go To Hellman Blog License Agreement. This blog uses StatCounter analytics; they set a tracking cookie that may spy on you. web counter