[HN Gopher] Running ArchiveTeam's Warrior in Kubernetes
       ___________________________________________________________________
        
       Running ArchiveTeam's Warrior in Kubernetes
        
       Author : gmemstr
       Score  : 55 points
       Date   : 2025-02-05 18:04 UTC (4 hours ago)
        
 (HTM) web link (gabrielsimmer.com)
 (TXT) w3m dump (gabrielsimmer.com)
        
       | badlibrarian wrote:
       | Many of these sites are already captured and archived by proper
       | entities as required by federal law. More is better, I guess,
       | except when it isn't. Duplication of effort is a huge problem in
       | the humanities in general and with archiving in particular.
       | 
       | The whole concept needs to be rethought. Captures from these
       | tools show up under "ArchiveTeam" which is currently pumping
       | thousands of copies of the Google Home Page into the Wayback
       | Machine every week. Or at least trying to.
       | 
       | https://web.archive.org/web/20250122000033/www.google.com
       | 
       | Like so many things about archive.org, when you dig in you start
       | to find wonder and craziness at every turn.
        
         | myself248 wrote:
         | > by proper entities as required by federal law.
         | 
         | What federal law do you suppose is guiding the mass deletions?
         | That doesn't look like archiving to me. Now that the foxes are
         | running the henhouse, how reliable do you suppose their own
         | archives are?
        
           | badlibrarian wrote:
           | Some of the mass deletions are merely a new administration
           | setting up shop. Policies from the previous administration
           | don't belong on the current whitehouse.gov. They wind up here
           | instead https://bidenwhitehouse.archives.gov/
           | 
           | We pay half a billion in tax dollars for the National
           | Archives, and nearly a billion to the Library of Congress to
           | preserve these records. Others are managed as part of
           | Presidential Libraries.
           | 
           | Thousands of employees, dozens of facilities, billions of
           | dollars.
           | 
           | Meanwhile archive.org doesn't have air conditioning and
           | preserves physical material within the blast radius of an oil
           | refinery. They let vagrants sleep on their steps yet seem
           | surprised when they set the utility pole outsides on fire.
           | 
           | I didn't say it didn't need to be done. I said the whole
           | process needs to be rethought with professional supervision.
           | Setting up more volunteer K8 clusters so that more copies of
           | the Google Home Page can be captured with the wrong user
           | agent isn't going to save democracy.
        
             | toomuchtodo wrote:
             | Archive.org is outside of the reach of the US government,
             | and is globally distributed. When the US government deletes
             | or darks data (as it has recently done across wide swaths
             | of the federal government website properties), you have no
             | recourse. This means your argument about the resources that
             | go into the US government as a data custodian are
             | meaningless: the outcome is what is material, which is the
             | archival and long term custody & availability of the data
             | sets in scope. Arguably, the Internet Archive has recently
             | proven _better_ at this job than the US government
             | (unsurprising).
             | 
             | You're angry at a high value non profit operating on a
             | limited budget. It's weird. I recommend focusing on more
             | important issues than "it is icky around the richmond
             | facility, the power goes out once in a while, and they use
             | ambient air and convection for system cooling which I don't
             | like."
             | 
             | If you want to save democracy, the Internet Archive doesn't
             | do that itself. It protects the historical record. If you
             | want to save democracy, that's a different conversation.
             | 
             | https://blog.archive.org/2024/05/08/end-of-term-web-
             | archive/
             | 
             | https://web.archive.org/collection-
             | search/EndOfTerm2024PreEl...
             | 
             | (no affiliation)
        
               | badlibrarian wrote:
               | I would classify the end of term web archive (which
               | archive.org is, in its typical fashion, taking far too
               | much credit for) as an example of entities doing things
               | right.
               | 
               | https://eotarchive.org/partners/
               | 
               | And saying "archive.org is outside the reach of the US
               | government" -- hell, it's not even outside the reach of
               | the RIAA or the book company with the little penguin on
               | the cover.
               | 
               | We should have proper supervised federal archiving and
               | archive.org should be far better run, too.
               | 
               | And I don't know what Archive Team is but maybe they
               | could update their site to provide some information on
               | the people involved. And perhaps update their
               | understanding of what's possible with docker containers
               | while they're at it.
               | 
               | Because the counterpoint to a radicalized Musk screwing
               | around with government databases isn't an opposing group
               | of anonymous radicals screwing around with commercial
               | databases.
        
               | toomuchtodo wrote:
               | Agree the US government should contribute in some
               | capacity. Agree they should be robustly funded to do
               | this. But, checks and balances are also important, and
               | when a node goes rouge or dark, the system must be fault
               | tolerant and operate when degradation occurs. I
               | previously said "I trust Brewster and the rest of the IA
               | gang more than the US government to safeguard the
               | Internet Archive." [1] I feel this assertion has been
               | proven out over the last few weeks.
               | 
               | ArchiveTeam stands on its own as an independent,
               | community driven volunteer digital archival and
               | preservation effort. If you don't understand why, what,
               | and how they operate, look closer and be more curious
               | [2].
               | 
               | [1] https://news.ycombinator.com/item?id=41984664
               | 
               | [2] https://en.wikipedia.org/wiki/Wikipedia:Chesterton%27
               | s_fence
        
               | badlibrarian wrote:
               | If the checks and balances of NARA and LOC (6,000
               | employees, $1.5 billion in annual funding) is Brewster
               | Kahle asking for $10 on pages serving pirated Nintendo
               | games, then we're in a bit of trouble, aren't we?
        
               | toomuchtodo wrote:
               | On the contrary, the fact that a single person's
               | charitable digital archive can stand toe to toe with a
               | global superpower's archival efforts is a sign that
               | success is possible. We may see things differently
               | though, and that's fine. Would I want to fund NARA and
               | LOC more? Or the Internet Archive? I prefer the latter.
               | Checks and balances. I have donated $10 on your behalf
               | (in addition to my annual donations).
               | 
               | (lots of good people at NARA and the LOC, but they are
               | subject to the whims of the US electorate, which is not
               | great; the Internet Archive is not)
        
               | badlibrarian wrote:
               | I think archive.org deserves more funding and I also
               | think they need to decide if they're an archive, a
               | library, or a pirate site. Since each has a different set
               | of costs, legal risk, and projected longevity.
               | 
               | For the record my opinion is that they need to focus on
               | archival and with a few tweaks could make it safe for
               | more users to upload more material. Going legit archive
               | (as their name implies) instead of hiding behind the DMCA
               | and playing high-stakes poker with copyright law would
               | also make it possible for more entities to provide direct
               | support.
               | 
               | I also disagree that NARA and LoC is subject to whims of
               | the electorate. The Library of Congress is set up to
               | serve, well, Congress. Who funds it. Lotta barriers to
               | cross there, even in these weird times.
               | 
               | I'll take that risk over one guy with limited governance
               | who seems genuinely surprised that he keeps gets hacked
               | and sued. There's a chance the whole thing goes away
               | because he couldn't resist serving up free Frank Sinatra
               | records and got hit with a $621 million lawsuit after he
               | thrice refused to take the stuff down.
        
               | gumball-amp wrote:
               | I'm interested in why you are saying that the Internet
               | Archive is taking too much credit for the end of term web
               | archive. The website you link to demonstrates that it's
               | run by the Internet Archive, although various partners
               | have joined it since it began.
               | 
               | Is that not correct?
               | 
               | > And I don't know what Archive Team is but maybe they
               | could update their site to provide some information on
               | the people involved.
               | 
               | You don't need to reveal your identity, but looking
               | through your comments, it looks like you originally spun
               | up this account to criticize the Internet Archive. I'll
               | just note that accusing others of being "anonymous
               | radicals" falls a little flatter when you're anonymous
               | yourself.
               | 
               | (Relevant disclosure: I've worked with IA and Brewster
               | Kahle, and defended him here before.)
        
               | badlibrarian wrote:
               | > Is that not correct?
               | 
               | It's not run by the Archive. It's a collaboration. They
               | didn't even do all the crawling, and the Library of
               | Congress keeps a copy.
               | 
               | https://eotarchive.org/about/
               | 
               | As for Archive Team, their site declares "Archive Team is
               | a loose collective of rogue archivists, programmers,
               | writers and loudmouths."
               | 
               | Dedication is great. And radicalization in response to
               | copyright and preservation certainly deserves some
               | leeway. But a little professionalism wouldn't hurt and
               | the 2600-era roleplay isn't fooling anyone.
        
             | taurknaut wrote:
             | > They let vagrants sleep on their steps yet seem surprised
             | when they set the utility pole outsides on fire.
             | 
             | Tbf I have let many people sleep on my doorstep and none of
             | them tried to set my building on fire. One of them even
             | sang for me; he had a killer baritone. Overall it seems
             | like a fairly harmless thing.
        
               | badlibrarian wrote:
               | I wasn't speaking metaphorically. Fire set to pole. Site
               | went down.
        
         | homebrewer wrote:
         | How do I as a non-US citizen get access to information from
         | those "proper entities"? Is it even possible for US citizens?
         | This is often a surprise for some visitors of this fine
         | website, but there's a large world outside the US where
         | "federal law" does not apply.
        
           | badlibrarian wrote:
           | We fund the Library of Congress (largest library in the
           | world) and the National Archives (NARA) who make all of this
           | stuff public. Other goverments do similar things. It's all on
           | the web.
           | 
           | https://www.archives.gov/presidential-
           | records/research/archi...
           | 
           | There are other agencies and data sources to be monitored of
           | course but I'm not seeing a lot of nuance in those efforts
           | yet.
        
       | ch71r22 wrote:
       | For anyone else interested in running this, it only took a couple
       | seconds to launch their docker-compose.yml
       | 
       | https://github.com/ArchiveTeam/warrior-dockerfile/blob/maste...
        
         | NortySpock wrote:
         | I noticed from the docker overlay filesystem that the container
         | was spraying files all over the disk. (Ephemeral, destroyed on
         | container shutdown, sure, but I wanted to reduce write-wear on
         | my ssd...)
         | 
         | I tried setting it up with /tmp as a tmpfs (ramdisk) but it
         | then refused to start...
         | 
         | Anyone know any broad-spectrum docker incantations to force all
         | overlay writes to RAM, for a container?
        
           | lopkeny12ko wrote:
           | > the container was spraying files all over the disk
           | 
           | Right, that's basically the point...the Warrior downloads
           | files, compresses them, and uploads them for archival. This
           | necessarily requires staging the files somewhere between
           | download and upload.
           | 
           | > Anyone know any broad-spectrum docker incantations to force
           | all overlay writes to RAM, for a container?
           | 
           | Why would you want this? This sounds like a terrible footgun.
        
       ___________________________________________________________________
       (page generated 2025-02-05 23:00 UTC)