[HN Gopher] Show HN: Self-host Reddit - 2.38B posts, works offli...
       ___________________________________________________________________
        
       Show HN: Self-host Reddit - 2.38B posts, works offline, yours
       forever
        
       Reddit's API is effectively dead for archival. Third-party apps are
       gone. Reddit has threatened to cut off access to the Pushshift
       dataset multiple times. But 3.28TB of Reddit history exists as a
       torrent right now, and I built a tool to turn it into something you
       can browse on your own hardware.  The key point: This doesn't touch
       Reddit's servers. Ever. Download the Pushshift dataset, run my tool
       locally, get a fully browsable archive. Works on an air-gapped
       machine. Works on a Raspberry Pi serving your LAN. Works on a USB
       drive you hand to someone.  What it does: Takes compressed data
       dumps from Reddit (.zst), Voat (SQL), and Ruqqus (.7z) and
       generates static HTML. No JavaScript, no external requests, no
       tracking. Open index.html and browse. Want search? Run the optional
       Docker stack with PostgreSQL - still entirely on your machine.  API
       & AI Integration: Full REST API with 30+ endpoints - posts,
       comments, users, subreddits, full-text search, aggregations. Also
       ships with an MCP server (29 tools) so you can query your archive
       directly from AI tools.  Self-hosting options: - USB drive / local
       folder (just open the HTML files) - Home server on your LAN - Tor
       hidden service (2 commands, no port forwarding needed) - VPS with
       HTTPS - GitHub Pages for small archives  Why this matters: Once you
       have the data, you own it. No API keys, no rate limits, no ToS
       changes can take it away.  Scale: Tens of millions of posts per
       instance. PostgreSQL backend keeps memory constant regardless of
       dataset size. For the full 2.38B post dataset, run multiple
       instances by topic.  How I built it: Python, PostgreSQL, Jinja2
       templates, Docker. Used Claude Code throughout as an experiment in
       AI-assisted development. Learned that the workflow is "trust but
       verify" - it accelerates the boring parts but you still own the
       architecture.  Live demo: https://online-archives.github.io/redd-
       archiver-example/  GitHub: https://github.com/19-84/redd-archiver
       (Public Domain)  Pushshift torrent:
       https://academictorrents.com/details/1614740ac8c94505e4ecb9d...
        
       Author : 19-84
       Score  : 175 points
       Date   : 2026-01-13 15:35 UTC (7 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | NickNaraghi wrote:
       | Data is available via torrent in this section:
       | https://github.com/19-84/redd-archiver?tab=readme-ov-file#-g...
        
         | 19-84 wrote:
         | I have also published sub statistics and profiling for each
         | platform. these can be used to help identify which subs to
         | prioritize for archiving.
         | 
         | reddit: https://github.com/19-84/redd-
         | archiver/blob/main/tools/subre...
         | 
         | voat: https://github.com/19-84/redd-
         | archiver/blob/main/tools/subve...
         | 
         | ruqqus: https://github.com/19-84/redd-
         | archiver/blob/main/tools/guild...
        
       | elSidCampeador wrote:
       | I wonder if this can be hooked up with the now-dead Apollo app in
       | some way, to get back a slice of time that is forever lost now?
        
         | 19-84 wrote:
         | the API should allow for a lot of different integrations
        
       | Aurornis wrote:
       | Cool way to self-host archives.
       | 
       | What I'd really like is a plugin that automatically pulls from
       | archives somewhere and replaces deleted comments and those bot-
       | overwritten comments with the original context.
       | 
       | Reddit is becoming maddening to use because half the old links I
       | click have comments overwritten with garbage out of protest for
       | something. Ironically the original content is available in these
       | archives (which are used for AI training) but now missing for
       | actual users like me just trying to figure out how someone fixed
       | their printer driver 2 years ago.
        
         | anonymous908213 wrote:
         | That would only really be ironic if the reason for people
         | overwriting their comments was out of protest for LLM training,
         | but the main reason that resulted in by far the biggest wave of
         | deletions was Reddit locking down their API. If the result of
         | their protest is that the site is less useful for you, the
         | user, then in fact it served its purpose, as the entire point
         | was an attempt to boycott Reddit, ie. get people to stop using
         | it by removing the user contributions that give the site its
         | only value in the first place.
        
           | Aurornis wrote:
           | > If the result of their protest is that the site is less
           | useful for you, the user, then in fact it served its purpose,
           | as the entire point was an attempt to boycott Reddit, ie. get
           | people to stop using it by removing the user contributions
           | that give the site its only value in the first place.
           | 
           | In practice I just give them more page views because I have
           | to view more threads before I find the answer.
           | 
           | Reddit's DAU numbers have only gone up since the protest.
        
             | anonymous908213 wrote:
             | I did phrase it as "an attempt". In the end the protest
             | probably wasn't as effective as protestors might have
             | hoped, and it didn't get Reddit to change course on their
             | enshittification decisions. I do think it was good that
             | there was an attempt at pushback, at least, when most
             | software users just accept enshittification as normal and
             | continue tolerating whatever abuse their masters throw at
             | them.
        
       | kylehotchkiss wrote:
       | _Hacker News collectively grabs the dataset to train their models
       | on how to become effective reddit trolls_
        
         | 19-84 wrote:
         | the API and MCP server is very powerful ;)
        
         | layer8 wrote:
         | Don't we have enough of those already? ;)
        
       | Jordan-117 wrote:
       | >Voat
       | 
       | Gross. Why would anyone want to have an archive of Reddit For
       | Neonazis?
        
         | 19-84 wrote:
         | thank you for your comment, I will support any platform that
         | has complete dataset available. I will take submissions for any
         | complete datasets through github issues.
         | https://github.com/19-84/redd-archiver/blob/main/.github/ISS...
        
         | diggyhole wrote:
         | Wat?
        
           | Jordan-117 wrote:
           | It sold itself as a healthier alternative to Reddit, but by
           | the end of its run virtually every post sitewide was some
           | flavor of virulently racist, misogynistic, anti-semitic,
           | fringe conspiratorial, etc.
        
         | devilsdata wrote:
         | Might be good for researchers to be able to perform studies on.
        
         | metaPushkin wrote:
         | It seems you have no understanding of the term neo-fascism, and
         | yes, it's not what your propaganda talks about.
        
           | nozzlegear wrote:
           | Can you explain for the class? Don't just say that and leave
           | us wondering.
        
         | apstls wrote:
         | There are certainly things to be learned from analysis of the
         | dataset. Keep your friends close but your enemies as JSON, or
         | something...
        
       | dvngnt_ wrote:
       | I want to do the same thing for tiktok. I have 5k videos starting
       | from the pandemic downloaded. want to find a way to use AI to tag
       | and categorize the videos to scroll locally.
        
       | syngrog66 wrote:
       | Did you pay all the people who created its content?
        
         | devilsdata wrote:
         | I have no problem with this being downloaded for personal use,
         | in fact that's a good thing. But of course we both know it'll
         | be used to train AI.
        
         | nullandvoid wrote:
         | Did anyone ever comment on reddit with an expectation of pay?
         | 
         | It's an open forum - similar to here, whatever I post I it's in
         | the public forum and therefore I expect it to be used / remixed
         | however anyone wants.
        
           | nozzlegear wrote:
           | > Did anyone ever comment on reddit with an expectation of
           | pay?
           | 
           | Maybe Gallowboob
        
             | Sohcahtoa82 wrote:
             | That's a name I haven't seen in a LONG time.
        
         | antisthenes wrote:
         | Reddit didn't pay me for posting either. Not that I posted in
         | the last decade.
        
       | alcroito wrote:
       | I tried spinning up the local approach with docker compose, but
       | it fails.
       | 
       | There's no `.env.example` file to copy from. And even if the env
       | vars are set manually, there are issues with the mentioned
       | volumes not existing locally.
       | 
       | Seems like this needs more polish.
        
         | 19-84 wrote:
         | thank you for your comment, some example dot files were not
         | copied in my original repo, they have now been added.
         | 
         | https://github.com/19-84/redd-archiver/commit/0bb103952195ae...
         | 
         | the docs have been updated with mkdir steps
         | 
         | https://github.com/19-84/redd-archiver/commit/c3754ea3a0238f...
        
           | alcroito wrote:
           | Cheers. I checked the updated steps.
           | 
           | This is still missing creating the `output/.postgres-data`
           | dir, without which docker compose refuses to start.
           | 
           | After creating that manually, going to http://localhost/
           | shows a 403 Forbidden page, which makes you believe that
           | something might have gone wrong.
           | 
           | This is before running `reddarchiver-builder python
           | reddarc.py` to generate the necessary DB from the input data.
        
             | 19-84 wrote:
             | I've updated the workflow and added a placeholder page that
             | will serve before archives are created. thanks again!
             | https://github.com/19-84/redd-
             | archiver/commit/0dfd505ca81cb2...
        
       | diggings wrote:
       | This is a neat project, nice work.
       | 
       | You've probably come across this already but there are
       | alternative archives to PushShift that may have differing sets of
       | posts and comments (perhaps depending on removal request
       | coverage?)
       | 
       | One is Arctic Shift:
       | https://github.com/ArthurHeitmann/arctic_shift/releases
       | 
       | Another is PullPush: https://pullpush.io/
        
       | bkovacev wrote:
       | Is there any way to check if a subreddit that was made private
       | (2-3 years ago) is in the data dump?
        
         | 19-84 wrote:
         | I included a metadata dump of every subreddit found in the
         | torrent. it includes a status field which will show of a
         | subreddit is private along with a much more details
         | 
         | data catalog readme: https://github.com/19-84/redd-
         | archiver/blob/main/tools/READM...
         | 
         | reddit data: https://github.com/19-84/redd-
         | archiver/blob/main/tools/subre...
        
       | m463 wrote:
       | I wonder if you could use this to "Seed" a new distributed social
       | media thing and just take over from there.
       | 
       | sort of like forking a project.
        
         | 19-84 wrote:
         | ive created tooling for an instance registry and team based
         | leaderboard. the API has function to support this as well, so
         | that we can collectively host archives in a decentralized and
         | distributed manner
         | 
         | registry readme: https://github.com/19-84/redd-
         | archiver/blob/main/docs/REGIST...
         | 
         | register instances: https://github.com/19-84/redd-
         | archiver/blob/main/.github/ISS...
        
       ___________________________________________________________________
       (page generated 2026-01-13 23:01 UTC)