[HN Gopher] Leaving the Basement
___________________________________________________________________
Leaving the Basement
Author : timf
Score : 66 points
Date : 2022-12-04 13:11 UTC (2 days ago)
(HTM) web link (community.hachyderm.io)
(TXT) w3m dump (community.hachyderm.io)
| [deleted]
| bluedino wrote:
| Why weren't the disks just replaced?
| robga wrote:
| Answer from the author
| https://news.ycombinator.com/item?id=33855250#33856184
| uniqueuid wrote:
| Wild speculation, but perhaps ZFS + postgres + potentially
| write-sensitive SSDs resulted in write amplification that would
| just occur again.
| rglullis wrote:
| As someone running a commercial provider for Mastodon (and
| Matrix, and XMPP...), I am somewhat envious of these posts. "Wow,
| 30000 users! If I had that many users on my service paying the
| $0.50/month I am charging, it would be enough to pay myself a
| full salary!".
|
| But then I realize that they are only getting these many people
| because they are not driven by commercial interests: even with
| donations, I can bet they are not collecting enough to keep
| things afloat and they only keep going because they don't mind
| spending all this time, money and resources of their own on this
| project. They can treat it as a (relatively expensive) hobby, and
| they can keep it running as long as it satisfies them.
|
| The problem is that I think that this is harmful in the long run.
| Yes, people now are finally seeing the issue with ad-funded
| social media. But if we want to have a healthy alternative, we
| need to understand TANSTAAFL, we need to accept that we need to
| give real money to the people working on this and to have the
| servers available 24/7 to store and distribute the hot takes and
| stupid memes that we so bizarrely crave every day.
|
| I worry that if we don't change the mindset quickly, the whole
| Twitter drama would be a wasted opportunity and Mastodon (and the
| Fediverse in general) will go back to the status quo, where
| surveillance capitalism is the norm and truly open systems are
| just a geeky curiosity.
|
| I wish I could fund a tech-equivalent of the "buy local and
| organic" campaign. I wish I had more people thinking "ok, I will
| pay $5/month to this guy and I will bring 10 people to this
| instance" because it is the _ethical thing to do_.
| buovjaga wrote:
| For financials, see
| https://community.hachyderm.io/blog/2022/12/04/growth-and-su...
| rglullis wrote:
| You lost me at "call for moderators and volunteers". Unless
| these people are actually paid for this taxing and stressful
| work, I don't believe it can be called "sustainable".
| abraae wrote:
| Wikipedia would be unsustainable using that test. Yet it's
| been around for two decades and shows no signs of going
| anywhere - if that's not sustainable then no web business
| is.
| rglullis wrote:
| When I make a edit on wikipedia, I am doing it for my own
| benefit and others and I am not dealing with _stressful_
| work. That is totally different of having to work as a
| moderator on an instance where some of the members might
| be reporting cases of racism, harassment and sometimes
| just petty uncivil people.
|
| Also, Wikipedia has donation drives, non-profit status
| and a _highly controversial history of how it spends the
| funds they collect_. Lots of high quality contributors
| already left because they were not recognized in any way
| and ended up feeling exploited.
| abraae wrote:
| I agree with all that. However you said "You lost me at
| call for moderators and volunteers".
|
| I do not agree that just calling for moderators and
| volunteers means a business is unsustainable.
|
| Perhaps it does in this case - I don't know enough about
| Hachyderm to know for sure. But it's possible that the
| roles of moderators and volunteers at Hachyderm might not
| be (or could be made not to be) so terrible, in which
| case relying on free labour is a proven and sustainable
| business model (for some businesses anyway).
|
| Also worth noting that a "highly controversial history of
| how it spends the funds they collect" at Wikipedia is not
| directly correlated to Wikipedia's sustainability at all
| - the history shows otherwise.
| rglullis wrote:
| > I do not agree that just calling for moderators and
| volunteers means a business is unsustainable.
|
| The problem is that example you gave (Wikipedia) is _not_
| a business.
|
| > I don't know enough about Hachyderm to know for sure.
|
| My point is not about Hachyderm in the particular. Like I
| said in the first comment, I think they have what it
| takes to continue operating and serving their community
| for the longer term.
|
| My issue is with the overall ecosystem and the
| expectations of the people coming in from "traditional
| social media sites". If we want to provide an ethical
| alternative to Twitter/Facebook/Instagram/WhatsApp, we
| need to find a way to serve hundreds of millions of
| users. How are these people going to be spread around the
| instances? For context, we would need ~15000 Hachyderms
| to replace Twitter and ~60000 pixelfeds to replace
| Instagram. Are they all going to be dependent on
| volunteers? Are all these instances be operated by highly
| paid professionals from SV who can sink a few hundred
| dollars every month? Or are they going to go to appeal to
| "your donation is very important to us" and expect that a
| few generous souls make up for the free-riders? Or are
| they all going to eventually cave in, start treating it
| as a business and start charging from their users
| something that can pay actual salaries for everyone
| involved?
| wmf wrote:
| I suspect Wikipedia has a very different reader:writer
| ratio than social media and thus Wikipedia needs
| proportionately less moderation. On the hardware side,
| Wikipedia is centralized and mostly static content so it
| should be cheaper to operate.
| cyberpunk wrote:
| > "We can then leverage Mastodon's S3 feature to write the "hot"
| data directly back to Digital Ocean using a reverse Nginx proxy."
|
| How does that work?
| josteink wrote:
| Digital Ocean offers S3 compatible storage.
|
| I guess they use Nginx to reroute traffic which by default is
| targeting aws.amazon.com?
| convolvatron wrote:
| "In other words, every ugly system is also a successful system.
| Every beautiful system, has never seen spontaneous adoption." -
| not only is this logically fallacious, its pretty offensive about
| the general notion of software quality
| nmyk wrote:
| Where's the fallacy? They're saying the only systems that are
| beautiful are the ones that haven't been forced by massive
| spontaneous adoption to scale faster than the developer(s) can
| come up with a beautiful design to meet the new requirements.
|
| Maybe you think if the original system were _really_ beautiful
| and of high quality, it would have scaled with the adoption on
| its own, with no need for ugly patches... but in that case the
| original system would have had the capacity to do a lot more
| than what was originally required. It would have been
| overengineered, in other words, and it would have been more
| beautiful if it had met its original requirements more cheaply.
|
| The notion that a sudden change in requirements that must be
| dealt with quickly results in an uglier system seems fairly
| straightforward to me, and certainly not offensive.
| watchdogtimer wrote:
| Kris described hachyderm's infrastructure operating in her
| basement on the Oxide and Friends podcast in mid-November. Kudos
| to her for being able to keep it going there so long!
| CharlesW wrote:
| As someone who'd hoped to start an industry-focused instance, I
| found Kris' Medium articles about Hachyderm's growth really
| interesting: https://medium.com/@kris-nova
|
| The latest post, _" Yelping: Action Through Criticism"_,
| includes links to additional Hachyderm-related content near the
| top, then talks about they handle an obnoxious online behavior
| they've experienced because of the recent popularity of
| Hachyderm.
| musk_micropenis wrote:
| I would like to understand why Mastodon requires such a huge
| amount of hardware for mediocre traffic volumes. Not just the
| lazy "it's Rails" answer - I know Rails is a resource hog, but
| that doesn't go far enough to explain the extreme requirements
| here.
|
| As a point of reference, look at what Stack Overflow is run on.
| As a caveat, SO is probably more read-heavy than Mastodon, but it
| also serves several orders of magnitude more volume (on a normal
| day in 2016 they would serve 209,420,973 HTTP requests[0]). They
| did this on 4 DB servers and 11 web servers. And in fact, it can
| (and has) worked serving this volume of traffic on only a single
| server.
|
| With this setup SO was not even close to maxing out their
| hardware (servers were under 10% load, approximately). SO also
| listed their server hardware[1] in 2016. I don't know enough
| about server hardware to assess the difference, but to my eye
| they look similar on the web tier with similar amounts of memory,
| similar disk, etc.
|
| I'm not saying Hachyderm is doing anything wrong, but it makes me
| wonder if there's a fundamental problem with the design of
| Mastodon. And to be clear I understand that this particular issue
| was caused by a disk failure, but that they even had this
| hardware in place running Hachyderm is surprising to me.
|
| [0] https://nickcraver.com/blog/2016/02/17/stack-overflow-the-
| ar...
|
| [1] https://nickcraver.com/blog/2016/03/29/stack-overflow-the-
| ha...
| rglullis wrote:
| Every post and every reply generates a request to all servers
| you are federating to:
| https://aeracode.org/2022/12/05/understanding-a-protocol/
| spankalee wrote:
| This ultimately polynomial direct-connection approach is
| clearly fatally unscalable.
|
| The fediverse needs to figure out a hubs-and-spokes or
| supernodes pattern so that service providers can scale up
| syncing, indexing etc.
|
| ie, my personal instance should be able to offload most of
| the message passing to an supernode intermediary that lots of
| other instances use for federation so that my instance only
| needs one connection, and the supernodes only need to connect
| to each other and their local network.
| rglullis wrote:
| They do have it for some things already. There is this
| concept of "relays", which you can use as a feed of the
| data from the larger instances. But AFAIK it's used only as
| a content source and it's not something that you can set up
| now as a way to help with scalability.
|
| I am also closely following https://github.com/nostr-
| protocol/nostr to see how they go along, because I am
| growing weary of the "tech elite" that is moving to
| Mastodon and is pushing for "moderation by committee". I've
| gotten myself with discussions already with people who
| actually want server operators that want only to open
| federation for those that abide by some "Covenant". This
| seems rooted in good intentions, but it reeks of something
| that might lead to a corporate copout of a network which is
| supposed to be open.
| musk_micropenis wrote:
| I guess this goes some way to answering my thought,
|
| > but it makes me wonder if there's a fundamental problem
| with the design of Mastodon.
|
| I also note the article says,
|
| > During the month of November we averaged 36.86 Mbps in
| traffic with samples taken every hour
|
| That seems like a large amount of bandwidth to service 30,000
| users (who knows what fraction of them are actually active at
| any given moment). But I guess there's going to be a lot of
| video and image content. I have tried searching all of their
| linked blog posts about scaling but can't find any number
| that might map to requests per second without making huge
| assumptions.
| bscphil wrote:
| I vouched for this comment because it's a good question,
| although I'm worried that your account is not long for this
| world with an inflammatory username like that.
|
| The problem probably starts with the inefficiency of RoR, as
| you've guessed. Mastodon is a very dynamic site which limits
| the amount of caching that can be done, and there are hot code
| paths like filtering streams using a user's block lists and
| word filters that are not particularly optimized - all this
| happens in Ruby.
|
| But there are other inefficiencies, compared to SO:
|
| 1. Mastodon is a media heavy site, with a lot of uploading by
| users. Mastodon has to convert user-uploaded media to
| standardized representations (e.g. JPEG and h.264), which takes
| a lot of CPU time.
|
| 2. Mastodon has a "firehose" feed which is available in the UI
| and actually used by many users. Filters apply to the firehose
| feed as well. Obviously this requires quite a lot of bandwidth
| and processing.
|
| 3. Federation is a weakness when it comes to traffic. If user X
| has an account on server A, and at least one user on 1000 other
| instances follow user X, server A has to immediately send any
| posts to _all 1000_ other instances, regardless of whether
| anyone on the other end will ever deliberately view them. (Of
| course, some users may view them in their instance 's firehose
| feed.) The instance then has to duplicate this traffic when
| sending it to the actual subscribed users. By this standard
| both large non-federated "servers" (like Twitter) and widely
| federated pull-only servers (think RSS) are more efficient than
| ActivityPub (the open standard Mastodon uses).
|
| 4. Federation is a weakness when it comes to trust. Instances
| do not (and must not) fully trust each other, except for things
| like "@x@thisinstance said 'P'". So for example, the little
| Open Graph based preview cards you're used to seeing on Twitter
| and elsewhere have to be generated for links _per instance_.
| The first time a Mastodon server sees a link, it must fetch
| that link and generate a preview card itself. Because new posts
| by popular accounts are syndicated immediately, this is a
| burden on websites as well.
| https://www.jwz.org/blog/2022/11/mastodon-stampede/ (note: copy
| link or disable sending referrers from HN for this site)
|
| 5. Scaling is not really a solved problem yet for Mastodon,
| because in practice it hasn't had to be. It's easy to pass the
| buck to instance operators, who end up needing a $20/month VPS
| to run a small instance rather than $5/month. Even the very
| biggest servers are scarcely larger than 1M users. At that kind
| of scale you can patch over performance problems by just
| throwing more hardware at the problem - and e.g.
| mastodon.social has the funds from Mastodon (the org) to do
| that. Note that Hachyderm, AFAIK, is an obvious example of
| this; it was started by a tech worker in Seattle with much
| better access to expensive hardware than most casual instance
| operators can dream of. It's not surprising that they can pull
| the funds together to scale up before they start seeing
| performance issues.
| kibwen wrote:
| In practice #3 is the only one that matters. For reducing
| dynamism/increasing caching potential, it would be fairly
| easy to run a fork of the site with the more dynamic features
| excised (donate $1 a month to get access to dynamic features
| like filtering). For media transcoding, that's a textbook
| case of a CPU-bound operation that you could offload to an
| isolated Rust component for a CPU savings of 99% compared to
| Ruby (not an exaggeration). But the exponential nature of the
| network scaling will still kill you despite all this, and
| needs to be addressed at the protocol level ASAP.
| Cyberdog wrote:
| Interesting username you have there.
|
| I don't see why you don't accept "it's Rails." There are other
| issues as sibling comments have pointed out, but by starting
| with an ecosystem known to have performance limitations, this
| sort of outcome is inevitable, is it not? I'm sure the Mastodon
| team were never expecting the degree of usage which has been
| thrust upon its larger instances, but now that it has happened
| and the limitations have become apparent, I'd encourage people
| who are interested in setting up fediverse/"Mastodon network"
| instances to consider the alternatives to Mastodon, however
| paltry they currently are.
|
| I know that the Pleroma front end and its forks are written in
| something called Elixir, which I have no idea about but I can't
| imagine it could be much worse than Ruby. What I'd really like
| to see is something written in a language known to be actually
| fast, though - PHP or Lua.
| BeetleB wrote:
| I would like to point out that you asked this same question a
| few days ago and got several answers:
|
| https://news.ycombinator.com/item?id=33855686
|
| I would recommend people read that thread before responding
| with the same answers.
| maxsilver wrote:
| > I would like to understand why Mastodon requires such a huge
| amount of hardware for mediocre traffic volumes.
|
| There's some inherent overhead in a federated model (vs a
| single-source one), and the ActivityPub protocol Mastodon
| happens to use, wasn't necessarily designed to be the lightest
| possible thing in all use-cases.
|
| Also, there's just a lot more traffic. My instance said, after
| Twitter's major struggles, they saw something like 30x more
| traffic and 20x more daily registrations. For instances that,
| prior to the influx, were running by volunteers in spare time
| out of people's bedrooms or small cheap VPS's and such.
|
| These instances weren't necessarily ideally performance-tuned
| prior to the influx (and even if _yours_ was, the remote ones
| your users might need to hit to fetch content from may not have
| been)
| imtringued wrote:
| That basement hardware didn't last long. If you don't know how
| big your userbase is going to be it would be better to avoid
| committing money to specific hardware.
| nix0n wrote:
| > committing money to specific hardware
|
| Note that Dell R620 and R630 servers have been discontinued for
| a couple of years now, were probably bought used, and can
| probably be re-sold.
| lantry wrote:
| hacker news: you don't need the cloud! you can just run a
| couple machines in your basement!
|
| also hacker news: why would you try to run something in your
| basement? Just use the cloud!
| CleverLikeAnOx wrote:
| There are multiple people here and their opinions vary.
| ocdtrekkie wrote:
| Note that Nova didn't feel an actual hardware capacity level
| was hit here. However, the setup lacked the redundancy to
| handle hardware outages for something like drive replacements
| without a significant outage. And I believe one of the main
| considerations in moving to a cloud service was actually
| limited connectivity options, because only so much fiber
| capacity was even available.
|
| > Our limiting factor in Hachyderm had almost nothing to do
| with the amount of users accessing the system as much as it did
| the amount of data we were federating. Our system would have
| flapped if we had 100 users, or if we had 1,000,000 users. We
| were nowhere close to hitting limits of DB size, storage size,
| or network capacity. We just had bad disks.
| dang wrote:
| Recent and related:
|
| _Post mortem on Mastodon outage with 30k users_ -
| https://news.ycombinator.com/item?id=33855250 - Dec 2022 (101
| comments)
|
| (Offtopic meta note: Alert users will note that that thread was
| posted later than this one. This is because the second-chance
| process (https://news.ycombinator.com/item?id=26998308) has a
| race condition: the events "story makes front page" and
| "moderator puts story in second-chance pool" sometimes diverge
| and can happen in any order.)
___________________________________________________________________
(page generated 2022-12-06 23:01 UTC)