[HN Gopher] NonRAID - fork of unRAID array kernel module
       ___________________________________________________________________
        
       NonRAID - fork of unRAID array kernel module
        
       Author : qvr
       Score  : 34 points
       Date   : 2025-07-22 20:16 UTC (2 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | miffe wrote:
       | What makes this different from regular md? I'm not familiar with
       | unRAID.
        
         | eddythompson80 wrote:
         | unRAID is geared towards homelab style deployments. Its main
         | advantages over typical RAID is it's flexibility
         | (https://www.snapraid.it/compare):
         | 
         | - It lets you throw JBODs (of ANY size) and you can create a
         | "RAID" over them.
         | 
         | - The biggest drive must be a parity drive(s).
         | 
         | - N parity = surviving N drive failures.
         | 
         | - You can expand your storage pool 1 drive at a time. You need
         | to recalculate parity for the full array.
         | 
         | The actual data is spread across drives. If a drive fails, you
         | rebuild it from the parity. This is another implementation
         | (using MergerFS + SnapRAID)
         | https://perfectmediaserver.com/02-tech-stack/snapraid/
         | 
         | It's a very simple model to think of compared to something like
         | ZFS. You can add/remove capacity _AND_ protection as you go.
         | 
         | Its perf is significantly less than ZFS of course.
        
           | somat wrote:
           | Those are the features I liked about ceph, with the benefit
           | that it uses the network to scale.
        
           | phoronixrly wrote:
           | I have an issue with this though... Won't you get a write on
           | the parity drive for each write on any other drive? Doesn't
           | seem well balanced... to be frank, looks like a good way to
           | shoot yourself in the foot. Have a parity drive fail, then
           | have another drive fail during the rebuild (a taxing process)
           | and congrats -- your data is now eaten, but at least you
           | saved a few hundred dollars by not buying drives of equal
           | size...
        
             | hammyhavoc wrote:
             | No, because you have a cache pool and calculate the parity
             | changes on a schedule, or when specific conditions are met,
             | e.g., remaining available storage on the cache pool.
             | 
             | The cache pool is recommended to be mirrored for this
             | reason (not many people see why I find this to be amusing).
        
               | phoronixrly wrote:
               | And let me guess, the cache pool is suggested to be on an
               | SSD?
               | 
               | > Increased perceived write speed: You will want a drive
               | that is as fast as possible. For the fastest possible
               | speed, you'll want an SSD
               | 
               | Great, now I have an SSD that is treated as a
               | consummative and will die and need to be replaced. Oh and
               | btw you are going to need two of them if you don't want
               | to accidentally your data.
               | 
               | The alternative? Have the cache on a pair of spinning
               | rust drives which will _again_ be overloaded and are
               | expected to fail earlier and need to be replaced while
               | also having the benefit of being slow... But at least you
               | won 't have to go through a full rebuild after a cache
               | drive failure.
               | 
               | Man, I am not sold on the cost savings of this approach
               | at all... Let alone the complexity and moving parts that
               | can fail...
        
               | dawnerd wrote:
               | But you'd have that problem on any system really.
        
               | hammyhavoc wrote:
               | Hit the comment depth limit again, but yes, SSDs!
               | 
               | Yes, Unraid can crash-and-burn in quite a lot of
               | different ways. Ask me how I know! Why I'm all-in on ZFS
               | now.
        
             | eddythompson80 wrote:
             | > Have a parity drive fail, then have another drive fail
             | during the rebuild (a taxing process) and congrats -- your
             | data is now eaten
             | 
             | That's just your drive failure tolerance. It's the same
             | risk/capacity trade as RAIDZ1, but with less performance
             | and more flexibility on expanding. Which is exactly what I
             | said.
             | 
             | If 1 drive failure isn't acceptable for you, you wouldn't
             | use RAIDZ1 and wouldn't use 1 parity drive.
             | 
             | You can use 2 parity drives for RAIDZ2-like protection.
             | 
             | You can use 3 drives for RAIDZ3-like protection.
             | 
             | You can use 4 drives, 10 drives. Add and remove as many
             | parity/capacity as you want. Can't do that with RAID/RAIDZ
             | easily.
             | 
             | You manage your own risk/reward ratio
        
               | phoronixrly wrote:
               | My issue is that due to uneven load balancing, the parity
               | drive _is_ going to fail more often than in a
               | configuration with distributed parity, thus you are going
               | to need to recalculate parity for the array more often,
               | which is a risky and taxing operation for all drives in
               | the array.
               | 
               | As hammyhavoc below noted, you can work around this by
               | having cache, and 'by deferring the inevitable parity
               | calculation until a later time (3:40 am server time, by
               | default)'.
               | 
               | Which seems like a hell of a bodge -- both risky, and
               | expensive -- now the unevenly balanced drive is the cache
               | one, it is also not parity protected. So you need
               | mirroring for it in case you don't want to lose your
               | data, and the cache drives are still expected to fail
               | before a drive in an evenly load-balanced array, so
               | you're going to have to buy new ones?
               | 
               | Oh and btw you are still at risk of bit flips and garbage
               | data due to cache not being checksum-protected.
        
               | eddythompson80 wrote:
               | You need to run frequent scrubs on the whole zfs array as
               | well.
               | 
               | On unraid/snapraid you need to spin 2 drives up (one of
               | then is always the parity)
               | 
               | On zfs, you are always spinnin up multiple drives too.
               | Sure the "parity" isn't always the same drives or at
               | least it's up to zfs to figure that out.
               | 
               | Nonetheless, this is all not really likely to have a
               | significant impact. Spinning disks failure rates don't
               | exactly correlate with their utilization[1][2]. Between
               | SSD cache, ZFS scrubs, general usage, I don't think the
               | parity drives are necessarily more at risk. This is
               | anectodal, but when I ran an unRAID box for few years
               | myself, I only had 1 failure and it was a non-parity
               | drive.
               | 
               | [1] Google study from 2007 for harddrive failure rates: h
               | ttps://static.googleusercontent.com/media/research.google
               | .c...
               | 
               | [2] "Utilization" in the paper is defined as:
               | The literature generally refers to utilization metrics by
               | employing the term duty cycle which unfortunately has no
               | consistent and precise definition, but can be roughly
               | characterized as the fraction of time a drive is active
               | out of the total powered-on time. What is widely reported
               | in the literature is that higher duty cycles affect disk
               | drives negatively
        
               | Dylan16807 wrote:
               | > due to uneven load balancing, the parity drive is going
               | to fail more often than in a configuration with
               | distributed parity
               | 
               | Good, it can be the canary.
               | 
               | > thus you are going to need to recalculate parity for
               | the array more often, which is a risky and taxing
               | operation for all drives in the array
               | 
               | This is not worth worrying about.
               | 
               | First off, if the risk is linear then your increased
               | parity failure is offset by decreased other-drive failure
               | and I don't think you'll have more rebuilds.
               | 
               | And even if you do get more rebuilds, it's significantly
               | less than one per year, and one extra full-drive read per
               | year is a negligible amount of load. If you're worried
               | about it all hitting at once then A) you should be
               | scrubbing more often and B) throttle the rebuild.
        
             | nodja wrote:
             | The wear on the parity drive is the same regardless of raid
             | technology you choose, unraid just lets you have mismatched
             | data drives. In fact you could argue that unraid is
             | healthier for the drives since a write doesn't trigger a
             | write on all drives, just 2. The situation you described is
             | true for any raid system.
        
             | dawnerd wrote:
             | Depends. If you use a cache like they recommend you'd only
             | get parity writes when it runs its mover command.
             | Definitely adds a lot of wear but so far i haven't had any
             | parity issues with two parity drives protecting 28 drives.
        
             | wongarsu wrote:
             | You want your drives to fail at different times! Which
             | means you want your load to be unbalanced, from a
             | reliability standpoint. If you put the same load on every
             | drive (like in a traditional RAID5/6) then the drives are
             | likely to fail at around the same time. Especially if you
             | don't go out of your way to get drives from different
             | manufacturing batches. But if you misbalance the amount of
             | work the drives get they accumulate wear and tear at
             | different rates and spend different amounts of time in
             | idle, leading them to fail at wildly different times,
             | giving you ample time to rebuild the raid.
             | 
             | I'd still recommend anyone to have two parity drives (which
             | unraid does support)
        
           | hammyhavoc wrote:
           | > If a drive fails, you rebuild it from the parity.
           | 
           | But if the file system is corrupt then you're hosed and end
           | up with a `lost+found`. It sounds great until it fails, and
           | then you realize why ZFS with replication makes sense. Unraid
           | doesn't do automatic repairs from replicated ZFS datasets yet
           | either even if you use individual ZFS disks within your
           | Unraid array.
        
             | Whatarethese wrote:
             | Hence why this is for home users that store media. ZFS is
             | for the enterprise where you have people to babysit storage
             | solutions.
        
         | wongarsu wrote:
         | md takes multiple partitions to make a virtual device you can
         | put a file system on, with striping and traditional RAID levels
         | 
         | unRaid takes multiple partitions, dedicates one or two of them
         | to parity, and hands the other partitions through. You can
         | handle those normally, putting different file systems on
         | different partitions in the same array and treating them as
         | completely separate file systems that happen to be protected by
         | the same parity drives
         | 
         | This enables you to easily mix drives of different sizes (as
         | long as the parity drives are at least as large as the largest
         | data partition), add, remove or upgrade drives with relative
         | ease, and means that every read operation only goes to one
         | drive, and writes to that drive plus the parity drives.
         | Depending on how you organize your files you can have drives
         | that are basically never on, while in an md array every drive
         | is used for basically every read or write.
         | 
         | The disadvantages are that you lose out on the performance
         | advantages of a RAID, and that the raid only really protects
         | against losing entire disks. You can't easily repair single
         | blocks the way a zfs RAID could. Also you have a number of file
         | systems you have to balance (which unRaid helps you with, but I
         | don't know how much of that is in this module)
        
           | phoronixrly wrote:
           | Not sure what you mean by 'easily repair single blocks the
           | way a zfs RAID could', but often the physical devices handle
           | bad blocks, and md has one safety layer on top of this - bad
           | blocks tracking. No relocation in md though, AFAIK.
        
             | hammyhavoc wrote:
             | If you have a redundant dataset (#1 reason to use ZFS
             | replication) then you can repair a ZFS dataset.
        
               | phoronixrly wrote:
               | I'm sorry, I still don't quite follow... If you have a
               | RAID5, you can repair a drive failure... Weren't we
               | talking about handling 'blocks'? Is it bad blocks or bad
               | block devices (a.k.a. dead drives)?
        
               | hammyhavoc wrote:
               | Hit the comment depth limit (so annoying), but the
               | comment about repairing blocks means that you can repair
               | bitrot/corruption/malicious changes/whatever down to the
               | block level of a ZFS dataset if you have a redundant
               | replicated dataset.
               | 
               | The magic of ZFS repairs isn't in RAID itself, IMO, it's
               | in being able to take your cold replicated dataset, e.g.,
               | from LTO, an external disk, remote server etc, and repair
               | any issues without needing to resilver, stress the whole
               | array, interrupt access, or hurt performance.
               | 
               | RAID can correct issues, yes, but ZFS as a filesystem can
               | repair itself from redundant datasets. Likewise, you can
               | mount the snapshots like Apple Time Machine and get back
               | specific versions of individual files.
               | 
               | I wish HN didn't limit comment depth as these are great
               | questions and this is heavily under-discussed, but it's
               | arguably the best reason to run ZFS, IMO.
               | 
               | Another way of putting this--you don't need a RAID array,
               | you can do individual ZFS disks and still replicate and
               | repair them. There's no limits to how many replicas or
               | mediums you use either. It's quite amazing for self-
               | healing problems with your datasets.
        
             | wongarsu wrote:
             | What I mean is that unraid, zfs and md all allow you to run
             | a scrub over your raid to check for bit rot. That might
             | happen for all kinds of issues, including cosmic rays just
             | flipping bits on the drive platter. The issue is that
             | unraid and md can't do much if they detect a block/stripe
             | where the parity doesn't match the data (because it doesn't
             | know which of the drives suffered a bit flip). Zfs on the
             | other hand can repair the data in that scenario because it
             | keeps checksums.
             | 
             | Now a fairly common scenario is to use unRaid with zfs as
             | the file system for each partition, having Y independent
             | zfs file systems. In that case in theory the information to
             | repair blocks exists: a zfs scrub will tell you which
             | blocks are bad, and you could repair those from parity. And
             | a unraid parity check will do the same for the parity
             | drives. But there is no system to repair single blocks. You
             | either have to dig in and do it yourself or just resilver
             | the whole disk
        
       | nullc wrote:
       | Are any filesystems offering file level FEC yet?
       | 
       | If a file has a hundred thousand blocks you could tack on a
       | thousands blocks of error correction for the cost of making it
       | just 1% larger. If the file is a seldom/never written archive
       | it's essentially free beyond the space it takes up.
       | 
       | The kind of massive data archives that you want to minimize
       | storage costs of tend to be read-mostly affairs.
       | 
       | It won't save you from a disk failure but I see bad blocks much
       | more often than whole disk failures these days... and raid/5/6
       | have rather high costs while being still quite vulnerable to the
       | possibility of an aligned fault on multiple disks.
       | 
       | Of course you could use par or similar tools, but that lacks nice
       | FS transparent integration and particularly doesn't benefit from
       | checksums already implemented in (some) FS (as you need half the
       | error correction data to recover from known-position errors, and-
       | or can use erasure only codes).
        
         | leptons wrote:
         | RAID and any other fault-tolerance scheme can not be the only
         | way you protect your data. I have two RAID 10 arrays, one is
         | for active data, one is for backup, and the backup system has
         | an LTO tape drive, where I also use PAR parity files on the
         | tape backups. Important stuff is backed-up to multiple tape
         | sets. Both systems are in different buildings, with tapes
         | stored in a third.
         | 
         | My point is, it doesn't much matter what your FS does, so long
         | as you have 3 or more of them.
        
         | oakwhiz wrote:
         | Rateless erasure FEC can go even further.
        
       | hebocon wrote:
       | This should be a "Show HN".
        
       ___________________________________________________________________
       (page generated 2025-07-22 23:00 UTC)