Posts by wollman@mastodon.social
(DIR) Post #B5SdYaMnddIqg8SiGG by wollman@mastodon.social
0 likes, 1 repeats
There's something that #ZFS #OpenZFS does at mount time (I'm told it's processing deferred frees) that sometimes takes an incredibly long time (hours), and I'm still struggling to understand: why must this be done at mount? (And why mount and not pool import?) Why can't it continue to be deferred like it was during operations?Still waiting for `zfs mount -a` after an unplanned power outage, 90 minutes after restore.
(DIR) Post #B5SdYajUHHS1oV4pcm by wollman@mastodon.social
0 likes, 0 repeats
The server eventually finished booting, after about three hours.
(DIR) Post #B5SjYykVaIkm1Z7edM by wollman@mastodon.social
0 likes, 0 repeats
@mwl There's never a situation in which that would be an option. The server has already rebooted at the point I'm being paged and the users want to know why their NFS isn't working.
(DIR) Post #B5SsZ1j3Pui6qLXXPc by wollman@mastodon.social
0 likes, 0 repeats
@mwl Maybe. These servers have 60+ drives in them; I'm not sure how big I would have to make the message buffer to be able to collect all of that. And I'd have to do it across all 22 servers, since I have no way of knowing a priori which one is going to have the problem, which can't happen until our June outage window anyway.
(DIR) Post #B5Suino8DcvCTthhp2 by wollman@mastodon.social
0 likes, 0 repeats
@mwl Oh god no. Our typical configuration is 28 mirrors, 4 spares. Some older machines have 7x8 RAID-Z2 but we've gotten away from that as the resilver times are too high (and some users need more than 7 virtual spindles). But most of our older machines have multipath in the mix as well, so there's 2 daN devices per physical disk. Biggest server is 102 drives: 2 enclosures with 51 drives each (51 mirrors and 6 spares total with each mirror split between the enclosures).
(DIR) Post #B5SxEVNgSo0IVVZtuS by wollman@mastodon.social
0 likes, 0 repeats
@mwl If the multipath were confusing things, that would show up during pool import and not mount.
(DIR) Post #B5VGkaAuShyw5TX31E by wollman@mastodon.social
0 likes, 0 repeats
@dexter @mwl I *know* what /etc/rc.d/zfs is doing: it's running `zfs mount -va`. What I don't know is what ZFS itself is doing, why it takes so long, or why it has to do it right then and there.
(DIR) Post #B5VGkaUlGtrT52otxg by wollman@mastodon.social
0 likes, 0 repeats
@dexter @mwl Five or fewer per filesystem.
(DIR) Post #B5VGkaf2eg5Javd6zA by wollman@mastodon.social
0 likes, 0 repeats
@dexter @mwl (Not seeing how that would cause the same server with the same filesystems to sometimes take a minute to mount 90 filesystems and other times take three hours, but OK, if someone can tell me a plausible story I can consider a different way to store that information.)
(DIR) Post #B5VGkb1NJdwuiC4wnQ by wollman@mastodon.social
0 likes, 0 repeats
@dexter @mwl Without some sort plausible mechanism it's hard to consider, especially for something that only shows up during recovery from an outage.
(DIR) Post #B5VGkbKWATGHfZ2EdM by wollman@mastodon.social
0 likes, 0 repeats
@mwl @dexter A scrub on one of our pools takes more than a day; it's hard to imagine any of Sun's customers BITD finding that acceptable either. We certainly do notice interrupted scrubs resuming immediately on boot, but not waiting for completion.
(DIR) Post #B7dxgWY8ixTCczDV5c by wollman@mastodon.social
0 likes, 0 repeats
@fribbledom Why would there be a need for "vim for people who hate vim" when the real thing ((n)vi) is right there?Signed,alias vi=nvi