[HN Gopher] PCIe 4.0 Card Hosts 21M.2 SSDs: Up To 168TB, 31 GB/s
___________________________________________________________________
PCIe 4.0 Card Hosts 21M.2 SSDs: Up To 168TB, 31 GB/s
Author : ohmyblock
Score : 84 points
Date : 2023-03-13 14:08 UTC (8 hours ago)
(HTM) web link (www.tomshardware.com)
(TXT) w3m dump (www.tomshardware.com)
| _joel wrote:
| I'll take two please
| [deleted]
| eqvinox wrote:
| (Pulling up from child comment)
|
| The chip this uses is likely a PM4x100 (x [?] {0, 1, 2}) from
| Microchip (formerly Microsemi (formerly PMC-Sierra)):
|
| https://www.microchip.com/en-us/product/PM40100
|
| ^ runs you $800 without bulk discounts
| [https://www2.mouser.com/ProductDetail/Microchip-Technology-A...]
| -- if you can get them, that is.
|
| https://www.microchip.com/en-us/product/PM41100
|
| https://www.microchip.com/en-us/product/PM42100
|
| ^ these latter two I don't see publicly listed prices for
| anywhere.
|
| The PCIe 5.0 equivalent is in "Samples available", i.e. not full
| production yet, which is likely why the card only does PCIe 4.0:
|
| https://www.microchip.com/en-us/product/PM50100
| rektide wrote:
| CXL will also make attaching drives & ram a much easier
| experience, much more regular.
| qwertox wrote:
| I wonder how the SSDs are exposed to the OS.
|
| While dealing with the Samsung Pro Firmware issue, I read that
| SSDs mounted on a hardware RAID controller need to be removed
| from the RAID in order to have their Firmware update applied,
| since Samsung's tool won't see the SSDs if they are placed on the
| controller.
| rasz wrote:
| If its indeed PM42100 based then you will see separate drives.
| h2odragon wrote:
| > the manufacturer confirmed that the X21 offers 100 PCIe lanes,
| suggesting the presence of a PCIe switch.
|
| Almost like it's custom designed for a particular application
| where money's no concern... Perhaps someone in Utah needs big
| rainbow tables?
| amluto wrote:
| This actually sounds like it could be a nice mid-range product.
| For lots of money, you can get a fancy motherboard and
| enclosure that routes a ton of CPU PCIe lanes direct to the
| NVMe drives. This ends up with a lot of performance per unit
| storage, which one might not want.
|
| With a card like this, one can get a ton of high-speed (much
| better than SATA but not as fast as direct NVMe) storage in a
| regular machine.
| bick_nyers wrote:
| Could one hypothetically install an M.2 -> PCIE x16 riser and
| install a GPU?
| amluto wrote:
| I don't see why not. OTOH, if you try this on an NVMe
| "hardware RAID" slot, you may get hilarious results.
| bick_nyers wrote:
| Stripe your GPU in a Raid 0 configuration for maximum
| performance, if your GPU doesn't have ECC VRAM, consider
| mirroring them :)
| jeffbee wrote:
| People like to throw the innuendo around but the NSA's pathetic
| little datacenter is something that you would lose in a corner
| of a real datacenter operated by a real hyperscale system like
| Amazon or Google.
| xen2xen1 wrote:
| But the NSA one is dedicated to invading everyone's privac...
| Hey, wait!
| zamnos wrote:
| Just like a cluster of Bitcoin miners will run absolute
| _circles_ around a similarly sized corner of an AWS data
| center, and the supercomputer at Oak Ridge will run circles
| around a similar sized corner of AWS of GPU EC2 instances
| connected via gigabit Ethernet, the NSA 's cluster's got a
| different use case than running web services for every SaaS
| company that wants to run in AWS. I imagine it's aimed at
| saving and analyzing/decrypting large amounts of data, and
| thus is architected and tuned towards that purpose, and thus
| runs circles around a similarly sized corner of AWS for that
| particular task.
|
| Unless you have experience with the NSA's cluster that you'd
| like to share with the rest of the class, that is.
| jeffbee wrote:
| Needs citation. I think the idea that the NSA has stronger
| data storage and analysis infrastructure than commercial
| operators is not even conjecture, it's something weaker, a
| fantasy. Commercial hyperscale operators claimed the
| ability to sort 50PB datasets at 600GB/s, eleven years ago.
| Storing and analyzing bulk data is the #1 thing these guys
| are good at.
| [deleted]
| rektide wrote:
| PCIe switches just shouldn't be so dammed expensive. A decade
| ago there was a lot more market competition but now there are,
| what, two companies with chips?
|
| It has gotten a good bit harder to build, especially with so
| many of the tricks & tight timings in PCIe 5 and 6, but the
| lack of market competition has made getting any kind of parts
| at all much much more expensive.
| bick_nyers wrote:
| With the cost of PLX PCIE Switches (allows you to e.g. PCIE
| 4.0 x16 -> PCIE 3.0 x32) it is actually worth considering
| just buying a second desktop and throwing in some high-
| bandwidth NIC and forming your own HPC. Or instead of using
| desktop parts just going EPYC/Xeon/Threadripper.
|
| Of course it all comes down to the fact that if you need
| those PCIE lanes, there's a very good chance that it's for
| your job, meaning that businesses are the target market, not
| the enthusiast building a homelab for tinkering with LLM off
| the clock.
| eqvinox wrote:
| > Almost like it's custom designed [...]
|
| https://www.microchip.com/en-us/product/PM42100
|
| It's a standard COTS part.
|
| Coincidentally, the PCIe 5.0 variant is in "Samples available",
| i.e. not full production yet, which is very likely the reason
| for this card only being PCIe 4.0.
|
| https://www.microchip.com/en-us/product/PM50100
| rasz wrote:
| Doesnt help that PM42100 is $7.5K and out of stock.
| eqvinox wrote:
| That's the price of the development/evaluation kit. Those
| are produced in small numbers, have provisions for
| everything and debugging the kitchen sink, and thus always
| this expensive.
|
| With the PM40100 being $800 (single unit, no bulk pricing),
| the PM41100 / PM42100 are probably < $1500. (They do seem
| to have more features, not quite clear without proper
| datasheet sadly.)
| mikece wrote:
| This is No Such Agency who would buy as many of these as could
| be produced...
| wdb wrote:
| Nice, I am currently in the jungle of SSD enclosures for my Apple
| M2 Pro device and it's pretty confusing. But 31 GB/s seems wild.
| vardump wrote:
| Same issue here, found anything good?
| wdb wrote:
| Not yet Orico seems pretty disappointing 600MB/s for their
| 20Gbps enclosures. I would thought you should already get
| that with their 10Gbps ones. So a bit wary to try out their
| 40Gbps enclosures.
| AceJohnny2 wrote:
| Note that a Thunderbolt 4 Hub can be a bottleneck:
|
| https://eclecticlight.co/2023/02/21/thunderbolt-4-hubs-can-s...
| formerly_proven wrote:
| Note 32 Gb/s, not GB/s.
| phonon wrote:
| No, it's GB/s. PCIe 4.0 x16 has a bandwidth of 32
| Gigabytes/s.
| [deleted]
| manav wrote:
| Might be able to eventually get 80Gbps out of USB4v2.
| Havoc wrote:
| Really hope they release a smaller version for home use.
|
| Something with say 8 slots would turn all those gen 4 pcie gaming
| motherboards retiring soon into a great NAS.
|
| Asus I think already makes a similar one but it isn't fanless
| toast0 wrote:
| You can do a passive x16 -> 4 x4, _iff_ your board supports
| pci-e bifurcation. Theoretically, you could do x16 - > 8 x2
| also passively, but I haven't seen bifurcation go down to x2.
| PCI-e switches are probably too expensive for anything active
| though.
| Havoc wrote:
| Yep - board supports it. Unfortunately even the passive cards
| seem to be minimum 150 bucks.
|
| I have a feeling that by the time I get round to this 8TBs
| may be so cheap that dual of those in the mobo ports may be
| enough haha
| toast0 wrote:
| Here's one for half that... https://www.amazon.com/ASUS-M-2
| -X16-V2-Threadripper/dp/B07NQ...
|
| This isn't an endorsement. Just an encouragement to shop
| more. This one says pci-e 3.0, fwiw, but I don't know how
| important pci-e 4.0 is to you?
| [deleted]
| throitallaway wrote:
| Does anyone else marvel at data throughputs nowadays? People talk
| about 5GB/s NVME cards as being "slow." Same with Internet
| speeds. It's unreal the progress that we've made (and continue to
| make.)
| forinti wrote:
| Sometimes you have to move 15TB about and then nothing is fast
| enough.
|
| At 5GB/s, that would take nearly an hour and 5GB/s would be the
| fastest part of the trip. If it has to land on tape or travel
| through the net, it's going to take days.
| jgalt212 wrote:
| sneakernet will always be with us.
| LeonM wrote:
| "Never underestimate the bandwidth of a station wagon full of
| magnetic tapes hurtling down the highway" - Andrew S.
| Tannenbaum
|
| But you are right, writing to and reading from tape will take
| a long time. Modern tape drives can do ~500MB/s, so 15TB will
| still take ~9 hours. Though that may still be faster than a 1
| gbit internet connection (depending on how far you must
| drive).
| godelski wrote:
| Yes and no. Throughputs are crazy high these days but the rate
| at which they increase is slower than the rate at which compute
| increases, by a lot. So if we're discussing in a relative
| sense, then there's a growing divergence and thus one can argue
| that I/O is getting "slower." This is actually one of the major
| topics in HPC discussion and why there have been so many crazy
| hacks. Things like flashbuffers are pretty much essential these
| days. Even if you're doing multi-node ML training you see
| pretty big differences using infiniband due to the frequency in
| which nodes need to communicate (there are regularization
| interval tricks too). In scientific computing this is a big
| limitation to our ability to visualize at high resolutions and
| is why in situ visualization is growing popular.
|
| As far as consumer hardware and consumer usages, yeah,
| everything just feels fast though.
| vinyl7 wrote:
| Its a shame all our data is sent/received over HTTP these days,
| otherwise I'd be excited about it
| organsnyder wrote:
| This is a very different HTTP than 1.0 or 1.1.
| kevin_thibedeau wrote:
| The protocol implementation doesn't matter if everything is
| serialized into ASCII. At some point there's going to be a
| web 4.0 where people figure out the performance advantages
| of binary data.
| Mountain_Skies wrote:
| It's starting to become difficult to comprehend much of it now.
| My ISP is pushing me to replace my 300 Mbps service with 2Gbps,
| but I never even saturate what I have. Maybe if I were a gamer
| with huge downloads it would make sense to upgrade.
| SketchySeaBeast wrote:
| I'm a gamer on 300 Mbps. Even 100 GB huge games aren't much
| more than a half hour away on steam. I don't really see the
| value in being able to download that much in 5 minutes.
| 0cf8612b2e1e wrote:
| I do not know if it is still true, but originally the
| PlayStation did not perform patch diffs, and any kind of
| update could be 10s of gigabytes. If I were routinely
| having to wait to start a frequently patched game, that
| would get old pretty quickly.
|
| Outside of that, yeah, I am not sure what use greater than
| gigabyte would be for 95% of the population.
| runnerup wrote:
| Blizzards servers never saturate my 1gbps even during extreme
| off hours.
| doubled112 wrote:
| Absolutely. I keep saying to people "it doesn't matter,
| everything is fast now".
|
| Gigabit fibre to the home, NVMe that is way faster than RAM was
| not that long ago, CPUs in phones that make old desktops look
| like toasters.
|
| The disconnect is that the numbers feel huge in comparison, and
| what my computer can do for me really, is not hugely different.
| tester756 wrote:
| Seems like we'll start using NVMe as RAM, so we'll be able to
| have higher RAM sizes
| rhn_mk1 wrote:
| It's called swap.
| doubled112 wrote:
| Could we go full circle and use that for a RAM disk?
| sidpatil wrote:
| IIRC tmpfs already does that, by swapping to disk.
| tester756 wrote:
| Conceptually? yes, but I meant using NVMe fully as RAM,
| without your RAM sticks.
| godelski wrote:
| Isn't the issue endurance? I'm pretty sure your drive
| would die before the year is over. Probably a month or
| two.
| Night_Thastus wrote:
| Isn't NVMe storage much higher latency than RAM, which is
| no good for the CPU? IIRC, NVMe is also poor at random
| access.
| bick_nyers wrote:
| Yup.
|
| Now Optane on the other hand...
| DaiPlusPlus wrote:
| > NVMe that is way faster than RAM was not that long ago
|
| Citation?
| doubled112 wrote:
| Depends on what you consider a long time, maybe.
|
| https://www.samsung.com/us/computing/memory-storage/solid-
| st...
|
| > Sequential read/write speeds up to 7,450/6,900 MB/s
|
| https://en.wikipedia.org/wiki/DDR2_SDRAM
|
| Lists DDR2-400 capable of 3200 MB/s of throughput.
| ciupicri wrote:
| Yeah, but there's a _random_ in RAM and SSDs aren 't that
| great yet.
| rektide wrote:
| DDR2-1066 was pretty rare (fast), and rated as PC2-8500,
| meaning 8.5GBps. DDR3 started here-about, in ~2007.
|
| Consumer PCIe 5.0 ssds will in some cases likely surpass
| that.
| jakogut wrote:
| DDR2-800 has a maximum theoretical bandwidth of 6,400 MB/s,
| and was in common use well after 2010.
| kiratp wrote:
| Lets all keep in mind that a very small portion of the global
| population has access to this. It is our responsibility to
| bring all of humanity forward with us as we write software.
|
| https://perfnow.nl/speakers#alex
| SketchySeaBeast wrote:
| While you're not wrong about the portions who can access
| it, I wonder at how many of us doing work that brings
| humanity forward at all. Sure, let's not use fast system as
| an excuse for bad code, but I'm not going to pretend my
| CRUD app is elevating humanity. Even the stuff that claims
| to feels like at best a lateral move a lot of the time.
| alfiedotwtf wrote:
| Talking size and not speed, to be honest without joking, I
| still marvel that I have a 64Gb USB stick. 64Gb! HUGE!
| DaiPlusPlus wrote:
| I marvel at 5GB/s mass-storage, but then I groan at the
| knowledge it will be used to run a denornalized Postgres
| database table containing JSON blobs for queries that will do
| full tablescans because apparently dropping $Lots on a high-end
| storage system is preferable to learning how to do things
| properly.
| whoomp12342 wrote:
| reasoning: engineers can be lazy too
| runnerup wrote:
| Moreso, because this solution costs $4,000 so it only needs
| to save one week of one persons engineering time to be a
| more cost effective solution than "better software
| engineering".
|
| It's laziness, but it's cost effective laziness.
| foepys wrote:
| The end result of this is something like Microsoft Teams
| where they are in the process of swapping out the entire
| underlying runtime because the team that wrote the
| application itself was so lazy (read: incompetent) that it
| is slow on literally every machine it runs on and this is
| apparently the only sane way to fix the whole mess - bar a
| whole rewrite of the entire application.
| 0cf8612b2e1e wrote:
| Is this true or confirmed somewhere official? Clearly the
| underlying architecture has major issues.
| LeifCarrotson wrote:
| I'm in the automation/manufacturing industry, and while I
| know there are big servers running ERPs and SCMs and MES and
| other TLAs with poorly-optimized giant databases I don't
| personally touch those. But still I suspect the same truths
| apply whether you're talking about "big iron" computers or
| literal metals: when evaluating the strength of an average
| weldment on an ordinary conveyor or machine, we like to say
| "Steel is cheap, engineering is expensive."
|
| Obviously there are situations when you have to employ more
| rigor and do the FEA, but typically, when choosing between a
| just-right solution and one that's obviously strong enough,
| just overbuilding it is a lot more efficient in terms of
| value.
|
| With 21 of the pictured $150 Samsung 1TB 990 Pro SSDs and,
| hypothetically, $1000 for this card, you're looking at $4,150
| for this storage solution. If that solves your problem and
| lets you apply off-the-shelf Postgres and JSON and
| unoptimized queries, do it! That money only buys a handful of
| site visits and maybe a week of engineering hours to change a
| system that may involve dozens of users, tens or hundreds of
| thousands of lines of code, and rigid requirements from
| upstream and downstream...maybe you can change those
| eventually, and it would definitely have been cheaper if all
| the stakeholders had a fundamental understanding of the
| compute requirements of full table scans and non-native blobs
| and designed their business around those mathematics, but
| that doesn't sound likely.
| kickaha wrote:
| Got a VIC-20 for my 13th birthday in 1983. The local TV shop
| sold Commodore hardware and hosted a BBS that my little nerd
| cronies and I connected to at 300 bps. All of us knew the owner
| who ran the store and mentored us and hired some of us when we
| got old enough. On one visit he took me back, behind an actual
| curtain, and showed me the BBS machine. If you know the 1541
| Disk Drive, you know. Well this beast had a Commodore branded 5
| MB hard disk connected to a C-64. In my memory it was 6" by 8"
| by 24" long, with an enormous power supply, and must have cost
| thousands of Reagan-first-term US dollars.
|
| Two or three years later another mentor hired me to put a 33 MB
| hard disk into his IBM PC. Not a clone. My memory tells me it
| was a DOS imposed limit, those 33 MB: the biggest drive
| available. I managed to plug the connector in upside down and
| released the magic smoke. That was a many-hundreds-of-dollars
| mistake. (And a good lesson in patient mentoring.)
|
| In 1991 I obtained a used 80 MB drive (half height!) to put
| into my own PC XT clone, via a local Usenet group. I set the
| volume name to $1_PER_MB because going under that threshold was
| so impressive.
|
| Those are my reference points for storage.
| kickaha wrote:
| I'm so old I forget that old memories can be found again with
| Wikipedia and Google.
|
| https://en.m.wikipedia.org/wiki/Commodore_D9060
|
| https://www.commodore-
| info.com/brochure/item/commodore_d9060...
| Mountain_Skies wrote:
| 300 baud was comfortable reading speed, at least for me as a
| child. It was also fun to pick up the phone and be able to
| distinguish individual bytes of data (though not the actual
| content of the byte). Once we got to 1200 baud, it just
| became a stream of warbling.
| Arrath wrote:
| > I set the volume name to $1_PER_MB because going under that
| threshold was so impressive.
|
| Hah! Too funny, my own personal memory for 'cheap' storage
| was keeping an eye on the Fry's print ads in the Sunday
| newspaper while saving all my allowance and summer job money,
| finally buying the outrageously large 200GB HDD for a mere $1
| a gig!
| jjoonathan wrote:
| What are the current strategies for leveraging NVMe speed &
| volume in a NAS?
|
| When I look at NAS offerings, I see lots of 2.5" and 3.5" bays
| and 1Gbe (maaaybe 10Gbe at the high end) which is a bit stifling.
| toast0 wrote:
| u.2 NVMe uses 2.5" bays and SATA Express connectors to offer up
| to 4 lanes of PCI-e or two lanes of SATA. That's probably where
| most of the enterprisy NAS is going.
| jeffbee wrote:
| In my humble opinion some home NAS use cases are beneficially
| converted to Thunderbolt. It's a lot more practical than the
| faster varieties of ethernet. If you have two hosts that need
| fast access to the NAS you can do it with TB4, and one of those
| hosts can re-export the stored resources for applications with
| lesser performance requirements, over SMB or iSCSI or whatever.
| rektide wrote:
| 40Gbps host-to-host networking in thunderbolt/usb4 is so epic
| & great to have.
| jjoonathan wrote:
| Does someone make a thunderbolt -> thunderbolt connector
| that exposes an ethernet PCIe endpoint to each host and
| forwards packets between them?
|
| 10 years ago I joked that someone should do this, but I
| thought high speed ethernet would trickle down and obviate
| the need. Evidently not, lol.
| jeffbee wrote:
| IP-over-Thunderbolt is a thing. It's the thing we are
| discussing! All major operating systems can do it.
| jjoonathan wrote:
| I didn't know you could just connect two thunderbolt
| hosts and get a network interface. Amazing!
|
| Ok, so now I need to find a small cheap computer with a
| thunderbolt port or two and lots of NVMe.
| jjoonathan wrote:
| Yeah, but then the primary host inherits the uptime
| requirements of a NAS and this gets really awkward when you
| have workflows in different operating systems that both want
| to use the fast storage. Both OSes can access thunderbolt, of
| course, but only one can be a good server. Now that I've
| experienced the separation of concerns that comes from an
| independent NAS I do not want to go back.
| jeffbee wrote:
| Which host have you nominated as "primary" in this
| scenario? I'm only thinking of the TB4 link as a relatively
| cheap and simple point-to-point networking link. I think
| it's a lot more practical, and cheaper, that 25gbps
| ethernet. A dual-port 25gb ethernet NIC costs hundreds of
| dollars, while you can find all kinds of cheap computers
| with 2 TB4 ports.
| jjoonathan wrote:
| The host connected to the NVMe storage.
| lyind wrote:
| This nice approach has at least these drawbacks:
|
| 1. Swapping drives is hard * may be overcome by
| declaring failure domain = node
|
| 2. No powerloss protection advertised to OS, ie. slow synchronous
| writes * may be overcome by software hacks and
| whole-system battery supply
|
| 3. Potential slowdown on continuous write load (weeks or months,
| depending on drive) * may be overcome by
| software in _some_ situations
|
| At least the last two points are a no-go for enterprise use-
| cases, if not addressed.
| ciupicri wrote:
| > No powerloss protection advertised to OS
|
| How can the OS (I'm interested in Linux) know about this
| feature?
| eqvinox wrote:
| I'm gonna claim that this card isn't aimed at enterprise use-
| cases that need this kind of service. I'd put it along the
| lines of "nearline SAS" HDDs, aimed at non-critical
| applications where you care more about bulk capacity than
| reliability.
|
| Relatedly, M.2 SSDs are inherently slower than the same pile of
| silicon in an U.2/2.5" form factor -- the power/heat budget is
| noticeably lower.
| rasz wrote:
| >Apex Storage doesn't reveal the inner details of the X21
|
| >In a single-card configuration, the X21 delivers sequential read
| and write speeds up to 30.5 GBps and 28.5 GBps, respectively.
|
| did they test it or reprint press release?
|
| >According to Apex Storage
|
| ah
|
| >The AIC has an average read and write access latency of 79us and
| 52us
|
| that doesnt make sense unless its additional latency of
| controller or they ship it populated with drieves.
|
| >However, Apex Storage didn't expose the type of RAID arrays. The
| X21 also flaunts "enterprise-grade reliability," NVMe 2.0
| support, advanced EEC, data protection, and error recovery. Apex
| Storage didn't reveal the pricing or availability for the X21.
|
| so only revealed performance figures and pictures
|
| Their previous product was a fancy looking bracket for holding 16
| SATA M.2 drives https://www.kickstarter.com/projects/storage-
| scaler/storage-....
|
| >A cross platform drop in card that can massively increase the
| amount of storage for your computer using cost effective m.2
| SSDs.
|
| You had to read fine print to realized its just a m.2 stand
| requiring proper 16 port SATA controller to function. It still
| hasnt shipped to this day. Im mildly optimistic.
___________________________________________________________________
(page generated 2023-03-13 23:02 UTC)