Post B5x26e64VDY1jfD3LM by azonenberg@ioc.exchange
(DIR) More posts by azonenberg@ioc.exchange
(DIR) Post #B5wtX8ZOvaJoBhU9tg by azonenberg@ioc.exchange
0 likes, 0 repeats
You know what would be really nice (but nobody is ever going to build)?Oscilloscope that replaces the ungodly slow USB3/1000baseT PC interface port with NVLink.Forget PCIe and Thunderbolt... 900 GB/s of bandwidth straight from the ADC to my GPU? Sign me up.
(DIR) Post #B5wtyMw7cbTOYz68fI by azonenberg@ioc.exchange
0 likes, 0 repeats
For comparison... my 16 GHz LeCroy oscilloscope puts out 40 Gsps * 4 channels * 8 bits of raw ADC samples, not counting the flatness corrections done in gateware/firmware.That's 160 GB/s or 1.28 Tbps of raw samples.That would even fit in NVLink 2.0 much less the current gen4/5 stuff.Imagine four channels of 16 GHz bandwidth waveform data straight into a (very large) GPU nonstop... We'd have to do a hell of a lot of optimization to ngscopeclient to keep up and probably add multi-GPU support but it would be so much fun lol.
(DIR) Post #B5wuxOMUw7MceHX7UO by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg I'm always kind of weary of silicon manufacturer's proprietary high-speed buses, because they teeeend to be slightly use-case-specific and don't deal well with edge cases outside that. Anyways, when I saw NVLink my first reaction was "wait is this HyperTransport, but with expensive modern transceivers?"; it isn't, but my guess is that from a system's perspective, you'd be better off going for AMD's InfinityFabric, which seems to make stronger coherence statements (not sure). But, you
(DIR) Post #B5wv3ViaUzEFCXmc2S by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg are mostly only optimizing for nvidia GPUs anyways, so that might be a moot point.
(DIR) Post #B5wv3W4ZBGoGIi4AIS by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab Well I mean I would *like* a ludicrously high bandwidth portable interface, but the vendors aren't building it.Realistically, I think the best you can do portably today is 100GbE with RoCE.
(DIR) Post #B5wvb3UJxpiTUd6aZs by 1div0@mastodon.social
0 likes, 0 repeats
@azonenberg https://tenstorrent.com/hardware/cards#compare ?
(DIR) Post #B5wwvnRWujdjlXFnRw by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg oh you *can* buy 800 Gb/s interfaces, don't know what their host sides look like, if any for non network-vendor stuff (this is mostly aggregated traffic equipment, i.e. linking racks or DC 1 to DC 2)
(DIR) Post #B5wx3Zxmunq6HA3MeG by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab yeah exactly. 100G with a normal pcie interface is available today, i have a 100G pipe to my desk and have saturated it with iperf in benchmarks.And the nic has RoCE offload capabilities although I'm not using it yet
(DIR) Post #B5wzSSvqOWBPcH8ehc by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab You can go all the way up to 800G if you have a host system with PCIe gen6 and sufficiently deep pockets (I do not)https://www.fs.com/products/346847.html?attribute=128158&id=4830037
(DIR) Post #B5x14z8PDLB0RuwgV6 by scribblesonnapkins@mastodon.social
0 likes, 0 repeats
@azonenberg so my tek 11801C with it's ability to connect to an external sampling head array might present a problem? Oh no wait it's sampling to slow.
(DIR) Post #B5x1DJYTAwci6B6mlU by azonenberg@ioc.exchange
0 likes, 0 repeats
@scribblesonnapkins well equivalent time sampling is easy to handle with today's tech because the number of actual samples acquired per second is low.equally, a scope that acquires high speed data and buffers it in memory before processing at a much slower rate is something we can handle today.But the vision is to be able to do real time or at least lower-dead-time processing at much higher data rates. ThunderScope almost maxes out 10GbE, my vision is to able to keep up with 25/40/100G eventually
(DIR) Post #B5x1QV6PYODjg2Gvb6 by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg I doubt your pockets will be deep enough for NVlink things involving anything but GPUs :)
(DIR) Post #B5x1VSvP4qQKN5UUCW by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab oh i know, nvlink doesnt even let you get the PHY chiplets (the protocol itself is undocumented) unless you have NDAs and a partnership with nvidia etc.but I can dream...
(DIR) Post #B5x1w5bHuAKPqnmEYy by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg I was assuming that you'd probably (assuming infinite money) could buy an Nvidia server platform that has network->VRAM piping (I assume this because I presume that's what nvidia bought mellanox for)
(DIR) Post #B5x213QdPIoaYxA4ie by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab That's where RoCE comes in.But Ethernet today tops out at 800 Gbps while the latest NVLink can do 14.4 Tbps
(DIR) Post #B5x26e64VDY1jfD3LM by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab NVLink is the fantasy, the actually achievable real world implementation is to make the scope speak RoCE, put a mellanox NIC in the client, and RDMA the incoming Ethernet frames straight into VRAM.But it still has to cross over PCIe and get bottlenecked on that bandwidth
(DIR) Post #B5x2idTvEqvSpbTqyW by scribblesonnapkins@mastodon.social
0 likes, 0 repeats
@azonenberg I was trying to make a joke with the 1st part "Oh no wait it's sampling to slow."But the second part about the thunderscope is cool.
(DIR) Post #B5x9fa5I72BZ6aRxbs by penguin42@mastodon.org.uk
0 likes, 0 repeats
@azonenberg How about CXL4? That claims 242GB/s and is at least designed for external connectivity.
(DIR) Post #B5x9jrBUTdbRIYO9Z2 by azonenberg@ioc.exchange
0 likes, 0 repeats
@penguin42 If somebody makes a GPU with CXL I'll be all over it.Until then I'm stuck with what I can get my hands on. Realistically, that's PCIe and RoCE
(DIR) Post #B5xDdvJYblXwBdaVlo by CliffsEsport@mastodon.social
0 likes, 0 repeats
@azonenberg @funkylab I am curious what fields use Oscilloscopes at level you build and test for? I am guessing radio & perhaps medical? I've only ever used them for basic electronics back in the 90s so the performance of your stuff is just stunning.
(DIR) Post #B5xE1WzegVVfhrXI0G by azonenberg@ioc.exchange
0 likes, 0 repeats
@CliffsEsport @funkylab My focus is mostly on the high speed digital side of things, so networking, high speed buses, etc. Modern digital interfaces are absurdly fast.Even USB 3.0 was 5 Gbps per pair and that's pretty slow compared to modern stuff. PCIe gen6 runs at 64 Gbps.DisplayPort goes up to 20 Gbps per lane now.But understanding complex issues around these buses involves recording a lot of data, processing it fast, looking at packet captures and physical layer signal quality, etc. There's always room to crunch more data faster.
(DIR) Post #B5xENVlxu7QiI6c6Oe by azonenberg@ioc.exchange
0 likes, 0 repeats
@CliffsEsport @funkylab Much lower speed stuff still benefits from better processing.Modern cars are switching from CAN bus to Ethernet for talking between modules. 100baseT1 is a common flavor that runs Ethernet frames at 100 Mbps bidirectionally over a single pair of wires.If you want to simultaneously look at a protocol decode and signal quality view of that signal in both directions you're looking at around a GB/s of data. Which ngscopeclient is able to crunch in real time (barely, this is a recent tech demo)
(DIR) Post #B5xEjfOknTGmW0Jxya by CliffsEsport@mastodon.social
0 likes, 0 repeats
@azonenberg @funkylab Oh, wow, I did not know cars were switching to Ethernet.
(DIR) Post #B5xHXk3fweQ9M6T248 by funkylab@mastodon.social
0 likes, 0 repeats
@CliffsEsport @azonenberg oh, that's not very similar to the Ethernet you know. The MAC mechanism is *completely* different.
(DIR) Post #B5xHXkFNF9mJwNwNIe by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab @CliffsEsport The layer 2 MAC is identical.PHY? Yeah, that's totally different.
(DIR) Post #B5xJEXayG2SR6n0rTM by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg @CliffsEsport waitwaitwait, you're referring to something that's not ×Base-T1S, right? Because that has PHY-Level Collision avoidance, with a master controlling transmission opportunities based on Node IDs (and methods to keep the system running should the master fail, by another node taking over that job).
(DIR) Post #B5xJT4MbDtJun8Dbf6 by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab @CliffsEsport Yes I am talking about 100/1000 baseT1 which is a normal full duplex switched fabric like any other modern Ethernet flavor, just running over a single twisted pair.T1S is a different animal outright.
(DIR) Post #B5xJkJa1twFCuLivVg by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg @CliffsEsport "Just" running full-duplex on a single fiber is still a bit cool to me, but yeah, that thing (maybe aside from the "master/slave" phase during link training) a lot more classical.100/1000-T1S is indeed a lot more like what happens when CAN, someone who hates CAN, and Ethernet raise a protocol child together (lovingly, but competitively). (aunt Firewire comes by once in a while to tell stories of address arbitration)
(DIR) Post #B5xKngVDEVc2p5doCO by f4grx@chaos.social
0 likes, 0 repeats
@funkylab @azonenberg assuming infinite money you buy nvidia and make all their protocols open source, plus you hire engineers to develop nvlink into a generic large bit pipe.
(DIR) Post #B5xKnghcUNXNRZRiXQ by azonenberg@ioc.exchange
0 likes, 0 repeats
@f4grx @funkylab lol fair
(DIR) Post #B5xMEn4f0WEQygi7NI by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg @f4grx I think that's not actually … hm quite true. NVLink isn't the protocol you want there, far as I can tell. It's a GPU-to-GPU-first low-coherency, huge-transfer bus. What it does is not that useful to streaming data that comes in in tiny quanta, far as I understand. Sure, it's fast on paper, but it solves a different set of problems than PCIe or Infinity Fabric, or, just to put something else out there, 200GBase-KR2 or -KR4.
(DIR) Post #B5xMPAyYzq0a0GoLHU by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab @f4grx I'm picturing oscilloscopes RDMAing 50M sample waveform blocks into VRAM. That sounds like something it would do really well at.
(DIR) Post #B5xOuiVGakrys9H7XU by funkylab@mastodon.social
0 likes, 0 repeats
@azonenberg @f4grx yep, but how does your scope know which part of the buffer can be overwritten? How does the GPU know what data the SMs (are they still called that way in modern nvidia parlance? last time I did raw CUDA was 2010?) need to fetch from VRAM, and what's still fresh in local or Warp memory? Sure, if you can restrict yourself to "here's a block of 50 MS, and even with a many-GS/s scope that gives us milliseconds to propagate that info", that'll be fine. But if that's fine,
(DIR) Post #B5xPWr7yyUUCcnZQ5g by azonenberg@ioc.exchange
0 likes, 0 repeats
@funkylab @f4grx There would be an Ethernet control-plane socket regardless of whether the data plane was RoCE, NVLink, or something else. The idea would be to do something similar to what we do now with normal waveform captures: data shows up over a socket (or nvlink) as raw ADC samples, is converted from native int8/16 format to float32, written to a FIFO of pending waveform objects., then the filter graph pops waveforms each trigger and evaluates them.With higher bandwidth streaming you'd need larger buffers, of course, but hopefully by the time a scope with nvlink or 100GbE becomes a thing, 32/64GB VRAM desktop GPUs will be as well
(DIR) Post #B5y00gsTFIfPs6qAGu by AMS@infosec.exchange
0 likes, 0 repeats
@azonenberg @funkylab ConnectX-7 can do 400GbE RoCE I think (Basically fill the 16x PCIe 5.0 to the GPU).
(DIR) Post #B5y0B9UTia6FVdbfTk by azonenberg@ioc.exchange
0 likes, 0 repeats
@AMS @funkylab Nice.I have a 100G CX6 in my desktop now but can't push >40G to it from anything external because the core router is on the only other 100G port on the switch and all the other ports are 40.I hope to play with RoCE in ngscopeclient in the next year or two perhaps as part of R&D for the next gen thunderscope
(DIR) Post #B5y1TRHUPbYBqtoaLw by AMS@infosec.exchange
0 likes, 0 repeats
@azonenberg @funkylab I'm holding on for the 10GbE Thunderscope. Figure it'll beat the absolute pants off the midrange (350MHz) fiber coupled probes because I can put a few channels (Gate, drain, a supply rail or two, maybe a current) on the thing swinging around at 70V/ns.