Post B5n0BCS4ZTMPxEweEC by Stormgren@obsidianmoon.com
 (DIR) More posts by Stormgren@obsidianmoon.com
 (DIR) Post #B5m2ONhc3fzDMXfPWK by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       Well this was not expected.The 100G optics I ordered for my desk came in, and swapping the pipe to my desk from 40 to 100G gave a huge improvement in Ceph performance.But I don't understand *why*.Baseline config from earlier in the year: client on 40G, cluster nodes on dual 10G, 1558 MB/s on linear readsMoving cluster nodes to 40G, client on 40G: 1864 MB/sMoving client to 100G, cluster nodes on 40G: 3787 MB/s.The confusing thing is, even 3787 MB/s is only about 30 Gbps, so after protocol overhead I would expect it to fit comfortably in 40G. Why can I get this performance with the client on 100G, but not on 40?
       
 (DIR) Post #B5m2jWKDC9F0BbnjYe by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @xabean Intel XL710 -> Mellanox ConnectX 6. Haven't checked specs but it's likely faster, I think the old NIC was a gen3?
       
 (DIR) Post #B5m34Ar6DI6LgH24Wm by whitequark@social.treehouse.systems
       0 likes, 0 repeats
       
       @azonenberg latency?
       
 (DIR) Post #B5m34B6LIcIKRYAFHs by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @whitequark That's kinda what I am wondering, is ceph super sensitive to latency at this level? I don't actually know but my gut feeling is that it shouldn't be they're plugged into the same switch. the difference in latency between 40 and 100G should be like a few microseconds and I'm done filesystem operations here not HFT
       
 (DIR) Post #B5m4UBBzgOfuD2lnZw by ChuckMcManis@chaos.social
       0 likes, 0 repeats
       
       @azonenberg What is the max op rate of the controller? Channel semantics being ops/sec dominate until packets get big enough and then bandwidth dominates.
       
 (DIR) Post #B5m50BDL8abrXket7o by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @ChuckMcManis No idea, and no idea how to find out either lol. This isn't a spec people usually publish in an obvious place like the bandwidth
       
 (DIR) Post #B5m58ZLpHUNAied0Xw by jpm@aus.social
       0 likes, 0 repeats
       
       @azonenberg @whitequark I’m also thinking latency, but more end to end latency of the entire transaction from the beginning of the ceph client issuing a read request to the end of read responses being received by the client, which is the point the client can issue another request. If there’s less buffering happening inside the switch going from 40G to 100G, with all the frames queued from the cluster exiting to the client faster, that could explain it.
       
 (DIR) Post #B5m5J4UsKtQMDErvX6 by ChuckMcManis@chaos.social
       0 likes, 0 repeats
       
       @azonenberg At the company doing SDRs with the Zynq Ultrascale+ we had it do 68 byte packets as fast as it could on a  100G link to figure out the ops limit. (it was actually about twice as good as the 40G link but note the bandwidth was nominally 2.5x more. The customer was happy with that though so we didn't do any deep dive optimization.
       
 (DIR) Post #B5mFVwFfInn3wjWpQO by tnt@chaos.social
       0 likes, 0 repeats
       
       @azonenberg Did you just change the optics ( like the QSFP modules ) and the NIC stayed the same (supported 100G all along) ?
       
 (DIR) Post #B5mFyNsy9Zm1lATbxA by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @tnt no i swapped an xl710 for a connectx6
       
 (DIR) Post #B5mGFB4f2Ufawsx9Ci by manawyrm@chaos.social
       0 likes, 0 repeats
       
       @azonenberg @xabean Just better drivers all around, then. Mellanox NICs have the very solid mlx5 kernel driver, with well working kernel offloading support for a _ton_ of features, saving a bunch of CPU and interrupts.Part of the reason why I like the ConnectX-4 and newer cards so much.
       
 (DIR) Post #B5mGW2CFIceqL1XSbo by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @manawyrm @xabean yeah I've been gradually rolling out cx5s and cx6s across my fleet as i got my hands on them. I'm actually in the middle of reconfiguring my VM server to use the one i just put in it
       
 (DIR) Post #B5mGfkSi5fWeEPs6qm by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @manawyrm @xabean (it had an x520-da2 before)
       
 (DIR) Post #B5mGm63UihSmwJ6XFQ by manawyrm@chaos.social
       0 likes, 0 repeats
       
       @azonenberg @xabean my usual rule of thumb is: "don't use anything other than mellanox for anything that has bridges or ip_forwarding enabled". intel seems mostly fine for TCP server stuff. broadcom is just awful all-around.both intel and broadcom have some weird offloading/driver bugs that will bite you in prod workloads, it's kinda crazy to me.
       
 (DIR) Post #B5n0BCS4ZTMPxEweEC by Stormgren@obsidianmoon.com
       0 likes, 0 repeats
       
       @jpm @azonenberg @whitequark I would tend to agree about the buffering, sometimes that throws timing of packet flows off a bit.  I've found that 4x10G lanes from a switch to a QSFP vs 4x25G can be treated very differently.  I've oft suspected that the chip and SERDES optimization on a lot of platforms is oriented for 100G and 40G is there just for backward compatibility but isn't performant.At the same time.however, I wouldn't be surprised if it really just did come down to the NIC.  The 710 series isn't a great card in general, and I actually had to go back to the 520-DA2s to get things like hardware timestamping back and the 520 also handled VFIO a heck of a lot better.  The second I get out of home improvement hell here and can scrape up some cash to go ConnectX of some kind, I am 100% going for it (along with the required switching upgrade, I'm 10G only for now).
       
 (DIR) Post #B5n0BCijZWgimujxCK by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @Stormgren @jpm @whitequark I had a 710 in my workstation and 520-DA2s in almost every box until recently. I've been gradually replacing them with CX5/CX6.The client CPU might have something to do with it too, now that I think about it... I unfortunately did not a get a benchmark after the mobo upgrade (2x Xeon 6144 -> 2x Xeon 8362, 192GB DDR4 2666 to 512GB DDR4 3200) with the old NIC since I did the board and NIC swap simultaneously.Nor did I get a benchmark with the CX6 lit up at 40G, since I only had it in that configuration for about 12 hours before the shipment from FS with the 100G optic I needed to light up the switch port arrived
       
 (DIR) Post #B5zcqKlUhBPblWac2i by evey@chaos.social
       0 likes, 0 repeats
       
       @azonenberg @tnt x710 is known to be performs not great and the connectx6 (dx) I assume does a lot more offloading / fancy things to improve performance. Plus I assume more modern PCIe version the x710 series is pushing PCIe bandwidth quite hard.