Post B6IbBQCr9T2ColYQSm by azonenberg@ioc.exchange
 (DIR) More posts by azonenberg@ioc.exchange
 (DIR) Post #B6Hy09bIeen1uzcjSa by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       Post fan/GPU upgrade, and some additional fan RPM tuning via IPMI: VM server is running a lot cooler, for the most part. CPU VRM temperatures during a big compile job are less than the *idle* temps previously.But I'm now seeing NIC temperature and it's concerningly hot. I'm not sure why it wasn't showing up before so I have no idea how toasty it was.I'm also seeing what appears to be poor / unstable network performance.The ConnectX6 is passively air cooled and sits just to the right of the new 80mm fans (as seen from the rear panel), and I suspect what is happening is that the negative pressure from the new fans is drawing front-to-back airflow slightly to the left and reducing airflow over its heatsink. Thermal engineering is hard.I have another PCIe slot exhaust fan on order coming tomorrow so hopefully things are tolerable between now and then.
       
 (DIR) Post #B6I0eQy9hJoxHHeIcK by karppinen@mastodon.online
       0 likes, 0 repeats
       
       @azonenberg I recently upgraded the firmware on twelve ConnectX-6 Dx cards, took something like 5 minutes per card (including one reboot), that 5 minutes was enough to make the heatsink too hot to touch. This was on a box with decent airflow too
       
 (DIR) Post #B6ID6NQJfnvBq4ua92 by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       Historical network traffic before and after the recent reconfiguration.I wonder why all my virtual desktops are so slow?
       
 (DIR) Post #B6IDGMbNaBZOBdqepM by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       The SSD is slightly cooler, other than a short spike right after boot that might be  before the fans spun up fully or something
       
 (DIR) Post #B6IDRUrXD6TizYiInw by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       We can also see the bad network performance in the CPU usage charts, showing up as increased dom0 iowait time due to CephFS operations lagging
       
 (DIR) Post #B6IDcnTKMQdXtTmMtc by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       This also shows up as spikes in load average, i think due to IO syscalls blocking more
       
 (DIR) Post #B6IDhMXLvr0hnuq4xc by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       Not a whole lot I can do about it until the new fan comes in though, so i'm just gonna have to deal with the poor performance until tomorrow.Right now I can cook the CPU VRMs or the NIC, there's no "neither" option
       
 (DIR) Post #B6IEYzWxgXM25ZEDTc by Mellivora@im-in.space
       0 likes, 0 repeats
       
       @azonenberg linux user here... just stick on top one of those mini fans. Works like a charm everytime.
       
 (DIR) Post #B6IXknxqPeVqov9Iwa by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       Set up this OEM approved cooling solution.Temps dropped by 10C or so but still getting 570 Mbps in / 138 out on iperf3. On a 40Gbase-SR4 pipe. Something is definitely wrong.
       
 (DIR) Post #B6IXyO1VenkKQ7L7YW by DausDD@dresden.network
       0 likes, 0 repeats
       
       @azonenberg Temporary solutions are always the most lasting
       
 (DIR) Post #B6IY3I8mxKCTcj5E6i by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @DausDD This one is going away in ~24 hours when the internally mounted fan I have on order comes in.But if this is having no measurable impact on performance, maybe something else is wrong
       
 (DIR) Post #B6IYJv5Dw6x12ELRBo by penguin42@mastodon.org.uk
       0 likes, 0 repeats
       
       @azonenberg Have you looked at the PCIe speeds using lspci -vvv?  A 'LnkSta' line I think on the device should show you speed and width, like 'LnkSta:Speed 8GT/s, Width x8'
       
 (DIR) Post #B6IYNFYmGEzbPXkwr2 by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @penguin42 Yeah it's linked at gen4 x16. Whatever the problem is, it's not that.
       
 (DIR) Post #B6IZxVzKl3zkk32B1s by bencc@morehammer.uk
       0 likes, 0 repeats
       
       @azonenberg Error counts going up on either side? Worth making sure your MTP connectors are fully pushed home, if they're a bit stiff and not quite engaged they can still bring up a link but might not work well. Multimode is too forgiving sometimes (I had this exact thing happen to me last week).
       
 (DIR) Post #B6IbBQCr9T2ColYQSm by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @bencc Unplugged and re-seated both the QSFPs in the cages, and the MTPs in the sockets. And cleaned the MTPs with a fiber cleaner.RX power reported by the switch is -1.95, -2.93, -2.62, -2.72 dBm which seems normal.
       
 (DIR) Post #B6IbOVwz6goaOgNDSS by bencc@morehammer.uk
       0 likes, 0 repeats
       
       @azonenberg ah yeah, those seem fine - went from -15 to -3 with our not-quite-seated connector!
       
 (DIR) Post #B6IbbEVdGQkI4Ke0rg by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @bencc one thing that looks linteresting is that I am seeing 452 million giant packets on the switch port perf counters which seems wrong, every endpoint should have a 1500 byte MTU (even though the switch supports up to 9600)
       
 (DIR) Post #B6IcMS0gsxlHYGge92 by bencc@morehammer.uk
       0 likes, 0 repeats
       
       @azonenberg Some sort of offloading now turned on? I was wondering if it's config that's referring to an old name, but looking at the history it doesn't look like you've move the card to a different slot so possibly not that. I assume both sides are reporting the same link speed?
       
 (DIR) Post #B6IcVEldUWY55OpoqO by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @bencc yeah this card hasn't moved since I installed it, offloading is all in the default mlx5 settings without me having touched anything by hand.Both sides have 40G optics in them; the switch port is 40G only and the NIC port is 100G capable. Both report linked at 40G.
       
 (DIR) Post #B6IcoSMLBGa2k8ARuK by bencc@morehammer.uk
       0 likes, 0 repeats
       
       @azonenberg Well that is very annoying.
       
 (DIR) Post #B6IdNE4qH1dmaCYuxc by azonenberg@ioc.exchange
       0 likes, 0 repeats
       
       @bencc Checking ethtool counters on the VM box, some that sound related:rx_crc_errors_phy = 0rx_out_of_buffer = 40526rx_packets_phy = 525122818rx_bytes_phy = 497156560480tx_bytes_phy = 805933109350rx_symbol_err_phy = 0rx_discards_phy = 265985tx_discards_phy = 577tx_pci_signal_integrity: 124374433rx_pci_signal_integrity: 0rx_err_lane*_phy: all 0tx_global_pause: 37106According to the docs, that's a bad sign and suggests PCIe communications problems between the host and card (despite being linked at gen4 x16).Unsure if that's due to thermal issues or something else, but it's likely part of the puzzle.
       
 (DIR) Post #B6IhI64ozUpA8gMWVE by bencc@morehammer.uk
       0 likes, 0 repeats
       
       @azonenberg Huh, I guess more testing after properly solving the thermal problem is a good place to start. Probably worth a card reseat too given how warm it was before.