[HN Gopher] Zml-smi: universal monitoring tool for GPUs, TPUs an...
___________________________________________________________________
Zml-smi: universal monitoring tool for GPUs, TPUs and NPUs
Author : steeve
Score : 74 points
Date : 2026-03-31 13:29 UTC (5 days ago)
(HTM) web link (zml.ai)
(TXT) w3m dump (zml.ai)
| mrflop wrote:
| Renaming fopen64 to intercept library calls feels like a brittle
| hack masquerading as "sandboxing." Why not just upstream this
| hardware support to nvtop instead of fragmenting the ecosystem?
| steeve wrote:
| sadly, sandboxing is something that can't be upstreamed. this
| way, sandboxing is kept in zml instead of patching mesa.
|
| as for nvtop, great program, but we missed a few features (such
| as sandboxing)
| pstuart wrote:
| It looks cool and I was excited to get monitoring for the NPU
| on my Ryzen AI 395+, unfortunately it does not show. NPU
| support in linux really seems to be an afterthought.
| steeve wrote:
| Weird, because we tried it. It doesn't show anything?
|
| We use the amdsmi to get metrics. I'll investigate.
| marwanet wrote:
| If this logic were pushed into nvtop, wouldn't the codebase
| become unmaintainable? Each vendor's interception method is
| going to be different.
| rdyro wrote:
| Looks cool!
|
| nvtop can actually support TPUs too via
| https://github.com/rdyro/libtpuinfo/
| https://github.com/Syllo/nvtop/blob/76890233d759199f50ad3bdb...
| 152334H wrote:
| "NPU" seems to refer to trainium only?
| serialx wrote:
| Look into all-smi https://github.com/lablup/all-smi It supports
| all GPUs thinkable including Apple Silicon and many AI
| accelerator cards.
| imcritic wrote:
| Is it capable of exposing metrics in Prometheus format?
| steeve wrote:
| consider it done
| synergy20 wrote:
| would be nice to have cpu usage added so I have all in one?
|
| currently I use btop which shows basic gpu load along with cpu,
| network, etc.
___________________________________________________________________
(page generated 2026-04-05 23:02 UTC)