[HN Gopher] SIMD programming in pure Rust
___________________________________________________________________
SIMD programming in pure Rust
Author : randomint64
Score : 29 points
Date : 2026-01-19 09:39 UTC (2 days ago)
(HTM) web link (kerkour.com)
(TXT) w3m dump (kerkour.com)
| crote wrote:
| What is the "nasty surprise" of Zen 4 AVX512? Sure, it's not
| quite the twice as fast you might initially assume, but (unlike
| Intel's downclocking) it's still a strict upgrade over AVX2, is
| it not?
| cogman10 wrote:
| It's splitting a 512 instruction into 2 256 instructions
| internally. That's the main nasty surpise.
|
| I suppose it saves on the decoding portion a little but it's
| ultimately no more effective than just issuing the 2 256
| instructions yourself.
| MobiusHorizons wrote:
| The benefit seems to be that we are one step closer to not
| needing to have the fallback path. This was probably a lot
| more relevant before Intel shit the bed with consumer avx-512
| with e-cores not having the feature
| convolvatron wrote:
| axv-512 for zen4 also includes a bunch of instructions that
| weren't in 256, including enhanced masking, 16 bit floats,
| bit instructions, double-sized double-width register file
| rwaksmunski wrote:
| Every Rust SIMD article should mention the .chunks_exact() auto
| vectorization trick by law.
| ChadNauseam wrote:
| Didn't know about this. Thanks!
|
| Not related, but I often want to see the next or previous
| element when I'm iterating. When that happens, I always have to
| switch to an index-based loop. Is there a function that returns
| Iter<Item=(T, Option<T>)> where the second element is a
| lookahead?
| formerly_proven wrote:
| Lazy man's "kinda good enough for some cases SIMD in pure Rust"
| is to simply target x86-64-v3 (RUSTFLAGS=-Ctarget-cpu=x86-64-v3),
| which is supported by all AMD Zen and Intel CPUs since Haswell;
| and for floating point code, which cannot be auto-vectorized due
| to the accuracy implications, "simply" write it with explicit
| four or eight-way lanes, and LLVM will do the rest. Usually.
| Loops may need explicit handling of head or tail to auto-
| vectorize (chunks_exact helps with this, it hands you the tail).
| dfajgljsldkjag wrote:
| The benchmarks on Zen 5 are absolutely insane for just a bit of
| extra work. I really hope the portable SIMD module stabilizes
| soon, so we do not have to keep rewriting the same logic for NEON
| and AVX every time we want to optimize something. That example
| about implementing ChaCha20 twice really hit home for me.
___________________________________________________________________
(page generated 2026-01-21 23:00 UTC)