[HN Gopher] Multiplatform Matrix Multiplication Kernels
___________________________________________________________________
Multiplatform Matrix Multiplication Kernels
Author : homarp
Score : 38 points
Date : 2025-07-18 19:59 UTC (3 hours ago)
(HTM) web link (burn.dev)
(TXT) w3m dump (burn.dev)
| raphaelty wrote:
| Very interesting, willing to try burn
| nathanielsimard wrote:
| One of the author here, don't hesitate if you have any question
| or comment!
| almostgotcaught wrote:
| I'm sorry this is a low brow comment but this is the dumbest
| thing you can do in this space:
|
| > Unit (thread in CUDA, invocation in Vulkan/Wgpu): the smallest
| execution entity performing computations.
|
| > Plane (warp in CUDA, subgroup in Vulkan/Wgpu): a group of
| (typically 32) units executing in lockstep and able to share data
| efficiently through registers.
|
| > Cube (thread block in CUDA, workgroup in Vulkan/Wgpu): a group
| of units that execute on the same SM, sharing memory and able to
| synchronize
|
| It's already bad enough that the vendors themselves insisted on
| different names but why in the bejesus would you rename these
| concepts and diverge from literally all existing naming
| conventions when you're providing middleware. Ie when using your
| tool I'm still going to reference NVIDIA's or AMD's docs to
| understand how the hardware actually works. Like do you really
| think otherwise - that your thing is gonna be end of the line???
|
| FYI the word warp isn't random techno babble but is actually a
| very clever pun that actually fits very well conceptually:
|
| https://en.m.wikipedia.org/wiki/Warp_and_weft
| nathanielsimard wrote:
| Using the naming from one of the existing API would put too
| much bias towards that API. It started as a WebGPU project
| early on, but some features are not present so mixing terms
| wasn't ideal. We're also working on extending CubeCL to CPU, so
| we want terms not only tied to the GPU word.
| almostgotcaught wrote:
| Thread, group, workgroup.
|
| There you go you've hit basically two of 3 completely (AMD
| and Vulkan) and are close enough to CUDA that people would
| get it.
|
| I have no idea what a plane connotes and a cube literally
| gives a distinct enough picture from block that I will be
| continuously reminding myself of the mapping.
|
| What you did was pointless - you assigned new words to
| objects that you don't own and now your conceptual framework
| is askew from the actual underlying (true) conceptual
| framework.
|
| > CubeCL to CPU
|
| There is zero affinity between GPU programing models and
| multicore CPU programing models. If you don't believe me go
| ask the OpenMP people how they're doing supporting GPUs.
| nathanielsimard wrote:
| Well we can agree to disagree, CubeCL also has the concept
| of instruction parallelism, which would be used to target
| simd instructions on CPU. Our algorithms are normally
| flexible on both the plane size and the line size, adapting
| to the hardware with comptime logique. You are free to
| dislike the naming, but imo a mix of multiple APIs is worse
| than something new.
| almostgotcaught wrote:
| > Our algorithms are normally flexible on both the plane
| size and the line size
|
| Congrats - I have no idea what this means lol.
| sroussey wrote:
| Why unit instead of point?
|
| Unit, plane (as vs train), and cube?
|
| Or point, plane, cube (1d, 2d, 3d)?
| nathanielsimard wrote:
| I don't recall the reason why, point is a valid name.
| airstrike wrote:
| burn is awesome
| Lerc wrote:
| Has there been much research into slightly flawed matrix
| multiplications?
|
| If you have a measure of correctness, and a measure of
| performance. Is there a maximum value of correctness per some
| unit of processing that exists below a full matrix multiply
|
| Obviously it can be done with precision, since that is what
| floating point is. But is there anything where you can save x% of
| computation and have fewer than x% incorrect values in a matrix
| multiplications?
|
| Gradient descent wouldn't really care about a few (Reliably) dud
| values.
___________________________________________________________________
(page generated 2025-07-18 23:00 UTC)