[HN Gopher] The death of thread per core
___________________________________________________________________
The death of thread per core
Author : ibobev
Score : 61 points
Date : 2025-10-20 21:19 UTC (1 days ago)
(HTM) web link (buttondown.com)
(TXT) w3m dump (buttondown.com)
| vacuity wrote:
| There are no hard rules; use principles flexibly.
|
| That being said, there are some things that are generally true
| for the long term: use a pinned thread per core, maximize
| locality (of data and code, wherever relevant), use asynchronous
| programming if performance is necessary. To incorporate the OP,
| give control where it's due to each entity (here, the scheduler).
| Cross-core data movement was never the enemy, but unprincipled
| cross-core data movement can be. If even distribution of work is
| important, work-stealing is excellent, as long as it's done
| carefully. Details like how concurrency is implemented (shared-
| state, here) or who controls the data are specific to the
| circumstances.
| AaronAPU wrote:
| I did mass scale performance benchmarking on highly optimized
| workloads using lockfree queues and fibers, and locking to a
| core almost never was faster. There were a few topologies where
| it was, but they were outliers.
|
| This was on a wide variety of intel, AMD, NUMA, ARM processors
| with different architectures, OSes and memory configurations.
|
| Part of the reason is hyper threading (or threadripper type
| archs) but even locking to groups wasn't usually faster.
|
| This was even moreso the case when you had competing workloads
| stealing cores from the OS scheduler.
| zamadatix wrote:
| I think workload might be as (if not more) the factor than
| the uniqueness of the topology itself for how much pinning
| matters. If your workload is purely computationally limited
| then it doesn't matter. Same if it's actually I/O limited. If
| it's memory bandwidth limited then it depends on things like
| how much fits in per core cache vs shared cache vs going to
| RAM, and how is RAM actually fed to the cores.
|
| A really interesting niche is all of the performance
| considerations around the design/use of VPP (Vector Packet
| Processing) in the networking context. It's just one example
| of a single niche, but it can give a good idea of how both
| "changing the way the computation works" and "changing the
| locality and pinning" can come together at the same time. I
| forget the username but the person behind VPP is actually on
| HN often, and a pretty cool guy to chat with.
|
| Or, as vacuity put it, "there are no hard rules; use
| principles flexibly".
| jandrewrogers wrote:
| Most high-performance workloads are limited by memory-
| bandwidth these days. Even in HPC that became the primary
| bottleneck for a large percentage of workloads in the 2000s.
| High-performance data infrastructure is largely the same. You
| can drive 200 GB/s of I/O on a server in real systems today.
|
| The memory-bandwidth bound cases is where thread-per-core
| tends to shine. It was the problem in HPC that thread-per-
| core was invented to solve and it empirically had significant
| performance benefits. Today we use it in high-scale databases
| and other I/O intensive infrastructure if performance and
| scalability are paramount.
|
| That said, it is an architecture that does not degrade
| gracefully. I've seen more thread-per-core implementations in
| the wild that were broken by design than ones that were
| implemented correctly. It requires a commitment to rigor and
| thoroughness in the architecture that most software devs are
| not used to.
| bob1029 wrote:
| I look at cross core communication as a 100x latency penalty.
| Everything follows from there. The dependencies in the workload
| ultimately determine how it should be spread across the cores (or
| not!). The real elephant in the room is that oftentimes it's much
| faster to just do the whole job on a single core even if you have
| 255 others available. Some workloads do not care what kind of
| clever scheduler you have in hand. If everything constantly
| depends on the prior action you will never get any uplift.
|
| You see this most obviously (visually) in places like game
| engines. In Unity, the difference between non-burst and burst-
| compiled code is very extreme. The difference between single and
| multi core for the job system is often irrelevant by comparison.
| If the amount of cpu time being spent on each job isn't high
| enough, the benefit of multicore evaporates. Sending a job to be
| ran on the fleet has a lot of overhead. It has to be worth that
| one time 100x latency cost both ways.
|
| The GPU is the ultimate example of this. There are some workloads
| that benefit dramatically from the incredible parallelism. Others
| are entirely infeasible by comparison. This is at the heart of my
| problem with the current machine learning research paradigm. Some
| ML techniques are _terrible_ at running on the GPU, but it seems
| as if we 've convinced ourselves that GPU is a prerequisite for
| any kind of ML work. It all boils down to the latency of the
| compute. Getting data in and out of a GPU takes an eternity
| compared to L1. There are other fundamental problems with GPUs
| (warp divergence) that preclude clever workarounds.
| bsenftner wrote:
| Astute points. I've worked on an extremely performant facial
| recognition system (tens of millions of face compares per
| second per core) that lives in L1 and does not use the GPU for
| the FR inference at all, only for the display of the video and
| the tracked people within. I rarely even bother telling
| ML/DL/AI people it does not use the GPU, because I'm just tired
| of the argument that "we're doing it wrong".
| kiitos wrote:
| > I look at cross core communication as a 100x latency penalty
|
| if your workload is majority cpu-bound then this is true,
| _sometimes_ , and _at best_
|
| most workloads are io (i.e. syscall) bound, and io/syscall
| overhead is >> cross-core communication overhead
| dist-epoch wrote:
| The thing with GPUs is that for many problems really dumb and
| simple algorithms (think bubble sort equivalent) are many times
| faster than very fancy CPU algorithms (think quick sort
| equivalent). Your typical non-neural-network GPU algorithm is
| rarely using more than 50% of it's power, yet still outperforms
| carefully written CPU algorithms.
| anyfoo wrote:
| > If everything constantly depends on the prior action you will
| never get any uplift.
|
| I mean... that's kind of a pathological case, no?
| pron wrote:
| That's fine, but a work-stealing scheduler doesn't redistribute
| work willy-nilly. Locally-submitted tasks are likely to remain
| local, and are generally stolen when stealing does pay off. If
| everything is more-or-less evenly distributed, you'll get
| little or no stealing.
|
| That's not to say it's perfect. The problem is in anticipating
| how much workload is about to arrive and deciding how many
| worker threads to spawn. If you overestimate and have too many
| worker threads running, you will get wasteful stealing; if
| you're overly conservative and slow to respond to growing
| workload (to avoid over-stealing), you'll wait for threads to
| spawn and hurt your latencies just as the workload begins to
| spike.
| vlovich123 wrote:
| There's secondary costs though - because you might run on any
| thread you have to sprinkle atomics and/or mutexes all over
| the place (in Rust parlance the tasks spawned must be Send)
| which have all sorts of implicit performance costs that stack
| up even if you never transfer the task.
|
| In other words, you could probably easily do 10m op/s per
| core on a thread per core design but struggle to get 1m op/s
| on a work stealing design. And the work stealing will be
| total throughput for the machine whereas the 10m op/s design
| will generally continue scaling with the number of CPUs.
| Archit3ch wrote:
| > If everything constantly depends on the prior action you will
| never get any uplift.
|
| Not always. For differential equations with large enough
| matrices, the independent work each core can do outperforms the
| communication overhead of core-to-core latency.
| josefrichter wrote:
| Isn't this what Erlang/Elixir BEAM is all about?
| ameliaquining wrote:
| How so? AFAIK BEAM is pretty much agnostic between work-
| stealing and work-sharding* architectures.
|
| * I prefer the term "work-sharding" over "thread-per-core",
| because work-stealing architectures usually also use one thread
| per core, so it tends to confuse people.
| adsharma wrote:
| Morsel driven parallelism is working great in DuckDB, KuzuDB and
| now Ladybug (fork of Kuzu after archival).
| jandrewrogers wrote:
| I've worked on several thread-per-core systems that were purpose-
| built for extreme dynamic data and load skew. They work
| beautifully at very high scales on the largest hardware. The
| mechanics of how you design thread-per-core systems that provide
| uniform distribution of load without work-stealing or high-touch
| thread coordination have idiomatic architectures at this point.
| People have been putting thread-per-core architectures in
| production for 15+ years now and the designs have evolved
| dramatically.
|
| The architectures from circa 2010 were a bit rough. While the
| article has some validity for architectures from 10+ years ago,
| the state-of-the-art for thread-per-core today looks nothing like
| those architectures and largely doesn't have the issues raised.
|
| News of thread-per-core's demise has been greatly exaggerated.
| The benefits have measurably increased in practice as the
| hardware has evolved, especially for ultra-scale data
| infrastructure.
| touisteur wrote:
| I feel I'm still doing it the old 2010 way, with all my hand-
| crafted dpdk-and-pipelines-and-lockless-queues-and-homemade-
| taskgraph-scheduler, any modern reference (apart from 'use
| seastar' ? ... which fair if it fills your needs) ?
| FridgeSeal wrote:
| Are there any resources/learning material about the more modern
| thread-per-core approaches? It's a particular area of interest
| for me, but I've had relatively little success finding more
| learning material, so I assume there's lots of tightly guarded
| institutional knowledge.
___________________________________________________________________
(page generated 2025-10-21 23:01 UTC)