[HN Gopher] The death of thread per core
       ___________________________________________________________________
        
       The death of thread per core
        
       Author : ibobev
       Score  : 61 points
       Date   : 2025-10-20 21:19 UTC (1 days ago)
        
 (HTM) web link (buttondown.com)
 (TXT) w3m dump (buttondown.com)
        
       | vacuity wrote:
       | There are no hard rules; use principles flexibly.
       | 
       | That being said, there are some things that are generally true
       | for the long term: use a pinned thread per core, maximize
       | locality (of data and code, wherever relevant), use asynchronous
       | programming if performance is necessary. To incorporate the OP,
       | give control where it's due to each entity (here, the scheduler).
       | Cross-core data movement was never the enemy, but unprincipled
       | cross-core data movement can be. If even distribution of work is
       | important, work-stealing is excellent, as long as it's done
       | carefully. Details like how concurrency is implemented (shared-
       | state, here) or who controls the data are specific to the
       | circumstances.
        
         | AaronAPU wrote:
         | I did mass scale performance benchmarking on highly optimized
         | workloads using lockfree queues and fibers, and locking to a
         | core almost never was faster. There were a few topologies where
         | it was, but they were outliers.
         | 
         | This was on a wide variety of intel, AMD, NUMA, ARM processors
         | with different architectures, OSes and memory configurations.
         | 
         | Part of the reason is hyper threading (or threadripper type
         | archs) but even locking to groups wasn't usually faster.
         | 
         | This was even moreso the case when you had competing workloads
         | stealing cores from the OS scheduler.
        
           | zamadatix wrote:
           | I think workload might be as (if not more) the factor than
           | the uniqueness of the topology itself for how much pinning
           | matters. If your workload is purely computationally limited
           | then it doesn't matter. Same if it's actually I/O limited. If
           | it's memory bandwidth limited then it depends on things like
           | how much fits in per core cache vs shared cache vs going to
           | RAM, and how is RAM actually fed to the cores.
           | 
           | A really interesting niche is all of the performance
           | considerations around the design/use of VPP (Vector Packet
           | Processing) in the networking context. It's just one example
           | of a single niche, but it can give a good idea of how both
           | "changing the way the computation works" and "changing the
           | locality and pinning" can come together at the same time. I
           | forget the username but the person behind VPP is actually on
           | HN often, and a pretty cool guy to chat with.
           | 
           | Or, as vacuity put it, "there are no hard rules; use
           | principles flexibly".
        
           | jandrewrogers wrote:
           | Most high-performance workloads are limited by memory-
           | bandwidth these days. Even in HPC that became the primary
           | bottleneck for a large percentage of workloads in the 2000s.
           | High-performance data infrastructure is largely the same. You
           | can drive 200 GB/s of I/O on a server in real systems today.
           | 
           | The memory-bandwidth bound cases is where thread-per-core
           | tends to shine. It was the problem in HPC that thread-per-
           | core was invented to solve and it empirically had significant
           | performance benefits. Today we use it in high-scale databases
           | and other I/O intensive infrastructure if performance and
           | scalability are paramount.
           | 
           | That said, it is an architecture that does not degrade
           | gracefully. I've seen more thread-per-core implementations in
           | the wild that were broken by design than ones that were
           | implemented correctly. It requires a commitment to rigor and
           | thoroughness in the architecture that most software devs are
           | not used to.
        
       | bob1029 wrote:
       | I look at cross core communication as a 100x latency penalty.
       | Everything follows from there. The dependencies in the workload
       | ultimately determine how it should be spread across the cores (or
       | not!). The real elephant in the room is that oftentimes it's much
       | faster to just do the whole job on a single core even if you have
       | 255 others available. Some workloads do not care what kind of
       | clever scheduler you have in hand. If everything constantly
       | depends on the prior action you will never get any uplift.
       | 
       | You see this most obviously (visually) in places like game
       | engines. In Unity, the difference between non-burst and burst-
       | compiled code is very extreme. The difference between single and
       | multi core for the job system is often irrelevant by comparison.
       | If the amount of cpu time being spent on each job isn't high
       | enough, the benefit of multicore evaporates. Sending a job to be
       | ran on the fleet has a lot of overhead. It has to be worth that
       | one time 100x latency cost both ways.
       | 
       | The GPU is the ultimate example of this. There are some workloads
       | that benefit dramatically from the incredible parallelism. Others
       | are entirely infeasible by comparison. This is at the heart of my
       | problem with the current machine learning research paradigm. Some
       | ML techniques are _terrible_ at running on the GPU, but it seems
       | as if we 've convinced ourselves that GPU is a prerequisite for
       | any kind of ML work. It all boils down to the latency of the
       | compute. Getting data in and out of a GPU takes an eternity
       | compared to L1. There are other fundamental problems with GPUs
       | (warp divergence) that preclude clever workarounds.
        
         | bsenftner wrote:
         | Astute points. I've worked on an extremely performant facial
         | recognition system (tens of millions of face compares per
         | second per core) that lives in L1 and does not use the GPU for
         | the FR inference at all, only for the display of the video and
         | the tracked people within. I rarely even bother telling
         | ML/DL/AI people it does not use the GPU, because I'm just tired
         | of the argument that "we're doing it wrong".
        
         | kiitos wrote:
         | > I look at cross core communication as a 100x latency penalty
         | 
         | if your workload is majority cpu-bound then this is true,
         | _sometimes_ , and _at best_
         | 
         | most workloads are io (i.e. syscall) bound, and io/syscall
         | overhead is >> cross-core communication overhead
        
         | dist-epoch wrote:
         | The thing with GPUs is that for many problems really dumb and
         | simple algorithms (think bubble sort equivalent) are many times
         | faster than very fancy CPU algorithms (think quick sort
         | equivalent). Your typical non-neural-network GPU algorithm is
         | rarely using more than 50% of it's power, yet still outperforms
         | carefully written CPU algorithms.
        
         | anyfoo wrote:
         | > If everything constantly depends on the prior action you will
         | never get any uplift.
         | 
         | I mean... that's kind of a pathological case, no?
        
         | pron wrote:
         | That's fine, but a work-stealing scheduler doesn't redistribute
         | work willy-nilly. Locally-submitted tasks are likely to remain
         | local, and are generally stolen when stealing does pay off. If
         | everything is more-or-less evenly distributed, you'll get
         | little or no stealing.
         | 
         | That's not to say it's perfect. The problem is in anticipating
         | how much workload is about to arrive and deciding how many
         | worker threads to spawn. If you overestimate and have too many
         | worker threads running, you will get wasteful stealing; if
         | you're overly conservative and slow to respond to growing
         | workload (to avoid over-stealing), you'll wait for threads to
         | spawn and hurt your latencies just as the workload begins to
         | spike.
        
           | vlovich123 wrote:
           | There's secondary costs though - because you might run on any
           | thread you have to sprinkle atomics and/or mutexes all over
           | the place (in Rust parlance the tasks spawned must be Send)
           | which have all sorts of implicit performance costs that stack
           | up even if you never transfer the task.
           | 
           | In other words, you could probably easily do 10m op/s per
           | core on a thread per core design but struggle to get 1m op/s
           | on a work stealing design. And the work stealing will be
           | total throughput for the machine whereas the 10m op/s design
           | will generally continue scaling with the number of CPUs.
        
         | Archit3ch wrote:
         | > If everything constantly depends on the prior action you will
         | never get any uplift.
         | 
         | Not always. For differential equations with large enough
         | matrices, the independent work each core can do outperforms the
         | communication overhead of core-to-core latency.
        
       | josefrichter wrote:
       | Isn't this what Erlang/Elixir BEAM is all about?
        
         | ameliaquining wrote:
         | How so? AFAIK BEAM is pretty much agnostic between work-
         | stealing and work-sharding* architectures.
         | 
         | * I prefer the term "work-sharding" over "thread-per-core",
         | because work-stealing architectures usually also use one thread
         | per core, so it tends to confuse people.
        
       | adsharma wrote:
       | Morsel driven parallelism is working great in DuckDB, KuzuDB and
       | now Ladybug (fork of Kuzu after archival).
        
       | jandrewrogers wrote:
       | I've worked on several thread-per-core systems that were purpose-
       | built for extreme dynamic data and load skew. They work
       | beautifully at very high scales on the largest hardware. The
       | mechanics of how you design thread-per-core systems that provide
       | uniform distribution of load without work-stealing or high-touch
       | thread coordination have idiomatic architectures at this point.
       | People have been putting thread-per-core architectures in
       | production for 15+ years now and the designs have evolved
       | dramatically.
       | 
       | The architectures from circa 2010 were a bit rough. While the
       | article has some validity for architectures from 10+ years ago,
       | the state-of-the-art for thread-per-core today looks nothing like
       | those architectures and largely doesn't have the issues raised.
       | 
       | News of thread-per-core's demise has been greatly exaggerated.
       | The benefits have measurably increased in practice as the
       | hardware has evolved, especially for ultra-scale data
       | infrastructure.
        
         | touisteur wrote:
         | I feel I'm still doing it the old 2010 way, with all my hand-
         | crafted dpdk-and-pipelines-and-lockless-queues-and-homemade-
         | taskgraph-scheduler, any modern reference (apart from 'use
         | seastar' ? ... which fair if it fills your needs) ?
        
         | FridgeSeal wrote:
         | Are there any resources/learning material about the more modern
         | thread-per-core approaches? It's a particular area of interest
         | for me, but I've had relatively little success finding more
         | learning material, so I assume there's lots of tightly guarded
         | institutional knowledge.
        
       ___________________________________________________________________
       (page generated 2025-10-21 23:01 UTC)