[HN Gopher] Building Burstables: CPU slicing with cgroups
       ___________________________________________________________________
        
       Building Burstables: CPU slicing with cgroups
        
       Author : msarnowicz
       Score  : 79 points
       Date   : 2025-05-02 16:45 UTC (6 hours ago)
        
 (HTM) web link (www.ubicloud.com)
 (TXT) w3m dump (www.ubicloud.com)
        
       | msarnowicz wrote:
       | Hey, author here. Please AMA.
       | 
       | I came into the Linux world via Postgres, and this was an
       | interesting project for me learning more about Linux internals.
       | While cgroups v2 do offer basic support for CPU bursting, the
       | bursts are short-lived, and credits don't persist beyond sub-
       | second intervals. If you've run into scenarios where more
       | adaptive or sustained bursting would help, we'd love to hear
       | about them. Knowing your use cases will help shape what we build
       | next.
        
         | parrit wrote:
         | Thanks! That was a pleasant read. I have been wanting to mess
         | with cgroups for a while, in order to hack together a "docker"
         | like many have done before to understand it better. This will
         | help!
         | 
         | Are there typical use cases where you reach for cgroups
         | directly instead of using the container abstraction?
        
       | jauntywundrkind wrote:
       | I'd also strongly recommend this view of how Kubernetes uses
       | cgroups, showing similar drill downs for how everything gets
       | managed. Lovely view of what's really happening!
       | https://martinheinz.dev/blog/91
       | 
       | I've been a bit apoplectic in the past that cgroups seemed not
       | super helpful in Kubernetes, but this really showed me how the
       | different Kubernetes QoS levels are driven by similar juggling of
       | different cgroups.
       | 
       | I'm not sure if this makes use of _cpu.max.burst_ or not. There
       | 's a fun article that monkeys with these cgroups directly, which
       | is neat to see. It also links to an ask that Kubernetes get
       | support for the new (5.14) CFS Burst system. Which is a whole
       | nother fun rabbit hole of fair share bursting to go down!
       | https://medium.com/@christian.cadieux/kubernetes-throttling-...
       | https://github.com/kubernetes/kubernetes/issues/104516
        
         | msarnowicz wrote:
         | Thank you, that is a good perspective, too!
        
         | __turbobrew__ wrote:
         | cpu.max.burst increases the chances of noisy neighbours
         | stealing CPU from other tenants.
         | 
         | I run multi-tenant k8s clusters with hundreds of tenants and it
         | fundamentally is a hard problem to balance workload performance
         | with efficiency. Sharing resources increases efficiency but in
         | most cases increases tail latencies.
        
           | jeffbee wrote:
           | If you use k8s qos levels "guaranteed" cpu resources will be
           | distinct -- via cpu sets -- from the ones used by the riff-
           | raff. This is a good way to segregate latency-sensitive apps
           | where you care about latency from throughtput-oriented stuff
           | where you don't.
        
           | hinkley wrote:
           | I suspect you can only really count on neighbors to take care
           | of their own. Anything else they see will be taken as an
           | entitlement.
           | 
           | So for instance if you run three processes for the same
           | customer, can you set them to use the same cpu slices and
           | deal with one of their apps occasionally needing a burst of
           | CPU?
        
         | msarnowicz wrote:
         | Reading through the description of how cgroups are used in
         | Kubernetes, I can see some similarities and some differences as
         | well. It is interesting to compare the approaches.
         | 
         | We chose not to use cpu.weight, and instead divide the host
         | explicitly using cgroups (slice in systemd). We put Standard
         | VMs in dedicated slices to keep them isolated and let several
         | Burstable VMs share a slice. This provides a trade off between
         | the price of the VM and resource guarantees.
         | 
         | We use cpu.max.burst to allow the VMs to "expand" a bit, while
         | we understand that this creates a "noisy neighbor" problem. At
         | the same time there is a minimum guarantee of the CPU. The
         | cgroups allow for all those knobs and give a lot of control.
         | Combining them in various ways is an interesting puzzle.
        
       ___________________________________________________________________
       (page generated 2025-05-02 23:00 UTC)