[HN Gopher] Building Burstables: CPU slicing with cgroups
___________________________________________________________________
Building Burstables: CPU slicing with cgroups
Author : msarnowicz
Score : 79 points
Date : 2025-05-02 16:45 UTC (6 hours ago)
(HTM) web link (www.ubicloud.com)
(TXT) w3m dump (www.ubicloud.com)
| msarnowicz wrote:
| Hey, author here. Please AMA.
|
| I came into the Linux world via Postgres, and this was an
| interesting project for me learning more about Linux internals.
| While cgroups v2 do offer basic support for CPU bursting, the
| bursts are short-lived, and credits don't persist beyond sub-
| second intervals. If you've run into scenarios where more
| adaptive or sustained bursting would help, we'd love to hear
| about them. Knowing your use cases will help shape what we build
| next.
| parrit wrote:
| Thanks! That was a pleasant read. I have been wanting to mess
| with cgroups for a while, in order to hack together a "docker"
| like many have done before to understand it better. This will
| help!
|
| Are there typical use cases where you reach for cgroups
| directly instead of using the container abstraction?
| jauntywundrkind wrote:
| I'd also strongly recommend this view of how Kubernetes uses
| cgroups, showing similar drill downs for how everything gets
| managed. Lovely view of what's really happening!
| https://martinheinz.dev/blog/91
|
| I've been a bit apoplectic in the past that cgroups seemed not
| super helpful in Kubernetes, but this really showed me how the
| different Kubernetes QoS levels are driven by similar juggling of
| different cgroups.
|
| I'm not sure if this makes use of _cpu.max.burst_ or not. There
| 's a fun article that monkeys with these cgroups directly, which
| is neat to see. It also links to an ask that Kubernetes get
| support for the new (5.14) CFS Burst system. Which is a whole
| nother fun rabbit hole of fair share bursting to go down!
| https://medium.com/@christian.cadieux/kubernetes-throttling-...
| https://github.com/kubernetes/kubernetes/issues/104516
| msarnowicz wrote:
| Thank you, that is a good perspective, too!
| __turbobrew__ wrote:
| cpu.max.burst increases the chances of noisy neighbours
| stealing CPU from other tenants.
|
| I run multi-tenant k8s clusters with hundreds of tenants and it
| fundamentally is a hard problem to balance workload performance
| with efficiency. Sharing resources increases efficiency but in
| most cases increases tail latencies.
| jeffbee wrote:
| If you use k8s qos levels "guaranteed" cpu resources will be
| distinct -- via cpu sets -- from the ones used by the riff-
| raff. This is a good way to segregate latency-sensitive apps
| where you care about latency from throughtput-oriented stuff
| where you don't.
| hinkley wrote:
| I suspect you can only really count on neighbors to take care
| of their own. Anything else they see will be taken as an
| entitlement.
|
| So for instance if you run three processes for the same
| customer, can you set them to use the same cpu slices and
| deal with one of their apps occasionally needing a burst of
| CPU?
| msarnowicz wrote:
| Reading through the description of how cgroups are used in
| Kubernetes, I can see some similarities and some differences as
| well. It is interesting to compare the approaches.
|
| We chose not to use cpu.weight, and instead divide the host
| explicitly using cgroups (slice in systemd). We put Standard
| VMs in dedicated slices to keep them isolated and let several
| Burstable VMs share a slice. This provides a trade off between
| the price of the VM and resource guarantees.
|
| We use cpu.max.burst to allow the VMs to "expand" a bit, while
| we understand that this creates a "noisy neighbor" problem. At
| the same time there is a minimum guarantee of the CPU. The
| cgroups allow for all those knobs and give a lot of control.
| Combining them in various ways is an interesting puzzle.
___________________________________________________________________
(page generated 2025-05-02 23:00 UTC)