For the love of god stop using CPU limits in Kubernetes
The analysis compares the same app running with and without a Kubernetes CPU limit, keeping CPU requests identical. It argues that CPU limits act as a hard ceiling that wastes idle node CPU and throttles the app many times per second, even when the node has free capacity. CPU requests, by contrast, guarantee a share of CPU and allow borrowing idle CPU without taking from anyone else. Memory limits are still recommended because they protect the node.
The underlying mechanism involves cgroup settings: cpu.weight controls the request (fair sharing) and cpu.max controls the limit (quota). When a limit is set, the kernel throttles the app's CPU usage, pausing it repeatedly. This throttling is what causes the observed performance degradation.
Measured results include: tail latency holds up under traffic peaks instead of collapsing, and CPU-bound startup work completes about 2x sooner after removing limits. Additionally, most clusters reserve far more CPU than they actually use at peak, so after removing limits and right-sizing requests, a meaningful number of nodes can be removed. An illustrative cost model estimates savings in the tens of thousands of dollars per year per cluster.
The analysis is structured into 12 sections covering requests vs limits, throttling mechanics, CFS fair sharing, noisy neighbors, rollout steps, and more. It links to longer supporting documents on theory, .NET runtime effects, Postgres specifics, counter-objections, cost modeling, and a staged rollout plan with a one-line rollback.