JOURNAL

Kubernetes Just Flipped a Default That Will Refuse to Boot Your Old Nodes

Kubelet's failCgroupV1 field defaulted to true in 1.35. A cgroup v1-only node won't start anymore unless you explicitly override it. Here's how to check today.

Read with AI

Choose content to copy and paste into your AI assistant. Nothing is sent automatically. CMS content is converted to Markdown; original Markdown is used when available.

A kubelet config field most people have never typed, `failCgroupV1`, flipped its default from false to true in Kubernetes 1.35. That’s not a removal. It’s quieter and, for a lot of clusters, more disruptive: a node still running cgroup v1 will now refuse to start unless you explicitly override the field. If your node images were built two or three years ago and nobody’s touched the base image since, this is the kind of change that shows up as a failed upgrade at 2am, not as a line item anyone budgeted time for.

The relevant proposal is KEP-5573, “Remove cgroup v1 support,” which builds on KEP-4569 (cgroup v1 moved to maintenance mode, GA since 1.31). I want to correct a detail that gets repeated inaccurately in a lot of secondary coverage: the KEP does not commit to removing v1 support in version 1.38. The actual text says removal “will be done no earlier than 1.38, to maintain the k8s deprecation policy.” There’s an unresolved marker in the KEP itself where the removal plan is still being debated among SIG Node maintainers. So the honest framing is a floor, not a date: nothing before 1.38, no committed target after it either.

Why the default flip matters more than the eventual removal

The uncertain removal date is almost a distraction from the thing that already shipped. `failCgroupV1` defaulting to true in 1.35 means that today, right now, upgrading kubelet on an unaudited cgroup v1 node can take that node out of your cluster. You don’t get a deprecation warning cycle for this one the way you might expect — the kubelet just won’t come up cleanly on v1-only hosts unless the override is set. Check this before your next kubelet upgrade, not after.

The check itself is one command, run on the node:

stat -fc %T /sys/fs/cgroup/

`cgroup2fs` means you’re on v2 and fine. `tmpfs` means the node is still on v1 and you need to plan a migration before you touch that kubelet version. Run this across your node fleet now, not as part of incident response later. It’s a five-minute audit that turns a surprise outage into a scheduled migration.

What the migration actually requires

Cgroup v2 needs a Linux kernel 5.8 or newer, containerd 1.4 or newer (or CRI-O 1.20 or newer), and the kubelet’s cgroup driver set to `systemd` rather than the older `cgroupfs` driver — Kubernetes’ own container runtime docs are explicit that cgroupfs is the wrong choice once you’re on v2. Most mainstream distributions shipped v2 as default years ago, so if your node images come from a current Ubuntu, Debian, or RHEL release, you’re probably already fine. The risk sits specifically with older, frozen golden images — the ones built once and never rebuilt because “it still works.”

Flow diagram: check cgroup version with stat, if the result is tmpfs the node is still cgroup v1, upgrade the container runtime to containerd 1.4+ or CRI-O 1.20+, set cgroupDriver to systemd, rebuild the node image on kernel 5.8 or newer, then verify Memory QoS and PSI are working since both require cgroup v2.
Kubelet’s failCgroupV1 field defaulted to true in Kubernetes 1.35 — a node that’s still cgroup v1-only will refuse to start unless you override it.

What you lose by staying on v1 even if the override keeps working

Even setting the override aside, cgroup v1 is quietly falling behind on features that matter for production stability. Memory QoS (KEP-2570, beta and default-on since 1.37) explicitly requires cgroup v2 — it’s unavailable on v1 nodes, full stop. Pressure Stall Information metrics need cgroup v2 as well; the feature gate for PSI has been locked to true since 1.36, but it produces nothing useful on a v1 host. And OOM handling genuinely improves under v2: it can kill an entire process group as one unit through `memory.oom.group`, where v1 can only kill individual processes, which tends to leave containers in a half-dead state instead of cleanly restarting.

Managed Kubernetes isn’t uniformly ahead of you here

If you’re on a managed service, check your specific provider rather than assuming they’ve handled this. GKE has published the most concrete timeline of the three major clouds: cgroup v2 became default for new node pools in 1.26, v1 was marked deprecated at 1.31, existing clusters got auto-migrated at 1.33, and GKE drops v1 support entirely at 1.35. EKS, by contrast, has not published a removal date — AWS’s own AMI deprecation FAQ says explicitly that a date for full removal hasn’t been announced yet, while still recommending migration to AL2023, Bottlerocket, RHEL9+, Ubuntu 22.04+, or Debian 11+. AKS confirms v2 became the default back in Kubernetes 1.25 and offers a rollback DaemonSet, but I couldn’t find a published forward-looking removal timeline from Microsoft to cite here, so I’m not going to invent one.

If you run self-managed nodes — kOps, kubeadm, bare metal — none of those managed-service timelines apply to you anyway, and the `failCgroupV1` default flip in 1.35 is the thing that actually bites first. Run the `stat` check across your fleet this week. It costs five minutes per node and the alternative is finding out during your next kubelet upgrade, which is a worse time to find out.

Discussion

Comments are reviewed before publication. Your email is kept private.

← Back to allĐọc tiếng Việt