Software

Kubernetes Shifts to cgroup v2 as Edge AI and Density Demands Reshape Cloud-Native Infrastructure

As Kubernetes deployments expand to edge environments and AI workloads, the ecosystem is moving away from deprecated Linux kernel interfaces while tackling memory constraints and self-hosted platform complexity.

4 min read
Kubernetes on cgroup v1 is dead. Here’s what comes next.

The cloud-native landscape continues to evolve ahead of KubeCon + CloudNativeCon North America 2026, scheduled for November 9-12 in Salt Lake City, Utah. This year's agenda reflects a fundamental shift: enterprises are pushing Kubernetes beyond its traditional data-center role into edge environments where AI inference runs closer to data sources, reducing latency and compliance risks.

Purpose-built hardware emerges as edge AI accelerator

Running AI inference at the edge offers tangible benefits—fewer calls back to centralized servers, lower data egress costs, and stronger data residency for regulated industries. Yet many organizations attempt to deploy AI workloads on generic infrastructure, creating operational friction. Aaron Lamond, from Hewlett Packard Enterprise, argues that "as intelligence becomes more distributed, purpose-built compute becomes increasingly critical to operational success." HPE's approach emphasizes hardware and software co-design, with ProLiant edge servers engineered for resource-constrained deployments that maintain enterprise-grade security.

Kubernetes on Edge Day addresses operational reality gap

Edge Kubernetes deployments have matured significantly, with CNCF projects like KubeEdge and numerous vendor platforms now supporting distributed architectures. This year's conference will feature a dedicated Kubernetes on Edge Day co-located event. Mars Toktonaliev and Katerina Arzhayev note on the CNCF blog that "the event was created to address the gap between traditional, data-center-centric cloud native approaches and the operational realities of edge computing." The schedule will feature operational guidance and case studies tailored for engineers managing Kubernetes across distributed, resource-constrained environments, with particular emphasis on observability and security.

Node swap feature triples cluster density for bursty workloads

Agentic AI workloads create unpredictable memory demands—sudden spikes followed by idle periods. This pattern strains Kubernetes clusters, where memory typically becomes the limiting factor before CPU. Ocean Xie and Yuan Wang explain on the Kubernetes blog: "Memory is often the first hard limit a Kubernetes cluster hits. Nodes run out of RAM long before they run out of CPU, and the new wave of agentic AI workloads makes this worse." The solution involves enabling node swap, a feature that reached general availability in Kubernetes v1.34. When backed by fast NVMe SSDs, node swap can deliver density improvements up to three times in certain scenarios, acting as a buffer during traffic surges while reducing overall infrastructure costs.

Self-hosted deployments demand pre-installation preparation

Fairwinds, a managed cloud-native infrastructure provider, cautions that shipping a working installer represents only the beginning of a self-hosted platform deployment. Munib Ali, director of engineering, emphasizes that vendors often overlook critical prerequisites on the customer side: establishing team responsibilities, configuring identity and access management, planning maintenance windows, ensuring compatibility with private cloud infrastructure, and validating DNS configurations. Ali states: "Self-hosted Kubernetes deployments for AI platforms often stall when customer prerequisites, environment restrictions, and cross-team handoffs are incomplete." Without addressing these foundational elements before go-live, installation timelines slip or projects fail to launch. Making self-hosted delivery seamless requires treating preparation as a shared responsibility between vendor and customer.

cgroup v2 adoption becomes mandatory for modern Kubernetes

A significant architectural shift is underway in Kubernetes' relationship with Linux kernel resource management. Paco Xu, open source team lead at DaoCloud, explains that cgroup v1, the traditional interface for managing CPU and memory resources, carries substantial limitations. Xu writes: "Compared with cgroup v1, cgroup v2 provides a single unified hierarchy, a more consistent interface, and a stronger foundation for resource isolation and modern resource-management features." The newer version enables memory quality of service updates, container-aware out-of-memory handling, rootless container support, and additional capabilities depending on Kubernetes version and configuration. Starting with v1.35, the kubelet refuses to start on cgroup v1 nodes by default. Operators running older releases must migrate all Linux nodes to cgroup v2 before upgrading, making this transition unavoidable for staying current.

Cilium addresses GPU and AI networking challenges

Cilium, the CNCF-graduated eBPF-based networking and security project, will host a co-located event at this year's conference on November 9. Joe Stringer from Isovalent at Cisco and Jordan Rife from Google note on the CNCF blog that "this year's agenda goes straight at the problems that show up when AI and GPU workloads push Kubernetes networking past what it was built for." The community is actively responding to GPU resource demands, AI-driven observability requirements, and emerging security vulnerabilities. Recent case studies from Splunk, Celonis, and Preferred Networks demonstrate eBPF's expanding role in production security and AI workloads. Bill Mulligan, Cilium and eBPF community pollinator, highlights BpfJailer, an experimental open-source eBPF security tool from Meta that represents a rewrite of Meta's internal closed-source security implementation. The community continues to find new production use cases for both Cilium and eBPF technology.

Road to KubeCon is an eight-part series presented by HPE. The New Stack readers can access a 10% discount on conference tickets using code KCNA26MED10 at registration.

Source: The New Stack · Reporting supplemented by The Silicon Ledger staff.