Persistent volume issues

On CFKE, you install and own the CSI drivers for persistent volumes (see storage). This page covers the volume problems that reach support most often.

Volume stuck attaching or mounting (Hetzner)

Symptoms:

  • Pod stuck in ContainerCreating with FailedAttachVolume: volume attachment is being deleted
  • FailedMount with a missing device path such as /dev/disk/by-id/scsi-0HC_Volume_VOLUME_ID
  • Multi-Attach error for volume when a Pod moves between nodes

Fixes:

  1. Keep the CSI driver current. Older hcloud-csi releases contain attach and mount races fixed upstream. Version v2.21.2 or later fixes the “device needs to be ready on node publish” failure behind the missing /dev/disk/by-id path:

    bash
    helm upgrade --install hcloud-csi hcloud/hcloud-csi -n kube-system
  2. Clear a stale VolumeAttachment. When a volume moves between nodes, a known upstream race can leave a VolumeAttachment object behind that blocks re-attachment. Deleting the stale object is non-destructive (it does not touch your data). It triggers a clean re-attach:

    bash
    kubectl get volumeattachments
    kubectl delete volumeattachment ATTACHMENT_NAME

Volumes and node placement

Block volumes are zonal: you can attach a volume only to nodes in the location where you created it. The scheduler and the node auto-provisioner respect volume topology automatically, as described in the storage overview. Keep topology in mind when you constrain workloads: if you pin a Pod away from its volume’s region, the Pod is unschedulable. If a StatefulSet must survive node consolidation without disruption, protect it with a PodDisruptionBudget and consider a conservative auto-provisioning profile.

AWS nodes with an availability zone ID as their zone

Symptom: on an AWS node, the EBS CSI node plugin crash-loops with:

RegisterPlugin error ... detected topology value collision: driver reported
"topology.kubernetes.io/zone":"AZ_NAME" but existing label is
"topology.kubernetes.io/zone":"AZ_ID"

and PVCs stay Pending with no topology key found for node.

Cause: earlier CFKE releases labelled auto-provisioned AWS nodes with the availability zone ID (for example euc1-az2) instead of its name (for example eu-central-1a). Current releases set topology.kubernetes.io/zone to the zone name and put the zone ID in topology.k8s.aws/zone-id.

Fix: CFKE replaces nodes with the old label automatically. To replace a node right away, delete its NodeClaim:

bash
kubectl get nodeclaims -o wide
kubectl delete nodeclaim NODECLAIM_NAME

Set up the driver as described in Amazon EBS.

Pods evicted with disk pressure

Symptom: Pods are evicted with node was low on resource: ephemeral-storage and left in Failed or ContainerStatusUnknown. The node reports the DiskPressure condition.

What to check:

  1. Set ephemeral-storage requests and limits on workloads that write significant local data (logs, caches, temporary files). The scheduler then places them on nodes with enough disk, and the kubelet can enforce a bound.
  2. Look for log-heavy containers without rotation and images with large writable layers.
  3. On auto-provisioned nodes, disk size follows the instance type. If your workloads need more local disk, steer them to larger instance types with resource requests or an instance family constraint.

Kubernetes does not clean up evicted Pod records automatically. After you resolve the pressure, remove them:

bash
kubectl delete pods --field-selector=status.phase=Failed -A
On this page