Local disks

The Kubernetes local volume type turns a disk of a node into a PersistentVolume. Local disks are fast and need no CSI driver. They fit self-managed and bare-metal nodes, scratch data, and apps that replicate their data themselves. A local volume belongs to one node, so the Pod that uses it always runs on that node.

Prerequisites

  • A CFKE cluster with self-managed nodes
  • A disk or partition mounted on the node, for example at /mnt/disks/data1
  • kubectl configured for the cluster

The examples use the node name NODE_NAME. Replace it with the name of your node, as kubectl get nodes shows it.

Step 1: Create a storage class

Create a storage class for local disks:

yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: local-disk
provisioner: kubernetes.io/no-provisioner
volumeBindingMode: WaitForFirstConsumer

The class creates no volumes. WaitForFirstConsumer delays binding until the scheduler places the Pod, so the claim binds to a volume on a node that fits the Pod.

Step 2: Create a PersistentVolume

Create one PersistentVolume for each disk:

yaml
apiVersion: v1
kind: PersistentVolume
metadata:
  name: NODE_NAME-data1
spec:
  capacity:
    storage: 10Gi
  accessModes: ["ReadWriteOnce"]
  persistentVolumeReclaimPolicy: Retain
  storageClassName: local-disk
  local:
    path: /mnt/disks/data1
  nodeAffinity:
    required:
      nodeSelectorTerms:
        - matchExpressions:
            - key: kubernetes.io/hostname
              operator: In
              values: ["NODE_NAME"]

The node affinity pins the volume, and every Pod that uses it, to the node. Retain keeps the data on the disk when you delete the claim.

Step 3: Use a volume

Create a StatefulSet that requests a local volume:

yaml
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: local-example
spec:
  serviceName: local-example
  replicas: 1
  selector:
    matchLabels:
      app: local-example
  template:
    metadata:
      labels:
        app: local-example
    spec:
      containers:
        - name: app
          image: busybox:1.37
          command: ["sh", "-c", "date >> /data/boots; cat /data/boots; exec sleep infinity"]
          resources:
            requests:
              cpu: 10m
              memory: 16Mi
          volumeMounts:
            - name: data
              mountPath: /data
  volumeClaimTemplates:
    - metadata:
        name: data
      spec:
        accessModes: ["ReadWriteOnce"]
        storageClassName: local-disk
        resources:
          requests:
            storage: 10Gi

The Pod needs no node selector, because the PersistentVolume already pins it to the node. Check that the claim is bound:

bash
kubectl get pvc data-local-example-0

Each start of the Pod appends a line to /data/boots. Delete the Pod and check the log of the new Pod. It lists the earlier starts, because the new Pod runs on the same node with the same disk.

What happens when the node is unavailable

The data of a local volume exists only on its node. When the node is down or cordoned, the Pod stays Pending until the node returns. It never starts with an empty volume on another node.

The node auto-provisioner may still launch a node for the Pending Pod. The Pod cannot use that node, and the node auto-provisioner removes it again.

To recover the Pod on another node, delete its PersistentVolumeClaim. Then restore the data from the replicas or backups of your app.

A released PersistentVolume keeps its data. Before you reuse the volume, delete the PersistentVolume, clean the directory on the node, and create the PersistentVolume again.

More disks

Creating one PersistentVolume for each disk does not scale to many nodes. Two projects automate it:

  • The local static provisioner discovers the disks mounted under a directory, for example /mnt/disks, and creates a PersistentVolume for each.
  • The local-path-provisioner creates a directory on one disk of the node for each claim. Volumes share the disk, and the provisioner does not enforce their size.

Do not use hostPath volumes for app data. A hostPath volume has no claim and no node affinity, so a rescheduled Pod starts with an empty directory on another node. The baseline Pod Security Standard also blocks hostPath volumes.

Next steps

On this page