GCP Persistent Disk

The Compute Engine Persistent Disk CSI driver provisions Persistent Disk volumes for Pods on the GCP nodes of your Fleet. The driver authenticates through Workload Identity Federation with your cluster’s OIDC tokens. You create no GCP service account keys and store no credentials in the cluster.

Prerequisites

  • A CFKE cluster with a GCP Fleet
  • Permission to create Workload Identity Pools and to grant IAM roles in the GCP project
  • kubectl configured for the cluster

The examples use the cluster ID CLUSTER_ID and the project ID PROJECT_ID. Replace them with your values.

Step 1: Grant the driver access to the project

The driver’s controller runs as the csi-gce-pd-controller-sa service account in the gce-pd-csi-driver namespace. Create a Workload Identity Pool that trusts your cluster’s tokens, and grant this service account’s federated identity access to Compute Engine. The grant uses direct resource access, so you do not create a GCP service account.

If you already created a Workload Identity Pool for the cluster, for example to access cloud APIs, reuse it and keep only the IAM bindings.

terraform
locals {
  cluster_id = "CLUSTER_ID"
  project_id = "PROJECT_ID"
  issuer     = "https://api.cloudfleet.ai/v1/clusters/${local.cluster_id}"
}

resource "google_iam_workload_identity_pool" "cfke" {
  project                   = local.project_id
  workload_identity_pool_id = "cfke-${substr(local.cluster_id, 0, 8)}"
  display_name              = "CFKE"
}

resource "google_iam_workload_identity_pool_provider" "cfke" {
  project                            = local.project_id
  workload_identity_pool_id          = google_iam_workload_identity_pool.cfke.workload_identity_pool_id
  workload_identity_pool_provider_id = "cfke"
  display_name                       = "CFKE"
  attribute_mapping                  = { "google.subject" = "assertion.sub" }
  oidc {
    issuer_uri        = local.issuer
    allowed_audiences = [local.issuer]
  }
}

resource "google_project_iam_member" "pd_csi" {
  for_each = toset(["roles/compute.storageAdmin", "roles/compute.instanceAdmin.v1"])
  project  = local.project_id
  role     = each.key
  member   = "principal://iam.googleapis.com/${google_iam_workload_identity_pool.cfke.name}/subject/system:serviceaccount:gce-pd-csi-driver:csi-gce-pd-controller-sa"
}

output "workload_identity_provider" {
  value = google_iam_workload_identity_pool_provider.cfke.name
}

The driver needs roles/compute.instanceAdmin.v1 to attach disks to instances. For tighter permissions, create a custom role from the permissions that the driver uses.

Step 2: Install the driver

The project publishes the driver as Kustomize manifests. By default, the controller reads a service account key from a secret. The following kustomization.yaml replaces the key with the Pod’s service account token, exchanged through Workload Identity Federation. It also runs the controller and the node plugin on GCP nodes, where they read the project and zone from the metadata server.

Create a directory named gce-pd with a kustomization.yaml file.

Important: Replace WORKLOAD_IDENTITY_PROVIDER in the file with the workload_identity_provider output from step 1. Print it with terraform output -raw workload_identity_provider. The value has the form projects/PROJECT_NUMBER/locations/global/workloadIdentityPools/POOL_ID/providers/cfke.

yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: gce-pd-csi-driver
resources:
  - https://github.com/kubernetes-sigs/gcp-compute-persistent-disk-csi-driver/deploy/kubernetes/overlays/stable-master?ref=v1.27.2
configMapGenerator:
  - name: gcp-federation
    options:
      disableNameSuffixHash: true
    literals:
      - >-
        credentials.json={"type": "external_account",
        "audience": "//iam.googleapis.com/WORKLOAD_IDENTITY_PROVIDER",
        "subject_token_type": "urn:ietf:params:oauth:token-type:jwt",
        "token_url": "https://sts.googleapis.com/v1/token",
        "credential_source": {"file": "/var/run/secrets/kubernetes.io/serviceaccount/token"}}        
patches:
  - patch: |-
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: csi-gce-pd-controller
      spec:
        template:
          spec:
            nodeSelector:
              cfke.io/provider: gcp
            containers:
              - name: gce-pd-driver
                env:
                  - name: GOOGLE_APPLICATION_CREDENTIALS
                    value: /etc/gcp/credentials.json
                volumeMounts:
                  - name: cloud-sa-volume
                    $patch: delete
                    mountPath: /etc/cloud-sa
                  - name: gcp-federation
                    mountPath: /etc/gcp
                    readOnly: true
              - name: taint-controller
                env:
                  - name: GOOGLE_APPLICATION_CREDENTIALS
                    value: /etc/gcp/credentials.json
                volumeMounts:
                  - name: cloud-sa-volume
                    $patch: delete
                    mountPath: /etc/cloud-sa
                  - name: gcp-federation
                    mountPath: /etc/gcp
                    readOnly: true
            volumes:
              - name: cloud-sa-volume
                $patch: delete
              - name: gcp-federation
                configMap:
                  name: gcp-federation      
  - patch: |-
      apiVersion: apps/v1
      kind: DaemonSet
      metadata:
        name: csi-gce-pd-node
      spec:
        template:
          spec:
            nodeSelector:
              cfke.io/provider: gcp      

Create the namespace and apply the configuration:

bash
kubectl create namespace gce-pd-csi-driver
kubectl apply -k gce-pd

Check that the controller and the node plugins are running:

bash
kubectl get pods -n gce-pd-csi-driver

The csi-gce-pd-controller Pod and one csi-gce-pd-node Pod per node run on GCP nodes. If the cluster has no GCP node yet, the node auto-provisioner adds one for the controller.

Step 3: Create a storage class

Create a storage class for Persistent Disk volumes:

yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: gce-pd
provisioner: pd.csi.storage.gke.io
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
parameters:
  type: pd-balanced

For other disk types, change parameters.type or add storage classes. See the driver’s storage class parameters.

Step 4: Use a volume

Create a StatefulSet that requests a Persistent Disk volume and runs on GCP nodes:

yaml
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: pd-example
spec:
  serviceName: pd-example
  replicas: 1
  selector:
    matchLabels:
      app: pd-example
  template:
    metadata:
      labels:
        app: pd-example
    spec:
      nodeSelector:
        cfke.io/provider: gcp
      containers:
        - name: app
          image: busybox:1.37
          command: ["sh", "-c", "date >> /data/boots; cat /data/boots; exec sleep infinity"]
          resources:
            requests:
              cpu: 10m
              memory: 16Mi
          volumeMounts:
            - name: data
              mountPath: /data
  volumeClaimTemplates:
    - metadata:
        name: data
      spec:
        accessModes: ["ReadWriteOnce"]
        storageClassName: gce-pd
        resources:
          requests:
            storage: 10Gi

Check that the claim is bound:

bash
kubectl get pvc data-pd-example-0

Each start of the Pod appends a line to /data/boots. Delete the Pod, or the node it runs on, and check the log of the new Pod. It lists the earlier starts, because the volume moved with the Pod to a node in the same zone.

Limitations

The controller lists disks and instances in the region of the GCP node it runs on. Use it with a GCP Fleet in a single region.

Next steps

On this page