Run PostgreSQL on Kubernetes with CloudNativePG

This guide walks you through running a highly available PostgreSQL cluster on Cloudfleet Kubernetes Engine (CFKE) using CloudNativePG, commonly shortened to CNPG, the CNCF operator for PostgreSQL. You will finish with a three-instance PostgreSQL 18 cluster on Hetzner Cloud block storage, with streaming replication, TLS-secured connections, controlled node placement, and automatic failover that completes in seconds. The last section adds a shared cluster that applications in other namespaces reach without any password, using certificates issued by cert-manager.

Hetzner is a good place to run this. Managed PostgreSQL is priced per gigabyte of memory and storage, and those are exactly the two things a database needs most of, which is why a self-managed cluster on Hetzner Cloud servers costs a fraction of the equivalent managed instance. CNPG is what makes running it yourself reasonable rather than a second job.

CNPG gives you essentially everything a managed PostgreSQL service does: high availability with automatic failover, streaming replication, backups and point-in-time recovery to object storage, connection pooling, rolling minor version upgrades, and TLS with certificates it issues and rotates itself. The difference is what you keep. The same manifests run on any Kubernetes cluster, on any provider, in any datacenter, so the database travels with your workloads instead of anchoring them to one vendor’s control plane. If you later want to move from a hyperscaler to European infrastructure, or from cloud to your own hardware, the database moves the same way the rest of your workloads do.

That portability is also what makes CNPG a practical answer to data sovereignty requirements. You decide exactly which country and which datacenter your data sits in, you can prove it from the cluster itself, and changing that decision later is a scheduling change rather than a migration project. Managed database services rarely offer that, because the control plane belongs to the provider even when the data is stored in-region.

CloudNativePG is a particularly good fit for CFKE. It has no external dependencies, it manages the full lifecycle of a PostgreSQL cluster through a single custom resource, and it relies on ordinary Kubernetes primitives for scheduling and storage. That means CFKE’s just-in-time node provisioner handles the compute side automatically: you declare three PostgreSQL instances, and the nodes to run them are provisioned for you.

Prerequisites

  • A running CFKE cluster with a fleet attached. See the getting started guide if you need one. This guide uses a Hetzner fleet named hetzner-fleet.
  • The Cloudfleet CLI, kubectl and helm on your local machine.

A few commands below take your cluster ID. Export it once:

bash
export CLUSTER_ID=<cluster-id>

CloudNativePG itself is provider-agnostic, and so is most of this guide. Fleets also support AWS, using an IAM role ARN, and Google Cloud, using a project ID, and self-managed nodes connect any other cloud, on-premises, or edge infrastructure. If you are on one of those, swap two things: the storage class in step 1 and the cfke.io/provider value in every nodeSelector. The rest of the manifests work unchanged. The parts that are genuinely Hetzner-specific are called out where they appear, and the region section is the main one, because Hetzner volumes are bound to a single location.

Step 1: Set up persistent storage

CloudNativePG gives every instance its own PersistentVolumeClaim, so your cluster needs a storage class that can provision block volumes before you deploy a database. Three properties matter:

  • ReadWriteOnce block storage. Each instance owns its volume, so no shared filesystem is needed.
  • WaitForFirstConsumer binding. Volumes are created where the pod actually lands, which matters on CFKE because nodes are provisioned just in time.
  • Volume expansion, so you can grow a database later without recreating it.

Which driver provides this depends on your fleet. On Hetzner, install the Hetzner Cloud CSI driver by following Hetzner Cloud Volumes. It gives you a default hcloud-volumes storage class with all three properties, and it reuses the Hetzner token CFKE already stored when you created the fleet, so no second API token is needed. On self-managed nodes you can use local disks, keeping in mind that local storage ties each instance to one node. For AWS and GCP, follow storage to install the provider’s CSI driver without keys.

Whichever driver you choose, set resource requests in its chart values. CFKE sizes and provisions nodes from pod resource requests, and CSI charts commonly ship resources: {}. Install one unchanged and CFKE’s admission policy tells you so:

Warning: Validation failed for ValidatingAdmissionPolicy 'resource-requests-are-not-set':
Resource requests are not set on one or more containers in pod template.

It is a warning, not a rejection: the pods still schedule and the driver still works. But requests are the only signal the auto-provisioner has for sizing nodes, so without them it works from bad information, and nodes can end up too small for what lands on them. Setting requests costs a few lines. This applies to every third-party chart you install on CFKE, including the CloudNativePG operator in the next step.

Confirm your storage class before continuing:

bash
kubectl get storageclass

The rest of this guide uses hcloud-volumes. If you are on a different class, substitute the name in the manifests below. Two Hetzner specifics do carry through: volumes start at 10 GB, and each volume lives in a single Hetzner location, which is what the region section later builds on.

Step 2: Install the CloudNativePG operator

Add the CloudNativePG chart repository:

bash
helm repo add cnpg https://cloudnative-pg.github.io/charts
helm repo update cnpg

The operator chart also defaults to empty resource requests, so set them explicitly. Create cnpg-operator-values.yaml:

yaml
replicaCount: 1

resources:
  requests:
    cpu: 100m
    memory: 200Mi
  limits:
    memory: 200Mi

Note the memory limit with no CPU limit. That is the recommended shape for CFKE workloads: a memory limit protects the node from a runaway container, while a CPU limit only causes unnecessary CFS throttling.

Install the operator:

bash
helm upgrade --install cnpg cnpg/cloudnative-pg --version 0.29.0 \
  -n cnpg-system --create-namespace --values cnpg-operator-values.yaml --wait

Confirm the operator is running and its custom resources are registered:

bash
kubectl get pods -n cnpg-system
kubectl get crd | grep cnpg

You should see clusters.postgresql.cnpg.io among the CRDs, along with resources for backups, poolers, databases, database roles, and publications.

Step 3: Deploy a PostgreSQL cluster

Create a namespace and a Cluster resource. Save this as pg-cluster.yaml:

yaml
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: pg-demo
  namespace: demo
spec:
  instances: 3
  imageName: ghcr.io/cloudnative-pg/postgresql:18.1

  primaryUpdateStrategy: unsupervised

  bootstrap:
    initdb:
      database: appdb
      owner: appuser

  storage:
    size: 10Gi
    storageClass: hcloud-volumes
  walStorage:
    size: 10Gi
    storageClass: hcloud-volumes

  resources:
    requests:
      cpu: 500m
      memory: 1Gi
    limits:
      memory: 1Gi

  affinity:
    nodeSelector:
      cfke.io/provider: hetzner
    enablePodAntiAffinity: true
    topologyKey: kubernetes.io/hostname
    podAntiAffinityType: required

  postgresql:
    parameters:
      max_connections: "200"

Several choices here are worth calling out.

walStorage puts the write-ahead log on its own volume. This is the recommended production layout: WAL writes are sequential and constant, data file writes are random and bursty, and separating them keeps one from starving the other. It also means a filling WAL cannot consume the space your data needs.

Hetzner Cloud volumes start at 10 GB, so 10Gi is the smallest size worth asking for. Smaller requests are rejected.

The memory request equals the memory limit, and there is no CPU limit, following the same reasoning as the operator values above. CloudNativePG derives shared_buffers from the memory available to the pod, so this also determines how much PostgreSQL will cache.

nodeSelector: cfke.io/provider: hetzner only does work on a cluster that spans providers. A Hetzner Cloud volume can attach only to a Hetzner server, so on a mixed cluster the selector is what stops an instance from being scheduled somewhere its storage cannot follow. If your fleet uses a single provider, every node already matches and you can leave the selector out: CFKE will use whatever infrastructure it provisions. What matters in that case is that the storage class you name works on the nodes you have, which is the point of step 1.

Keeping instances on separate nodes

The affinity block is the part that makes this cluster genuinely highly available, so it deserves a closer look:

yaml
  affinity:
    enablePodAntiAffinity: true
    topologyKey: kubernetes.io/hostname
    podAntiAffinityType: required

enablePodAntiAffinity tells the operator to generate a pod anti-affinity rule across all instances of this PostgreSQL cluster. topologyKey: kubernetes.io/hostname makes the unit of separation a node. podAntiAffinityType: required makes it a hard constraint rather than a preference.

Without this, nothing stops the scheduler from packing all three instances onto one node, and a single node failure would take down the whole database along with every replica meant to survive it. CloudNativePG defaults enablePodAntiAffinity to true, but the default anti-affinity type is preferred, which the scheduler will happily violate when capacity is tight. On a database, make it required.

On CFKE this constraint also drives node provisioning. As each replica is created, the scheduler finds no eligible node, and CFKE provisions a new one to satisfy the rule. You will watch the cluster grow from one node to three as the replicas join, which is the behaviour you want: capacity follows the availability requirement instead of you sizing a node pool up front.

Apply it:

bash
kubectl create namespace demo
kubectl apply -f pg-cluster.yaml

Watch the cluster build itself:

bash
kubectl get cluster -n demo pg-demo -w

The status moves through Setting up primary, then Creating a new replica once for each replica, and finally settles on Cluster in healthy state. Expect the whole sequence to take five to ten minutes on an empty cluster, most of it spent provisioning Hetzner nodes.

When it is done you have three pods on three separate nodes:

bash
kubectl get pods -n demo -l cnpg.io/cluster=pg-demo \
  -o custom-columns='POD:.metadata.name,STATUS:.status.phase,NODE:.spec.nodeName'
POD         STATUS    NODE
pg-demo-1   Running   model-husky-436990372
pg-demo-2   Running   pumped-ibex-3909944675
pg-demo-3   Running   cute-crab-640389611

You can now see the anti-affinity rule the operator generated from those three fields:

bash
kubectl get pod -n demo pg-demo-1 -o jsonpath='{.spec.affinity.podAntiAffinity}' | jq
json
{
  "requiredDuringSchedulingIgnoredDuringExecution": [
    {
      "labelSelector": {
        "matchExpressions": [
          { "key": "cnpg.io/cluster", "operator": "In", "values": ["pg-demo"] },
          { "key": "cnpg.io/podRole", "operator": "In", "values": ["instance"] }
        ]
      },
      "topologyKey": "kubernetes.io/hostname"
    }
  ]
}

The rule selects on cnpg.io/podRole: instance, so it separates PostgreSQL instances from each other without interfering with the operator’s transient bootstrap and join jobs.

Confirm replication is streaming:

bash
kubectl exec -n demo pg-demo-1 -c postgres -- \
  psql -U postgres -x -c "SELECT application_name, state, sync_state FROM pg_stat_replication;"
-[ RECORD 1 ]----+----------
application_name | pg-demo-2
state            | streaming
sync_state       | async
-[ RECORD 2 ]----+----------
application_name | pg-demo-3
state            | streaming
sync_state       | async

The operator also created two PodDisruptionBudgets for you, one that keeps at least one instance available during voluntary disruptions and one that protects the primary specifically. This is what makes node consolidation and cluster upgrades safe.

Step 4: Connect to the database

CloudNativePG creates three services in the namespace:

ServiceRoutes toUse for
pg-demo-rwthe current primarywrites and read-write transactions
pg-demo-roreplicas onlyread-only queries you want kept off the primary
pg-demo-rany instanceread-only queries where either is fine

Applications should connect to pg-demo-rw. The service follows the primary automatically, so a failover does not require any application change.

Credentials for the application user live in the pg-demo-app secret, which contains ready-made uri, jdbc-uri, and pgpass entries alongside the individual fields. The cluster’s CA is in pg-demo-ca.

Here is a job that writes through the read-write endpoint and reads back through the read-only endpoint, with full TLS verification. Save it as pg-smoke-test.yaml:

yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: pg-smoke-test
  namespace: demo
spec:
  backoffLimit: 2
  template:
    spec:
      restartPolicy: Never
      nodeSelector:
        cfke.io/provider: hetzner
      containers:
        - name: psql
          image: ghcr.io/cloudnative-pg/postgresql:18.1
          resources:
            requests: {cpu: 50m, memory: 64Mi}
            limits: {memory: 64Mi}
          env:
            - name: PGHOST
              value: pg-demo-rw
            - name: PGDATABASE
              value: appdb
            - name: PGUSER
              valueFrom: {secretKeyRef: {name: pg-demo-app, key: username}}
            - name: PGPASSWORD
              valueFrom: {secretKeyRef: {name: pg-demo-app, key: password}}
            - name: PGSSLMODE
              value: verify-full
            - name: PGSSLROOTCERT
              value: /etc/pg-ca/ca.crt
          volumeMounts:
            - name: ca
              mountPath: /etc/pg-ca
              readOnly: true
          command:
            - bash
            - -c
            - |
              set -e
              psql -c "CREATE TABLE IF NOT EXISTS demo (id serial primary key, note text, at timestamptz default now());"
              psql -c "INSERT INTO demo (note) VALUES ('written via pg-demo-rw');"
              psql -c "SELECT count(*) AS rows FROM demo;"
              echo "--- read-only endpoint ---"
              PGHOST=pg-demo-ro psql -c "SELECT pg_is_in_recovery() AS on_replica, count(*) FROM demo;"              
      volumes:
        - name: ca
          secret:
            secretName: pg-demo-ca
            items: [{key: ca.crt, path: ca.crt}]
bash
kubectl apply -f pg-smoke-test.yaml
kubectl logs -n demo job/pg-smoke-test
CREATE TABLE
INSERT 0 1
 rows
------
    1
(1 row)

--- read-only endpoint ---
 on_replica | count
------------+-------
 t          |     1
(1 row)

The row written through the primary is immediately visible on a replica, and pg_is_in_recovery() returns true, confirming the read-only endpoint really did route to a standby. PGSSLMODE=verify-full means the client verified the server certificate against the cluster CA, so this is an encrypted and authenticated connection, not just an encrypted one.

Setting PGSSLMODE=verify-full is worth doing in your own applications too. CloudNativePG issues server certificates automatically, so the only cost is mounting the CA.

Step 5: Test failover

The point of three instances is surviving the loss of one. Delete the primary pod and watch what happens. The PRIMARY column shows which instance is serving writes:

bash
kubectl get cluster -n demo pg-demo
kubectl delete pod -n demo pg-demo-1
kubectl get cluster -n demo pg-demo

The operator detects the loss, promotes the most advanced replica, and repoints the pg-demo-rw service at it. In this run the primary moved from pg-demo-1 to pg-demo-2 in six seconds. Applications connected to pg-demo-rw see a dropped connection and reconnect to the new primary.

The old instance is not discarded. CloudNativePG restarts it, reattaches its Hetzner volume, and rejoins it as a replica of the new primary:

bash
kubectl get cluster -n demo pg-demo
NAME      INSTANCES   READY   STATUS                     PRIMARY
pg-demo   3           3       Cluster in healthy state   pg-demo-2

Verify the data survived:

bash
kubectl exec -n demo pg-demo-2 -c postgres -- \
  psql -U postgres -d appdb -c "SELECT count(*), max(at) FROM demo;"

Controlling which region the database runs in

A CFKE cluster is scoped neither to a region nor to a provider. A single fleet can span Hetzner, AWS, and Google Cloud and provision nodes in any location of any of them, and self-managed nodes can join the same cluster from any other cloud, on-premises, or edge site. By default the provisioner picks whichever location is cheapest, and the scheduler places pods on any node that fits. For most workloads that is exactly right. For a database it usually is not, and the reason is storage.

Hetzner Cloud volumes are bound to a single location. The CSI driver reflects this by labelling nodes with csi.hetzner.cloud/location and stamping a matching node affinity onto every PersistentVolume it creates:

bash
kubectl get pv -o jsonpath='{.items[0].spec.nodeAffinity}' | jq
json
{
  "required": {
    "nodeSelectorTerms": [
      {
        "matchExpressions": [
          { "key": "csi.hetzner.cloud/location", "operator": "In", "values": ["nbg1"] }
        ]
      }
    ]
  }
}

A PostgreSQL instance whose volume lives in nbg1 can therefore only ever be scheduled onto a node in nbg1. That is what you want, and it is why the failed instance in the previous step recovered cleanly. But it also means that if the provisioner scatters your instances across locations, each one is pinned to the location it landed in: replication traffic crosses the public network, and an instance can only ever recover where its volume already is.

CFKE labels every node so you can be explicit about this. A Hetzner node in Nuremberg carries:

cfke.io/provider:                 hetzner
cfke.io/region:                   europe
cfke.io/subregion:                central
topology.kubernetes.io/region:    nbg1
topology.kubernetes.io/zone:      nbg1
csi.hetzner.cloud/location:       nbg1
node.kubernetes.io/instance-type: cx23
karpenter.sh/capacity-type:       on-demand

Note that cfke.io/region is coarse, at continent level, while topology.kubernetes.io/region is the Hetzner location. Use the latter when you mean a specific datacenter.

There are three ways to take control, depending on whether the whole cluster should be regional, only the database should be, or the database should deliberately span locations.

Option A: Lock the entire fleet to one location

If everything on this cluster should stay in one place, constrain the fleet itself. Fleet constraints move placement policy out of every pod spec and onto the fleet, where a missing selector cannot quietly put a workload in the wrong datacenter:

bash
cloudfleet clusters fleets create $CLUSTER_ID -f - <<EOF
{
  "id": "hetzner-fleet",
  "hetzner": { "enabled": true, "apiKey": "<your-hetzner-api-token>" },
  "limits": { "cpu": 16 },
  "constraints": {
    "topology.kubernetes.io/region": ["nbg1"],
    "kubernetes.io/arch": ["amd64", "arm64"]
  }
}
EOF

The provisioner will now only ever create nodes in nbg1, and no workload can accidentally end up elsewhere. Constraints compose, so you can pin architecture, instance family, and purchase type in the same block. This is the simplest option and a good default for a single-region product.

The architecture constraint above is worth setting deliberately. A fleet created without a constraints block defaults to kubernetes.io/arch: [amd64], which quietly excludes Hetzner’s arm64 CAX line, often the cheapest memory you can buy there and a good match for a database. CloudNativePG publishes multi-architecture images, so listing both lets the provisioner pick on price.

Two caveats when applying this to a fleet that already has workloads. Constraining an existing fleet can strand pods that no longer match, leaving them Pending because no fleet can satisfy them, so check what is currently scheduled first. And fleet updates are full replacements: cloudfleet clusters fleets update resets any field you leave out, so read the current fleet, edit it, and pipe it back rather than sending a partial document.

Option B: Keep the cluster multi-region, pin the database

More often you want the opposite: stateless services spread across locations for latency and resilience, with the database deliberately kept together. Leave the fleet unconstrained and pin the database instead, using affinity.nodeSelector on the Cluster resource:

yaml
  affinity:
    nodeSelector:
      cfke.io/provider: hetzner
      topology.kubernetes.io/region: nbg1
    enablePodAntiAffinity: true
    topologyKey: kubernetes.io/hostname
    podAntiAffinityType: required

This is the combination worth understanding: the nodeSelector keeps all three instances inside one Hetzner location, while the anti-affinity rule keeps them on three different nodes inside it. You get node-level redundancy and local replication, and the rest of the cluster remains free to schedule anywhere. Anything else that needs to sit next to the database, such as a PgBouncer pooler or a batch job, should carry the same nodeSelector.

The operator applies the selector to every instance pod it manages, so you can verify it took effect without inspecting each manifest:

bash
kubectl get pods -n demo -l cnpg.io/cluster=pg-demo \
  -o custom-columns='POD:.metadata.name,SELECTOR:.spec.nodeSelector,NODE:.spec.nodeName'
POD         SELECTOR                                                              NODE
pg-demo-1   map[cfke.io/provider:hetzner topology.kubernetes.io/region:nbg1]      set-panther-1265738465
pg-demo-2   map[cfke.io/provider:hetzner topology.kubernetes.io/region:nbg1]      pumped-ibex-3909944675
pg-demo-3   map[cfke.io/provider:hetzner topology.kubernetes.io/region:nbg1]      cute-crab-640389611

Changing affinity on a running cluster triggers a rolling restart, replicas first and the primary last, so it is safe to apply to an existing database. Instances whose volumes already live in the selected location stay where they are.

Option C: Spread deliberately across locations

If you want the cluster to survive the loss of an entire Hetzner datacenter, separate the instances by location rather than by node. CloudNativePG exposes this through the same affinity block, by changing the anti-affinity topologyKey:

yaml
  affinity:
    nodeSelector:
      cfke.io/provider: hetzner
    enablePodAntiAffinity: true
    topologyKey: topology.kubernetes.io/zone
    podAntiAffinityType: required

This replaces the kubernetes.io/hostname key from step 3. You do not need both: if every instance is in a different Hetzner location, they are necessarily on different nodes.

Because the rule is required, an instance cannot be placed in a location that already holds one. Once the existing locations are used up, the next instance is unschedulable everywhere, and CFKE provisions a node in a location it has not used yet. On a three-instance cluster this produces one instance per location, each with its volume created alongside it:

bash
kubectl get pods -n demo -l cnpg.io/cluster=pg-demo \
  -o custom-columns='POD:.metadata.name,NODE:.spec.nodeName'
kubectl get nodes -L topology.kubernetes.io/zone

Reading the two outputs together gives one instance per location:

POD         NODE                       ZONE
pg-demo-1   bright-muskox-2248219584   fsn1
pg-demo-2   model-husky-436990372      nbg1
pg-demo-3   supreme-tahr-384143606     hel1

Two limits are worth knowing before you rely on this.

You need at least as many locations as instances. Three instances need three Hetzner locations. Ask for more instances than there are available locations and the surplus stays Pending indefinitely, because no node can ever satisfy the rule.

It only works on a new cluster, or on instances that have not been created yet. A volume is bound to the location it was created in, so applying this to a running single-location cluster will not move anything. The existing instances stay where their data is.

The cost of spreading is latency on WAL shipping and cross-location traffic counted against your Hetzner server allowances. With CloudNativePG’s default asynchronous replication that latency is absorbed by the replicas rather than by your writes, so for most workloads this is a reasonable trade. If you also enable synchronous replication, measure it first.

Passwordless connections with mutual TLS

The smoke test in step 4 authenticates with a password from the pg-demo-app secret. A password is a long-lived credential: it gets copied into CI variables and local shells, it works from anywhere the database is reachable, and rotating it means touching every client at once.

PostgreSQL has supported certificate authentication for over 15 years, yet few teams use it. Outside Kubernetes, someone has to run a CA, hand every application a certificate, track expiry dates, and redeploy before anything expires. That operational burden is why passwords win by default.

Kubernetes removes that burden. cert-manager runs the CA and renews every certificate on its own, CloudNativePG reloads new certificates without a restart, and an admission policy decides which namespace may hold which identity. You declare the certificates once, and the cluster keeps them valid. This section builds a second cluster, pg-shared, that is passwordless from the start, and connects applications from other namespaces to it.

Target architecture

One PostgreSQL cluster serves several applications. Each application belongs to a team and runs in the team’s own namespace, here orders and billing. Each team gets its own database and role on the shared cluster, so you run one highly available cluster instead of one per application.

flowchart LR
    CA["cert-manager<br/>ClusterIssuer postgres-ca"]
    subgraph ns-orders["namespace orders"]
        OA["orders app<br/>certificate CN=orders"]
    end
    subgraph ns-billing["namespace billing"]
        BA["billing app<br/>certificate CN=billing"]
    end
    subgraph ns-databases["namespace databases"]
        subgraph pg["pg-shared: 3 PostgreSQL instances"]
            DO[("database orders<br/>role orders")]
            DB[("database billing<br/>role billing")]
        end
    end
    CA -. issues .-> OA
    CA -. issues .-> BA
    CA -. issues .-> pg
    OA -- "mutual TLS" --> DO
    BA -- "mutual TLS" --> DB

A private CA run by cert-manager issues one certificate to the database and one to each team. The certificate’s common name (CN) is the team’s database role, and it always equals the team’s namespace. The roles have no password at all, so there is nothing to leak, store in a secrets manager, or rotate by hand, and cert-manager renews every certificate before it expires.

How PostgreSQL decides who reaches which database

Three checks run on every connection, in this order:

  1. TLS handshake. The server accepts only client certificates signed by the CA in the cluster’s clientCASecret.
  2. pg_hba rules. PostgreSQL takes the first rule that matches the connection type, database, and role. The rule hostssl orders orders all cert lets role orders into database orders over TLS with certificate authentication. The cert method also requires the certificate’s CN to equal the role, so a certificate for orders cannot log in as billing.
  3. Fallback to passwords. Every other combination falls through to CloudNativePG’s default scram-sha-256 rule. The team roles have no password, so that rule always refuses them.

The certificate is therefore the credential, and its CN decides which database it opens. A Kubernetes admission policy in this section makes sure each namespace can only obtain a certificate for its own name.

Install cert-manager

If cert-manager already runs in your cluster, for example from the NGINX Ingress and cert-manager tutorial, skip this step. Otherwise create cert-manager-values.yaml:

yaml
crds:
  enabled: true

resources:
  requests: {cpu: 10m, memory: 64Mi}
  limits: {memory: 128Mi}
webhook:
  resources:
    requests: {cpu: 10m, memory: 32Mi}
    limits: {memory: 64Mi}
cainjector:
  resources:
    requests: {cpu: 10m, memory: 64Mi}
    limits: {memory: 128Mi}
bash
helm repo add jetstack https://charts.jetstack.io
helm repo update jetstack
helm upgrade --install cert-manager jetstack/cert-manager --version v1.21.2 \
  -n cert-manager --create-namespace --values cert-manager-values.yaml --wait

Create a private CA

These certificates do not come from a public CA such as Let’s Encrypt. Public CAs only sign names they can verify on the internet, never .svc Service names or client identities like orders. A private CA inside the cluster signs both, and no external service is involved.

Save postgres-ca.yaml. A self-signed issuer creates the CA once, and a ClusterIssuer signs certificates with it in every namespace:

yaml
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: selfsigned
spec:
  selfSigned: {}
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: postgres-ca
  namespace: cert-manager
spec:
  isCA: true
  commonName: postgres-ca
  secretName: postgres-ca
  duration: 87600h # 10 years
  privateKey:
    algorithm: ECDSA
    size: 256
  issuerRef:
    name: selfsigned
    kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: postgres-ca
spec:
  ca:
    secretName: postgres-ca
bash
kubectl apply -f postgres-ca.yaml

If your organization already runs a PKI, put your intermediate CA in the postgres-ca secret instead, for example with External Secrets Operator syncing it from your secrets manager. Everything else stays the same.

Bind each namespace to its own identity

The ClusterIssuer signs certificates in every namespace, and cert-manager signs whatever CN a Certificate asks for. Without a guard, anyone who can create a Certificate in orders could request CN=billing and log in as the billing team. A ValidatingAdmissionPolicy, built into Kubernetes, closes that gap without any change to cert-manager. Save postgres-ca-policy.yaml:

yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: postgres-ca-identities
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
      - apiGroups: [cert-manager.io]
        apiVersions: ["*"]
        operations: [CREATE, UPDATE]
        resources: [certificates, certificaterequests]
  matchConditions:
    - name: issued-by-postgres-ca
      expression: >-
        object.spec.issuerRef.name == 'postgres-ca' &&
        object.spec.issuerRef.kind == 'ClusterIssuer'        
  validations:
    - expression: >-
        object.kind != 'Certificate' || request.namespace == 'databases' ||
        (has(object.spec.commonName) && object.spec.commonName == request.namespace &&
         !has(object.spec.dnsNames) && !has(object.spec.ipAddresses) &&
         !has(object.spec.uris) && !has(object.spec.emailAddresses) &&
         !has(object.spec.literalSubject) &&
         !(has(object.spec.isCA) && object.spec.isCA))        
      message: outside the databases namespace, the common name must equal the namespace and no other names are allowed
    - expression: >-
        object.kind != 'CertificateRequest' ||
        request.userInfo.username == 'system:serviceaccount:cert-manager:cert-manager'        
      message: create a Certificate instead of a CertificateRequest
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: postgres-ca-identities
spec:
  policyName: postgres-ca-identities
  validationActions: [Deny]

The first rule applies to every Certificate from postgres-ca outside the databases namespace: its CN must equal the namespace, and it may carry no DNS names, IP addresses, or CA flag, so no other namespace can impersonate a team or the database server. The second rule blocks hand-written CertificateRequest objects. Only cert-manager creates them, and only from a Certificate that already passed the first rule. The policy checks updates too, so an existing certificate cannot be renamed.

bash
kubectl apply -f postgres-ca-policy.yaml

A namespace that asks for someone else’s identity is now rejected when the manifest is applied:

The certificates "spoof" is invalid: : ValidatingAdmissionPolicy 'postgres-ca-identities' with binding 'postgres-ca-identities' denied request: outside the databases namespace, the common name must equal the namespace and no other names are allowed

Access to a team’s database now follows Kubernetes RBAC: whoever can create resources in the orders namespace can reach the orders database, and nobody else can.

Issue the server and replication certificates

CloudNativePG needs a server certificate for its Services and a client certificate for its streaming_replica user. Issuing both from postgres-ca means clients in any namespace verify the server with the same ca.crt they receive with their own certificate. Save pg-shared-certificates.yaml:

yaml
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: pg-shared-cert-server
  namespace: databases
spec:
  secretName: pg-shared-cert-server
  usages: [server auth]
  dnsNames:
    - pg-shared-rw
    - pg-shared-rw.databases
    - pg-shared-rw.databases.svc
    - pg-shared-rw.databases.svc.cluster.local
    - pg-shared-ro
    - pg-shared-ro.databases
    - pg-shared-ro.databases.svc
    - pg-shared-ro.databases.svc.cluster.local
    - pg-shared-r
    - pg-shared-r.databases
    - pg-shared-r.databases.svc
    - pg-shared-r.databases.svc.cluster.local
  privateKey:
    algorithm: ECDSA
    size: 256
  issuerRef:
    name: postgres-ca
    kind: ClusterIssuer
  secretTemplate:
    labels:
      cnpg.io/reload: ""
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: pg-shared-cert-replication
  namespace: databases
spec:
  secretName: pg-shared-cert-replication
  usages: [client auth]
  commonName: streaming_replica
  privateKey:
    algorithm: ECDSA
    size: 256
  issuerRef:
    name: postgres-ca
    kind: ClusterIssuer
  secretTemplate:
    labels:
      cnpg.io/reload: ""

The DNS names match the ones CloudNativePG puts in its own certificates. The cnpg.io/reload label makes CloudNativePG reload the certificates whenever cert-manager renews them, without restarting PostgreSQL.

Do not name these secrets pg-shared-server or pg-shared-replication. CloudNativePG uses <cluster>-server and <cluster>-replication for the certificates it manages itself. A cert-manager secret with the same name silently replaces the operator’s certificate, and replicas then fail to join with connection requires a valid client certificate.

bash
kubectl create namespace databases
kubectl apply -f pg-shared-certificates.yaml

Create the shared cluster

Save pg-shared.yaml. It is the cluster from step 3 with three additions, certificates, pg_hba, and a team database in initdb:

yaml
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: pg-shared
  namespace: databases
spec:
  instances: 3
  imageName: ghcr.io/cloudnative-pg/postgresql:18.1

  primaryUpdateStrategy: unsupervised

  bootstrap:
    initdb:
      database: orders
      owner: orders

  certificates:
    serverTLSSecret: pg-shared-cert-server
    serverCASecret: pg-shared-cert-server
    replicationTLSSecret: pg-shared-cert-replication
    clientCASecret: pg-shared-cert-replication

  postgresql:
    parameters:
      max_connections: "200"
    pg_hba:
      - hostssl orders orders all cert
      - hostssl billing billing all cert

  storage:
    size: 10Gi
    storageClass: hcloud-volumes
  walStorage:
    size: 10Gi
    storageClass: hcloud-volumes

  resources:
    requests:
      cpu: 500m
      memory: 1Gi
    limits:
      memory: 1Gi

  affinity:
    nodeSelector:
      cfke.io/provider: hetzner
    enablePodAntiAffinity: true
    topologyKey: kubernetes.io/hostname
    podAntiAffinityType: required

certificates points CloudNativePG at the cert-manager secrets. cert-manager stores the issuing CA as ca.crt in every certificate secret, so the same secrets also serve as CA secrets. With all four fields set, CloudNativePG runs in user-provided certificates mode and trusts only client certificates from postgres-ca.

pg_hba holds one rule per team. CloudNativePG places them before its final password rule, which gives the evaluation order described above.

bash
kubectl apply -f pg-shared.yaml
kubectl get cluster -n databases pg-shared -w

The cluster bootstraps exactly like pg-demo, but its replicas join with the cert-manager replication certificate. It ends in:

NAME        AGE     INSTANCES   READY   STATUS                     PRIMARY
pg-shared   4m42s   3           3       Cluster in healthy state   pg-shared-1

Create a role and database for each team

DatabaseRole and Database resources manage PostgreSQL objects declaratively. They live next to the cluster in the databases namespace, so the cluster’s owners decide which teams exist. Save teams.yaml:

yaml
apiVersion: postgresql.cnpg.io/v1
kind: DatabaseRole
metadata:
  name: orders
  namespace: databases
spec:
  cluster: {name: pg-shared}
  name: orders
  login: true
  disablePassword: true
---
apiVersion: postgresql.cnpg.io/v1
kind: Database
metadata:
  name: orders
  namespace: databases
spec:
  cluster: {name: pg-shared}
  name: orders
  owner: orders
---
apiVersion: postgresql.cnpg.io/v1
kind: DatabaseRole
metadata:
  name: billing
  namespace: databases
spec:
  cluster: {name: pg-shared}
  name: billing
  login: true
  disablePassword: true
---
apiVersion: postgresql.cnpg.io/v1
kind: Database
metadata:
  name: billing
  namespace: databases
spec:
  cluster: {name: pg-shared}
  name: billing
  owner: billing

disablePassword: true sets the role’s password to NULL, so password authentication can never succeed. This also applies to the orders role that initdb created, and the password CloudNativePG stored for it in the pg-shared-app secret stops working. DatabaseRole is available from CloudNativePG 1.30; on older versions, declare the roles under spec.managed.roles in the Cluster. Adding a team later means one role, one database, and one pg_hba line.

bash
kubectl apply -f teams.yaml
kubectl get databaserole,database -n databases
NAME                                      AGE   CLUSTER     PG NAME   APPLIED
databaserole.postgresql.cnpg.io/billing   15s   pg-shared   billing   true
databaserole.postgresql.cnpg.io/orders    15s   pg-shared   orders    true

NAME                                  AGE   CLUSTER     PG NAME   APPLIED
database.postgresql.cnpg.io/billing   16s   pg-shared   billing   true
database.postgresql.cnpg.io/orders    16s   pg-shared   orders    true

Connect from the team namespaces

Each team requests its own client certificate in its own namespace and mounts it into its Pods. Save orders-client.yaml:

yaml
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: postgres-client
  namespace: orders
spec:
  secretName: postgres-client
  commonName: orders
  usages: [client auth]
  privateKey:
    algorithm: ECDSA
    size: 256
  issuerRef:
    name: postgres-ca
    kind: ClusterIssuer
---
apiVersion: v1
kind: Pod
metadata:
  name: psql
  namespace: orders
spec:
  nodeSelector:
    cfke.io/provider: hetzner
  securityContext:
    runAsUser: 1000
    fsGroup: 1000
  containers:
    - name: psql
      image: ghcr.io/cloudnative-pg/postgresql:18.1
      command: [sleep, infinity]
      resources:
        requests: {cpu: 50m, memory: 64Mi}
        limits: {memory: 64Mi}
      env:
        - {name: PGHOST, value: pg-shared-rw.databases.svc}
        - {name: PGUSER, value: orders}
        - {name: PGDATABASE, value: orders}
        - {name: PGSSLMODE, value: verify-full}
        - {name: PGSSLROOTCERT, value: /etc/pg-tls/ca.crt}
        - {name: PGSSLCERT, value: /etc/pg-tls/tls.crt}
        - {name: PGSSLKEY, value: /etc/pg-tls/tls.key}
      volumeMounts:
        - name: pg-tls
          mountPath: /etc/pg-tls
          readOnly: true
  volumes:
    - name: pg-tls
      secret:
        secretName: postgres-client
        defaultMode: 0440

The PG* variables are read by libpq, so most PostgreSQL clients pick them up without code changes. PGSSLMODE=verify-full checks the server certificate and host name against ca.crt, so the client verifies the database while the database verifies the client. defaultMode: 0440 with fsGroup keeps the private key readable only by the Pod’s group, because libpq refuses keys that other users can read.

The billing client is the same manifest with orders replaced by billing. To see CFKE’s multi-cloud networking at work, also change its selector to cfke.io/provider: aws, so the application runs on a different provider than the database.

bash
kubectl create namespace orders
kubectl create namespace billing
kubectl apply -f orders-client.yaml -f billing-client.yaml

Write through the read-write endpoint and read back through the read-only endpoint:

bash
kubectl exec -n orders psql -- psql \
  -c "CREATE TABLE demo (id serial PRIMARY KEY, note text);" \
  -c "INSERT INTO demo (note) VALUES ('written via pg-shared-rw');" \
  -c "SELECT current_user, version AS tls FROM pg_stat_ssl WHERE pid = pg_backend_pid();"
kubectl exec -n orders psql -- env PGHOST=pg-shared-ro.databases.svc psql \
  -c "SELECT pg_is_in_recovery() AS on_replica, count(*) FROM demo;"
CREATE TABLE
INSERT 0 1
 current_user |   tls
--------------+---------
 orders       | TLSv1.3
(1 row)

 on_replica | count
------------+-------
 t          |     1
(1 row)

No password was sent or stored. The connection is encrypted, the client verified the server, and the server verified the client.

The billing team connects the same way:

bash
kubectl exec -n billing psql -- psql -c "SELECT current_user, current_database();"
 current_user | current_database
--------------+------------------
 billing      | billing
(1 row)

In this run the billing Pod ran on AWS in eu-central-1 and the database on Hetzner in hel1. CFKE connected them over its WireGuard® encrypted node network, with no VPC peering, VPN, or public database endpoint. Mutual TLS adds workload identity on top: the network proves which node sent a packet, the certificate proves which application holds the connection. It is the same principle CFKE applies to cloud APIs, where Pods authenticate with workload identity instead of stored keys.

Verify the isolation

Run these from the orders Pod. Each one is refused:

bash
# Another team's database
kubectl exec -n orders psql -- env PGPASSWORD=guess psql -d billing -c "SELECT 1"
# Another team's role
kubectl exec -n orders psql -- psql -U billing -d billing -c "SELECT 1"
# No client certificate
kubectl exec -n orders psql -- env PGSSLCERT=/nonexistent PGSSLKEY=/nonexistent psql -c "SELECT 1"
FATAL:  password authentication failed for user "orders"
FATAL:  certificate authentication failed for user "billing"
FATAL:  connection requires a valid client certificate

No pg_hba rule lets orders into billing, so the first attempt falls through to the password rule, and the role has no password. The second presents CN=orders for role billing. The third has no certificate at all.

To keep other namespaces from even opening a connection to the database, add a network policy such as the one in Isolate a sensitive namespace.

Cleaning up

Deleting the CloudNativePG cluster removes its pods and PVCs, and the CSI driver deletes the underlying Hetzner volumes:

bash
kubectl delete -f pg-cluster.yaml

If you followed the mutual TLS section, also delete its namespaces, which removes pg-shared and its volumes:

bash
kubectl delete namespace databases orders billing

To tear down everything, delete the Cloudfleet cluster. CFKE deprovisions all fleet nodes with it:

bash
cloudfleet clusters delete $CLUSTER_ID

Next steps

Backups and point-in-time recovery. CloudNativePG 1.26 and later handle object storage backups through the Barman Cloud plugin, which you install alongside the operator and configure with an ObjectStore resource. Any S3-compatible endpoint works, including Hetzner Object Storage. Follow the official backup documentation and the plugin repository for setup. Note that the plugin requires cert-manager in the cluster, which the mutual TLS section installs.

Connection pooling. PostgreSQL handles connections with one backend process each, which becomes expensive well before max_connections is reached. CloudNativePG’s Pooler resource deploys PgBouncer in front of your cluster as a managed resource.

Synchronous replication. The setup above uses asynchronous replication, which can lose recently committed transactions on failover. If your workload cannot tolerate that, configure synchronous replication so commits wait for a replica to acknowledge.

Monitoring. Every CloudNativePG instance exposes Prometheus metrics. Set monitoring.enablePodMonitor: true on the cluster once you have the Prometheus Operator installed, and import the project’s Grafana dashboard.

Exposing PostgreSQL outside the cluster. If you need external access, front the -rw service with a LoadBalancer service and set externalTrafficPolicy: Local. On CFKE the default Cluster policy provisions a load balancer in every location where nodes exist, which costs more and routes less efficiently. See the Cloudfleet load balancing documentation.

For the full range of what the operator can do, see the CloudNativePG documentation.

On this page