Run PostgreSQL on Kubernetes with CloudNativePG
This guide walks you through running a highly available PostgreSQL cluster on Cloudfleet Kubernetes Engine (CFKE) using CloudNativePG, commonly shortened to CNPG, the CNCF operator for PostgreSQL. You will finish with a three-instance PostgreSQL 18 cluster on Hetzner Cloud block storage, with streaming replication, TLS-secured connections, controlled node placement, and automatic failover that completes in seconds. The last section adds a shared cluster that applications in other namespaces reach without any password, using certificates issued by cert-manager.
Hetzner is a good place to run this. Managed PostgreSQL is priced per gigabyte of memory and storage, and those are exactly the two things a database needs most of, which is why a self-managed cluster on Hetzner Cloud servers costs a fraction of the equivalent managed instance. CNPG is what makes running it yourself reasonable rather than a second job.
CNPG gives you essentially everything a managed PostgreSQL service does: high availability with automatic failover, streaming replication, backups and point-in-time recovery to object storage, connection pooling, rolling minor version upgrades, and TLS with certificates it issues and rotates itself. The difference is what you keep. The same manifests run on any Kubernetes cluster, on any provider, in any datacenter, so the database travels with your workloads instead of anchoring them to one vendor’s control plane. If you later want to move from a hyperscaler to European infrastructure, or from cloud to your own hardware, the database moves the same way the rest of your workloads do.
That portability is also what makes CNPG a practical answer to data sovereignty requirements. You decide exactly which country and which datacenter your data sits in, you can prove it from the cluster itself, and changing that decision later is a scheduling change rather than a migration project. Managed database services rarely offer that, because the control plane belongs to the provider even when the data is stored in-region.
CloudNativePG is a particularly good fit for CFKE. It has no external dependencies, it manages the full lifecycle of a PostgreSQL cluster through a single custom resource, and it relies on ordinary Kubernetes primitives for scheduling and storage. That means CFKE’s just-in-time node provisioner handles the compute side automatically: you declare three PostgreSQL instances, and the nodes to run them are provisioned for you.
Prerequisites
- A running CFKE cluster with a fleet attached. See the getting started guide if you need one. This guide uses a Hetzner fleet named
hetzner-fleet. - The Cloudfleet CLI,
kubectlandhelmon your local machine.
A few commands below take your cluster ID. Export it once:
export CLUSTER_ID=<cluster-id>CloudNativePG itself is provider-agnostic, and so is most of this guide. Fleets also support AWS, using an IAM role ARN, and Google Cloud, using a project ID, and self-managed nodes connect any other cloud, on-premises, or edge infrastructure. If you are on one of those, swap two things: the storage class in step 1 and the cfke.io/provider value in every nodeSelector. The rest of the manifests work unchanged. The parts that are genuinely Hetzner-specific are called out where they appear, and the region section is the main one, because Hetzner volumes are bound to a single location.
Step 1: Set up persistent storage
CloudNativePG gives every instance its own PersistentVolumeClaim, so your cluster needs a storage class that can provision block volumes before you deploy a database. Three properties matter:
- ReadWriteOnce block storage. Each instance owns its volume, so no shared filesystem is needed.
WaitForFirstConsumerbinding. Volumes are created where the pod actually lands, which matters on CFKE because nodes are provisioned just in time.- Volume expansion, so you can grow a database later without recreating it.
Which driver provides this depends on your fleet. On Hetzner, install the Hetzner Cloud CSI driver by following Hetzner Cloud Volumes. It gives you a default hcloud-volumes storage class with all three properties, and it reuses the Hetzner token CFKE already stored when you created the fleet, so no second API token is needed. On self-managed nodes you can use local disks, keeping in mind that local storage ties each instance to one node. For AWS and GCP, follow storage to install the provider’s CSI driver without keys.
Whichever driver you choose, set resource requests in its chart values. CFKE sizes and provisions nodes from pod resource requests, and CSI charts commonly ship resources: {}. Install one unchanged and CFKE’s admission policy tells you so:
Warning: Validation failed for ValidatingAdmissionPolicy 'resource-requests-are-not-set':
Resource requests are not set on one or more containers in pod template.It is a warning, not a rejection: the pods still schedule and the driver still works. But requests are the only signal the auto-provisioner has for sizing nodes, so without them it works from bad information, and nodes can end up too small for what lands on them. Setting requests costs a few lines. This applies to every third-party chart you install on CFKE, including the CloudNativePG operator in the next step.
Confirm your storage class before continuing:
kubectl get storageclassThe rest of this guide uses hcloud-volumes. If you are on a different class, substitute the name in the manifests below. Two Hetzner specifics do carry through: volumes start at 10 GB, and each volume lives in a single Hetzner location, which is what the region section later builds on.
Step 2: Install the CloudNativePG operator
Add the CloudNativePG chart repository:
helm repo add cnpg https://cloudnative-pg.github.io/charts
helm repo update cnpgThe operator chart also defaults to empty resource requests, so set them explicitly. Create cnpg-operator-values.yaml:
replicaCount: 1
resources:
requests:
cpu: 100m
memory: 200Mi
limits:
memory: 200MiNote the memory limit with no CPU limit. That is the recommended shape for CFKE workloads: a memory limit protects the node from a runaway container, while a CPU limit only causes unnecessary CFS throttling.
Install the operator:
helm upgrade --install cnpg cnpg/cloudnative-pg --version 0.29.0 \
-n cnpg-system --create-namespace --values cnpg-operator-values.yaml --waitConfirm the operator is running and its custom resources are registered:
kubectl get pods -n cnpg-system
kubectl get crd | grep cnpgYou should see clusters.postgresql.cnpg.io among the CRDs, along with resources for backups, poolers, databases, database roles, and publications.
Step 3: Deploy a PostgreSQL cluster
Create a namespace and a Cluster resource. Save this as pg-cluster.yaml:
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: pg-demo
namespace: demo
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:18.1
primaryUpdateStrategy: unsupervised
bootstrap:
initdb:
database: appdb
owner: appuser
storage:
size: 10Gi
storageClass: hcloud-volumes
walStorage:
size: 10Gi
storageClass: hcloud-volumes
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
memory: 1Gi
affinity:
nodeSelector:
cfke.io/provider: hetzner
enablePodAntiAffinity: true
topologyKey: kubernetes.io/hostname
podAntiAffinityType: required
postgresql:
parameters:
max_connections: "200"Several choices here are worth calling out.
walStorage puts the write-ahead log on its own volume. This is the recommended production layout: WAL writes are sequential and constant, data file writes are random and bursty, and separating them keeps one from starving the other. It also means a filling WAL cannot consume the space your data needs.
Hetzner Cloud volumes start at 10 GB, so 10Gi is the smallest size worth asking for. Smaller requests are rejected.
The memory request equals the memory limit, and there is no CPU limit, following the same reasoning as the operator values above. CloudNativePG derives shared_buffers from the memory available to the pod, so this also determines how much PostgreSQL will cache.
nodeSelector: cfke.io/provider: hetzner only does work on a cluster that spans providers. A Hetzner Cloud volume can attach only to a Hetzner server, so on a mixed cluster the selector is what stops an instance from being scheduled somewhere its storage cannot follow. If your fleet uses a single provider, every node already matches and you can leave the selector out: CFKE will use whatever infrastructure it provisions. What matters in that case is that the storage class you name works on the nodes you have, which is the point of step 1.
Keeping instances on separate nodes
The affinity block is the part that makes this cluster genuinely highly available, so it deserves a closer look:
affinity:
enablePodAntiAffinity: true
topologyKey: kubernetes.io/hostname
podAntiAffinityType: requiredenablePodAntiAffinity tells the operator to generate a pod anti-affinity rule across all instances of this PostgreSQL cluster. topologyKey: kubernetes.io/hostname makes the unit of separation a node. podAntiAffinityType: required makes it a hard constraint rather than a preference.
Without this, nothing stops the scheduler from packing all three instances onto one node, and a single node failure would take down the whole database along with every replica meant to survive it. CloudNativePG defaults enablePodAntiAffinity to true, but the default anti-affinity type is preferred, which the scheduler will happily violate when capacity is tight. On a database, make it required.
On CFKE this constraint also drives node provisioning. As each replica is created, the scheduler finds no eligible node, and CFKE provisions a new one to satisfy the rule. You will watch the cluster grow from one node to three as the replicas join, which is the behaviour you want: capacity follows the availability requirement instead of you sizing a node pool up front.
Apply it:
kubectl create namespace demo
kubectl apply -f pg-cluster.yamlWatch the cluster build itself:
kubectl get cluster -n demo pg-demo -wThe status moves through Setting up primary, then Creating a new replica once for each replica, and finally settles on Cluster in healthy state. Expect the whole sequence to take five to ten minutes on an empty cluster, most of it spent provisioning Hetzner nodes.
When it is done you have three pods on three separate nodes:
kubectl get pods -n demo -l cnpg.io/cluster=pg-demo \
-o custom-columns='POD:.metadata.name,STATUS:.status.phase,NODE:.spec.nodeName'POD STATUS NODE
pg-demo-1 Running model-husky-436990372
pg-demo-2 Running pumped-ibex-3909944675
pg-demo-3 Running cute-crab-640389611You can now see the anti-affinity rule the operator generated from those three fields:
kubectl get pod -n demo pg-demo-1 -o jsonpath='{.spec.affinity.podAntiAffinity}' | jq{
"requiredDuringSchedulingIgnoredDuringExecution": [
{
"labelSelector": {
"matchExpressions": [
{ "key": "cnpg.io/cluster", "operator": "In", "values": ["pg-demo"] },
{ "key": "cnpg.io/podRole", "operator": "In", "values": ["instance"] }
]
},
"topologyKey": "kubernetes.io/hostname"
}
]
}The rule selects on cnpg.io/podRole: instance, so it separates PostgreSQL instances from each other without interfering with the operator’s transient bootstrap and join jobs.
Confirm replication is streaming:
kubectl exec -n demo pg-demo-1 -c postgres -- \
psql -U postgres -x -c "SELECT application_name, state, sync_state FROM pg_stat_replication;"-[ RECORD 1 ]----+----------
application_name | pg-demo-2
state | streaming
sync_state | async
-[ RECORD 2 ]----+----------
application_name | pg-demo-3
state | streaming
sync_state | asyncThe operator also created two PodDisruptionBudgets for you, one that keeps at least one instance available during voluntary disruptions and one that protects the primary specifically. This is what makes node consolidation and cluster upgrades safe.
Step 4: Connect to the database
CloudNativePG creates three services in the namespace:
| Service | Routes to | Use for |
|---|---|---|
pg-demo-rw | the current primary | writes and read-write transactions |
pg-demo-ro | replicas only | read-only queries you want kept off the primary |
pg-demo-r | any instance | read-only queries where either is fine |
Applications should connect to pg-demo-rw. The service follows the primary automatically, so a failover does not require any application change.
Credentials for the application user live in the pg-demo-app secret, which contains ready-made uri, jdbc-uri, and pgpass entries alongside the individual fields. The cluster’s CA is in pg-demo-ca.
Here is a job that writes through the read-write endpoint and reads back through the read-only endpoint, with full TLS verification. Save it as pg-smoke-test.yaml:
apiVersion: batch/v1
kind: Job
metadata:
name: pg-smoke-test
namespace: demo
spec:
backoffLimit: 2
template:
spec:
restartPolicy: Never
nodeSelector:
cfke.io/provider: hetzner
containers:
- name: psql
image: ghcr.io/cloudnative-pg/postgresql:18.1
resources:
requests: {cpu: 50m, memory: 64Mi}
limits: {memory: 64Mi}
env:
- name: PGHOST
value: pg-demo-rw
- name: PGDATABASE
value: appdb
- name: PGUSER
valueFrom: {secretKeyRef: {name: pg-demo-app, key: username}}
- name: PGPASSWORD
valueFrom: {secretKeyRef: {name: pg-demo-app, key: password}}
- name: PGSSLMODE
value: verify-full
- name: PGSSLROOTCERT
value: /etc/pg-ca/ca.crt
volumeMounts:
- name: ca
mountPath: /etc/pg-ca
readOnly: true
command:
- bash
- -c
- |
set -e
psql -c "CREATE TABLE IF NOT EXISTS demo (id serial primary key, note text, at timestamptz default now());"
psql -c "INSERT INTO demo (note) VALUES ('written via pg-demo-rw');"
psql -c "SELECT count(*) AS rows FROM demo;"
echo "--- read-only endpoint ---"
PGHOST=pg-demo-ro psql -c "SELECT pg_is_in_recovery() AS on_replica, count(*) FROM demo;"
volumes:
- name: ca
secret:
secretName: pg-demo-ca
items: [{key: ca.crt, path: ca.crt}]kubectl apply -f pg-smoke-test.yaml
kubectl logs -n demo job/pg-smoke-testCREATE TABLE
INSERT 0 1
rows
------
1
(1 row)
--- read-only endpoint ---
on_replica | count
------------+-------
t | 1
(1 row)The row written through the primary is immediately visible on a replica, and pg_is_in_recovery() returns true, confirming the read-only endpoint really did route to a standby. PGSSLMODE=verify-full means the client verified the server certificate against the cluster CA, so this is an encrypted and authenticated connection, not just an encrypted one.
Setting PGSSLMODE=verify-full is worth doing in your own applications too. CloudNativePG issues server certificates automatically, so the only cost is mounting the CA.
Step 5: Test failover
The point of three instances is surviving the loss of one. Delete the primary pod and watch what happens. The PRIMARY column shows which instance is serving writes:
kubectl get cluster -n demo pg-demo
kubectl delete pod -n demo pg-demo-1
kubectl get cluster -n demo pg-demoThe operator detects the loss, promotes the most advanced replica, and repoints the pg-demo-rw service at it. In this run the primary moved from pg-demo-1 to pg-demo-2 in six seconds. Applications connected to pg-demo-rw see a dropped connection and reconnect to the new primary.
The old instance is not discarded. CloudNativePG restarts it, reattaches its Hetzner volume, and rejoins it as a replica of the new primary:
kubectl get cluster -n demo pg-demoNAME INSTANCES READY STATUS PRIMARY
pg-demo 3 3 Cluster in healthy state pg-demo-2Verify the data survived:
kubectl exec -n demo pg-demo-2 -c postgres -- \
psql -U postgres -d appdb -c "SELECT count(*), max(at) FROM demo;"Controlling which region the database runs in
A CFKE cluster is scoped neither to a region nor to a provider. A single fleet can span Hetzner, AWS, and Google Cloud and provision nodes in any location of any of them, and self-managed nodes can join the same cluster from any other cloud, on-premises, or edge site. By default the provisioner picks whichever location is cheapest, and the scheduler places pods on any node that fits. For most workloads that is exactly right. For a database it usually is not, and the reason is storage.
Hetzner Cloud volumes are bound to a single location. The CSI driver reflects this by labelling nodes with csi.hetzner.cloud/location and stamping a matching node affinity onto every PersistentVolume it creates:
kubectl get pv -o jsonpath='{.items[0].spec.nodeAffinity}' | jq{
"required": {
"nodeSelectorTerms": [
{
"matchExpressions": [
{ "key": "csi.hetzner.cloud/location", "operator": "In", "values": ["nbg1"] }
]
}
]
}
}A PostgreSQL instance whose volume lives in nbg1 can therefore only ever be scheduled onto a node in nbg1. That is what you want, and it is why the failed instance in the previous step recovered cleanly. But it also means that if the provisioner scatters your instances across locations, each one is pinned to the location it landed in: replication traffic crosses the public network, and an instance can only ever recover where its volume already is.
CFKE labels every node so you can be explicit about this. A Hetzner node in Nuremberg carries:
cfke.io/provider: hetzner
cfke.io/region: europe
cfke.io/subregion: central
topology.kubernetes.io/region: nbg1
topology.kubernetes.io/zone: nbg1
csi.hetzner.cloud/location: nbg1
node.kubernetes.io/instance-type: cx23
karpenter.sh/capacity-type: on-demandNote that cfke.io/region is coarse, at continent level, while topology.kubernetes.io/region is the Hetzner location. Use the latter when you mean a specific datacenter.
There are three ways to take control, depending on whether the whole cluster should be regional, only the database should be, or the database should deliberately span locations.
Option A: Lock the entire fleet to one location
If everything on this cluster should stay in one place, constrain the fleet itself. Fleet constraints move placement policy out of every pod spec and onto the fleet, where a missing selector cannot quietly put a workload in the wrong datacenter:
cloudfleet clusters fleets create $CLUSTER_ID -f - <<EOF
{
"id": "hetzner-fleet",
"hetzner": { "enabled": true, "apiKey": "<your-hetzner-api-token>" },
"limits": { "cpu": 16 },
"constraints": {
"topology.kubernetes.io/region": ["nbg1"],
"kubernetes.io/arch": ["amd64", "arm64"]
}
}
EOFThe provisioner will now only ever create nodes in nbg1, and no workload can accidentally end up elsewhere. Constraints compose, so you can pin architecture, instance family, and purchase type in the same block. This is the simplest option and a good default for a single-region product.
The architecture constraint above is worth setting deliberately. A fleet created without a constraints block defaults to kubernetes.io/arch: [amd64], which quietly excludes Hetzner’s arm64 CAX line, often the cheapest memory you can buy there and a good match for a database. CloudNativePG publishes multi-architecture images, so listing both lets the provisioner pick on price.
Two caveats when applying this to a fleet that already has workloads. Constraining an existing fleet can strand pods that no longer match, leaving them Pending because no fleet can satisfy them, so check what is currently scheduled first. And fleet updates are full replacements: cloudfleet clusters fleets update resets any field you leave out, so read the current fleet, edit it, and pipe it back rather than sending a partial document.
Option B: Keep the cluster multi-region, pin the database
More often you want the opposite: stateless services spread across locations for latency and resilience, with the database deliberately kept together. Leave the fleet unconstrained and pin the database instead, using affinity.nodeSelector on the Cluster resource:
affinity:
nodeSelector:
cfke.io/provider: hetzner
topology.kubernetes.io/region: nbg1
enablePodAntiAffinity: true
topologyKey: kubernetes.io/hostname
podAntiAffinityType: requiredThis is the combination worth understanding: the nodeSelector keeps all three instances inside one Hetzner location, while the anti-affinity rule keeps them on three different nodes inside it. You get node-level redundancy and local replication, and the rest of the cluster remains free to schedule anywhere. Anything else that needs to sit next to the database, such as a PgBouncer pooler or a batch job, should carry the same nodeSelector.
The operator applies the selector to every instance pod it manages, so you can verify it took effect without inspecting each manifest:
kubectl get pods -n demo -l cnpg.io/cluster=pg-demo \
-o custom-columns='POD:.metadata.name,SELECTOR:.spec.nodeSelector,NODE:.spec.nodeName'POD SELECTOR NODE
pg-demo-1 map[cfke.io/provider:hetzner topology.kubernetes.io/region:nbg1] set-panther-1265738465
pg-demo-2 map[cfke.io/provider:hetzner topology.kubernetes.io/region:nbg1] pumped-ibex-3909944675
pg-demo-3 map[cfke.io/provider:hetzner topology.kubernetes.io/region:nbg1] cute-crab-640389611Changing affinity on a running cluster triggers a rolling restart, replicas first and the primary last, so it is safe to apply to an existing database. Instances whose volumes already live in the selected location stay where they are.
Option C: Spread deliberately across locations
If you want the cluster to survive the loss of an entire Hetzner datacenter, separate the instances by location rather than by node. CloudNativePG exposes this through the same affinity block, by changing the anti-affinity topologyKey:
affinity:
nodeSelector:
cfke.io/provider: hetzner
enablePodAntiAffinity: true
topologyKey: topology.kubernetes.io/zone
podAntiAffinityType: requiredThis replaces the kubernetes.io/hostname key from step 3. You do not need both: if every instance is in a different Hetzner location, they are necessarily on different nodes.
Because the rule is required, an instance cannot be placed in a location that already holds one. Once the existing locations are used up, the next instance is unschedulable everywhere, and CFKE provisions a node in a location it has not used yet. On a three-instance cluster this produces one instance per location, each with its volume created alongside it:
kubectl get pods -n demo -l cnpg.io/cluster=pg-demo \
-o custom-columns='POD:.metadata.name,NODE:.spec.nodeName'
kubectl get nodes -L topology.kubernetes.io/zoneReading the two outputs together gives one instance per location:
POD NODE ZONE
pg-demo-1 bright-muskox-2248219584 fsn1
pg-demo-2 model-husky-436990372 nbg1
pg-demo-3 supreme-tahr-384143606 hel1Two limits are worth knowing before you rely on this.
You need at least as many locations as instances. Three instances need three Hetzner locations. Ask for more instances than there are available locations and the surplus stays Pending indefinitely, because no node can ever satisfy the rule.
It only works on a new cluster, or on instances that have not been created yet. A volume is bound to the location it was created in, so applying this to a running single-location cluster will not move anything. The existing instances stay where their data is.
The cost of spreading is latency on WAL shipping and cross-location traffic counted against your Hetzner server allowances. With CloudNativePG’s default asynchronous replication that latency is absorbed by the replicas rather than by your writes, so for most workloads this is a reasonable trade. If you also enable synchronous replication, measure it first.
Passwordless connections with mutual TLS
The smoke test in step 4 authenticates with a password from the pg-demo-app secret. A password is a long-lived credential: it gets copied into CI variables and local shells, it works from anywhere the database is reachable, and rotating it means touching every client at once.
PostgreSQL has supported certificate authentication for over 15 years, yet few teams use it. Outside Kubernetes, someone has to run a CA, hand every application a certificate, track expiry dates, and redeploy before anything expires. That operational burden is why passwords win by default.
Kubernetes removes that burden. cert-manager runs the CA and renews every certificate on its own, CloudNativePG reloads new certificates without a restart, and an admission policy decides which namespace may hold which identity. You declare the certificates once, and the cluster keeps them valid. This section builds a second cluster, pg-shared, that is passwordless from the start, and connects applications from other namespaces to it.
Target architecture
One PostgreSQL cluster serves several applications. Each application belongs to a team and runs in the team’s own namespace, here orders and billing. Each team gets its own database and role on the shared cluster, so you run one highly available cluster instead of one per application.
flowchart LR
CA["cert-manager<br/>ClusterIssuer postgres-ca"]
subgraph ns-orders["namespace orders"]
OA["orders app<br/>certificate CN=orders"]
end
subgraph ns-billing["namespace billing"]
BA["billing app<br/>certificate CN=billing"]
end
subgraph ns-databases["namespace databases"]
subgraph pg["pg-shared: 3 PostgreSQL instances"]
DO[("database orders<br/>role orders")]
DB[("database billing<br/>role billing")]
end
end
CA -. issues .-> OA
CA -. issues .-> BA
CA -. issues .-> pg
OA -- "mutual TLS" --> DO
BA -- "mutual TLS" --> DB
A private CA run by cert-manager issues one certificate to the database and one to each team. The certificate’s common name (CN) is the team’s database role, and it always equals the team’s namespace. The roles have no password at all, so there is nothing to leak, store in a secrets manager, or rotate by hand, and cert-manager renews every certificate before it expires.
How PostgreSQL decides who reaches which database
Three checks run on every connection, in this order:
- TLS handshake. The server accepts only client certificates signed by the CA in the cluster’s
clientCASecret. pg_hbarules. PostgreSQL takes the first rule that matches the connection type, database, and role. The rulehostssl orders orders all certlets roleordersinto databaseordersover TLS with certificate authentication. Thecertmethod also requires the certificate’s CN to equal the role, so a certificate fororderscannot log in asbilling.- Fallback to passwords. Every other combination falls through to CloudNativePG’s default
scram-sha-256rule. The team roles have no password, so that rule always refuses them.
The certificate is therefore the credential, and its CN decides which database it opens. A Kubernetes admission policy in this section makes sure each namespace can only obtain a certificate for its own name.
Install cert-manager
If cert-manager already runs in your cluster, for example from the NGINX Ingress and cert-manager tutorial, skip this step. Otherwise create cert-manager-values.yaml:
crds:
enabled: true
resources:
requests: {cpu: 10m, memory: 64Mi}
limits: {memory: 128Mi}
webhook:
resources:
requests: {cpu: 10m, memory: 32Mi}
limits: {memory: 64Mi}
cainjector:
resources:
requests: {cpu: 10m, memory: 64Mi}
limits: {memory: 128Mi}helm repo add jetstack https://charts.jetstack.io
helm repo update jetstack
helm upgrade --install cert-manager jetstack/cert-manager --version v1.21.2 \
-n cert-manager --create-namespace --values cert-manager-values.yaml --waitCreate a private CA
These certificates do not come from a public CA such as Let’s Encrypt. Public CAs only sign names they can verify on the internet, never .svc Service names or client identities like orders. A private CA inside the cluster signs both, and no external service is involved.
Save postgres-ca.yaml. A self-signed issuer creates the CA once, and a ClusterIssuer signs certificates with it in every namespace:
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: selfsigned
spec:
selfSigned: {}
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: postgres-ca
namespace: cert-manager
spec:
isCA: true
commonName: postgres-ca
secretName: postgres-ca
duration: 87600h # 10 years
privateKey:
algorithm: ECDSA
size: 256
issuerRef:
name: selfsigned
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: postgres-ca
spec:
ca:
secretName: postgres-cakubectl apply -f postgres-ca.yamlIf your organization already runs a PKI, put your intermediate CA in the postgres-ca secret instead, for example with External Secrets Operator syncing it from your secrets manager. Everything else stays the same.
Bind each namespace to its own identity
The ClusterIssuer signs certificates in every namespace, and cert-manager signs whatever CN a Certificate asks for. Without a guard, anyone who can create a Certificate in orders could request CN=billing and log in as the billing team. A ValidatingAdmissionPolicy, built into Kubernetes, closes that gap without any change to cert-manager. Save postgres-ca-policy.yaml:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: postgres-ca-identities
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: [cert-manager.io]
apiVersions: ["*"]
operations: [CREATE, UPDATE]
resources: [certificates, certificaterequests]
matchConditions:
- name: issued-by-postgres-ca
expression: >-
object.spec.issuerRef.name == 'postgres-ca' &&
object.spec.issuerRef.kind == 'ClusterIssuer'
validations:
- expression: >-
object.kind != 'Certificate' || request.namespace == 'databases' ||
(has(object.spec.commonName) && object.spec.commonName == request.namespace &&
!has(object.spec.dnsNames) && !has(object.spec.ipAddresses) &&
!has(object.spec.uris) && !has(object.spec.emailAddresses) &&
!has(object.spec.literalSubject) &&
!(has(object.spec.isCA) && object.spec.isCA))
message: outside the databases namespace, the common name must equal the namespace and no other names are allowed
- expression: >-
object.kind != 'CertificateRequest' ||
request.userInfo.username == 'system:serviceaccount:cert-manager:cert-manager'
message: create a Certificate instead of a CertificateRequest
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: postgres-ca-identities
spec:
policyName: postgres-ca-identities
validationActions: [Deny]The first rule applies to every Certificate from postgres-ca outside the databases namespace: its CN must equal the namespace, and it may carry no DNS names, IP addresses, or CA flag, so no other namespace can impersonate a team or the database server. The second rule blocks hand-written CertificateRequest objects. Only cert-manager creates them, and only from a Certificate that already passed the first rule. The policy checks updates too, so an existing certificate cannot be renamed.
kubectl apply -f postgres-ca-policy.yamlA namespace that asks for someone else’s identity is now rejected when the manifest is applied:
The certificates "spoof" is invalid: : ValidatingAdmissionPolicy 'postgres-ca-identities' with binding 'postgres-ca-identities' denied request: outside the databases namespace, the common name must equal the namespace and no other names are allowedAccess to a team’s database now follows Kubernetes RBAC: whoever can create resources in the orders namespace can reach the orders database, and nobody else can.
Issue the server and replication certificates
CloudNativePG needs a server certificate for its Services and a client certificate for its streaming_replica user. Issuing both from postgres-ca means clients in any namespace verify the server with the same ca.crt they receive with their own certificate. Save pg-shared-certificates.yaml:
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: pg-shared-cert-server
namespace: databases
spec:
secretName: pg-shared-cert-server
usages: [server auth]
dnsNames:
- pg-shared-rw
- pg-shared-rw.databases
- pg-shared-rw.databases.svc
- pg-shared-rw.databases.svc.cluster.local
- pg-shared-ro
- pg-shared-ro.databases
- pg-shared-ro.databases.svc
- pg-shared-ro.databases.svc.cluster.local
- pg-shared-r
- pg-shared-r.databases
- pg-shared-r.databases.svc
- pg-shared-r.databases.svc.cluster.local
privateKey:
algorithm: ECDSA
size: 256
issuerRef:
name: postgres-ca
kind: ClusterIssuer
secretTemplate:
labels:
cnpg.io/reload: ""
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: pg-shared-cert-replication
namespace: databases
spec:
secretName: pg-shared-cert-replication
usages: [client auth]
commonName: streaming_replica
privateKey:
algorithm: ECDSA
size: 256
issuerRef:
name: postgres-ca
kind: ClusterIssuer
secretTemplate:
labels:
cnpg.io/reload: ""The DNS names match the ones CloudNativePG puts in its own certificates. The cnpg.io/reload label makes CloudNativePG reload the certificates whenever cert-manager renews them, without restarting PostgreSQL.
Do not name these secrets pg-shared-server or pg-shared-replication. CloudNativePG uses <cluster>-server and <cluster>-replication for the certificates it manages itself. A cert-manager secret with the same name silently replaces the operator’s certificate, and replicas then fail to join with connection requires a valid client certificate.
kubectl create namespace databases
kubectl apply -f pg-shared-certificates.yamlCreate the shared cluster
Save pg-shared.yaml. It is the cluster from step 3 with three additions, certificates, pg_hba, and a team database in initdb:
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: pg-shared
namespace: databases
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:18.1
primaryUpdateStrategy: unsupervised
bootstrap:
initdb:
database: orders
owner: orders
certificates:
serverTLSSecret: pg-shared-cert-server
serverCASecret: pg-shared-cert-server
replicationTLSSecret: pg-shared-cert-replication
clientCASecret: pg-shared-cert-replication
postgresql:
parameters:
max_connections: "200"
pg_hba:
- hostssl orders orders all cert
- hostssl billing billing all cert
storage:
size: 10Gi
storageClass: hcloud-volumes
walStorage:
size: 10Gi
storageClass: hcloud-volumes
resources:
requests:
cpu: 500m
memory: 1Gi
limits:
memory: 1Gi
affinity:
nodeSelector:
cfke.io/provider: hetzner
enablePodAntiAffinity: true
topologyKey: kubernetes.io/hostname
podAntiAffinityType: requiredcertificates points CloudNativePG at the cert-manager secrets. cert-manager stores the issuing CA as ca.crt in every certificate secret, so the same secrets also serve as CA secrets. With all four fields set, CloudNativePG runs in user-provided certificates mode and trusts only client certificates from postgres-ca.
pg_hba holds one rule per team. CloudNativePG places them before its final password rule, which gives the evaluation order described above.
kubectl apply -f pg-shared.yaml
kubectl get cluster -n databases pg-shared -wThe cluster bootstraps exactly like pg-demo, but its replicas join with the cert-manager replication certificate. It ends in:
NAME AGE INSTANCES READY STATUS PRIMARY
pg-shared 4m42s 3 3 Cluster in healthy state pg-shared-1Create a role and database for each team
DatabaseRole and Database resources manage PostgreSQL objects declaratively. They live next to the cluster in the databases namespace, so the cluster’s owners decide which teams exist. Save teams.yaml:
apiVersion: postgresql.cnpg.io/v1
kind: DatabaseRole
metadata:
name: orders
namespace: databases
spec:
cluster: {name: pg-shared}
name: orders
login: true
disablePassword: true
---
apiVersion: postgresql.cnpg.io/v1
kind: Database
metadata:
name: orders
namespace: databases
spec:
cluster: {name: pg-shared}
name: orders
owner: orders
---
apiVersion: postgresql.cnpg.io/v1
kind: DatabaseRole
metadata:
name: billing
namespace: databases
spec:
cluster: {name: pg-shared}
name: billing
login: true
disablePassword: true
---
apiVersion: postgresql.cnpg.io/v1
kind: Database
metadata:
name: billing
namespace: databases
spec:
cluster: {name: pg-shared}
name: billing
owner: billingdisablePassword: true sets the role’s password to NULL, so password authentication can never succeed. This also applies to the orders role that initdb created, and the password CloudNativePG stored for it in the pg-shared-app secret stops working. DatabaseRole is available from CloudNativePG 1.30; on older versions, declare the roles under spec.managed.roles in the Cluster. Adding a team later means one role, one database, and one pg_hba line.
kubectl apply -f teams.yaml
kubectl get databaserole,database -n databasesNAME AGE CLUSTER PG NAME APPLIED
databaserole.postgresql.cnpg.io/billing 15s pg-shared billing true
databaserole.postgresql.cnpg.io/orders 15s pg-shared orders true
NAME AGE CLUSTER PG NAME APPLIED
database.postgresql.cnpg.io/billing 16s pg-shared billing true
database.postgresql.cnpg.io/orders 16s pg-shared orders trueConnect from the team namespaces
Each team requests its own client certificate in its own namespace and mounts it into its Pods. Save orders-client.yaml:
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: postgres-client
namespace: orders
spec:
secretName: postgres-client
commonName: orders
usages: [client auth]
privateKey:
algorithm: ECDSA
size: 256
issuerRef:
name: postgres-ca
kind: ClusterIssuer
---
apiVersion: v1
kind: Pod
metadata:
name: psql
namespace: orders
spec:
nodeSelector:
cfke.io/provider: hetzner
securityContext:
runAsUser: 1000
fsGroup: 1000
containers:
- name: psql
image: ghcr.io/cloudnative-pg/postgresql:18.1
command: [sleep, infinity]
resources:
requests: {cpu: 50m, memory: 64Mi}
limits: {memory: 64Mi}
env:
- {name: PGHOST, value: pg-shared-rw.databases.svc}
- {name: PGUSER, value: orders}
- {name: PGDATABASE, value: orders}
- {name: PGSSLMODE, value: verify-full}
- {name: PGSSLROOTCERT, value: /etc/pg-tls/ca.crt}
- {name: PGSSLCERT, value: /etc/pg-tls/tls.crt}
- {name: PGSSLKEY, value: /etc/pg-tls/tls.key}
volumeMounts:
- name: pg-tls
mountPath: /etc/pg-tls
readOnly: true
volumes:
- name: pg-tls
secret:
secretName: postgres-client
defaultMode: 0440The PG* variables are read by libpq, so most PostgreSQL clients pick them up without code changes. PGSSLMODE=verify-full checks the server certificate and host name against ca.crt, so the client verifies the database while the database verifies the client. defaultMode: 0440 with fsGroup keeps the private key readable only by the Pod’s group, because libpq refuses keys that other users can read.
The billing client is the same manifest with orders replaced by billing. To see CFKE’s multi-cloud networking at work, also change its selector to cfke.io/provider: aws, so the application runs on a different provider than the database.
kubectl create namespace orders
kubectl create namespace billing
kubectl apply -f orders-client.yaml -f billing-client.yamlWrite through the read-write endpoint and read back through the read-only endpoint:
kubectl exec -n orders psql -- psql \
-c "CREATE TABLE demo (id serial PRIMARY KEY, note text);" \
-c "INSERT INTO demo (note) VALUES ('written via pg-shared-rw');" \
-c "SELECT current_user, version AS tls FROM pg_stat_ssl WHERE pid = pg_backend_pid();"
kubectl exec -n orders psql -- env PGHOST=pg-shared-ro.databases.svc psql \
-c "SELECT pg_is_in_recovery() AS on_replica, count(*) FROM demo;"CREATE TABLE
INSERT 0 1
current_user | tls
--------------+---------
orders | TLSv1.3
(1 row)
on_replica | count
------------+-------
t | 1
(1 row)No password was sent or stored. The connection is encrypted, the client verified the server, and the server verified the client.
The billing team connects the same way:
kubectl exec -n billing psql -- psql -c "SELECT current_user, current_database();" current_user | current_database
--------------+------------------
billing | billing
(1 row)In this run the billing Pod ran on AWS in eu-central-1 and the database on Hetzner in hel1. CFKE connected them over its WireGuard® encrypted node network, with no VPC peering, VPN, or public database endpoint. Mutual TLS adds workload identity on top: the network proves which node sent a packet, the certificate proves which application holds the connection. It is the same principle CFKE applies to cloud APIs, where Pods authenticate with workload identity instead of stored keys.
Verify the isolation
Run these from the orders Pod. Each one is refused:
# Another team's database
kubectl exec -n orders psql -- env PGPASSWORD=guess psql -d billing -c "SELECT 1"
# Another team's role
kubectl exec -n orders psql -- psql -U billing -d billing -c "SELECT 1"
# No client certificate
kubectl exec -n orders psql -- env PGSSLCERT=/nonexistent PGSSLKEY=/nonexistent psql -c "SELECT 1"FATAL: password authentication failed for user "orders"
FATAL: certificate authentication failed for user "billing"
FATAL: connection requires a valid client certificateNo pg_hba rule lets orders into billing, so the first attempt falls through to the password rule, and the role has no password. The second presents CN=orders for role billing. The third has no certificate at all.
To keep other namespaces from even opening a connection to the database, add a network policy such as the one in Isolate a sensitive namespace.
Cleaning up
Deleting the CloudNativePG cluster removes its pods and PVCs, and the CSI driver deletes the underlying Hetzner volumes:
kubectl delete -f pg-cluster.yamlIf you followed the mutual TLS section, also delete its namespaces, which removes pg-shared and its volumes:
kubectl delete namespace databases orders billingTo tear down everything, delete the Cloudfleet cluster. CFKE deprovisions all fleet nodes with it:
cloudfleet clusters delete $CLUSTER_IDNext steps
Backups and point-in-time recovery. CloudNativePG 1.26 and later handle object storage backups through the Barman Cloud plugin, which you install alongside the operator and configure with an ObjectStore resource. Any S3-compatible endpoint works, including Hetzner Object Storage. Follow the official backup documentation and the plugin repository for setup. Note that the plugin requires cert-manager in the cluster, which the mutual TLS section installs.
Connection pooling. PostgreSQL handles connections with one backend process each, which becomes expensive well before max_connections is reached. CloudNativePG’s Pooler resource deploys PgBouncer in front of your cluster as a managed resource.
Synchronous replication. The setup above uses asynchronous replication, which can lose recently committed transactions on failover. If your workload cannot tolerate that, configure synchronous replication so commits wait for a replica to acknowledge.
Monitoring. Every CloudNativePG instance exposes Prometheus metrics. Set monitoring.enablePodMonitor: true on the cluster once you have the Prometheus Operator installed, and import the project’s Grafana dashboard.
Exposing PostgreSQL outside the cluster. If you need external access, front the -rw service with a LoadBalancer service and set externalTrafficPolicy: Local. On CFKE the default Cluster policy provisions a load balancer in every location where nodes exist, which costs more and routes less efficiently. See the Cloudfleet load balancing documentation.
For the full range of what the operator can do, see the CloudNativePG documentation.