Self-managed node issues
Self-managed nodes let you join any Linux server to a CFKE cluster. Unlike auto-provisioned nodes, you own the machine, its operating system, and its network. This page covers the failures we see most often and the responsibility boundary between you and Cloudfleet. For setup, see Self-managed nodes.
The join command fails
cloudfleet clusters add-self-managed-node connects to the machine over SSH, configures it, and waits for it to join the cluster. Run the command with --verbose to see every command it executes on the node. Adding a node is safe to repeat on the same machine. After you fix the cause, run the same command again (or re-apply your Terraform configuration). The most common failures:
The SSH connection fails. The CLI reports the specific cause:
- Nothing listens for SSH on the address and port. Usually a newly created machine is still booting. Check
--hostand--ssh-port. - The host cannot be reached or does not resolve. Check the address and any firewall between you and the node.
- The login is refused. The CLI names the user and key it tried. Check
--ssh-username,--ssh-key, and that the public key is in that user’sauthorized_keys. The CLI tries the key file from--ssh-keyand the keys loaded into your SSH agent. - The host key does not match. A reinstalled or newly created machine reuses the address. Remove the stale entry with
ssh-keygen -R HOSTand run the command again.
Unsupported operating system. The CLI checks the operating system before changing anything and refuses releases it does not support. Self-managed nodes must run Ubuntu 22.04 or 24.04, Debian 12 or 13, or the RHEL family (RHEL, Rocky Linux, AlmaLinux, CentOS Stream) 9 or 10.
Package installation fails. The node needs outbound internet access to download the Kubernetes packages. See Network requirements. If the command fails during a download, test IPv4 and IPv6 reachability of the host in the error message separately from the node. For example, for pkgs.k8s.io:
curl -4 -sSL -o /dev/null -w '%{http_code}\n' https://pkgs.k8s.io/core:/stable:/v1.33/rpm/repodata/repomd.xml.key
curl -6 -sSL -o /dev/null -w '%{http_code}\n' https://pkgs.k8s.io/core:/stable:/v1.33/rpm/repodata/repomd.xml.keyA 403 only on IPv6 means the CDN that serves the packages blocks your provider’s IPv6 range, and the package manager prefers IPv6 for downloads. This problem recurs on some VPS providers. Prefer a machine with a routed public IPv4 address. An IPv6-only or unrouted-IPv4 server also breaks image pulls from IPv4-only registries (such as ghcr.io) and Pod egress later. Symptoms include dial tcp IP:443: i/o timeout, even though kubectl still works.
After fixing the network, run the join command again.
The node is configured but never becomes Ready. The CLI waits up to three minutes (--wait-timeout) for the node to register and report Ready. If it times out:
- Run
kubectl describe node NODE_NAME(orkubectl get nodesif the node has not appeared yet) to see whether the node has registered and what its conditions report. - On the machine, run
sudo journalctl -u kubeletto see why kubelet cannot register with the control plane. - Check that the machine can reach the internet and the cluster endpoint (
cloudfleet clusters describe CLUSTER_IDshows it). Also check that the outbound destinations listed under Network requirements are allowed. - After you fix the cause, run the join command again.
SSH access is not possible. If you cannot reach the machine over SSH from your workstation, use the Terraform provider to generate the node’s join configuration as cloud-init data. Apply it yourself when the machine boots:
resource "cloudfleet_cfke_node_join_information" "node" {
cluster_id = CLUSTER_ID
region = "datacenter-1"
zone = "rack-a"
}See the Terraform documentation and the provider-specific guides for examples.
The node freezes when workloads start
Symptom: the machine becomes unreachable over SSH whenever kubelet and the container runtime start, or shortly after a specific workload is scheduled.
Cause seen in practice: a workload without CPU limits (for example CPU-based LLM inference) consumed every core and starved kubelet, the container runtime, and SSH itself. The node appears dead, but the hardware is fine.
Fix: set CPU limits on heavy workloads so cores remain for the system (for example, cap a workload at 12 of 16 cores). Set realistic requests everywhere.
This failure mode is not specific to self-managed nodes. A workload without resource requests and limits can starve the system components on any node and destabilize the whole cluster. Missing resource requests are the most common root cause of the unstable clusters we see in support. On self-managed nodes, the impact is worse because no provider replaces a starved machine automatically. On auto-provisioned nodes, missing requests also break node sizing (see Pods stuck in Pending).
Pods stop starting, existing Pods go Unknown
Symptom: the node reports healthy, but no new container starts and existing Pods drift into Unknown.
Cause seen in practice: /etc/resolv.conf disappeared on the host (systemd-resolved stopped or the symlink was removed). The container runtime needs this file to create every Pod sandbox.
Fix: restore DNS resolution on the host, then re-add the node:
sudo systemctl restart systemd-resolved
# or recreate the symlink:
sudo ln -sf /run/systemd/resolve/stub-resolv.conf /etc/resolv.confThe node is unreachable and logs fail with “No agent available”
Symptom: the node is NotReady, kubectl logs for Pods on it fails with No agent available, and SSH times out.
Cause: the machine itself is down or cut off, often a provider incident on the physical host.
What to know: Cloudfleet cannot see, escalate, or resolve incidents in your provider account. Check the provider’s status page and open a ticket with them. This is the key operational difference from auto-provisioned nodes: Cloudfleet replaces those automatically when a host fails.
Do not install a cloud provider integration
Do not install a cloud controller manager (for example the Hetzner cloud-controller-manager) or a provider node agent on CFKE nodes. CFKE ships its own cloud integration that manages load balancers (including PROXY protocol support) with its own annotations. A second controller conflicts with it and has caused cluster-wide networking outages. For persistent volumes, follow the provider-specific CSI guidance instead (see Persistent volume issues).
Also keep the host firewall permissive for the cluster’s overlay traffic. Custom firewall rules that block this traffic cut Pods off from the API server. Pods on the Pod network then fail with dial tcp 10.96.0.1:443: connect: no route to host, while host-network Pods still work.
Upgrades
To synchronize self-managed nodes with the control plane version, re-add them. See Kubernetes versions and upgrades.