Pods stuck in Pending

When you deploy a workload and no suitable node exists, Cloudfleet’s node auto-provisioner creates one. This normally takes 2-4 minutes. If your Pods stay in Pending longer than that, work through this checklist from top to bottom. Each step matches a failure mode we see regularly in support tickets.

Step 1: Read the scheduling events

Start with the events on the Pod:

bash
kubectl describe pod POD_NAME

Look at the FailedScheduling message. It usually names the blocking condition, such as Insufficient memory, an unsatisfiable node selector, or no nodes available to schedule pods.

Then check if the node auto-provisioner created a node request. A NodeClaim object represents each node request:

bash
kubectl get nodeclaims -o wide

Three outcomes are possible:

  • No NodeClaim appears: the node auto-provisioner cannot find an instance type that satisfies your Pod’s constraints and the Fleet’s configuration. Continue with step 2.
  • A NodeClaim exists but stays Registered=Unknown with reason NodeNotFound: the provider received a request to create a machine, but no machine joined the cluster. Continue with step 3.
  • A NodeClaim becomes ready but the Pod still does not fit: the Pod’s resource requests are the problem. Continue with step 4.

Step 2: No instance type matches your constraints

If the cluster events or the console Events tab show messages like no instance type met the scheduling requirements, your Pod’s scheduling constraints and the Fleet configuration exclude every available machine. Common causes:

  • An exact instance type pin that the location does not offer. Cloud providers do not offer every instance type in every zone or region. A nodeSelector such as node.kubernetes.io/instance-type: cpx31 with a specific zone fails permanently if the zone does not sell that model. Pin the instance family instead, and let Cloudfleet pick a concrete size:

    yaml
    nodeSelector:
      cfke.io/instance-family: cpx

    In general, avoid pinning exact instance types. Set realistic resource requests and let the node auto-provisioner select the most cost-effective match.

  • Invalid region values in the Fleet configuration. Region names are provider-specific. For example, AWS-style names such as eu-central-1 match nothing on a Hetzner Fleet, whose locations are fsn1, nbg1, or hel1. A Fleet constrained to nonexistent regions provisions nothing. Review your Fleet constraints and the valid values in Node regions.

  • Over-constrained combinations. An instance type pin, a zone pin, and Pod anti-affinity can each be satisfiable alone but impossible together. Relax one dimension at a time to find the conflict.

  • The Fleet’s vCPU limit is reached. Fleets can cap the total vCPU count as a cost control. When the Fleet reaches the cap, new Pods stay Pending until you raise the limit, add another Fleet, or remove workloads.

Step 3: The provider cannot create machines

If NodeClaims exist but no machine registers, the provider rejects or fails the launch. Check the Events tab in the Cloudfleet Console for provider errors, then verify:

  • Fleet credentials are valid and have write permission. This is the most common cause. If a Hetzner API token was revoked, regenerated, or created with read-only permission, every node creation fails with Unauthorized (401) or permission denied (403). Existing nodes keep serving traffic and can hide the problem, sometimes for weeks. Create a new token with “Read & Write” permission and update it on the Fleet in the Cloudfleet Console. The fix takes effect within minutes and does not affect running workloads.

  • The provider account is in good standing. Providers restrict accounts to read-only or suspend network access because of unpaid invoices, pending identity verification, or abuse reviews. In that state, Cloudfleet’s API calls fail even with a valid token. Check your provider console for notices. After the provider lifts the restriction, the node auto-provisioner recovers automatically.

  • Required provider APIs are enabled. On GCP, enable the Compute Engine API in the project before Cloudfleet can create any machine:

    bash
    gcloud services enable compute.googleapis.com --project PROJECT_ID
  • Provider quotas are sufficient. New cloud accounts often have low default quotas for vCPUs or instances per region. On Hetzner, when you reach the account’s server limit, the provisioning events show server limit reached. Raise the limit in the Hetzner console. Other providers return similar quota errors. Request an increase in the provider console. Provisioning then resumes automatically.

Step 4: Check your resource requests

On CFKE, resource requests are more than a scheduling detail. They drive node sizing, workload placement, and node stability. Workloads without realistic requests and limits are the most common root cause of unstable clusters, from scheduling loops to frozen nodes. The node auto-provisioner sizes nodes from your Pods’ resource requests, so implausible requests cause confusing failures:

  • Missing or very small requests make Cloudfleet pick the smallest available instance. System components can use nearly all of a small node and leave no room for your Pod. Nodes are then added and removed in a loop. Set realistic CPU and memory requests on all workloads, or constrain the Fleet to a minimum instance size.

  • Impossibly large requests have the opposite effect. For example, a misconfigured VerticalPodAutoscaler can set a container’s memory request to an effectively unbounded value. The scheduler then reports Insufficient memory on a nearly empty node. Compare the Pod’s requests with the node’s allocatable resources:

    bash
    kubectl get pod POD_NAME -o jsonpath='{.spec.containers[*].resources.requests}'
    kubectl describe node NODE_NAME | grep -A 6 Allocatable

    Fix the request (and any VPA maxAllowed policy), then recreate the Pod. If you restart the Pod without fixing the source, the bad value comes back.

  • Single-architecture container images crash-loop when Cloudfleet provisions a cheaper ARM node. See Handle single-platform container images.

Things that make it worse

  • Never rename or modify Cloudfleet-provisioned machines in the provider console. Cloudfleet tracks machines by their creation name. Cloudfleet treats a renamed server as lost, which causes phantom node churn and orphaned machines. Make all capacity changes through Cloudfleet.
  • Never delete NodeClaims manually. When you delete a NodeClaim, its node is removed. The system Pods on that node become unschedulable and trigger replacement provisioning, which can look like runaway node creation.

Still stuck?

If the checklist does not show the cause, check the Events tab in the Cloudfleet Console for provisioning errors and contact support with your cluster ID and the Pod’s FailedScheduling events.

On this page