Control plane pressure and API throttling

The Kubernetes API server protects itself under load with API Priority and Fairness. It queues requests per flow and rejects them when queues overflow. On CFKE, the most common source of control plane pressure is platform tooling, not application traffic:

  • GitOps controllers reconciling too often
  • operators re-listing large object sets
  • tools that store large objects in the cluster

This page shows how to recognize the condition and the fixes that resolved real support cases.

For hard object size and data store limits, see Kubernetes API server quotas.

Recognizing control plane pressure

Typical symptoms, often appearing together:

  • controllers receive HTTP 429 (throttled) or intermittent 500/504 responses
  • multiple controllers on different nodes lose leader election at the same time and restart: Failed to renew lease ... context deadline exceeded
  • kubectl feels slow while /readyz reports healthy
  • watch connections close repeatedly and controllers log reconnects

Confirm with the API server’s flow control metrics:

bash
kubectl get --raw /metrics | grep apiserver_flowcontrol_rejected_requests_total
kubectl get --raw /metrics | grep apiserver_flowcontrol_current_inqueue_requests

Non-zero rejections with reason="time-out", typically in the service-accounts or workload-low priority levels, mean requests time out in the queue. The control plane is saturated by request volume, not broken.

Fix reconcile storms from GitOps tooling

One misconfiguration can generate thousands of API requests per hour. Check each tool you run:

Flux:

  • For stable releases, raise spec.interval on HelmReleases and Kustomizations to 10 minutes or more. One-minute intervals re-list everything they manage every 60 seconds.

  • Fix permanently failing reconciliations. A Kustomization that fails every minute re-applies the entire repository every minute. A common example is a SOPS-encrypted secret without a decryption block on the root Kustomization (Secret/... is SOPS encrypted, configuring decryption is required):

    yaml
    spec:
      decryption:
        provider: sops
        secretRef:
          name: sops-age
  • After committing fixes, resume anything suspended: flux resume kustomization NAME -n flux-system.

Argo CD:

  • Remove empty parameters: [] blocks under spec.source.helm. Argo CD treats them as a permanent diff and does a full hard refresh (git re-clone and complete state read) on every sync cycle.
  • Make exactly one Application own cluster-scoped CRDs. If multiple Applications install the same chart at different versions, they rewrite the CRDs on every sync. This generates tens of thousands of updates and closes everyone else’s watch connections.

Operators:

  • Set sensible reconcile periods instead of the default of about one minute, for example the helm.sdk.operatorframework.io/reconcile-period annotation or the operator’s own interval setting.
  • Avoid seconds-level spec.interval on custom resources. 28 resources re-checked every 3 seconds is a request storm.
  • Scope kube-state-metrics with --resources so it does not watch secrets, configmaps, and leases it does not need.
  • If controllers lose leases under load, widen the leader election settings (leaseDuration, renewDeadline, retryPeriod). Also stagger resync intervals so operators do not re-list together.

Fix object bloat

The API server keeps stored objects in memory. Tools that write many large objects slow down the whole control plane:

  • Trivy Operator: SBOM generation creates one large report object per workload revision. If you do not consume SBOMs, disable it (OPERATOR_SBOM_GENERATION_ENABLED=false), delete existing sbomreports and clustersbomreports, and set scannerReportTTL. Vulnerability scanning is unaffected.
  • Velero: frequent full-cluster filesystem backups with long retention collect thousands of backup objects. Reduce schedule frequency and retention.
  • Log collectors: do not pull logs through the Kubernetes API from every node. Configure agents such as Grafana Alloy to discover only Pods on their own node and to tail log files from the host (/var/log/pods) instead of the API.

“Forbidden: this resource is crucial” errors

Symptom: a reconciler loops with errors such as:

daemonsets.apps "cilium" is forbidden: ... This resource is crucial for correct
functioning of your cluster. You are not allowed to create, update, or delete
this resource in CFKE.

Cause: CFKE write-protects the platform components it manages: the CNI, metrics-server, cloud controller, node provisioning objects, and admission webhooks. By design, CFKE rejects tools like Rancher, Flux, or Helm that try to adopt or reinstall them.

Fix:

Remove the duplicate installation from your GitOps configuration. The platform-managed copy already provides the function. For example, kubectl top and HPA work out of the box with the managed metrics-server. Do not attempt to grant the tool write access.

The errors are harmless to the cluster but cause steady API load and log noise. If you cannot remove the source quickly, filter the messages in your log pipeline.

The same applies to node provisioning objects. Do not manage NodePool or NodeClass objects with your own charts. Define Fleets instead.

When capacity is the answer

Basic clusters share control plane resources with other customers. Heavy controller stacks (GitOps, policy engines, monitoring, backup, service catalogs) can outgrow them even after tuning. Pro clusters run dedicated, multi-replica API servers that handle significantly higher request volume. If flow control rejections continue after the fixes above, consider upgrading the cluster tier or splitting operator-heavy stacks across clusters.

On this page