Skip to content

feat(MHC): add MachineHealthChecks to the templates - #31

Open
Herbaert wants to merge 2 commits into
mainfrom
feat/machine-health-checks
Open

Herbaert wants to merge 2 commits into
mainfrom
feat/machine-health-checks

Conversation

@Herbaert

Copy link
Copy Markdown
Collaborator

Closes #7

What

Templates / ClusterClass

  • A MachineHealthCheck for control-plane and worker Machines in every flavor (cluster-template*.yaml), and healthCheck for both classes in clusterclass.yaml:
    • nodeStartupTimeoutSeconds: 600
    • Ready=False and Ready=Unknown after 600s
    • workers only: remediation.triggerIf.unhealthyLessThanOrEqualTo: 40%
  • MachineDeployment spec.remediation.maxInFlight: 2.
  • nodeDrainTimeoutSeconds: 900 for control-plane and worker Machines.
  • Docs: cluster-template.md, clusterclass.md and cni.md. The CNI has to be installed within 10 min of the Nodes joining.

Provider

  • Non-ACTIVE servers used to report Provisioning: server state is X. They now get a reason that matches the state, e.g. InstanceStopped, InstanceFailed, InstanceStarting or InstanceBusy. This applies to StackitMachines and the bastion.
  • The server powerStatus is now read. An ACTIVE server with power status CRASHED or ERROR is reported as not ready (InstanceCrashed / InstancePowerError).
  • A Warning event is emitted when a server enters a state that does not resolve on its own (ERROR, INACTIVE, DEALLOCATED, PAUSED, RESCUE, crashed). Events are deduplicated by condition reason.

Why these values

A manual run on a real STACKIT cluster (1 control plane, 3 workers) tested each failure without and with a Ready=Unknown check:

Scenario Server state without Unknown with Unknown
containerd stopped (Ready=False) ACTIVE replaced after 10 min, but the drain hung 36 min –
VM deleted out of band 404 replaced after ~1 min (NodeDeleted) –
kubelet dead ACTIVE never replaced replaced after 10 min
VM stopped INACTIVE never replaced; the CCM does not delete the Node replaced after 10 min
hung VM (kernel panic) ACTIVE never replaced replaced after 10 min
1 worker cut off from the API server ACTIVE – healthy VM replaced (false positive)
all workers cut off ACTIVE – blocked during the outage; 2 healthy VMs replaced on recovery
  • Ready=Unknown: a dead kubelet or a hung VM produces no signal other than Ready=Unknown. The STACKIT server state stays ACTIVE. Losing the connection between the management and workload cluster does not cause remediation, because the MHC stops before it evaluates any targets.
  • Drain timeout: without it, remediating a Ready=False node hangs for as long as any pod cannot terminate. That happens, for example, when the container runtime is down.
  • triggerIf: it limits mass remediation during network outages, but does not prevent it entirely once nodes recover one after another.

Behaviour to be aware of

  • CAPI rounds 40% down. A MachineDeployment with fewer than 3 workers is never remediated, including Ready=False and deleted VMs. The docs recommend WORKER_MACHINE_COUNT ≥ 3.
  • nodeDrainTimeoutSeconds applies to every Machine deletion, including upgrades and scale-down. Pods still protected by a PodDisruptionBudget after 15 min are no longer waited for.

Testing

  • make test (new table test for the state mapping; envtest cases for STOPPING → INACTIVE with exactly one warning, ACTIVE + CRASHED, and a failed bastion) and make lint
  • Templates and ClusterClass applied with kubectl apply --dry-run=server --validate=strict against the CAPI v1.13.2 CRDs; all new fields are kept
  • make test-e2e-workload-topology passes with the final changes
  • The manual remediation run described above

@Herbaert
Herbaert requested review from a team and tuunit September 30, 2026 17:20

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add MachineHealthChecks to the templates

1 participant