Skip to main content

Control plane machines

A tenant cluster's control plane usually runs as a workload on an existing Kubernetes cluster. It can instead run directly on machines that vMetal provisions. vCluster calls this a standalone control plane, and the platform installs it on the server as a systemd service rather than as pods. For what a standalone control plane is and how vCluster runs on the server it lands on, see vCluster Standalone.

This is the path for an operator whose hardware is the starting point. You do not need a Kubernetes cluster to host tenant control planes, only the control plane cluster that runs Metal3 itself.

Control plane machines are still Machines. They use the same node types, the same server pool, and the same lifecycle as any other. See Machine and server concepts.

Modify the following with your specific values to replace them across the whole page:

Before you begin​

Work through Set up your hardware steps 1 through 5 first, so you have a NodeProvider and servers in available state. Step 6 there is a branch point that routes to one of three guides. This is the one for servers that carry a tenant cluster's control plane. It picks up where step 5 ends.

The NodeProvider, tenant cluster configuration, and Machine commands in this guide use the management cluster context. SSH and systemd checks run on the provisioned server itself.

You also need a node type whose servers are suitable for a control plane. Control planes want reliable, modestly sized machines rather than your largest GPU hosts. Select them by label the same way you would any other node type. See Configuration.

Decide this before you create the tenant cluster

controlPlane is immutable on a Machine. You cannot convert a worker Machine into a control plane Machine, or the reverse. The platform creates control plane Machines from the tenant cluster's own configuration, so the decision is made when you create the cluster.

1. Define a control plane node type​

Add a node type for the servers you want control planes to land on.

nodeTypes:
- name: "control-plane-node"
displayName: "Control Plane Node"
resources:
cpu: "8"
memory: 32Gi
bareMetalHosts:
selector:
matchLabels:
role: control-plane
properties:
vcluster.com/os-image: ubuntu-noble
vcluster.com/ssh-keys: ops-key

The install script needs a Linux distribution with systemd, plus curl, iptables, and root access. The node also needs raised inotify watcher limits. Use an OS image that covers this, and see Node requirements for the authoritative list. The script detects ARM and x86 hosts on its own.

Include SSH keys. A control plane machine has no kubelet joining it to anything, so if the install fails you reach the server over SSH to read the logs. See Register an SSH key.

2. Create the tenant cluster in standalone mode​

Point the tenant cluster's control plane at your NodeProvider. The platform creates one Machine per requested control plane node and installs vCluster on each.

controlPlane:
standalone:
enabled: true
autoNodes:
provider: metal3-provider
quantity: 3
nodeTypeSelector:
- property: vcluster.com/node-type
value: control-plane-node

The nesting is a fixed path the platform reads. Four values sit under it:

  • enabled selects standalone control plane mode. Without it, the platform does not create control plane Machines from autoNodes.
  • provider names the NodeProvider to claim servers from. The platform reports NoControlPlaneNodeProvider on the cluster and provisions nothing if you leave it out.
  • quantity is how many control plane machines to claim. It must be at least one. A quantity of zero is rejected with NoControlPlaneNodeQuantity.
  • nodeTypeSelector narrows which node type the machines come from. Without it, the platform picks from anything the provider offers.

Set quantity to three or more for a highly available control plane. The platform keeps the peer set and its PKI in sync across every control plane machine it creates, so members find each other without manual configuration.

3. Watch the control plane come up​

The platform names these Machines after the cluster, with a -cp- infix, and labels them loft.sh/vcluster-control-plane with the cluster name.

kubectl --context <management-context> get nodeclaims -n loft-p-my-project \
-l loft.sh/vcluster-control-plane=my-cluster

Control plane machines provision before the tenant cluster is reachable, which is the reverse of a worker. A worker Machine waits for the cluster to come online before it provisions, because it needs a live API server to join. A control plane Machine does not wait, because it is what brings the cluster online.

For the same reason, a control plane Machine only needs its network environment's infrastructure to be ready. It does not wait for Kubernetes to be provisioned there, the way a worker does.

The cluster controller reports VirtualClusterSynced and VirtualClusterReady after it has created the requested control plane Machine objects. Those conditions do not mean the servers have finished provisioning. Keep watching the Machines until they reach Available, then verify that you can connect to the cluster.

4. Add workers​

A standalone control plane serves workers the same way any tenant cluster does. Add them from the same fleet or a different node type, through the same privateNodes configuration. See Tenant cluster nodes.

Workers wait for the control plane to come online before they provision, so expect them to follow rather than arrive together.

Lifecycle​

Control plane Machines are owned by the tenant cluster that created them. Deleting the cluster deletes its control plane Machines, and the driver cleans those servers and returns them to the pool.

Reducing quantity does not remove existing control plane machines. The platform creates machines to reach the requested count and does not delete them to shrink it. Removing a control plane member is a deliberate operation on the cluster, not a value change.

The underlying server still follows the normal BareMetalHost lifecycle. See BareMetalHost states.

Troubleshooting​

The install runs from cloud-init at first boot, so failures show up on the server rather than in the platform.

  • The Machine provisions but the cluster never comes online. Follow Control plane Machine is available but the cluster is offline. The script checks for curl, systemd, and root access up front and exits with a clear message when one is missing. It does not check the rest, so a missing iptables or an inotify limit left at the distribution default fails later and less obviously.
  • The cluster reports NoControlPlaneNodeProvider or NoControlPlaneNodeQuantity. The controlPlane.standalone.autoNodes section is missing a required field. Nothing is provisioned until you fix it.
  • The Machine never leaves Pending. No server matched the node type selector. Check that servers carrying the right labels are in available state. See Troubleshooting.