From the makers of

Turn Your Racks Into EC2

Bare metal machines and VMs for every cluster you run, through one stable API. Powered by the provisioning stacks you need today, and any you’ll need tomorrow.

Provision and PXE boot bare metal servers, plus VMs
Manage declaratively or through a flexible UI
Runs on Metal3, KubeVirt, NVIDIA NICo, and more
Commercial support behind the open-source drivers
Automate networking, DNS, and machine lifecycle
Hard tenant isolation down at the hardware layer
Trusted by the fastest-growing AI cloud providers

One machine layer under every cluster you run.

vMetal provisions the bare metal and VMs. vCluster turns them into products you ship.

Clusters

Ship any cluster as a product

Kubernetes Clusters

Nested Clusters

Oss

Slurm Clusters

Beta

Run:AI Clusters

Ray Clusters

Inference Clusters

Dynamo · llm-d · + more

Soon

Agent Sandbox Clusters

Soon

+ more clusters

Platform for Tenant Management

Operate every tenant at scale

Cluster Templates

Capacity Management

Observability

Node Autohealing

Soon

Billing

Alpha
One stable API to consume
Machines

Provision hardware like a cloud

Bare Metal Machines

Virtual Machines

Node provisioning & lifecycle management
Drivers

Future-proof your node provisioning

Bare Metal Provisioning Drivers

+ more

Network Automation

VM Provisioning Drivers

+ more
YOUR GPU INFRASTRUCTURE
Bare metal GPU & CPU servers · Networking · Storage

Not another provisioning tool. A provisioning orchestrator.

vMetal isn't a provisioning tool itself. It's the orchestration layer that manages your provisioning, network, storage, and virtualization technologies, and solves their integration under one consistent API.
Future-proof your stack

Run several systems at once or swap one out. Whatever the future brings, nothing above has to change.

One consistent, EC2-like API

A stable interface upward, no matter what you change underneath.

Integrate OS, network, and storage

Bring your compute, networking, and storage layers together with ease.

Commercial support for all OSS

We stand behind every open-source technology in the stack, in production.

Proven architecture

Deep, battle-tested experience across every possible integration.

Orchestrate any technology

Metal3, KubeVirt, NVIDIA NICo, VMware, OpenStack, and more, under one layer.

Building the Machine Layer Means Rebuilding EC2

Each capability below is months of specialized engineering, and every one has to be rebuilt whenever your hardware or provisioning stack changes.

Machine-layer capability

Typical engineering effort

Time to Build

Business impact if you DIY

Bare metal provisioning & PXE automation

2–3 infrastructure engineers

3–4 months

GPUs wait on a provisioning layer you are still building

Multi-driver abstraction (Metal3, NICo, KubeVirt)

2–3 platform engineers

4–6 months

Vendor lock-in and a re-architecture when hardware changes

VM provisioning alongside bare metal

1–2 engineers

2–3 months

Half your estate left unsupported

Network automation & tenant isolation

1–2 networking engineers

2–3 months

Complex networking increases operational risk

Machine lifecycle & autohealing

1–2 platform engineers

2–3 months

Manual ops and longer time to recover failed nodes

What a Solid Machine Layer Gets You

Launch on Your Hardware Faster

Turn racked GPUs into an on-demand machine API in weeks.

Open, No Lock-In

Open drivers, standard API, your hardware. Never re-architect to escape a vendor.

Offer Any Cluster on One Foundation

Kubernetes, Slurm, Ray, and inference all run on the same machine layer.

Your Cloud, Your Margins

An EC2-class experience at your margins, with data residency at the infra layer.

Run It All With One Team

Every stack and GPU supplier from one control plane. Scale without scaling headcount.

A Team That’s on the Call

Commercial support behind every driver, from a team that runs the biggest AI clouds.

100K+
GPUs powered
50+
GPU clouds & Fortune 500s
40+
Hardcore infra engineers
1 msg
Away on Slack or a call

Guides, Reference Architectures, and Solutions

Everything you need to stand up the machine layer for a GPU cloud on your own infrastructure.

EBOOK
NVIDIA DGX Reference Architecture

A blueprint for bringing cloud-grade elasticity and automation to NVIDIA DGX systems.

Solution
Automate Network Isolation for Hard Tenant Isolation

vMetal and Netris integrate machine provisioning with network automation for clean per-tenant isolation.

Guide
Achieve a ClusterMAX™ Platinum Rating

Deliver enterprise-grade infrastructure for AI workloads and improve your ClusterMAX™ rating.

Give Every Cluster Its Machines, On Your Own Hardware

One stable API for bare metal and VMs. Any provisioning stack. Fully supported in production.