Owning GPU servers doesn't make a GPU cloud. The business begins when a customer can request capacity, receive an isolated environment, run work at a predictable level, see usage and get support without an engineer rebuilding the service by hand each time.
That distinction catches new operators out. They budget for accelerators and rack space, then discover that provisioning, identity, tenant networks, image hygiene, metering, billing, storage, abuse controls and failed-node recovery consume the team's time. The hardware may be fast while the service remains impossible to operate profitably.
Start with a narrow product and a measurable service promise. A provider that can deliver one dependable bare-metal GPU product to a well-defined customer usually has a stronger starting point than one advertising bare metal, virtual machines, fractional GPUs, managed Kubernetes and serverless inference before any of them has an operating model.
Choose what customers can actually buy
“GPU cloud” can describe several products that share hardware but demand different controls.
| Product | Customer receives | Operator burden | Suitable first offer? | | --- | --- | --- | --- | | Dedicated bare metal | A whole physical server for one tenant | Provisioning, network, storage, health and remote access | Often yes | | Dedicated VM | One or more whole GPUs passed to a tenant VM | Hypervisor, image lifecycle, isolation, network and storage | Yes, with mature virtualisation | | Fractional GPU | A partition or time-sliced share | Scheduling, noisy-neighbour control, compatibility and detailed metering | Usually later | | Managed Kubernetes | A tenant cluster or namespace plus GPU workers | Control plane, upgrades, policy, observability and support | Only with Kubernetes operations staff | | Managed inference | An endpoint, model and service objective | Model lifecycle, routing, autoscaling, latency, token accounting and application support | A separate platform business |
NVIDIA's Cloud Partner software reference separates the stack into Infrastructure as a Service, Container as a Service and AI Platform as a Service. That separation is commercially useful. Each layer adds customer convenience, but it also moves more responsibility onto the provider.
A new operator should write one sentence that names the sellable unit. Examples include “a dedicated eight-GPU server billed monthly”, “a single H100-class GPU passed through to a VM”, or “an isolated Kubernetes worker pool billed per reserved GPU-hour”. If the engineering and sales teams describe different units, the invoice will expose the disagreement.
Calculate billable capacity before revenue
Installed GPU-hours are not billable GPU-hours. A simple capacity model is:
Billable GPU-hours = installed GPUs x period hours x health rate x occupancy x billable allocation rate
Suppose an operator installs 64 GPUs for a 30-day month:
- Installed capacity:
64 x 720 = 46,080 GPU-hours - At 97% hardware and platform health:
44,697.6 GPU-hours - At 70% paid occupancy:
31,288.3 GPU-hours - After reserving 5% of occupied time for testing, migrations and service work: about
29,724 billable GPU-hours
The example doesn't predict a real fleet. It shows why a revenue model built on 46,080 saleable hours overstates capacity by more than a third under quite respectable operating assumptions.
Now divide the full monthly cost by billable hours:
Cost per billable GPU-hour = (capital recovery + facility + network + storage + software + staff + support + finance) / billable GPU-hours + variable usage cost
Run the model at several occupancy levels. A service that works only at 90% paid utilisation has little room for customer churn, delivery delays or maintenance. The go/no-go case should survive a weak first year, not just the steady state presented to investors.
Hardware must match the product
Dedicated bare metal favours complete nodes with predictable GPU topology, local NVMe and enough network capacity for the advertised workload. Shared VM or Kubernetes products need an equally deliberate CPU, RAM and storage ratio because tenants consume host resources alongside the GPU.
GPU choice follows the service sold:
- Large training customers may require eight-GPU HGX nodes, NVLink/NVSwitch and a non-blocking multi-node fabric.
- Fine-tuning, visual computing and independent inference jobs may fit PCIe GPU servers better, especially where customers rent one or two devices.
- Fractional inference needs a GPU and software combination that supports the intended partitioning method. NVIDIA MIG provides isolated GPU instances on supported accelerators, but instance profiles, memory sizes and software compatibility vary by generation.
- A mixed fleet increases sales coverage while multiplying images, drivers, firmware, spare parts and scheduler rules. Add a second platform only when customer demand pays for that complexity.
Use the GPU server configurator to inspect node options, then model the surrounding fabric and storage before fixing a retail price. A cheap server with the wrong topology can become the most expensive node in the fleet.
Separate tenant, operator and hardware control planes
NVIDIA's NCP reference describes a tenant view and an operator view. Keep that division visible in the design.
The tenant needs a stable API or portal for capacity requests, credentials, images, networks, volumes, console access, usage and support. The operator needs inventory, topology, health, provisioning, firmware, telemetry, billing events and break-fix controls. The baseboard management network sits below both and must never become a tenant shortcut.
At minimum, define separate paths for:
- East-west GPU traffic between nodes.
- Tenant north-south traffic and public access.
- Storage.
- Provisioning and in-band management.
- Out-of-band management for BMCs, switches and power systems.
Combining those functions without bandwidth and security boundaries makes failures hard to diagnose. A storage burst shouldn't block BMC access; a tenant network change shouldn't expose the provisioning plane.
Tenant isolation is a product promise
Dedicated physical nodes provide the clearest boundary and usually the best performance. They still need secure wipe, image provenance, network separation, credential rotation and a documented hand-back process.
Virtualised or shared hosts need more. NVIDIA's workload-isolation guidance warns that multiple tenants on one physical host carry greater security risk than single-tenant placement and calls for hardened hypervisors plus protection against contention across GPU, CPU, memory, network, storage and address space. BlueField DPUs and hardware-accelerated network controls can help, but equipment doesn't replace policy.
Write down the isolation boundary for each service:
| Layer | Question the service must answer | | --- | --- | | GPU | Whole device, MIG partition or time slice? What can another tenant contend for? | | CPU and RAM | Dedicated NUMA resources or shared host pool? | | Network | Separate VRF/VPC, VLAN, security policy and public IP behaviour? | | Storage | Dedicated volume, encrypted tenant pool or local ephemeral media? | | Control plane | Dedicated tenant control plane or shared service with enforced quotas? | | Operations | Which staff can access hosts, consoles, logs and tenant data? |
Security claims belong in the contract only after engineering has tested them. “Isolated” is too vague unless the provider can point to the mechanism and the remaining shared resources.
Provisioning must be repeatable
Phone calls and spreadsheets don't scale into cloud operations. NVIDIA's current requirements for AI clouds call for programmatic capacity discovery, stable node identifiers, health state, allocation, region and topology-aware reservation. Those are sensible minimums even when the fleet starts small.
A practical provisioning flow should:
1. Select capacity that matches the requested GPU, topology, network and locality. 2. Reserve the complete allocation atomically so half a multi-node cluster can't be sold elsewhere. 3. Apply a known firmware and host-image baseline. 4. Create tenant network and storage attachments. 5. Inject credentials through a controlled path. 6. Run health and performance checks before hand-off. 7. Start metering only after the resource reaches the contracted state.
Deprovisioning needs the same care in reverse. Stop billing at the defined point, revoke credentials, preserve agreed logs, sanitise storage, release network addresses, test the node and return it to the available pool.
Meter the thing named on the invoice
Capacity state and billing state must agree. Track at least delivered, healthy, reserved and active capacity, together with tenant, project, node, GPU, region, start time and end time. Record why capacity became unavailable and whether the event counts against the service objective.
For reserved bare metal, billing may use server-months with overage for network or storage. VM products may use GPU-hours, vCPU, RAM, volumes, snapshots and egress. Managed inference might charge for reserved capacity, tokens or endpoint time, but that requires application-layer measurement and a much more involved support boundary.
Avoid pricing a metric that operations can't reproduce. A customer should be able to compare the invoice with portal usage, and support should trace both back to immutable events.
Storage is part of the service, not an accessory
GPU tenants need somewhere to keep container images, datasets, checkpoints, models and output. Local NVMe works well for cache and temporary data but complicates migration and recovery. External storage needs tenant quotas, access controls, snapshots, failure handling and enough throughput to prevent paid GPUs waiting for data.
Decide which classes the service will offer:
- Ephemeral local scratch for maximum node-local performance.
- Persistent block volumes for VM and database workloads.
- Shared file storage for training datasets and checkpoints.
- Object storage for artefacts, archives and data pipelines.
State durability and backup separately. A persistent volume isn't automatically backed up, and replication doesn't protect against every deletion or tenant mistake.
Break-fix determines whether the SLA is credible
NVIDIA's NCP break-fix architecture describes a cycle of detecting, triaging, remediating and validating unhealthy nodes. Build that loop before the first contractual uptime promise.
Health collection should cover GPU errors, thermals, memory, NIC links, storage, host state and the services needed to provision the node. Automation can cordon a node, preserve diagnostic evidence, move or stop workloads, run field diagnostics and decide whether the node returns to service or enters RMA.
Keep two states separate: a server may respond to monitoring while one GPU, NIC rail or NVMe device has failed. If the sellable product requires the whole topology, that node isn't healthy capacity.
Spare strategy follows the repair model. Same-day onsite support still leaves diagnosis, approval and replacement time. A provider selling immediate replacement capacity needs warm spare headroom or another pool that can accept the workload.
Price support and people honestly
The platform needs on-call coverage, hardware diagnosis, network operations, storage operations, security response, image maintenance, billing support and customer communication. One engineer may cover several jobs during launch, but the cost model still needs those hours.
Support tiers should say what the clock measures. “24/7 support” might mean ticket intake, first human response, service restoration or hardware replacement; those are different promises with different staffing and spare requirements.
Include sales engineering and onboarding. Multi-node customers often need help with images, drivers, SSH, VPNs, storage mounts, scheduler integration and collective tests before productive use begins. If that work isn't priced, high-touch customers can consume the margin from the hardware rental.
A sensible launch sequence
Start with design partners whose workloads resemble the intended service. Run paid or clearly bounded pilots; free, open-ended trials produce noisy utilisation data and support habits that don't survive commercial launch.
Measure provisioning time, healthy capacity, occupancy, failed allocations, time to detect faults, time to restore service, support hours per customer and invoice disputes. Those figures tell the operator which automation to build next.
Don't add fractional GPUs because spare capacity looks idle. First find out why it is idle. Sales mismatch, inconvenient contract terms, weak onboarding, missing regions or unreliable service won't be repaired by a finer scheduler.
Go/no-go questions
The business isn't ready to launch until it can answer these:
- What exact resource does the customer reserve, and when does billing start and stop?
- How does the operator prove isolation and sanitise capacity between tenants?
- Which workloads fit the fabric, storage and GPU topology being sold?
- How are unhealthy GPUs removed from sale and returned after validation?
- What occupancy keeps the service cash-positive after staff, finance and support?
- Can every invoice line be traced to usage or reservation records?
- What happens when a customer needs capacity that the current pool can't supply?
Where GPUMachines fits
GPUMachines can help specify and source the physical layer: PCIe GPU servers, HGX systems, high-speed networking, storage and rack integration. The Buy & Host service may suit operators that want dedicated hardware without building the first facility footprint themselves.
The software, security and commercial service still need an owner. A hardware quote should sit beneath the product definition, isolation model, capacity calculation and operating plan described above.
FAQ
How many GPUs does a new provider need?
There is no responsible universal minimum. Start from the smallest fleet that can serve the chosen product, preserve maintenance headroom and meet customer topology requirements. Eight GPUs may support one dedicated-node offer; a training cloud with multi-node reservations needs enough homogeneous nodes to allocate complete blocks.
Is Kubernetes required?
No. Bare-metal and VM services can operate without exposing Kubernetes to customers. Managed Kubernetes and inference platforms benefit from it, but they also require control-plane, upgrade and policy skills.
Should a provider offer fractional GPUs first?
Usually not. Whole-GPU or whole-node products simplify performance, isolation and billing. Fractional services make sense after the operator has a clear workload, supported profiles and scheduling controls.
What utilisation target should go into the plan?
Model several occupancy levels and separate paid occupancy from hardware health. Don't call maintenance, internal testing or unsold reserved capacity “utilisation” if it produces no revenue.
Can GPUMachines operate the cloud platform?
Confirm the required scope directly. GPUMachines can design, configure, supply and host infrastructure; managed tenant portals, billing, platform software and day-two operations must be named explicitly in the proposal.
Sources
- NVIDIA Cloud Partner software reference guide
- NVIDIA workload-isolation guidance
- NVIDIA break-fix architecture
- NVIDIA requirements for AI clouds
- NVIDIA Multi-Instance GPU user guide
Discuss a GPU cloud hardware and hosting plan with GPUMachines.
.jpg)