GPUaaS onboarding · platform powered by hosted.ai

Bring the hardware.
We run the cloud.

This guide is for operators connecting GPU servers to a managed GPU cloud. The control plane runs on the hosted.ai GPUaaS platform and is operated for you. Your job is to stand the nodes up to spec. They're onboarded as workers, start serving tenants, and the platform scales as you add nodes. The whole thing can be white-labelled under your own brand.

CONTROL PLANELive · fully managed
YOUR ROLEGPUaaS worker nodes
BRANDINGWhite-label ready
BASE OSUbuntu 24.04
WORKLOADSContainerised
Architectural reference

One control plane.
Nodes that scale out.

The managed Admin/User panel sits separately from the GPU nodes and connects to each region as hardware is onboarded. Within a region, one node can run the controller, worker and storage roles together; additional nodes join as workers. The diagram below is a reference pattern. Final network design is aligned with your standards during deployment.

ADMIN / USER PANEL MANAGED CONTROL PLANE · YOUR BRAND tenancy · quota · metering · access · workload lifecycle REGION eu-west-4 GPUaaS Node 1 … (multi-region) MANAGEMENT private · health · scheduling · metrics REGION eu-west-2 GPUaaS NODE 1 Ubuntu 24.04 ROLES · GPUaaS Controller · GPUaaS Worker · GPUaaS Storage Service SERVICE GATEWAY · PUBLIC 122.111.222.111:32456 GPU WORKLOAD · containerised 2 vCPU · 64GB · 2 vGPU · vLLM 2TB persistent · 200GB ephemeral PHYSICAL GPUs 192.168.5.151 GPUaaS NODE 2 Ubuntu 24.04 ROLE · GPUaaS Worker 192.168.5.152 GPUaaS NODE 3 Ubuntu 24.04 ROLE · GPUaaS Worker 192.168.5.153 NODE NETWORK · 192.168.5.0/24 POD OVERLAY · 192.168.100.0/24 · IP-in-IP tunnel Service Gateway 122.111.222.111 Node Network 192.168.5.0/24 Pod Network 192.168.100.0/24 Cluster CIDR 10.96.0.0/12

Reference pattern adapted from the hosted.ai GPUaaS architectural reference. Addresses shown are illustrative; your node, pod and service ranges are confirmed during deployment.

Scope & node roles

What your hardware
becomes on the platform.

A node can carry one or more roles. Your first node in a region typically carries all three; every node after it joins as a worker and adds GPU capacity to the pool.

Controller

Runs the in-region control functions and talks to the Admin/User panel over the management network, handling node management, scheduling and health reporting.

Worker

Hosts the actual containerised GPU workloads. Every additional node is a worker, so more nodes means more GPU capacity for tenants. Workers are added as demand grows.

Storage Service

Presents persistent block storage to GPU workloads. Storage can be local to the node or presented over the network, depending on your architecture and target workloads.

Networking

Two networks in,
one network out.

The platform needs connectivity between the control plane and your GPU nodes, plus external connectivity for customer workloads. These are typical reference patterns; final design is aligned with your standards during deployment.

Management network

Private connectivity between the control plane and your GPU nodes, used for node management, scheduling and health reporting. Typically a standard VLAN or equivalent segmentation.

Service / public network

Carries customer access to workloads and exposed services. Can be backed by public IPs, NAT, load balancers or private connectivity. Outbound access for image and model pulls follows your security policies.

High-performance fabric · optional

For multi-node or latency-sensitive workloads, high-bandwidth interconnects such as NVLink, InfiniBand or RoCE can be supported where available.

Node network
Private range, e.g. 192.168.X.X/24
Pod network
192.168.100.0/24 overlay (predefined)
Service CIDR
10.96.0.0/12 (predefined)
Service gateway
Public IP, e.g. node 122.111.222.111
Storage layout

Three tiers
per node.

Persistent storage may be local or presented over the network, depending on your architecture and the workloads you intend to run.

Root disk

1 TB or larger NVMe preferred (SSD optional) for the OS and platform services.

Ephemeral

GPU workload scratch disk: fast, transient space for the running container.

Persistent

GPU persistent block storage, local to the node or network-presented.

Scaling

Add nodes
as demand grows.

Depending on the deployment model and target workloads, the platform scales in several directions, all under the same control plane.

01 / MORE WORKERS

GPU nodes on demand.

GPUaaS worker nodes are added as demand for GPUs grows. Each new node to spec joins the pool and starts serving tenants.

02 / HORIZONTAL + MULTI-REGION

Scale out, scale across.

Horizontal scaling and multiple regions are supported. The same panel manages nodes in eu-west-2, eu-west-4 and beyond.

03 / HPC INTERCONNECT

Low-latency multi-GPU.

Low-latency, multi-node, multi-GPU workloads are supported over NVLink and InfiniBand where the fabric is in place.

04 / DEDICATED TENANTS

Hardware tenant clusters.

Dedicated hardware tenant clusters can be carved out for customers that need isolated, reserved capacity.

Sizing & oversubscription

Pick the path
that fits your hardware.

GPU sizing and oversubscription work two ways. The right one depends on whether your hardware specs are still flexible or already fixed. Either way, the final ratios are agreed jointly before production tenants go live.

Approach A

Target-driven planning

Define a target oversubscription strategy upfront, from conservative to aggressive GPU sharing, and we work backward from it.

  • Advise on the CPU, memory, storage and network characteristics needed to support the target ratio
  • Identify bottlenecks and constraints at different oversubscription levels
  • Recommend guardrails for predictable performance across tenants

Use when specs are still being finalised, or you're designing a new GPU platform from scratch.

Approach B

Hardware-first validation

Supply the final node specs and we validate against them.

  • Evaluate the node design against real-world GPUaaS usage
  • Recommend realistic, sustainable oversubscription ranges
  • Confirm operational limits before enabling production tenants

Use when hardware is already procured or in delivery and the goal is to maximise utilisation without compromising service quality.

How onboarding works

From rack
to serving tenants.

The control plane is already in place. Onboarding a node is a collaborative sequence between your team and ours; we align the network design to your standards as we go.

01

Prep the node

Ubuntu 24.04, root NVMe, GPUs in place, and bonded networking (LACP) with private and public VLANs ready.

02

Share the specs

Send node details and target use cases. We confirm IP ranges, the oversubscription path (A or B) and the role mix.

03

We connect it

We attach the node to the control plane over the management network and bring up the worker and storage roles.

04

Validate & go live

We confirm operational limits, then enable tenants. Ongoing node- and cluster-level operations run with our support.

Start onboarding

Have the hardware?
Let's connect it.

  • Node specs: GPU model, GPUs/node, CPU, memory, storage, fabric
  • Target region and rough capacity plan
  • Your networking standards for management and service VLANs

Send those over and the onboarding team comes back with confirmed network ranges, a sizing path, and a deployment window for your region.