Operator guide · Private cloud setup

Set up your private cloud.

This page takes a datacenter operator from nothing to a running, branded, multi-tenant GPU cloud — including where every byte of your data lives. Choose a deployment model, create your operator account, and bring your fleet online.

Step 0 · Choose your deployment model

Managed, sovereign, or on-prem.

Managed LIVE

Fastest launch · Rackify operates

Rackify deploys, hosts, and operates the full stack. You get an operator invite and start enrolling servers the same day.

  • Zero infrastructure on your side beyond the GPU servers
  • Pure 5% revenue share — no license, no seats
  • Rackify runs upgrades, backups, and the drills

Sovereign / dedicated LIVE

Your cloud account · your jurisdiction

The entire platform deploys into your own cloud account and region from Terraform and documented setup — control data, metrics, and logs never leave your account or chosen jurisdiction.

  • Every cloud resource defined in Terraform; the console is code
  • Region-pinned data stores; no shared control plane
  • Operated by Rackify’s forward-deployed engineers, or by your team from the runbook

On-prem / air-gapped ROADMAP

Your metal · no public cloud

The architecture is built for it: agents speak plain HTTPS to one ingest API that owns all store access, so the managed metric and log backends swap for self-hosted equivalents behind the same contract — agents never change.

  • Designed-in backend swap (self-hosted metrics + log stores)
  • Outbound-443-only agents already fit strict egress policies
  • Scoping now with design partners — talk to us

Honest labeling, as everywhere: LIVE means running in production today and documented on this page; ROADMAP means the architecture supports it and it is being scheduled with design partners — never a silent gap.

Steps 1–12 · From invite to live fleet

Create your account, then your cloud.

Create your operator account

An operator is one independent datacenter — yours. It owns its own admin accounts, its own customer tenants, its own servers and its own capacity pool, and no other operator on the deployment can see any of it. Operator accounts are code-gated: Rackify issues your first admin invite code during onboarding, and that code names your operator, so redeeming it creates your datacenter and makes you its first admin. Single-use, valid 14 days, shown exactly once.

  1. Open /register on your console URL.
  2. Enter your email, a password of at least 12 characters, and paste the invite code.
  3. Click Create account — you land signed in, at the operator console (/admin), scoped to your operator.

Security model: passwords are argon2id-hashed; sessions are opaque 30-day tokens stored only as hashes; every agent, operator, and customer holds its own revocable identity. There are no shared credentials anywhere in the platform — and in sovereign deployments, no long-lived cloud keys either (the console reaches your cloud through short-lived, identity-federated credentials).

Add the rest of your team

More seats at your operator come from Rackify: ask for another operator invite naming your datacenter, and whoever redeems it at /register becomes an admin alongside you — same tenants, same fleet, same pool. Codes are single-use, valid 14 days, and shown exactly once.

  • Scoped to you — an invite names the operator it joins, so a colleague never lands in a different datacenter, and never lands in one of their own by accident.
  • Revocable one at a time — removing a teammate’s account kills their sessions first, then the account, and revokes every invite and reset code it handed out. Their enrollment tokens belong to the operator rather than to their seat, so revoke those from the token list on /admin/onboard (it names who issued each one). The servers they installed keep running either way: each node holds its own credential, not theirs.
  • Traceable — every invite minted and redeemed is recorded.

Inside your datacenter, an operator account is root. Anyone you add sees every one of your tenants’ machines and telemetry, your whole capacity pool, and can reassign nodes, issue and revoke the operator’s enrollment tokens (yours included) and create tenants. What they cannot see, ever, is another operator — that boundary is enforced on the server from the session and proven by an automated cross-operator test on every release. There is no narrower role within an operator yet, so keep your own list short.

Tour the operator console

Signing in as an admin lands on Fleet — every machine in your datacenter, and only yours:

  • Stat tiles — total GPUs, online nodes, live power draw (real PromQL against your telemetry), GPUs by state.
  • GPU node map — one cell per node, colored by state (never color alone: offline cells carry ✕, warnings carry !). Hover for detail, click to open the node.
  • Utilization & power charts by site over 24h / 7d / 30d — every chart has a table-view toggle for keyboard and screen-reader use.
  • Events feed & offline banner — enrolls, offline/online flaps, GPU-missing and hardware-swap events surface here.

Everything refreshes on a 15-second cadence without page reloads. Sidebar items marked SOON (Launch AI Factory, Workloads, Incidents, Tickets, Revenue) are roadmap previews.

Launch NeoCloud in the sidebar is a real page and an honest one: four sample storefront designs for renting out your GPUs — bare-metal benchmark, enterprise, price-transparent marketplace, developer-console — so you can pick a direction. Every preview is a mockup: figures are masked, regions are lettered, nothing is wired to your fleet or your rates, and the choice is not saved anywhere. Launching a storefront is not available yet and the button says so — what ships today is the decision, not the store.

Put your own name and logo on your console

Your console can carry your identity LIVE. Open Manage profile (/profile): a Custom product name field and a Console logo upload sit there, and saving either changes the name and mark in the top-left of your console on your next page load. Scope, precisely: this dresses your operator console (/admin and everything under it) and nothing else — your customers’ console, the master console and this website are untouched, because one datacenter putting its mark on its own console must never repaint another’s.

  • Custom product name: up to 64 characters — what you call the console you sell, which is not the same thing as the display name on your own account a card further down the page.
  • Logo: PNG, JPEG or WebP, up to 96 KB. The file’s actual bytes decide what it is — never its name or what the browser declares — so a renamed file is refused rather than trusted.
  • SVG is refused, deliberately: an SVG is a document that can carry script, and your logo is served back from your console’s own origin. Export a PNG at 2× instead.
  • Nothing is invented for you — clear the custom product name and the console falls back to your datacenter’s name, then to the platform’s. Remove the logo and the platform mark returns.
  • Yours alone — the bytes are served behind your session and scoped to your operator; another operator asking for them gets the same answer as for a logo that does not exist.

The name your customers see is a different setting, and it is not yours to change. Your custom product name dresses your console; the platform brand is what their console and every page title says. It is data, not code, so a rename takes effect everywhere immediately — but it is one row for the whole deployment: the separate Branding card (on /admin/tenants, not your profile) renders only for the deployment’s super admin, and the API answers an operator account a plain 403, because a rename by one operator would rebrand every other operator’s customers.

  • Managed: tell Rackify the name you want; we set it, and you see it on your next page load.
  • Sovereign / dedicated: the master console runs inside your own deployment — whoever holds its super-admin account sets the name from that same card.

Per-operator branding of your customers’ console — your name and mark where they sign in, rather than only where you work — is SOON. Until that row is per-operator, one deployment carries one customer-facing name; that is the honest limit, not a setting we have hidden from you.

Issue an enrollment token

Add GPU servers on your fleet page opens /admin/onboard — the guided walkthrough, and the place tokens live. Step 1 there is Get an enrollment token: switch to Create new token, give it a name (and a description if it helps), press Create token, and the token is revealed exactly once, at that moment — together with the install command to paste. Only its hash is stored, so reopening the page never reveals it again.

Hold as many as you find useful — one per datacenter, one per bootstrap script, one for a contractor you revoke afterwards. They belong to your operator, not to your seat: every admin on your team sees the same list, with each token’s name, who issued it and when it last enrolled a machine, and can revoke any of them. Revoking deletes it — the token stops enrolling immediately, and every server already installed with it keeps running, because each machine holds its own credential rather than yours.

Where servers land: every server enrolled with one of your tokens joins your operator’s capacity pool — yours alone, never visible to customers and never to another operator. Your customers rent whole nodes out of your pool from their own console (next steps), and you can reassign any node manually at any time.

Prepare a GPU server

Per box: Ubuntu 22.04 or 24.04+, an installed NVIDIA driver (nvidia-smi works), outbound HTTPS to your console URL (port 443 — no inbound access is ever needed), and a sane clock.

Install the agent — fleet-wide, same command

ubuntu@gpu-node-01
curl -sSL https://<your-console-url>/install.sh | sudo bash -s -- --token <TOKEN>

A token is reusable, so this exact line rolls out a whole rack (add --site <name> per site, e.g. dal-01). The installer preflights the OS, driver, connectivity, and clock with loud, actionable failures; downloads the checksum-verified agent; installs a hardened systemd unit; enrolls; and starts streaming. Each node and its GPUs appear in Fleet — in your capacity pool — within two minutes. The token resolves to your operator server-side, so a server can only ever enroll into the datacenter whose token installed it.

On a shared box, hand the token over in the environment instead — arguments are world-readable through ps: export CHAMBERD_ENROLL_TOKEN with a leading space (keeps it out of shell history), then pipe to sudo -E bash with no --token. The installer forwards it to the agent the same way, so the secret never reaches any process’s argv.

verify on the box
systemctl status chamberd     # active (running), Restart=always + watchdog
journalctl -u chamberd -f     # live agent log

Removal: curl -sSL https://<your-console-url>/install.sh | sudo bash -s -- --uninstall. The installer is idempotent — re-running it is safe. Agents buffer 24 hours of telemetry on disk and backfill with original timestamps after outages (up to a 55-minute replay cap; older gaps are shown honestly).

Put your sites on the map

Every node reports the --site label it was installed with, and the fleet map turns those into markers sized by GPU count and coloured by health. The card is the map: a title, one sentence, the markers and a compact legend. Everything about a site — its nodes, GPU models, aggregate GPU memory, newest heartbeat, online/offline split, open conditions, its coordinates and the button that changes them — is in the box that opens when you hover a marker. Tab to a marker for the same box, press E on a focused one to open its editor, and the box stays open while your pointer or focus is inside it.

A marker comes from one of two places, and the map always says which.

  • Coordinates you setEdit location in a marker’s hover box (or PUT /api/ops/sites/<site>). Solid marker. Yours always win: a position you typed is never replaced or moved by anything below, and clearing it is one click.
  • Where that site’s own agents connect from — for a site nobody has placed, we fall back to the approximate location of the machines’ own connections to this console. Drawn hollow and dashed, prefixed , and the hover box says the position came from the agents rather than from your team.
  • Unplaced — no coordinates and nothing reporting a location, so there is no marker to hover. These are the line under the map: “N site locations don’t have coordinates yet”, which opens a list of exactly those sites, each with its node and GPU counts intact and each editable there. An empty map means nothing placed yet, never no hardware.

Nothing is guessed from the name. dal-01 is a string you typed, not a claim about Dallas — there is no geocoder and no name-to-city table anywhere in the product. And a derived marker is always one real machine’s location, never a point averaged between machines: averaging two nodes on different continents would draw a marker in open water, where no hardware is.

What we record, plainly: every request an agent makes reaches us with an approximate location for that connection, and we keep the latest one per node (city-level at best) so the map can place a site you have not placed yourself. It is the location of the machine’s network egress — for a datacenter usually the building; behind a VPN, corporate tunnel or cloud NAT it can be another country — which is exactly why your own coordinates outrank it and why the two never look alike. It is visible only to your operator, and only rolled up to the site; simulated demo nodes and anything running off our platform get no position at all.

Create a tenant per customer

A tenant is the isolation boundary inside your datacenter — one per customer company, and every tenant you create belongs to your operator. Use Create tenant on /admin/tenants. Your pooled servers reach a tenant two ways: the customer rents whole nodes themselves from your live availability (by GPU model), or you reassign a node from its detail page. The per-tenant logs visible to members toggle (default on) controls whether that customer’s users can read their own machines’ logs.

Isolation is enforced server-side at both levels, always. Every metric and log line is attributed to its tenant and its operator by the platform — never by anything the agent or browser sends — and every query has its scope injected from the session: a customer’s to their tenant, yours to your operator. Nobody, including you, can hand the platform a raw query to widen it. Two automated leak tests run in CI on every change: one proves a customer cannot read another customer, the other that an operator cannot read another operator.

Invite your team and your customers

Invite member on a customer’s tenant mints their user accounts — send the single-use code; they follow the customer guide and can start renting your pool capacity the moment they sign in. Operator seats for your own team are a different door: ask Rackify for an invite naming your operator (above). The operator console mints member invites only, deliberately — an account that decides which datacenter a new operator joins is a decision that belongs one level up.

Read a node like an SRE

Click any node cell to open its detail page:

  • Inventory — driver & CUDA versions, every GPU’s UUID (the identity anchor), VRAM, MIG mode, and the interconnect topology matrix captured at enrollment.
  • Per-GPU live charts — utilization, VRAM, temperature, power; ECC single/double-bit and Xid lifetime counters.
  • Conditions — agent-reported health checks (NVML availability, clock skew, buffer pressure) with reasons.
  • Log viewer — search by time range, level, source, or free text; Follow mode tails live; Xid/NVRM lines are highlighted.
  • Hardware changes — a vanished GPU UUID raises a gpu_missing event; a new UUID is flagged for your explicit swap acknowledgment.

Day-2 operations

  • Reassign a node — between your tenants or back to your capacity pool — from its detail page; customers release their own rented nodes too. Attribution flips atomically; historical data keeps its original tenant (documented, never rewritten). A node never leaves your operator: there is no destination outside it.
  • Offline nodes — flagged in the UI after 2 missed minutes of heartbeats; a banner lists them. Agents self-heal: fix the network and telemetry backfills itself.
  • Demo mode — a simulated fleet can be run against the live console for sales demos; everything simulated is badged SIMULATED in both the data and the UI. Real telemetry is never mixed with it silently: simulated machines are hidden from every dashboard by default and only appear for accounts that switch Show simulated data on in the profile menu — and any view that hid something says how many it hid, so an empty dashboard is never ambiguous.
  • Operational rehearsals — agent-kill, 10-minute network cut, credential rotation, deploy rollback, and backup-restore drills are documented in the operator runbook with recorded results.

Data residency

Where your bytes live.

In a sovereign deployment every component below runs inside your cloud account, pinned to your chosen region. Nothing is shared with other operators; there is no cross-operator control plane.

ComponentWhat it holdsWhere it runs
Console & APIsThe operator and customer consoles, ingest and query APIsYour deployment, as code
Metrics storeAll GPU/node telemetry, 15s resolutionYour account, your region
Log storePer-node driver/dmesg/agent logs (30-day retention)Your account, your region
Control dataTenants, users, nodes, tokens (hashed), events, auditYour account, your region — point-in-time recovery on, nightly exports to your own bucket
Agent binariesSigned, checksum-verified releasesYour account’s artifact bucket
GPU serversYour metal — agents connect outbound-443 onlyYour datacenters

Self-hosting walkthrough: the platform is designed so a competent stranger can stand the whole thing up from the repository docs in about an hour — all cloud resources as Terraform, the console as code, CI with a required cross-tenant test, and a rehearsed backup/restore path. Ask us for it, or for the on-prem design-partner program.