Provisioning · Kubernetes & Slinky/Slurm · HPC help desk

Boutique Kubernetes & Slurm Cluster Integration

We deploy and manage Kubernetes and Slinky clusters (Slurm on Kubernetes) using hardware you own or a cloud account you already have. You keep the metal, the accounts, and the data. We do the architecture and run the system.

Your researchers get a production HPC cluster with the Slurm interface they already know, built on Kubernetes and running on infrastructure you own or rent directly. We do the build and the day-to-day operations, and we answer your users' tickets. We are your cluster team.

We are a small shop, and we design each cluster around the workloads it has to run, the hardware and network you already have, and the people who will use it. Then we build it as code so it stays that way.

What you get

A cluster that's stood up correctly, with Kubernetes and Slinky/Slurm kept running, and a help desk backing your users.

Provisioning

We stand up your cluster from a tested, repeatable build: control plane, GPU compute, storage, and networking. Nodes go through burn-in before they run a single job, so bad GPUs and NICs turn up before your users find them.

Kubernetes management

We keep Kubernetes healthy. Upgrades, patching, scaling nodes in and out, networking, storage, and monitoring all run through GitOps, so the cluster's state stays versioned and reproducible instead of accumulating hand-edits.

Slinky / Slurm management

We run the scheduler layer: partitions, QoS, preemption, accounting, and the Slinky operators behind Slurm on Kubernetes. Jobs land where they should, and usage is tracked.

HPC help desk

Direct support for your users and researchers on job submission, partitions and QoS, containers, software stacks, and workflow tuning. The engineers answering tickets run HPC systems for a living.

What that looks like in a rack

A single-site reference layout: what each tier is physically made of, how many of them there are, and which networks each one attaches to.

Site networkHead / provisioning node×1PXE, DHCP and TFTPNode image serviceRouting to the site networkTailscale overlayOPTIONALIdentity-based SSHAdmin access to UIs and the APINo public IPs, no new cablingOut-of-band / BMC1 GbEProvisioning + cluster25 GbEHigh-speed fabric200 Gb RDMAcompute and storage onlyControl plane×3HARDWAREcp-11U · CPUcp-21U · CPUcp-31U · CPURUNSKubernetes API + etcdGitOps reconcilerslurmctld + slurmdbdSlinky operatorVirtual IP shared by all threeInfrastructure×2+HARDWAREinfra-11U · CPUinfra-21U · CPUinfra-nas users growRUNSLogin nodesIdentity + POSIX accountsMetrics, logs, alertingIngress + certificatesTainted, so no user jobs land hereStorage×2+HARDWAREhome-12U · disk shelfscratch-12U · NVMescratch-nadd targetsRUNSShared homeParallel scratchSnapshots + backupStorage driversSized for job I/O, not just capacityCompute×NHARDWAREgpu-14U · 8 GPUgpu-24U · 8 GPUgpu-nidentical buildRUNSslurmdGPU device pluginsJob containersRDMA adapterSame image on every nodeREADING THE DRAWINGJunction dot: the tier attaches to that networkCrossing with no dot: the line passes by, unconnectedDashed unit: the tier scales out from hereDashed box: optional, layered on when you want it
One rack at one site, drawn to show which tier lands on which network. Control plane and infrastructure sit on the management and cluster networks only, while storage and compute also terminate on the high-speed fabric. That split decides how many fabric ports end up on your bill of materials. Switch models, power, and exact rack positions are site specific and get settled during design.

Runs on your infrastructure

Bare metal, cloud, rented GPU capacity, or a mix. The same reference architecture, adapted to where your compute lives.

We work with cloud providers and colocation facilities regularly, and we know what their quotes tend to leave out. If you are still deciding where to put the cluster, we can help you collect and compare bids from several of them, then size the architecture around what you actually buy.

The same holds for the systems you already run. Plenty of organizations have an identity provider they are not going to give up, and building the integration between it and the cluster's identity and account layer is part of the boutique work. We develop that integration with you, so your people reach the cluster with the accounts they already have.

Your bare metal

On-prem racks, a colocation cage, or a lab room in the basement. We handle OS images, network and fabric bring-up, and node lifecycle with Packer, MAAS, Warewulf, and Terraform. A rebuilt node comes back identical to the one it replaced.

Your cloud account

GPU instances from the cloud provider you already buy from, inside your project and under your billing and security controls. We bring the reference architecture. You keep the account, the commitments, and the data.

Rented GPU capacity

Bare-metal GPU capacity rented from a specialty provider works the same way. Give us nodes and a network, and we will build the cluster on top of them.

Hybrid and bursting

Keep a steady-state cluster on your own metal and burst to cloud when a deadline hits. One Slurm interface, consistent images and software stacks, and accounting that covers both.

What's under the hood

A modern, GitOps-native stack, configured, tested, and operated by practitioners. These are the parts that have to come together, and how they stack up.

Your git repository

Every box below is defined here, reviewed as a pull request, and applied as code.

reconciled continuously

Control plane

Runs the cluster and the scheduler. No user workloads land here.

Kubernetes control plane

RKE2

GitOps reconciler

Argo CD

Slinky operator

SchedMD slurm-operator

slurmctld

Slurm controller

slurmdbd

Slurm accounting on MariaDB

slurm-bridge

Slurm places Kubernetes workloads

Cluster networking

Cilium, kube-proxy replacement

Secrets and certificates

External Secrets Operator, cert-manager

schedules jobs and applies policy

Infrastructure services

The shared services behind every job and every login.

Login nodes

Slinky loginset

Identity management

Authentik, SCIM provisioning

POSIX account resolution

nsscache

Shared home and scratch

NFS on ZFS, BeeGFS

Storage drivers

csi-driver-nfs, local-path-provisioner

Metrics and alerting

Prometheus, Grafana, Alertmanager

Log aggregation

Loki, Alloy

Tailscale operatoroptional

tailnet access to UIs and the API

storage, identity, and telemetry

Compute nodes

Where jobs run. Identical from one node to the next.

slurmd

Slinky nodeset

GPU device plugin

NVIDIA and AMD

GPU telemetry

DCGM, AMD device metrics exporter

Host telemetry

node-exporter

Job container runtime

Enroot, Pyxis

High-speed fabric

RDMA over InfiniBand or Ethernet

Shared storage mounts

NFS and BeeGFS clients

Tailscale node agentoptional

tailscaled, Tailscale SSH

This is one arrangement of the parts, not the only one. Boxes get swapped to fit the hardware, the network, and the policies you already have. Standing up any single one of them is an afternoon. Keeping all of them working through a node reboot, an expired credential, and a driver upgrade in the same week is the engagement.

Slinky (Slurm-on-Kubernetes)

Your researchers get the Slurm they already know, with srun, sbatch, and sinfo running as Kubernetes workloads underneath. Familiar HPC ergonomics on modern orchestration.

Kubernetes, tuned for HPC

RKE2 with a high-performance CNI, topology-aware scheduling, and device plugins configured for GPU and fabric-heavy workloads.

GPU compute, ready to run

NVIDIA CUDA or AMD ROCm drivers, container toolkits, and RDMA fabric, configured and validated. Enroot and Pyxis for container jobs out of the box.

Shared storage & identity

Shared home and scratch storage, plus centralized user identity and POSIX accounts that keep working when the identity provider hiccups.

Monitoring & accounting

Prometheus, Grafana, and GPU health metrics, plus Slurm accounting backed by an in-cluster database. Usage and utilization stay visible.

Secure remote access with Tailscale

Add Tailscale for identity-based SSH and admin access without public IPs or firewall wrangling. Optional, and easy to layer on later.

slurm-bridge

Kubernetes workloads placed by Slurm instead of by a second scheduler. Containerized services and ordinary sbatch jobs draw from one pool of nodes under one set of priorities, QOS limits, and fairshare.

Self-service apps, one click

Launch a vLLM server, a Jupyter notebook, or an isolated VM from a button. Each lands as a Slurm job under your accounting and preemption rules, so an ordinary sbatch can take the node back.

Managed as code, across two repositories

Nothing about your cluster is configured by hand. Everything is defined as code and changed through reviewed pull requests, split across a platform repository we maintain and an apps repository you own.

The platform repository holds the reference architecture and the integration work that makes the pieces fit: Kubernetes and Slinky configuration, storage, identity, monitoring, and the tooling that glues them together. We maintain it and it stays with us. That is deliberate, and it is what you are buying. A fix we make on another cluster this month lands on yours, because you are getting a product we keep sharpening rather than a one-off build that starts aging the day it ships.

The apps repository is yours. It comes in as a submodule, and it holds the workloads your users actually run. A vLLM server behind an OpenAI-compatible endpoint, a Jupyter notebook, and an isolated VM are already built and tested, so putting those on your cluster is configuration rather than a project. Anything specific to you that deploys through slurm-bridge, we build with you in your repository, and you keep it.

Both sides work the same way. Changes are proposed, reviewed, and applied as code, so the cluster's state is versioned, auditable, and reproducible rather than hidden in someone's terminal history. Open a pull request against your apps and work through it with us.

Agent skills, shipped with the cluster

Your admins and researchers get skills that answer cluster questions in plain language, wired to the same read-only interfaces we use.

Your cluster ships with agent skills your team runs from Claude Code or another MCP-capable client. They reach the cluster through read-only tools over Slurm, Prometheus, Loki, and the Kubernetes API. Ask what is idle, what is blocking the queue, who is over quota, or why last night's run failed, and the answer is built from the cluster's own telemetry instead of guessed.

Read-only is the default, and it is the point. A skill that finds a quota problem writes the fix as a diff against your policy file for someone to review and merge. Nothing an agent concludes reaches the live cluster except through the same pull request every other change goes through.

We write skills for your workflows the same way we integrate apps. If your team keeps answering the same question by hand, it should be a skill.

How we engage

Take the whole thing or just the parts you're short on. Engagements move between these as your team grows.

Integration

A defined build. We design the architecture, provision the cluster on your infrastructure, validate it with real workloads, and hand over a documented, running system with your team trained to operate it.

Co-managed

Your admins keep the keys and the day-to-day work. We cover the deep end: upgrades, scheduler tuning, incident response, and the changes nobody on staff has done before.

Fully managed

We operate the cluster and staff the help desk. You keep the hardware, the accounts, and your apps, and you keep a support relationship instead of a hiring problem.

Pricing

Simple, cluster-size-based pricing for our services: provisioning, Kubernetes and Slinky/Slurm management, and the HPC help desk. Tell us your GPU count and where the nodes live, and we'll size a quote.

Hardware, colocation, and cloud compute are purchased on your own accounts and contracts. We don't resell capacity or mark it up. You pay your provider directly, and you pay us for the engineering, including the work of sourcing and comparing quotes on your behalf.