· Charlie Holland · DevOps  · 10 min read

MicroK8s in the Field: ML at the Edge When the Cloud Isn't an Option

Not everyone can 'just use the cloud.' Healthcare data sovereignty, air-gapped networks, public sector paranoia — sometimes you're deploying ML inference on a rack server in a basement. Here's what that actually looks like with MicroK8s.

Not everyone can 'just use the cloud.' Healthcare data sovereignty, air-gapped networks, public sector paranoia — sometimes you're deploying ML inference on a rack server in a basement. Here's what that actually looks like with MicroK8s.

“Just use the cloud.”

I hear this approximately once a day. Usually from someone who has never had to explain to a healthcare organisation’s information governance team why patient data should leave the building. Or told a public sector client that their classified workloads should run on infrastructure owned by an American technology company. Or discovered, mid-deployment, that the factory floor has no outbound internet connection and nobody thought to mention it.

Not everyone can use the cloud. And for those organisations, you still need to run ML inference workloads somewhere. For the past eighteen months, my somewhere has been MicroK8s on whatever hardware the client happens to have in a rack somewhere.

This is what that actually looks like.

Why MicroK8s

The short answer: it’s the easiest way to get a working Kubernetes cluster on a single machine or a small cluster of machines, with GPU support, without requiring a Kubernetes expert on staff.

The longer answer involves comparing the lightweight Kubernetes options, and there are several. But for our specific use cases — ML inference at the edge, on Ubuntu servers, often with NVIDIA GPUs — MicroK8s has been the best fit.

Installation

sudo snap install microk8s --classic --channel=1.21/stable
sudo microk8s status --wait-ready
sudo microk8s enable dns storage gpu

Three commands. That’s a working Kubernetes cluster with DNS, local storage, and GPU support. Try doing that with kubeadm. I’ll wait.

The snap-based distribution is MicroK8s’s killer feature. It handles the Kubernetes binary lifecycle, the container runtime, the CNI plugin, and the core add-ons in a single, upgradable package. snap refresh microk8s --channel=1.21/stable and you’ve upgraded Kubernetes. No draining nodes, no etcd backup dance, no “is the API server back yet” anxiety.

The add-on ecosystem

MicroK8s ships with a set of built-in add-ons that you enable with a single command:

  • gpu: Installs the NVIDIA GPU operator. Essential for ML workloads.
  • dns: CoreDNS. You need this for basically everything.
  • storage: Local path provisioner. Not production-grade, but functional.
  • metallb: Bare-metal load balancer. Because there’s no cloud LB on-prem.
  • istio: Service mesh. We use this for traffic management on model serving.
  • prometheus: Monitoring stack. More on this later.
  • registry: Local container registry. Critical for air-gapped deployments.

Each of these would be a separate Helm chart or operator installation on a vanilla Kubernetes cluster. On MicroK8s, it’s microk8s enable <addon>. I cannot overstate how much time this saves when you’re deploying to a new client site every few weeks.

MicroK8s vs k3s

I get asked this constantly, so let me be direct.

k3s is lighter. Smaller binary, lower memory footprint, faster startup. If you’re running on a Raspberry Pi or a genuinely resource-constrained device, k3s is probably the better choice. It’s excellent at what it does.

But for our use cases, MicroK8s wins on two points:

GPU support: MicroK8s’s GPU add-on integrates the NVIDIA GPU operator cleanly. k3s can run GPU workloads, but the setup is more manual — you’re installing the NVIDIA container toolkit and device plugin yourself, managing driver compatibility, and debugging containerd runtime configuration. The MicroK8s add-on handles most of this. Not all of it, mind — GPU driver management is still painful regardless — but most of it.

High availability: MicroK8s uses Dqlite as its distributed datastore for HA clusters. Three nodes, automatic leader election, no external etcd cluster to manage. k3s uses embedded SQLite for single-node and requires an external datastore (etcd, MySQL, or PostgreSQL) for HA. For edge deployments where we want a small HA cluster without additional infrastructure, Dqlite is a genuine advantage.

k3s has a larger community and more mindshare, and I wouldn’t argue against choosing it. But for ML-at-the-edge specifically, MicroK8s has been the better tool.

The real challenges

Here’s where the blog post stops being an advert for Canonical and starts being honest about what goes wrong.

Hardware variability

Every client is different. Every server rack is different. I have deployed MicroK8s on:

  • Dell PowerEdge R740s with dual NVIDIA T4s
  • HPE ProLiant DL380s with a single V100
  • Lenovo ThinkSystem SR650s with no GPU at all (CPU inference only)
  • One memorable occasion: a repurposed gaming PC under someone’s desk with a consumer RTX 2080

Each of these has different BIOS settings, different driver requirements, different network interface configurations, different storage layouts. The MicroK8s installation is the same, but everything around it is bespoke.

There is no Terraform for “whatever hardware Dave ordered six months ago.” You SSH in, you assess what you’ve got, and you make it work.

No cloud load balancer

In the cloud, you create a Service of type LoadBalancer and your cloud provider gives you an IP address. On-prem, that Service sits in Pending forever.

MetalLB solves this by assigning IPs from a pool you configure. It works. But:

  • You need to coordinate with the client’s network team to get a range of IPs allocated. This can take weeks in enterprise environments.
  • Layer 2 mode has failover delays. BGP mode requires router configuration that most on-prem network teams have never done for Kubernetes.
  • If the network team gives you a range that conflicts with something else, you find out at the worst possible time.

The alternative is NodePort, which works everywhere but means your services are on high-numbered ports (30000-32767), which confuses users, breaks assumptions in application configs, and looks amateurish. We use MetalLB wherever possible and accept the coordination overhead.

Storage

Local-path provisioner gives you ReadWriteOnce PVCs backed by the node’s local disk. It’s fine for development and testing. It’s not fine for anything that needs durability or shared access.

For production edge deployments, we’ve used:

  • Longhorn: Distributed block storage. Works well in small clusters (3-5 nodes). Adds resource overhead, but gives you replicated volumes and snapshot/backup capabilities. Our default choice when we have multiple nodes.
  • NFS: When the client already has NFS infrastructure. Simple, well-understood, terrible performance for random I/O. Fine for model artifact storage, bad for training data that needs fast random access.

We tried Rook/Ceph once on a three-node edge cluster. The resource overhead was enormous — Ceph wanted more RAM than the actual workloads. Rook is designed for large clusters with dedicated storage nodes, not edge deployments on constrained hardware. Lesson learned.

GPU driver management

This is the single most painful part of the entire stack. It’s not MicroK8s’s fault — it’s NVIDIA’s, and Ubuntu’s, and the kernel’s, and the immutable laws of driver compatibility.

The GPU operator add-on in MicroK8s installs the NVIDIA GPU Operator, which manages the driver lifecycle in-cluster. In theory. In practice:

  • The GPU operator version must be compatible with the GPU model.
  • The driver version must be compatible with the CUDA version your ML framework needs.
  • The CUDA version must be compatible with your model’s training environment (because a model trained on CUDA 11.4 may not run on CUDA 11.1 without recompilation).
  • The kernel version must be compatible with the driver version.
  • If the client has pre-installed NVIDIA drivers (which they often have, for their own reasons), the GPU operator will conflict with them.

I have spent entire days on GPU driver issues at client sites. Entire. Days. The nvidia-smi output becomes the most important diagnostic in your toolkit. If nvidia-smi shows your GPUs, you’re halfway there. If it doesn’t, clear your afternoon.

# The command you'll run 400 times
nvidia-smi
# The output you pray for
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 470.57.02    Driver Version: 470.57.02    CUDA Version: 11.4     |
+-----------------------------------------------------------------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|=============================================================================|
|   0  Tesla T4            On   | 00000000:3B:00.0 Off |                    0 |
| N/A   38C    P8     9W /  70W |      0MiB / 15109MiB |      0%      Default |
+-----------------------------------------------------------------------------+

Air-gapped installation

The most challenging deployments are air-gapped environments — networks with no outbound internet access. These are common in healthcare, defence, and heavy industry.

MicroK8s installs via snap, which normally pulls from the Snap Store. In an air-gapped environment, you need a Snap Store Proxy or you need to download the snap on an internet-connected machine and snap ack / snap install it from a local file.

But MicroK8s is just the orchestrator. You also need:

  • Container images for your workloads (pre-pulled to a local registry)
  • Container images for MicroK8s add-ons (CoreDNS, GPU operator, MetalLB, etc.)
  • Model artifacts (downloaded from wherever they were trained)
  • Python packages (if anything does pip install at runtime — and something always does)

We built an “air-gap preparation” script that runs on an internet-connected machine, pulls everything into a tarball, and transfers it via USB or secure file transfer. It’s about 400 lines of bash and it handles maybe 80% of cases. The other 20% is discovering at the client site that your model needs a Python package you forgot to include, and the data scientist who trained it forgot to mention it because it “just worked on my laptop.”

Monitoring on constrained hardware

You need monitoring. Prometheus + Grafana is the standard Kubernetes monitoring stack. But on edge hardware with 32-64GB of RAM shared between your actual workloads and the platform, the monitoring stack competes for resources.

We run Prometheus with aggressive retention settings (24-48 hours of local data), minimal scrape targets, and remote_write to a central monitoring system — when there is network connectivity to write to. In fully air-gapped environments, it’s local Prometheus and Grafana only, with manual checks.

The MicroK8s prometheus add-on deploys the full kube-prometheus-stack, which includes Alertmanager, node-exporter, kube-state-metrics, and Grafana. On a single-node deployment with 32GB RAM, this stack consumes about 2-3GB. It’s not nothing.

ML-specific issues

Model serving on limited resources

We use TensorFlow Serving for TensorFlow models and NVIDIA Triton Inference Server for everything else. Both work on edge hardware, but the resource constraints change your architecture.

In the cloud, you’d run multiple replicas of your serving container behind a load balancer, with autoscaling based on request queue depth. At the edge, you might have one GPU and one serving instance, and your scaling strategy is “hope it’s fast enough.”

Batching becomes critical. Triton’s dynamic batching feature is essential — it collects individual inference requests and batches them for GPU execution, significantly improving throughput. The configuration is fiddly (max batch size, max queue delay, preferred batch sizes), but the difference between batched and unbatched inference on a T4 can be 5-10x throughput.

Model updates without downtime

In the cloud, you blue-green deploy a new model version. At the edge, with intermittent connectivity and a single serving instance, it’s more complicated.

Our pattern: download the new model artifact during a maintenance window (or whenever connectivity is available), load it into a staging path, run validation inference against a test dataset, then swap the model path. Triton supports model versioning natively, which helps. TensorFlow Serving supports model version labels, which also helps.

What doesn’t help is when the new model is 4GB, the download connection is 10Mbps, and the maintenance window is two hours. You learn to be very intentional about model size optimisation, quantisation, and distillation when the deployment path involves transferring files over a constrained link.

The lesson

Kubernetes at the edge works. MicroK8s makes it considerably less painful than the alternatives. But you are absolutely, unequivocally trading cloud convenience for control.

The cloud gives you managed Kubernetes, managed model serving, managed storage, managed networking, managed monitoring, managed everything. The edge gives you hardware under a desk and an SSH connection. Everything between those two points is your responsibility.

Only do this if you genuinely can’t use the cloud. If data sovereignty, air-gap requirements, latency constraints, or regulatory compliance mean the data cannot leave the building, then MicroK8s on edge hardware is a legitimate, production-viable approach. We have deployments that have been running for over a year with minimal intervention.

But if someone suggests edge Kubernetes because it might be cheaper than cloud, or because they “don’t trust the cloud,” or because a vendor sold them hardware they now need to justify — push back. The cloud is easier. The cloud is faster. The cloud has a team of engineers who wake up at 3am when the storage system breaks, and that team isn’t you.

Edge Kubernetes is a last resort that works. Treat it accordingly.

Back to Blog

Related Posts

View All Posts »
The DevOps Bubble Is Bursting — Good

The DevOps Bubble Is Bursting — Good

We never really needed ten thousand DevOps engineers. The market's not dying — it's clearing out the noise. When the dust settles, the builders will still be here.

Ops in a Frock: Why Most SRE is Still Theatre

Ops in a Frock: Why Most SRE is Still Theatre

Pipeline jockeys cranking YAML. Firefighting teams on endless rota. Developers throwing features over the wall. Call it SRE, call it Platform Engineering — if nobody owns production, it's just ops in a frock.

Kubernetes Operators: The Abstraction That Actually Worked

Kubernetes Operators: The Abstraction That Actually Worked

Most Kubernetes abstractions add complexity for complexity's sake. Operators are the exception — they encode operational knowledge into software and turn 'file a ticket and wait three days' into 'apply a YAML file and get a database in two minutes.'

Shift-Left Security Sounds Great Until You See the Pipeline

Shift-Left Security Sounds Great Until You See the Pipeline

Everyone preaches 'find vulnerabilities earlier.' The concept is sound. The execution is usually a disaster — every build fails on day one, exception lists grow longer than vulnerability lists, and security becomes a rubber stamp that makes everyone feel safe without actually being safe.