· Charlie Holland · Architecture  · 12 min read

Fifteen Hyperdisks a Node, and the Arithmetic Nobody Will Explain

You size a GKE cluster for thirty small stateful workloads, each with a modest PVC. You expect them to pack onto the nodes neatly. Instead half the nodes are idling, a quarter of your pods are Pending, and the error is about 'volume attach limits.' The number, when you dig into it, is fifteen. Why fifteen? Good question.

You size a GKE cluster for thirty small stateful workloads, each with a modest PVC. You expect them to pack onto the nodes neatly. Instead half the nodes are idling, a quarter of your pods are Pending, and the error is about 'volume attach limits.' The number, when you dig into it, is fifteen. Why fifteen? Good question.

You have a GKE cluster. A few dozen small stateful workloads — batch workers that checkpoint their progress to local disk, a fleet of build agents that need a warm cache between jobs, a handful of ETL jobs that spool intermediate files to disk before flushing to object storage. Each pod doesn’t need much — a modest PVC of maybe 20 GiB — but it does need one of its own, because scratch that survives a pod restart is cheaper than rebuilding the state from scratch every time.

You choose Hyperdisk Balanced as the storage class because the Google Cloud blog has been telling you for two years that Hyperdisk is the future. Hyperdisk Balanced is the default on N4 machines. The performance is good, the pricing is reasonable, the IOPS are tunable per volume. Modern stuff.

You deploy. Thirty workloads, one PVC each. Kubernetes starts scheduling pods. And then it stops scheduling pods. Half your nodes are sitting at 25% CPU and 40% memory. The rest of the pods are Pending, and when you describe them you get this:

Events:
  Warning  FailedScheduling  35s  default-scheduler
  0/4 nodes are available:
    4 node(s) exceed max volume count.
    preemption: 0/4 nodes are available:
    4 Preemption is not helpful for scheduling.

“Max volume count.” On a node with 16 GiB free.

Welcome to the part of the Hyperdisk story that doesn’t make the blog posts.

The number is fifteen, and it’s in the source code

Not a hypervisor-level ceiling. Not a Compute Engine limit you can look up in a docs table. Not something the CSI driver asks the VM about at runtime. It’s a Go constant in the PD CSI driver source tree, typed by a human, checked into the master branch. Here it is, straight out of pkg/gce-pd-csi-driver/node.go:

// These constants are all the documented attach limit minus one because the
// node boot disk is considered an attachable disk so effective attach limit is
// documented limit minus one.
const (
    volumeLimitSmall int64 = 15
    volumeLimitBig   int64 = 127

    x4HyperdiskLimit       int64 = 39
    a4HyperdiskLimit       int64 = 127
    a4xMetalHyperdiskLimit int64 = 31
    c3MetalHyperdiskLimit  int64 = 15
)

That’s it. That’s the answer to “why fifteen.” Somebody typed it in, added a comment explaining it’s sixteen minus the boot disk, and moved on with their day.

The driver doesn’t consult the running VM. It doesn’t ask Compute Engine what the actual attach ceiling for this instance is. It takes the machine type string — n4-standard-8, c4-standard-4, whatever — and looks it up against hardcoded tables in pkg/constants/constants.go. For Hyperdisk Balanced on an N4 machine, the whole table is three lines:

var N4MachineHyperdiskAttachLimitMap = []MachineHyperdiskLimit{
    {Max: 8,  Value: 15},
    {Max: 80, Value: 31},
}

Eight vCPUs or fewer? You get fifteen. Nine to eighty? You get thirty-one. Above that? You get a bigger default from the fallback. The table doesn’t go any finer than that. An n4-standard-2 and an n4-standard-8 both get fifteen slots, despite one being four times the machine. A c4-standard-4 and a c4-standard-2 both round to the same bucket.

C4D is its own private flavour of this:

var C4DMachineHyperdiskAttachLimitMap = []MachineHyperdiskLimit{
    {Max: 2,   Value: 3},
    {Max: 4,   Value: 7},
    {Max: 8,   Value: 15},
    {Max: 96,  Value: 31},
    {Max: 192, Value: 63},
    {Max: 384, Value: 127},
}

Three disks on a c4d-standard-2. Yes, three. A StatefulSet with four replicas against Hyperdisk Balanced on that machine size will never schedule its fourth pod anywhere, ever, because the math doesn’t work and nobody told you.

Why is the number fifteen and not sixteen or twenty? Because the author picked fifteen, in Go, with a comment. Why does it apply uniformly across an entire eight-vCPU range for N4 but a two-vCPU range for C4D? Because the tables are shaped differently. Why does volumeLimitSmall exist as a top-level constant at all, as the fallback for any machine the driver doesn’t recognise? Because someone needed a conservative default, and fifteen is what they picked.

There’s no runtime check. There’s no “let me query the Compute Engine API to see what this VM can actually take.” There’s a match on the machine type prefix, an index into a slice, and out pops your ceiling.

What the scheduler actually sees

When the driver starts up on a node, it runs through that logic and reports the resulting number to the kubelet via a CSINode object. The scheduler then uses that number as a hard ceiling when deciding where to place pods that bring PVCs with them.

You can see exactly what your nodes are advertising:

kubectl get csinodes -o yaml \
  | yq '.items[] | {
      "node": .metadata.name,
      "drivers": .spec.drivers
    }'

On a small N4 node with Hyperdisk Balanced you’ll see:

node: gke-prod-default-pool-abc123-xyz
drivers:
  - name: pd.csi.storage.gke.io
    nodeID: projects/my-project/zones/europe-west2-a/instances/...
    allocatable:
      count: 15
    topologyKeys:
      - topology.gke.io/zone

allocatable.count: 15. That’s volumeLimitSmall. It has travelled, unchanged, from a Go const block in a GitHub repo to a hard constraint on how densely you can pack your cluster.

The default scheduler is obedient. It will refuse to place a sixteenth PVC on that node even if the node is otherwise empty. It doesn’t know that the underlying machine might be able to take more attachments. It only knows the number the driver put in the CSINode object, which is the number the Go code said to put there.

The “this is what you wanted, right?” part

The comment on the constants — “documented attach limit minus one because the node boot disk is considered an attachable disk” — refers to the Compute Engine documentation that says an n4-standard-8 supports up to sixteen Hyperdisk Balanced attachments. Minus one for boot, equals fifteen data disks, equals the number you see.

Which would be fine — almost reasonable — except:

  • It’s hardcoded, so when Google extends the supported attachment limits for a machine generation (and they do, frequently), the driver keeps advertising the old number until somebody opens a PR to update a table.
  • The bucket ranges are coarse. An n4-standard-4 gets the same fifteen that the upstream docs say applies up to n4-standard-8, but in reality smaller N4 shapes may support fewer. Everyone in the same bucket gets the most conservative number in the bucket.
  • There’s no per-disk-type differentiation in what gets reported. The driver looks up the Hyperdisk Balanced table whether you’re using Hyperdisk Balanced or Hyperdisk Throughput, because the CSINode allocatable.count is one number per driver, not one per storage class.
  • You can’t move it by changing the StorageClass, the PVC, or any API on the cluster. The only user-facing lever is the node label.

A GCP cluster with c3-standard-2, c3-standard-4, n4-standard-2, n4-standard-4 nodes can, as the driver’s own README acknowledges, “erroneously exceed the maximum attachable disk number, which should be 16.” Erroneously, here, means “in ways that surprise both the user and the driver.” The number is fifteen until it isn’t, and it isn’t until you hit a code path the table doesn’t cover.

Reproducing it at home

If you want to see this for yourself — and it’s worth seeing, because once you’ve watched a StatefulSet stall at replica sixteen you remember it — here’s the minimum reproduction.

A StorageClass for Hyperdisk Balanced:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: hyperdisk-balanced
provisioner: pd.csi.storage.gke.io
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
parameters:
  type: hyperdisk-balanced
  provisioned-iops-on-create: "3000"
  provisioned-throughput-on-create: "150Mi"

A StatefulSet that needs one small PVC per replica:

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: disky
spec:
  serviceName: disky
  replicas: 30
  selector:
    matchLabels: { app: disky }
  template:
    metadata:
      labels: { app: disky }
    spec:
      containers:
        - name: app
          image: busybox
          command: ["sh", "-c", "sleep 36000"]
          volumeMounts:
            - name: data
              mountPath: /data
  volumeClaimTemplates:
    - metadata:
        name: data
      spec:
        accessModes: ["ReadWriteOnce"]
        storageClassName: hyperdisk-balanced
        resources:
          requests:
            storage: 10Gi

Deploy that onto a three-node pool of n4-standard-8 and watch the first fifteen pods per node land cleanly, then the rest sit Pending with the volume-count error while the nodes themselves yawn at 15% utilisation. Your cluster can’t fill up, because it’s out of a resource that nobody thinks of as a resource.

Working around it

There are four options, roughly in order of how much you’ll enjoy them.

Buy bigger nodes. The simplest and least satisfying. Move to n4-standard-32 and the attach limit climbs to thirty-two. Your 30 workloads now fit, but you’re paying for 128 vCPUs across four nodes you weren’t using. Cloud economics: you solved a disk limit with CPU spend.

Override the driver’s made-up number with the node label. GKE 1.32.4-gke.1698000 and later support node-restriction.kubernetes.io/gke-volume-attach-limit-override, which tells the scheduler to use a number of your choosing instead of the one the driver hands it:

kubectl label node gke-prod-default-pool-abc123-xyz \
  node-restriction.kubernetes.io/gke-volume-attach-limit-override=31

This is the cleanest way out of the situation when you know the hardcoded number is under-representing what your VM can actually take. If the Compute Engine docs say your machine supports 32 attachments and the driver advertises 15, the label lets you reconcile them. What it does not do is create capacity that doesn’t exist — set it to 100 on an n4-standard-2 and the sixteenth attachment still fails at the Compute Engine layer, which moves the failure from FailedScheduling to FailedAttachVolume. Worse. The label is a correction tool, not a capacity multiplier. Read the actual Compute Engine limits for your machine shape before you set it.

Hyperdisk Storage Pools. Storage pools let you thin-provision capacity and IOPS across multiple volumes that share a pool. They help with cost and performance allocation, and they’re genuinely useful when you have a lot of volumes that are individually small but collectively need decent IOPS headroom. What they do not do is change the per-VM attachment limit. Fifteen disks from a storage pool and fifteen disks outside one look identical to the CSI driver’s constant table. Pools are a provisioning abstraction, not a driver one.

Go back to Persistent Disk, or go sideways to Filestore. Standard pd-balanced has the same 128-total ceiling but doesn’t have the per-Hyperdisk-type sub-ceiling. If you don’t need Hyperdisk’s per-volume IOPS tuning, the older Persistent Disk types let you pack more volumes per node. Or — radical idea — if your workload doesn’t genuinely need block storage per pod, use Filestore or a single shared ReadWriteMany volume and stop multiplying PVCs for no good reason. A lot of “every pod needs its own PVC” designs turn out, on inspection, to be “every pod needs its own path,” which is a problem that pre-dates Kubernetes by forty years and has been solved by NFS the whole time.

Why fifteen

Because volumeLimitSmall int64 = 15. That’s the whole answer. Not a PCIe slot budget, not a hypervisor-level attachment ceiling, not the outcome of a careful engineering study of NVMe throughput on Google’s custom silicon. A constant, in a Go file, in a publicly-visible GitHub repo, with a one-line comment.

Why not sixteen? Because the comment says “documented limit minus one for the boot disk,” and the most common documented limit for small Hyperdisk-capable VMs is sixteen. Why not twenty? Because nobody wrote a table entry that would produce twenty. Why not read it from the VM at runtime? Because that would be an API call, and this way is a slice lookup. The tables exist because somebody decided a slice lookup was good enough.

The more interesting question is the one underneath. Cloud providers love to tell you storage is elastic — and it is, right up until you need to attach it to a specific VM. At that point it becomes one of the most inelastic parts of the stack, rate-limited by constants that the CSI driver authors may or may not have got right for your machine type, in the version of the driver that GKE happens to have bundled with your cluster version. AWS has the equivalent with EBS volume attachment limits per instance type. Azure has it with VM-size-dependent data disk caps. None of them are negotiable, and most of them lose the per-instance nuance by the time they reach the CSI layer.

The broader point

Block storage is not a continuous resource the way CPU and memory are. It’s discrete and slot-based, and the slot count is lower than you think and further from where you look than you’d expect. For a Kubernetes architect designing a workload with a lot of small stateful pods, the attach limit is the most important number nobody’s going to tell you about. It won’t be in the GKE pricing calculator. It won’t be in the node-sizing guide. It will be in a Go file that the driver running on your nodes was compiled from, and you will find out about it when your cluster stops being able to schedule pods for reasons that have nothing to do with capacity as you understand the word.

If you’re designing for a PVC-per-pod pattern — which plenty of legitimate workloads are — check the driver’s actual numbers for your machine type before you commit. Open the repo. Read the table. Assume you’ll hit whichever number is in it, because you will.

And if anyone at Google is reading this: please surface these numbers somewhere other than a Go constant. A canonical doc page listing machine-type → Hyperdisk-type → attach-limit would save every GKE architect in your ecosystem a week of investigation apiece. Right now it takes two searches, a GitHub commit history, a kubectl command, and a strong coffee to find out that the answer, for your shape, is fifteen.

I’ve opened issue #2307 on the CSI driver repo proposing that the attach limit should be discovered dynamically from the VM rather than baked into the binary. If you’ve hit this and have an opinion, pile on.

Why fifteen? Because somebody typed it. That’s it. Now you know.

Back to Blog

Related Posts

View All Posts »
EKS vs AKS vs GKE: I've Run Production on All Three

EKS vs AKS vs GKE: I've Run Production on All Three

Most Kubernetes comparisons are feature matrices written by people who've read the docs. This is what the three major managed Kubernetes services actually feel like when you're running production workloads, handling upgrades at 2am, and arguing with IAM policies.

Nature Spent 60 Million Years Removing Complexity. We Add It Every Quarter.

Nature Spent 60 Million Years Removing Complexity. We Add It Every Quarter.

There's a lump of granite off the Ayrshire coast that used to be a volcano. Sixty million years of weather wore it down to the one part hard enough to matter, and we've made the world's curling stones out of it ever since. Software does the opposite. Nothing ever erodes, every layer you've ever shipped is still down there needing somebody to mind it, and you pay for all of it in headcount and meetings. Oh, Fred Brooks, save us all.

Istio: The Service Mesh Nobody Asked For

Istio: The Service Mesh Nobody Asked For

Everyone adopted Istio because everyone else was adopting Istio. Complex, resource-hungry, and solving problems most teams didn't have — it was the poster child for resume-driven development.