· Charlie Holland · DevOps · 9 min read
Kubernetes Operators: The Abstraction That Actually Worked
Most Kubernetes abstractions add complexity for complexity's sake. Operators are the exception — they encode operational knowledge into software and turn 'file a ticket and wait three days' into 'apply a YAML file and get a database in two minutes.'
The Kubernetes ecosystem has a complexity problem. I don’t think that’s controversial — even the people who love Kubernetes will admit, after a beer or two, that the tooling around it has gone somewhat mad.
Helm charts that are harder to read than the YAML they replace. Kustomize overlays stacked three deep that nobody can follow. Custom admission webhooks that silently mutate resources in ways that break deployments and baffle anyone debugging at 2am. Service meshes that add a sidecar proxy to every pod because someone read a blog post about Istio and got excited.
The Kubernetes ecosystem suffers from what I’d call abstraction addiction — the compulsive need to put another layer of indirection between the engineer and the thing they’re trying to do. Every layer promises simplicity and delivers complexity. Every tool claims to make Kubernetes “easier” and makes the debugging surface larger.
Operators are the exception. They’re the one abstraction in the Kubernetes world that genuinely reduces cognitive load for the people who consume them. And I say this as someone who has spent the last couple of years building and deploying them for two very different organisations — a major pharma company and a global marketing agency.
What operators actually are
For anyone who hasn’t encountered them, the Operator pattern is simple in concept: you extend the Kubernetes API with custom resources that represent the things your organisation cares about, and you write a controller that watches those resources and takes action when they change.
The idea came from CoreOS in 2016. The insight was that Kubernetes already has a powerful reconciliation loop — the control plane constantly watches the declared state of resources and works to make reality match. Operators let you plug your own operational logic into that same loop.
A concrete example. Say your development teams need PostgreSQL databases for testing. Without an operator, the workflow looks like this:
- Developer files a Jira ticket requesting a database.
- Ticket sits in the DBA team’s backlog for somewhere between two hours and two weeks, depending on mood and backlog depth.
- DBA provisions the database manually, or runs a script, or clicks through a cloud console.
- DBA sends the connection string back to the developer via email, Slack, or — in one memorable case — a Post-it note stuck to a monitor.
- Developer finally starts the work they needed the database for.
With an operator, the workflow is:
apiVersion: databases.internal/v1
kind: PostgresInstance
metadata:
name: my-test-db
namespace: team-alpha
spec:
version: "15"
storage: 10Gi
backup: falseApply that file. Wait two minutes. The operator provisions the database, creates the Kubernetes Secret with connection details, and the developer gets on with their life. When they’re done, they delete the resource and the operator cleans up.
That’s not a theoretical example — I built exactly that system for a pharma company where the average turnaround time for a test database had been four days. Four days of a developer sitting idle (or more likely context-switching to something else and losing momentum) because provisioning a database required human intervention.
Why this abstraction works when others don’t
I’ve been thinking about why operators succeed where so many Kubernetes abstractions fail, and I think it comes down to three things.
They reduce cognitive load for the consumer
Most Kubernetes tooling shifts complexity sideways rather than eliminating it. Helm doesn’t make Kubernetes simpler — it makes it differently complicated. Instead of writing YAML, you’re writing Go templates that generate YAML, with values files that override other values files, and chart dependencies that pull in more templates. The total complexity hasn’t decreased; it’s been repackaged.
Operators genuinely reduce cognitive load for the person consuming the service. The developer applying that PostgreSQL resource doesn’t need to know how the database is provisioned, what cloud API calls are involved, what networking configuration is required, or what backup policies apply. They specify what they want, and the operator handles how.
This is the correct direction for an abstraction — it hides complexity that the consumer doesn’t need to see, while keeping that complexity accessible to the people who maintain the operator. Joel Spolsky’s Law of Leaky Abstractions applies, of course — abstractions leak. But operators leak less than most because the boundary is clean: the consumer’s interface is a Kubernetes resource spec, and everything else is the operator’s problem.
They keep complexity where it belongs
The complexity doesn’t disappear when you build an operator — it moves into the operator code, where it’s maintained by the platform team. This is the right trade-off. The platform team has the expertise to manage cloud provisioning, networking, security policies, and backup configuration. The product developers have the expertise to build features.
Operators encode the platform team’s operational knowledge into software. Every decision that would previously have been made by a human — “what instance size should this database be?”, “what backup schedule?”, “what network security group rules?” — gets encoded as logic in the operator, with sensible defaults and overridable options.
This is what Google calls encoding operational knowledge in code — taking the things that experienced engineers know and making that knowledge available as software. It’s the same principle behind SRE: turn operations into engineering.
They use Kubernetes’ own machinery
Here’s the subtle bit. Operators don’t bolt new concepts onto Kubernetes — they extend the concepts Kubernetes already has. Custom resources look and behave like built-in resources. You create them with kubectl apply. You inspect them with kubectl get. You delete them with kubectl delete. They have status conditions, events, and finalizers. They participate in Kubernetes RBAC.
For a developer who already knows Kubernetes (and in 2024, that’s most developers at organisations running Kubernetes), operators feel native. There’s no new CLI to learn, no new configuration language, no new deployment model. It’s just Kubernetes, with more resource types.
Contrast this with Helm, which introduces its own packaging format, its own template language, its own release management concept, and its own CLI. Or Kustomize, which introduces overlays, patches, strategic merge patches, and a configuration model that’s different enough from raw YAML to be confusing but similar enough to be deceptive. These tools all live alongside Kubernetes rather than within it.
What I built and what I learned
At the pharma company, we built operators for three things: database provisioning (PostgreSQL and SQL Server instances for development and testing), cloud storage buckets (with appropriate access policies baked in), and ephemeral environments (complete application stacks that developers could spin up for feature branch testing and tear down when they were done).
The ephemeral environment operator was the most impactful. Before it existed, testing a feature branch against a full environment meant either sharing a staging environment (with all the conflicts and queuing that implies) or waiting for the ops team to provision one manually. After the operator was in place, any developer could run:
apiVersion: environments.internal/v1
kind: EphemeralEnvironment
metadata:
name: feature-payment-refactor
spec:
branch: feature/payment-refactor
services:
- api
- worker
- web
ttl: 48hForty-eight hours later, the environment would be automatically destroyed. No cleanup tickets. No forgotten environments running up cloud bills for months. The operator handled the entire lifecycle.
At the marketing agency, the use case was different but the pattern was the same. The agency ran campaigns for dozens of clients, each needing isolated infrastructure — separate databases, separate storage, separate API endpoints. The operator watched for new client onboarding resources and provisioned the entire stack automatically. What had been a two-week manual process became a fifteen-minute automated one.
Lessons from building them
Start with Kubebuilder or Operator SDK. Don’t write the scaffolding yourself. Both frameworks generate the boilerplate for custom resource definitions, controllers, and RBAC configuration. Kubebuilder is maintained by the Kubernetes SIG and is my preference, but Operator SDK (from Red Hat) is solid too. The choice doesn’t matter much — pick one and move on.
Design your CRD API carefully. The custom resource spec is the interface your consumers will use. Treat it like a product API. Sensible defaults for everything. Clear documentation. Validation that catches mistakes early rather than producing cryptic errors during reconciliation. Version it from day one — you will need to evolve it, and CRD versioning is easier if you start with it rather than retrofit it.
Reconciliation must be idempotent. The controller’s reconcile function will be called multiple times for the same resource — on creation, on update, on periodic resyncs, and on controller restart. If your reconcile function isn’t idempotent, you’ll create duplicate resources, trigger unnecessary updates, or — in the worst case — delete things that shouldn’t be deleted. This is the most common source of operator bugs I’ve encountered.
Status subresources are essential. Your operator should report its status through the custom resource’s status subresource. Is the database provisioning? Ready? Failed? What’s the connection string? What’s the last error? Consumers shouldn’t have to read logs to understand what’s happening — the status should tell them everything they need to know.
Test with envtest, not a real cluster. The Kubebuilder test framework includes envtest, which runs a local API server and etcd instance for testing. It’s fast, reliable, and doesn’t require a running Kubernetes cluster. Use it for unit and integration tests. Save real cluster testing for end-to-end validation.
The broader point about Kubernetes complexity
The Kubernetes ecosystem is guilty of accidental complexity on a grand scale. The CNCF landscape has over 1,000 projects. Every problem has seventeen solutions. Every solution introduces three new problems. The cognitive load on platform teams is enormous, and it’s growing.
Operators don’t fix this problem. But they demonstrate what good Kubernetes abstractions look like: they extend the platform rather than competing with it. They reduce cognitive load for consumers. They encode expertise in code. They use the machinery that’s already there.
The next time someone proposes adding a new tool to your Kubernetes stack — another service mesh, another policy engine, another GitOps controller, another abstraction layer — ask a simple question: does this reduce the total complexity of our system, or does it just move it somewhere else?
If the answer is “it moves it somewhere else,” it’s probably not worth the cost. The Kubernetes ecosystem has enough layers. What it needs is fewer, better abstractions that genuinely make things simpler for the people who use them.
Operators got this right. I wish more of the ecosystem would take the hint.
