Aakash
Parmar
← All posts

Case study · 2 Aug 2026 · 3 min read

Auditing a live Kubernetes cluster without taking it down

The checklist I follow when a cluster already runs real workloads and needs to be made production ready: CIS checks, RBAC cleanup, audit logs and network policies.

Building a secure cluster from scratch is the easy case. The harder one is when a cluster has been running for a year, has 30 apps on it, and someone asks “is this production ready?” You can’t rebuild it, and you can’t break anything while you look.

This is roughly how I approach that.

1. Find out what’s actually there

Before fixing anything, I want a clear picture:

kubectl get nodes -o wide
kubectl get ns
kubectl get pods -A -o wide
kubectl get clusterrolebindings -o wide

I also pull the API server, kubelet and etcd flags from the control plane nodes (or the managed provider’s settings). A lot of problems show up just from reading the flags: anonymous auth left on, audit logging off, no encryption for secrets at rest.

2. Run the CIS Benchmark

kube-bench checks nodes against the CIS Kubernetes Benchmark. I run it as a Job on each node type:

kubectl apply -f https://raw.githubusercontent.com/aquasecurity/kube-bench/main/job.yaml
kubectl logs job/kube-bench

The output is long. I sort the failures into three buckets:

3. Clean up RBAC

This is usually where the biggest risk is. Over time, people hand out cluster-admin to get something working and never take it back.

The first thing I look for is every binding to cluster-admin:

kubectl get clusterrolebindings -o json \
  | jq -r '.items[] | select(.roleRef.name=="cluster-admin") | .metadata.name + " -> " + ([.subjects[]?.name] | join(", "))'

Then I check what each service account can really do:

kubectl auth can-i --list --as=system:serviceaccount:payments:default -n payments

Some common things I end up fixing:

I don’t remove permissions blindly. I turn on audit logs first, watch which permissions are actually used for a week or so, and then trim to that.

4. Turn on audit logging

If audit logs are off, you have no way to answer “who deleted that deployment?” A simple policy is enough to start:

apiVersion: audit.k8s.io/v1
kind: Policy
rules:
  - level: None
    resources:
      - group: ""
        resources: ["events"]
  - level: Metadata
    resources:
      - group: ""
        resources: ["secrets", "configmaps"]
  - level: RequestResponse
    verbs: ["create", "update", "patch", "delete"]
  - level: Metadata

Note that secrets are logged at Metadata level only, so their values never end up in the logs. Ship the logs somewhere outside the cluster so they survive if the cluster has a bad day.

5. Default deny network policies

Out of the box, every pod can talk to every other pod. I add a default deny policy per namespace and then open only the traffic each app needs. With Cilium you also get Hubble, which shows you the real traffic flows, so you can write policies based on what is happening instead of guessing.

Roll this out one namespace at a time. Start with the least critical one and keep Hubble open while you do it.

6. Make it stick

An audit is only useful if the cluster doesn’t drift back. A few things that help:

What I learned

The technical fixes are rarely the hard part. The hard part is changing things without surprising the teams who run apps on the cluster. Telling them what’s changing and when, and giving them a way to test in staging first, matters more than any single CIS check.

Contact

Let's ship
something secure.

Looking for DevOps, DevSecOps or platform roles. Based in India, happy to work remote.