When I joined, every project deployed a little differently. Some used scripts on a VM, some ran kubectl apply from CI, and a couple were updated by hand. It worked, but nobody could tell you for sure what version was running where.
Now 25+ projects follow the same flow. Here’s how it fits together.
The big picture
There are two kinds of repos:
- App repos. The code, a Dockerfile and a CI workflow.
- A config repo. Helm values and Kubernetes manifests for every app and every environment.
CI builds and checks the image, then opens a change in the config repo. ArgoCD watches the config repo and makes the cluster match it. CI never talks to the cluster directly.
That last part is important. The CI runners don’t hold cluster credentials at all, so a leaked CI token can’t be used to change production.
Step 1: Build and check in CI
The workflow in each app repo is short because the real logic lives in a shared reusable workflow:
name: build
on:
push:
branches: [main]
jobs:
build:
uses: our-org/ci-templates/.github/workflows/build-and-scan.yml@v3
with:
image: ghcr.io/our-org/payments-api
secrets: inherit
The shared workflow does these steps:
- SAST on the source code to catch things like SQL injection or hardcoded secrets.
- Dependency scan (SCA) for known vulnerable libraries.
- Build the image with a multi-stage Dockerfile.
- Image scan with Trivy. High or critical issues that have a fix available fail the build.
- Push the image, tagged with the commit SHA. Never
latest.
Keeping this in one shared workflow means a fix or new check gets rolled out to every project by bumping one version tag.
Step 2: Update the config repo
If everything passes, CI updates the image tag in the config repo:
yq -i '.image.tag = "'"$GITHUB_SHA"'"' apps/payments-api/values-staging.yaml
git commit -am "payments-api: deploy ${GITHUB_SHA::7} to staging"
git push
For staging this is a direct commit. For production it opens a pull request, so a person has to approve it. Every deploy ends up as a commit, which gives us a full history and makes rollbacks a git revert.
Step 3: Policy checks before anything is applied
Before ArgoCD applies a change, Kyverno checks the manifests. Some of the rules we enforce:
- Containers must run as non root and drop all capabilities.
- Images must come from our own registry.
- Every container needs CPU and memory requests and limits.
- No
hostPathmounts or privileged pods.
The same policies also run in CI against the config repo using the Kyverno CLI, so most problems get caught on the pull request instead of at sync time:
kyverno apply ./policies --resource ./rendered/
Step 4: ArgoCD syncs it
ArgoCD notices the new commit and syncs the app. We use the app of apps pattern, so adding a new project is mostly adding a folder to the config repo.
A few settings that made a difference:
- Auto sync with self heal on staging, so manual changes in the cluster get reverted.
- Manual sync on production, for the extra control.
- Prune turned on, so resources deleted from Git get deleted from the cluster too.
Secrets
Secrets never go into Git in plain text. Depending on the project we use Sealed Secrets, SOPS or pull them from Vault at runtime. The config repo only ever holds encrypted values or references.
What changed for the team
Deploys went from something a couple of people knew how to do to something anyone on the team can follow. Once a commit lands, the rollout itself takes seconds. And when someone asks “what’s running in production right now?”, the answer is just whatever is in the config repo.
If I were starting over, I’d set up the shared CI workflow first. Getting every project onto the same pipeline was the hardest part, and having one template from day one would have saved a lot of migration work.