CASE STUDY 06Kubernetes · GitOpsReference build, not client work

GitOps platform

A cluster nobody can describe is a cluster nobody can rebuild

A production-shaped GitOps platform, built end to end as a reference. Ten polyglot microservices across three environments, with Git as the only source of truth and ArgoCD reconciling the cluster back to it.

Role
Architecture and build
Services
10, in Go, Java, Node.js and Python
Environments
dev, staging, prod
The problem

Clusters drift. Someone patches a deployment by hand at two in the morning, it works, and now the running state and the repository disagree with no record of which is right. Multiply that across three environments and ten services and rebuilding becomes archaeology.

The second problem is release coupling. When one pipeline deploys everything, a change to one service waits on every other service being ready.

Architecture
CI writes image tags to Git; ArgoCD reconciles the cluster to Git every three minutesGitHub Actionsbuilds, tags by SHAnever deployscommitsGitOps repoHelm charts + valuesthe only source of truthArgoCDApp of Appsauto-sync, self healCluster, 3 environments10 services, 4 languagesHPA on CPU + custom metricsExternal Secrets from AWSnginx ingress, cert-managerkube-prometheus-stackdrift check every 3 minutes, hand patches revertedA per-service tag override means one service rolls without waiting on the other nine.
What I did

Git decides, ArgoCD enforces, CI never touches the cluster

A root ArgoCD application owns child applications, App of Apps, so the whole platform is one object to install and one place to look. Each service is a Helm chart with per environment values, and ArgoCD checks for drift every three minutes with auto-sync and self healing, so a hand patch is reverted rather than silently kept.

CI never deploys. GitHub Actions builds each service, tags the image with the commit SHA, pushes it, and commits the per-service tag override back into the GitOps repository. ArgoCD picks that up and rolls that one service, which is what decouples the releases.

Around that: HPA on CPU and custom metrics, External Secrets Operator pulling from AWS Secrets Manager so no secret lives in Git, nginx ingress with cert-manager and NetworkPolicies, and RBAC at both the Kubernetes and ArgoCD layers. Terraform provisions the EKS side.

Observability is part of the platform, not bolted on: kube-prometheus-stack delivered through ArgoCD as a multi-source application and tuned for an 8 GB node, ServiceMonitors per service, and the cart and payment services instrumented with prom-client counters and gauges, verified end to end in Grafana through PromQL.

Two fixes worth keeping: the upstream Dockerfiles referenced an openjdk:8-jdk base that had been deleted, which I replaced with a maven:3.8-eclipse-temurin-8 multi-stage build on amazoncorretto:8, and the Go dispatch service needed Go 1.22 or later to clear a pprof dependency break.

Results

What the platform guarantees

MeasureBeforeAfterChange
Drift detection intervalmanual3 minutesauto-sync, self heal
Services deployable independently0 of 1010 of 10per-service SHA tags
Secrets stored in GitsomenoneExternal Secrets Operator
Environments from one manifest setn/a3Helm values per env
All workNext: GPU inference platform