GitOps platform
A cluster nobody can describe is a cluster nobody can rebuild
A production-shaped GitOps platform, built end to end as a reference. Ten polyglot microservices across three environments, with Git as the only source of truth and ArgoCD reconciling the cluster back to it.
- Role
- Architecture and build
- Services
- 10, in Go, Java, Node.js and Python
- Environments
- dev, staging, prod
Clusters drift. Someone patches a deployment by hand at two in the morning, it works, and now the running state and the repository disagree with no record of which is right. Multiply that across three environments and ten services and rebuilding becomes archaeology.
The second problem is release coupling. When one pipeline deploys everything, a change to one service waits on every other service being ready.
Git decides, ArgoCD enforces, CI never touches the cluster
A root ArgoCD application owns child applications, App of Apps, so the whole platform is one object to install and one place to look. Each service is a Helm chart with per environment values, and ArgoCD checks for drift every three minutes with auto-sync and self healing, so a hand patch is reverted rather than silently kept.
CI never deploys. GitHub Actions builds each service, tags the image with the commit SHA, pushes it, and commits the per-service tag override back into the GitOps repository. ArgoCD picks that up and rolls that one service, which is what decouples the releases.
Around that: HPA on CPU and custom metrics, External Secrets Operator pulling from AWS Secrets Manager so no secret lives in Git, nginx ingress with cert-manager and NetworkPolicies, and RBAC at both the Kubernetes and ArgoCD layers. Terraform provisions the EKS side.
Observability is part of the platform, not bolted on: kube-prometheus-stack delivered through ArgoCD as a multi-source application and tuned for an 8 GB node, ServiceMonitors per service, and the cart and payment services instrumented with prom-client counters and gauges, verified end to end in Grafana through PromQL.
Two fixes worth keeping: the upstream Dockerfiles referenced an openjdk:8-jdk base that had been deleted, which I replaced with a maven:3.8-eclipse-temurin-8 multi-stage build on amazoncorretto:8, and the Go dispatch service needed Go 1.22 or later to clear a pprof dependency break.
What the platform guarantees
| Measure | Before | After | Change |
|---|---|---|---|
| Drift detection interval | manual | 3 minutes | auto-sync, self heal |
| Services deployable independently | 0 of 10 | 10 of 10 | per-service SHA tags |
| Secrets stored in Git | some | none | External Secrets Operator |
| Environments from one manifest set | n/a | 3 | Helm values per env |