Firefly.online, off the cloud
Every managed service was a line item and a dependency
Firefly.online is an IoT SaaS platform built on Node.js, React and PostgreSQL. It ran entirely on AWS managed services. I led the move to self-managed bare metal, end to end, without taking the product offline.
- Role
- DevOps Engineer, migration lead
- Timeline
- Dec 2023 to Oct 2025
- Downtime
- None, customer facing
- Outcome
- ~60% lower spend
EC2, ECS, RDS, ALB, S3, CloudFront, Lambda and Redis. Convenient, and expensive at the platform’s size, with limited control over placement and tuning. The business wanted the spend down and the control back.
The constraint was that customers were live on it. A migration that needed a maintenance window long enough to move a production PostgreSQL database and its object storage was not acceptable.
Stand the new one up, prove it, then move the traffic
Every managed service got a self-hosted equivalent provisioned and running in parallel: PostgreSQL for RDS, nginx for the load balancer, self-hosted object storage for S3, Redis on our own hardware. The new stack ran alongside the old one until it was demonstrably correct.
The cutover was blue green: data migrated with replication catching up to the live database, then traffic switched. PostgreSQL, Redis persistence and object storage all moved with no customer-facing downtime, and infrastructure spend fell by roughly sixty percent.
The platform was simultaneously moving from Node.js to Java Spring Boot, so I designed and provisioned the development, staging and production environments for the new stack: Ubuntu hosts with Java, Maven, Docker, PostgreSQL, Redis, Apache Kafka and nginx upstream load balancing across multiple backend instances.
Then I made it observable and hardened it: Grafana, Prometheus, Loki and Node Exporter for metrics, centralised logging and alerting, with Let’s Encrypt, UFW, Fail2ban and SSH key restrictions across every environment. GitHub Actions and SonarQube handled code quality gates, image builds and per-environment deploys.
What the move produced
| Measure | Before | After | Change |
|---|---|---|---|
| Infrastructure spend | baseline | ~60% lower | same platform |
| Customer facing downtime | n/a | zero | blue green cutover |
| AWS managed services replaced | 8 | 0 | all self hosted |
| Environments rebuilt | ad hoc | 3 | dev, staging, prod |