CASE STUDY 04Cloud cost · reliabilityLive in production

Mediatiz Foundation

A bill sized for the peak, paid at every hour of the day

Mediatiz runs a media literacy programme at national scale: a learning platform and an Android app used by hundreds of thousands of students and teachers. I owned the AWS infrastructure underneath it, the release pipelines that shipped to it, and the reporting that ran on top.

Role
DevOps Engineer
Timeline
Oct 2025 to Apr 2026
Scale
1M+ users, 20K to 30K concurrent
The problem

The platform had to survive classroom-hour spikes of twenty to thirty thousand concurrent requests, so everything had been provisioned for the spike and left there. Outside those hours the same capacity sat idle and still billed.

Separately, an AI tutor feature was planned on a third party LLM API. That meant a per-token bill that scaled with student numbers, and student questions leaving the country, which the programme could not accept.

Architecture
Autoscaling follows real load; the LLM tutor runs self hosted1M+ users20K-30K concurrentALB + ECStask definitions tunedto measured usage30% less spendCloudWatch drives scalingRDS + S3right-sizedPyBot, self hostedDjango + Ollama, Gemma 2Bno external API billstudent PII never leavesServerless reportingLambda, API Gateway, EventBridge8 departments, was manualFour codebases shipped through GitHub Actions: Android, Laravel LMS, Django, Next.js admin. Zero-downtime ECS rollouts.
What I did

Follow the actual load, and bring the model in house

I put the scaling decisions on real signals: CloudWatch driven auto-scaling, ECS task definitions tuned to what the containers actually used rather than what had been guessed, and workload-aware right-sizing across staging and production. Thirty percent of the bill came off, and uptime stayed at 99.9%.

For the tutor I built PyBot instead: Django in front of Ollama running Gemma 2B on self-hosted infrastructure. No external API bill, and student data never leaves the estate.

I also built the release pipelines for four codebases at once, an Android app, a Laravel LMS, a Django service and a Next.js admin dashboard, with zero-downtime ECS rollouts, Slack deploy alerts and Play Console automation. Cross-departmental reporting moved onto a serverless Lambda, API Gateway, EventBridge and S3 path that replaced a manual workflow for eight departments.

And I ran the security testing in house against staging, brute force, SQL injection and CSRF, with SQLMap and SonarQube, then handed the development team a remediation report they adopted.

Results

Measured over the engagement

MeasureBeforeAfterChange
AWS monthly spendbaseline30% lowersame workload
Uptime at peak loadn/a99.9%sustained
External LLM API costper tokenzeroself hosted Gemma 2B
Departments on manual reporting80serverless pipeline
See it
All workNext: NextLab