
Taming a 3am Pager: SLOs and Error Budgets That Stuck
- Reliability
- DevOps
- 29 Aug 2026
- 07 Mins read
A platform team was losing people to burnout, and the cause was the pager. On-call meant a phone that went off day and night with alerts about CPU, memory, and pod restarts, the overwhelming majority

Multi-Region Active-Active for a Payments API
- Architecture
- Cloud Computing
- 16 Aug 2026
- 06 Mins read
A payments API that moves real money had been running comfortably in a single AWS region for years. It was reliable until the day it was not: a regional control-plane incident took the whole service o