Sheet ⁨20⁩ · ⁨Series⁩Surveyed ⁨2026⁩
Series

Reliability & Resilience Engineering

Back to Series

Reliability & Resilience Engineering

3parts, in order. Start at the top.

Blog Posts

2 items
Chaos Engineering: Testing Resiliency with Chaos Monkey and Gremlin

Chaos Engineering: Testing Resiliency with Chaos Monkey and Gremlin

Modern software systems are incredibly complex. They're spread across massive networks with countless moving parts. Because of this complexity, unexpected failures are inevitable. Servers crash. Netwo

Designing SLOs and Error Budgets: Your Blueprint for Sustainable Reliability

Designing SLOs and Error Budgets: Your Blueprint for Sustainable Reliability

Every business shipping software is trying to balance two big things: getting new features out fast and keeping their services super reliable. Every tech team deals with this, pushing for new ideas wh

Case Studies

1 item
Taming a 3am Pager: SLOs and Error Budgets That Stuck

Taming a 3am Pager: SLOs and Error Budgets That Stuck

A platform team was losing people to burnout, and the cause was the pager. On-call meant a phone that went off day and night with alerts about CPU, memory, and pod restarts, the overwhelming majority