What you will own
- Reliability automation across deployment and regional infrastructure.
- Observability standards, incident tooling, and capacity signals.
- Resilience testing and failure-mode exercises.
- Service-level objectives and production readiness reviews.
What we look for
- Hands-on experience with Linux, networking, and automation.
- Comfort debugging distributed systems from telemetry.
- A pragmatic approach to operational risk.
- Strong incident communication and post-incident learning habits.
First 90 days
Map the highest-risk failure paths, improve one operational control, and run a resilience exercise with the team.
Apply to reliability engineering
Include an example of an incident you helped resolve and what you changed afterward.
