Help ensure uptime, disaster recovery, and automated incident response across our clients' mission-critical clusters.
Requirements
- 4+ years of SRE or production infrastructure experience
- Deep knowledge of Linux internals and networking
- Hands-on monitoring and observability with OpenTelemetry/Grafana