Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and implement observability instrumentation, telemetry pipelines, internal platforms, and tooling. You will operationalize service-level indicators, objectives, and alerting; improve incident root-cause analysis; create actionable dashboards; and balance telemetry signal, cost, noise, and performance impact.
Requirements
- Backend or systems software engineering
- Go, C++, Rust, Java, or Python
- Distributed systems
- Networking
- Concurrency
- Performance optimization
- Metrics
- Logging
- Distributed tracing
- Production monitoring
- Alerting
- OpenTelemetry
- Prometheus
- Grafana
- Datadog, Elastic, Jaeger, Tempo, or similar tools
- Telemetry pipeline
- Service-level indicator
- Service-level objective
Responsibilities
- Design and implement observability instrumentation across services and platforms
- Build and maintain telemetry pipelines for metrics, logs, and traces
- Develop internal observability platforms, libraries, and tooling
- Define and operationalize service-level indicators, objectives, and alerting strategies
- Make systems debuggable by design
- Reduce incident resolution time through root-cause analysis
- Create actionable dashboards and alerts
- Balance telemetry signal against cost, noise, and performance impact
- Improve the developer experience for observability and debugging