About the Role
You will own the architecture, deployment, and operation of production infrastructure across cloud services, Kubernetes, databases, CI/CD, observability, security, and blockchain systems. You will automate operations, design resilient systems, investigate incidents, and improve availability, recovery, and performance.
Requirements
- 5+ years of DevOps, SRE, platform, or infrastructure engineering experience
- Experience building and operating highly available production infrastructure
- Hands-on Kubernetes, AWS, and Linux experience
- Experience with Terraform, GitOps, and CI/CD pipeline design
- Networking knowledge including DNS, load balancing, TLS, and service discovery
- PostgreSQL production experience including replication, backups, recovery, and performance
- Experience with high-availability design, incident response, root-cause analysis, and capacity planning
- Hands-on blockchain infrastructure experience and knowledge of distributed blockchain network failures
Responsibilities
- Own infrastructure architecture and operations for production systems
- Design, build, and operate highly available AWS cloud infrastructure
- Manage infrastructure as code with Terraform, GitOps, and CI/CD pipelines
- Operate PostgreSQL for availability, replication, backups, performance, and recovery
- Build observability and early incident alerting across infrastructure, applications, databases, and blockchain nodes
- Run and monitor blockchain nodes, RPC services, and indexers
- Design disaster recovery, failover, and security hardening for critical systems
- Diagnose full-stack production incidents and automate operational processes
Benefits