About the Role
fomo is seeking a hands-on Staff Distributed Systems Engineer to own critical shared infrastructure such as datastores, caches, messaging systems, and regional application services. The role focuses on building predictable, high-throughput systems and improving observability, resilience, and data locality during failures and traffic surges.
Requirements
- 8 or more years of backend, platform, or infrastructure engineering experience, or equivalent practical experience.
- Experience designing and debugging distributed, high-throughput production systems.
- Strong PostgreSQL experience, including query performance, indexing, connection pooling, replication, transaction contention, and failure modes.
- Strong experience with Redis-compatible systems such as Redis, Valkey, Dragonfly, or KeyDB, including sharding, replication, memory management, hot keys, and failure handling.
- Experience operating services on AWS, ideally using ECS, RDS, and ElastiCache.
- Experience with infrastructure as code, preferably Terraform.
- Proficiency in Go, TypeScript/Node.js, or a comparable systems-oriented language.
- Hands-on experience designing and testing failover and disaster-recovery systems, including backup restoration, replication, regional failover, and RTO/RPO validation.
- Experience with NATS JetStream, Kafka, or another durable messaging system is a plus.
- Familiarity with Datadog APM and AWS Performance Insights is a plus.
- Experience performing live datastore or cache topology migrations is a plus.
- Experience operating systems with bursty or unpredictable traffic is a plus.
- Experience with financial, trading, cryptocurrency, gaming, or other high-throughput systems is a plus.
Responsibilities
- Design and operate high-throughput, multi-region services.
- Improve datastore and cache performance, capacity, replication, and failure handling.
- Implement backpressure, concurrency limits, load shedding, rate limiting, circuit breakers, and bounded retries.
- Reduce cross-region latency and improve data locality.
- Design and test service, datastore, and regional failover procedures.
- Help architect new features to operate at scale from day one.
- Level up the team on how to think about scale.