About The Role
The role focuses on maintaining high availability, low latency, and robust security across large-scale distributed systems supporting millions of active users.
The team works closely with software engineers to automate infrastructure provisioning, enhance observability, and drive incident response protocols.
Key Responsibilities
- Design and manage cloud infrastructure on AWS using Terraform and Kubernetes, ensuring high availability and zero-downtime deployments
- Build and maintain comprehensive monitoring, logging, and alerting systems using Prometheus, Grafana, and Datadog
- Automate deployment pipelines and release processes with CI/CD tools such as GitHub Actions and ArgoCD
- Lead incident response, conduct thorough post-mortem analyses, and implement preventive measures to eliminate recurring failure modes
- Optimize infrastructure costs, resource utilization, and database performance across multiple production environments
- Participate in an on-call rotation and collaborate with development teams to ensure production readiness for new services
What We Are Looking For
- 3–6 years of experience in Site Reliability Engineering, DevOps, or systems administration within cloud-native environments
- Strong proficiency in Linux systems administration, networking fundamentals (TCP/IP, DNS, TLS), and containerization technologies (Docker, Kubernetes)
- Hands-on experience with Infrastructure as Code (Terraform, CloudFormation) and configuration management tools
- Proficiency in at least one scripting or programming language: Python, Go, or Bash
- Solid understanding of distributed systems architecture, microservices patterns, and resilience engineering principles
- Bonus: Bachelor's degree in Computer Science, certified Kubernetes administrator (CKA), or experience with service mesh technologies like Istio