About The Role
The role owns the reliability, scalability, and security of production infrastructure supporting high-traffic applications serving millions of users globally.
You will work alongside software engineering teams to design resilient systems, automate deployment pipelines, and establish rigorous observability standards.
Key Responsibilities
- Design, provision, and manage cloud infrastructure on AWS using Terraform and infrastructure-as-code best practices
- Build and maintain robust CI/CD pipelines using GitHub Actions or GitLab CI for seamless, automated deployments
- Manage Kubernetes clusters in production, ensuring high availability, optimal resource utilization, and secure cluster configuration
- Implement comprehensive monitoring, logging, and alerting systems using Prometheus, Grafana, and Datadog
- Participate in an on-call rotation to troubleshoot and resolve production incidents, conducting thorough post-mortems to prevent recurrence
- Enforce security baselines, manage IAM roles, and participate in compliance and vulnerability remediation efforts
What We Are Looking For
- 3–6 years of experience in DevOps, Site Reliability Engineering, or systems engineering roles in high-growth environments
- Strong proficiency with AWS core services (EKS, EC2, RDS, IAM, VPC) and Infrastructure-as-Code tools like Terraform
- Deep operational experience with Kubernetes, Docker containerization, and service mesh architectures
- Solid programming and scripting skills in Python, Go, or Bash for automation and tooling
- Deep understanding of networking concepts (TCP/IP, DNS, TLS, load balancing) and Linux system administration
- Bonus: Experience with service mesh technologies like Istio, FinOps cost optimization practices, or holding active AWS/CKA certifications