About The Role
The role owns the reliability, scalability, and security of production infrastructure supporting millions of daily users across distributed cloud environments.
You will collaborate closely with software engineering teams to build automated deployment pipelines, enhance observability, and ensure robust disaster recovery protocols.
Key Responsibilities
- Architect and maintain cloud infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform and Ansible
- Design and manage CI/CD pipelines using GitHub Actions, GitLab CI, or ArgoCD to ensure rapid and safe code deployments
- Implement comprehensive monitoring, logging, and alerting systems using Prometheus, Grafana, and Datadog
- Drive containerization efforts using Docker and orchestrate workloads with Kubernetes across multi-region environments
- Participate in an on-call rotation to troubleshoot and resolve critical production incidents efficiently
- Enforce security best practices, access controls, and compliance standards across all cloud resources
What We Are Looking For
- 3–6 years of experience in DevOps, site reliability engineering, or infrastructure engineering roles
- Strong proficiency in cloud platforms (AWS, GCP, or Azure) and infrastructure-as-code frameworks (Terraform)
- Hands-on experience with Kubernetes, Docker, and service mesh architectures in high-throughput production environments
- Solid scripting skills in Python, Bash, or Go for automation and tooling
- Deep understanding of networking fundamentals, TCP/IP, DNS, load balancing, and TLS configurations
- Bonus: Certified Kubernetes Administrator (CKA), experience with service meshes like Istio, or background in FinOps cost optimization