About The Role
The role owns the reliability, scalability, and security of distributed infrastructure powering high-traffic production systems.
You will work alongside software engineering teams to build modern platform capabilities, automate deployments, and ensure zero-downtime operations.
Key Responsibilities
- Design and maintain cloud infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform or OpenTofu
- Manage Kubernetes clusters in production, ensuring high availability, resource optimization, and secure multi-tenancy
- Build and optimize CI/CD pipelines using GitHub Actions, GitLab CI, or ArgoCD for fast, reliable software delivery
- Implement comprehensive observability stacks using Prometheus, Grafana, Datadog, or OpenTelemetry for monitoring and alerting
- Enforce security best practices, IAM policies, network segmentation, and compliance standards across all cloud environments
- Participate in an on-call rotation to triage and resolve production incidents, conducting root cause analyses to prevent recurrence
What We Are Looking For
- 3–7 years of experience in DevOps, Site Reliability Engineering, or platform engineering roles within cloud-native environments
- Strong proficiency in Infrastructure as Code (Terraform) and container orchestration (Kubernetes, Docker)
- Deep familiarity with CI/CD tooling, GitOps workflows, and Linux system administration
- Programming and scripting skills in Python, Go, or Bash for automation and tooling development
- Bonus: Experience with service mesh technologies like Istio, FinOps cost optimization, or large-scale data migrations