About The Role
The role is responsible for the availability, latency, performance, efficiency, and capacity management of a high-throughput cloud-native platform. This involves designing and implementing automated infrastructure pipelines that support rapid deployment cycles while maintaining strict uptime and reliability SLAs.
The engineer will work closely with software engineering teams to containerize workloads, manage stateful and stateless services in Kubernetes, and build comprehensive observability frameworks that enable proactive system monitoring and rapid incident response.
Key Responsibilities
- Design, provision, and maintain multi-region cloud infrastructure on AWS or GCP using Terraform and Infrastructure-as-Code best practices
- Manage and scale production-grade Kubernetes (EKS/GKE) clusters, including networking, ingress controllers, and IAM integrations
- Develop and optimize CI/CD pipelines using GitHub Actions, GitLab CI, or ArgoCD to automate build, test, and deployment processes
- Implement comprehensive observability stacks using Prometheus, Grafana, OpenTelemetry, and ELK/Datadog to monitor system health and latency
- Participate in a blameless on-call rotation, conducting post-mortems and implementing preventative automation to reduce operational toil
- Collaborate with security teams to enforce IAM roles, network security policies, and vulnerability scanning within build and runtime environments
What We Are Looking For
- 3–6 years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering supporting high-traffic production environments
- Strong proficiency with Infrastructure as Code (IaC) tools, specifically Terraform, and container orchestration with Kubernetes
- Solid software engineering skills in at least one scripting or programming language, such as Python, Go, or Bash
- Deep understanding of Linux systems administration, networking fundamentals (TCP/IP, DNS, VPCs), and cloud security best practices
- Experience configuring and managing CI/CD tools and modern observability pipelines (Prometheus, Grafana, Datadog)
- Bonus: Experience with service meshes (Istio/Linkerd), GitOps workflows (ArgoCD/Flux), or managing distributed databases at scale