About The Role
The role owns the availability, latency, performance, efficiency, and capacity management of core infrastructure serving millions of global users.
The team builds and maintains the platform automation, CI/CD pipelines, and observability systems that allow engineering groups to ship code rapidly and safely.
Key Responsibilities
- Design, build, and maintain production infrastructure on AWS using Terraform, Kubernetes, and Helm charts
- Manage and scale observability pipelines utilizing Prometheus, Grafana, and Datadog for real-time alerting and metrics collection
- Automate deployment workflows and CI/CD pipelines with GitHub Actions to streamline releases and minimize downtime
- Participate in an on-call rotation to troubleshoot and resolve infrastructure incidents, conducting thorough root cause analyses
- Implement security hardening, IAM policies, and compliance standards across all cloud resources
- Write infrastructure as code and internal tooling in Go or Python to eliminate manual operational toil
What We Are Looking For
- 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering roles in high-scale environments
- Deep expertise in Kubernetes administration, containerization, and service mesh architectures
- Strong proficiency in infrastructure-as-code tooling, specifically Terraform and CloudFormation
- Solid programming skills in Python, Go, or Bash for automation and tooling development
- Extensive experience with major cloud providers, preferably AWS, including networking, IAM, and compute primitives
- Bonus: Experience migrating legacy architectures to cloud-native stacks or contributing to open-source infrastructure projects