About The Role
The DevOps Engineer builds and operates the infrastructure that runs production services at scale. The role spans cloud architecture, deployment automation, observability, incident response, and reliability engineering across containerized workloads and distributed systems.
You will partner with software engineers and security teams to improve delivery speed without compromising stability. The team is looking for an engineer who can turn operational needs into repeatable platforms, clear runbooks, and measurable improvements in system availability and performance.
Key Responsibilities
- Design and manage highly available infrastructure on AWS, including VPCs, IAM, EC2, EKS, RDS, S3, and CloudWatch
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or comparable tooling to automate testing, releases, and rollback procedures
- Provision and configure infrastructure with Terraform, Helm, and Kubernetes while maintaining reusable, version-controlled modules
- Develop observability systems using Prometheus, Grafana, OpenTelemetry, Datadog, or similar tools to track availability, latency, capacity, and error budgets
- Lead incident response for production issues, coordinate technical remediation, and improve runbooks and post-incident action plans
- Harden production environments through access controls, secrets management, vulnerability remediation, patching, and infrastructure-as-code reviews
- Collaborate with application teams to improve service reliability, deployment safety, scalability, and operational readiness before launch
What We Are Looking For
- 3–8 years of experience in DevOps, site reliability engineering, platform engineering, or a closely related infrastructure role
- Hands-on experience operating production workloads in AWS or another major cloud provider, including networking, compute, storage, identity, and managed databases
- Strong Kubernetes experience, including deployments, services, ingress, resource management, troubleshooting, and Helm-based releases
- Proficiency with Terraform or an equivalent infrastructure-as-code tool and practical experience designing maintainable CI/CD pipelines
- Working knowledge of Linux administration, Bash or Python scripting, Git workflows, and common networking concepts such as DNS, TLS, HTTP, and load balancing
- Experience with monitoring, logging, tracing, on-call operations, incident management, and reliability practices such as SLIs, SLOs, and error budgets
- Bachelor’s degree in computer science, engineering, information technology, or a related field, or equivalent practical experience; Bonus: experience with Go, service meshes, ArgoCD, Vault, Kafka, multi-cloud environments, or compliance-focused infrastructure