About The Role
We are seeking an experienced and technically strong AWS DevOps Engineer with 4+ years of hands-on experience to join our engineering team. You will play a critical role in architecting and managing scalable AWS cloud platforms, driving enterprise-scale CI/CD automation, implementing Infrastructure as Code (IaC), and optimizing Kubernetes orchestration.
In this role, you will collaborate closely with development, QA, security, and infrastructure teams in a fast-paced Agile environment to ensure operational stability, optimize cloud efficiency, and strengthen security and observability across multiple environments.
What You’ll Do
- Cloud Infrastructure & Architecture: Design, build, and manage highly scalable, secure, and resilient AWS infrastructure (EC2, EKS, RDS, S3, VPC). Execute modernization initiatives, multi-account AWS IAM setups, and disaster recovery strategies.
- DevOps & CI/CD Engineering: Design and maintain enterprise-grade CI/CD pipelines using GitHub Actions, Jenkins, CircleCI, or AWS CodePipeline. Build automated deployment workflows supporting microservices with blue-green, canary, and rolling strategies.
- Infrastructure as Code & Automation: Lead IaC implementations using Terraform, Terragrunt, AWS CloudFormation, or CDK. Develop reusable modules, automate environment provisioning, configuration management, and enforce GitOps best practices.
- Containerization & Kubernetes: Manage and optimize containerized workloads using Docker and Kubernetes (EKS). Deploy and maintain clusters with Helm charts, auto-scaling, ingress controllers, and troubleshoot cluster networking and performance.
- Monitoring, Observability & Reliability: Implement enterprise observability solutions using Prometheus, Grafana, CloudWatch, OpenSearch, or ELK stack. Define alerting standards and perform root cause analysis (RCA) for critical incidents.
- Production Operations & On-Call Support: Participate in on-call support rotations to respond to high-severity production incidents, perform thorough root cause analysis (RCA), and implement long-term remediation to minimize Mean Time to Repair (MTTR).
- Cloud Cost Optimization & FinOps: Proactively monitor and optimize AWS cloud expenditure by applying resource right-sizing, cost-allocation tagging, spot/reserved instance strategies, and lifecycle management policies.
- Security & Compliance: Implement DevSecOps practices, managing IAM policies, secrets management, encryption, and integrating security scanners (Trivy, SonarQube, Checkov) into CI/CD pipelines.
- Collaboration & Leadership: Mentor junior DevOps engineers, promote an automation culture, participate in architecture reviews, capacity planning, and drive continuous delivery improvements.
Qualifications & Requirements
- Experience: 4+ years of dedicated hands-on experience in DevOps, Cloud Infrastructure, or Site Reliability Engineering (SRE).
- AWS Expertise: Deep experience with AWS core services including EC2, ECS, EKS, Lambda, VPC, Route53, RDS, S3, IAM, and CloudWatch.
- Infrastructure as Code: Strong hands-on expertise in Terraform, Terragrunt, Ansible, and automated infrastructure lifecycle management.
- CI/CD & Automation: Proven track record of building complex CI/CD workflows using GitHub Actions, Jenkins, CircleCI, or AWS DevOps tools.
- Containers & Kubernetes: Solid proficiency in Docker, Kubernetes (EKS), Helm charts, auto-scaling, and container networking.
- Operations & Reliability: Hands-on experience in production on-call support, incident troubleshooting, RCA execution, and MTTR reduction.
- Scripting & Systems: Strong Linux system administration and scripting skills in Bash or Python (Golang is a plus).
- Observability: Experience setting up monitoring, logging, and alerting using Prometheus, Grafana, CloudWatch, and ELK/OpenSearch.
- Cost Optimization: Practical knowledge of FinOps practices, resource right-sizing, cost tagging, and AWS billing optimization tools.
Preferred Skills (Nice-to-Have)
- Certifications: AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect.
- GitOps & Service Mesh: Hands-on experience with ArgoCD, FluxCD, or service mesh technologies (Istio/Linkerd).
- Data & Messaging: Familiarity with caching and messaging systems (Redis, Kafka, RabbitMQ) and databases (PostgreSQL, DynamoDB, MongoDB).
- Practices: Exposure to SRE principles (SLIs/SLOs, error budgets) and security tools (Artifactory, Checkov).
Information Security & ISO 27001 Compliance Security is at the heart of everything we do. In this role, you will be strictly required to uphold our Information Security Management System (ISMS) policies in alignment with ISO/IEC 27001 standards. Responsibilities include safeguarding sensitive asset data, completing mandatory security awareness training, reporting potential security incidents or vulnerabilities immediately, and ensuring daily operations comply with rigorous data protection protocols.