Job Title: Senior DevOps Engineer
Experience: 10+ Years
Location: Bangalore
Employment Type: Full-Time
Job Summary
We are looking for an experienced Senior DevOps Engineer with 10+ years of hands-on experience in DevOps, cloud infrastructure, automation, CI/CD, Infrastructure as Code, and cloud-native technologies.
The ideal candidate should have strong expertise in AWS, Terraform, Kubernetes/EKS, Docker, GitLab CI/CD, ArgoCD, serverless technologies, and observability tools such as Datadog/Splunk. The candidate will be responsible for designing, implementing, securing, and supporting highly scalable and reliable cloud platforms and production deployment environments.
Key Responsibilities
- Design, build, and support cloud-native platforms, deployment pipelines, and automation frameworks in production.
- Develop and maintain CI/CD pipelines using GitLab CI/CD, Jenkins, or equivalent tools.
- Implement Infrastructure as Code using Terraform, including reusable modules, environment consistency, peer review, and controlled deployments.
- Design and manage AWS cloud environments across Production, Development, QA, and lower environments.
- Work extensively with AWS services such as:
- Lambda
- API Gateway
- IAM
- EventBridge
- DynamoDB
- CloudWatch
- VPC/Networking
- Security services
- Implement and support serverless applications and APIs, including authentication, integrations, security, and operational support.
- Manage containerized workloads using Docker and Kubernetes, preferably AWS EKS.
- Implement GitOps-based deployment practices using ArgoCD or similar tools.
- Design and improve observability solutions using Datadog, Splunk, CloudWatch, Prometheus, or similar tools.
- Build actionable monitoring and alerting strategies to reduce alert noise and improve incident response.
- Implement resiliency, fault tolerance, disaster recovery, backup, and rollback strategies for business-critical applications.
- Support safe production releases through automation, validation, canary/phased deployments, and controlled rollback mechanisms.
- Perform performance tuning, capacity planning, cost optimization, system analysis, and Root Cause Analysis (RCA).
- Integrate and support third-party/vendor APIs and managed AWS services.
- Apply cloud security best practices, IAM controls, privileged access management, and secure deployment practices.
- Use scripting and automation with Python, Bash, PowerShell, or similar technologies.
- Explore and implement AI-driven solutions to improve automation, observability, operational efficiency, and incident management.
Required Technical Skills Cloud & AWS
- Strong hands-on experience with AWS
- AWS Lambda
- API Gateway
- IAM
- EventBridge
- DynamoDB
- CloudWatch
- Networking & Security
Infrastructure as Code
- Terraform
- CloudFormation or equivalent
- Reusable Terraform modules
- Infrastructure automation and environment management
CI/CD & GitOps
- GitLab CI/CD
- Jenkins or equivalent
- ArgoCD
- GitOps practices
- Automated deployment and release management
Containers & Orchestration
Monitoring & Observability
- Datadog
- Splunk
- CloudWatch
- Prometheus or similar monitoring platforms
- Alerting, logging, metrics, and incident response
Scripting
- Python
- Bash
- PowerShell or similar scripting languages
Additional / Preferred Skills
- Kafka / event streaming
- Kong API Gateway
- Akamai / CDN technologies
- Serverless architecture
- Distributed systems and system design
- Scalability and performance engineering
- Disaster Recovery and business continuity
- eCommerce, retail, or inventory management platforms
- AI/ML technologies and AI-driven DevOps automation
Required Qualifications
- 10+ years of experience in DevOps, Systems Engineering, Cloud Engineering, or related roles.
- Bachelor's degree in Computer Science, Engineering, or equivalent experience.
- Strong hands-on experience with AWS cloud platforms.
- Proven experience with CI/CD tools such as GitLab CI/CD or Jenkins.
- Hands-on experience with Terraform / Infrastructure as Code.
- Experience with GitOps and ArgoCD or similar tools.
- Strong understanding of Kubernetes and containerization.
- Strong scripting/programming skills using Python, Bash, PowerShell, or similar.
- Good understanding of system design, scalability, distributed systems, security, and reliability engineering.
- Strong production troubleshooting and incident management experience.