Location: Bangalore
Experience: 8–10 years
Role Summary:
As a Sr. DevOps Engineer / SRE, you will own CI/CD pipelines, multi-cloud infrastructure,
Kubernetes/EKS operations, and production reliability across multiple high-throughput product pods. This is a hands-on and leadership role you will design automation, reduce operational toil, enforce security/compliance, lead incident response, mentor junior engineers, and oversee complex release management.
Key Responsibilities
CI/CD & Release Management Design, implement, and maintain Jenkins declarative pipelines for 5–10 microservice pods.
Lead release management across multiple environments, including production; target <10
mins deployment and fast rollback.
Architect blue/green, canary, and rolling deployment strategies for high-availability systems.
Infrastructure & IaC Author and maintain reusable Terraform/Terragrunt modules.
Provision and manage AWS cloud infrastructure: VPC, ALB/NLB, Route53, S3, RDS/EFS, EC2,
and EKS.
Implement infrastructure automation and enforce IaC best practices.
Kubernetes & Containerization
Operate and scale EKS/Kubernetes clusters; manage upgrades, Helm charts, service mesh,
and resource tuning.
Ensure containerized applications follow best practices for reliability and performance.
Observability, SLA & Reliability
Implement monitoring and alerting (Datadog, ELK, Prometheus/Grafana).
Define SLOs, SLIs, error budgets, and proactively drive reliability improvements.
Security & Compliance
Integrate SAST/DAST into CI pipelines (SonarQube, Checkmarx).
Enforce IAM, secrets management, and cloud security best practices.
Cost & Performance
Run FinOps initiatives: optimize instance types, reserved vs. spot usage, and propose measurable cost savings.
Conduct load/stress testing and tuning to maximize performance.
Support & On-call
Lead P0/P1 incidents; mentor L1/L2 engineers during incident response.
Own escalation processes and post-mortems; track MTTR improvements.
Must-Haves
● 8–10 years experience in production AWS environments (EKS/ECS, networking, VPC,
ALB/NLB).
● Hands-on Jenkins (declarative pipelines) and Terraform/Terragrunt.
● Kubernetes operational expertise (cluster upgrades, Helm, service mesh,
troubleshooting).
● Strong Linux system administration and scripting (Bash, Python, Perl).
● Experience with monitoring/observability tools (Datadog, ELK, Prometheus/Grafana).
● Proven history of on-call ownership and incident leadership.
● CI/CD implementation experience and release management across multiple
environments.