Position Overview
We are looking for an experienced Technical Lead – DevOps to lead and drive DevOps engineering, automation, CI/CD, cloud infrastructure, Kubernetes, and platform engineering initiatives. The candidate should have strong hands-on experience across the complete DevOps lifecycle and be capable of designing, implementing, troubleshooting, and optimizing enterprise-grade DevOps platforms.
The role requires advanced expertise in AWS Cloud, Kubernetes, CI/CD, Infrastructure as Code, configuration management, automation, monitoring, security, and release management, AWS expertise along with the ability to provide technical leadership to the DevOps team.
Key Responsibilities
- DevOps & CI/CD:
- Own and drive the complete DevOps lifecycle, from source-code management through build, testing, security scanning, deployment, monitoring, and release.
- Design, implement, and maintain enterprise-grade CI/CD pipelines.
- Develop reusable pipeline frameworks and automation for multiple applications and environments.
- Implement CI/CD best practices including automated testing, quality gates, artifact management, approvals, rollback, and release strategies.
- Integrate tools such as Jenkins, GitLab, Git, SonarQube, Fortify, Trivy and artifact repositories.
- Troubleshoot complex pipeline, build, deployment, and release issues.
- Establish DevOps standards, processes, branching strategies, and deployment methodologies.
- AWS Cloud – Advanced:
Strong hands-on and architectural knowledge of AWS services is required, including:
- VPC – subnets, route tables, NAT Gateway, Internet Gateway, VPC peering, Transit Gateway, security groups, NACLs and VPC endpoints.
- EKS – cluster architecture, node groups, networking, IAM/RBAC, autoscaling, ingress, upgrades, troubleshooting and production operations.
- EC2 – instance sizing, AMIs, autoscaling, lifecycle management and troubleshooting.
- IAM – roles, policies, permissions, least-privilege access and service accounts.
- CloudWatch – metrics, logs, dashboards, alarms, monitoring and alerting.
- ELB/ALB/NLB – load balancing, target groups, health checks and integrations with Kubernetes.
- S3 – storage, lifecycle policies, encryption and access control.
- Route 53 – DNS management, routing policies and health checks.
- EBS/EFS – storage provisioning, performance and lifecycle management.
- CloudFormation/Terraform – infrastructure provisioning and automation.
- AWS security, networking, monitoring, cost optimization and high-availability architecture.
- Kubernetes & Containerization:
- Design, deploy, and manage production-grade Kubernetes/EKS clusters.
- Strong knowledge of Kubernetes architecture, networking, storage, RBAC, scheduling, autoscaling and security.
- Hands-on experience with:
- Deployments and StatefulSets
- Services and Ingress
- ConfigMaps and Secrets
- Persistent Volumes/PVCs
- HPA/VPA
- PodDisruptionBudgets
- Affinity/anti-affinity
- Taints and tolerations
- Network policies
- Helm
- Containerization using Docker/OCI standards.
- Troubleshoot application, networking, resource, storage, and cluster-level issues.
- Plan and execute Kubernetes version upgrades with minimal production impact.
- Ansible & Automation:
- Advanced hands-on experience with Ansible.
- Develop reusable Ansible roles, playbooks, inventories, templates, and automation frameworks.
- Infrastructure as Code:
- Strong experience with Terraform and/or CloudFormation.
- Design reusable infrastructure modules.
- Manage infrastructure lifecycle through version-controlled IaC.
- Implement remote state management, state locking, workspaces, variables, modules, and environment-specific configurations.
- Establish Infrastructure-as-Code standards across development, QA, staging, and production environments.
- Monitoring & Observability:
- Design and implement enterprise monitoring and observability solutions.
- Strong knowledge of:
- AWS CloudWatch
- Prometheus
- Grafana
- ELK/EFK
- Fluent Bit/Fluentd
- Define application and infrastructure KPIs, dashboards, alerts, and operational thresholds.
- Implement proactive monitoring and automated alerting.
- Analyze logs, metrics, and performance trends to identify bottlenecks and production issues.
- DevSecOps & Security:
- Integrate security into the CI/CD lifecycle.
- Implement SAST, DAST, container scanning, dependency scanning, and vulnerability management.
- Experience with tools such as Trivy, SonarQube, Fortify, Snyk, or equivalent.
- Implement secure secrets management and IAM best practices.
- Ensure container, Kubernetes, AWS, and pipeline security best practices.
- Promote shift-left security practices across development and DevOps teams.
- Production Operations & Troubleshooting:
- Provide technical leadership for critical production incidents.
- Perform root-cause analysis and implement permanent corrective actions.
- Troubleshoot complex issues across: AWS, kubernetes, networking
- Participate in incident management, problem management, capacity planning, and disaster recovery activities.
- Drive automation to eliminate repetitive operational activities.
Mandatory Technical Skills
| Area | Required Expertise |
| ------------------------ | -------------------------------------------------------------------- |
| DevOps | Advanced |
| CI/CD | Jenkins, GitLab CI/CD |
| Cloud | AWS – Advanced |
| AWS Networking | VPC, Subnets, Route Tables, NAT, IGW, TGW, Security Groups, NACL |
| Kubernetes | Advanced / EKS |
| Configuration Management | Ansible – Advanced |
| IaC | Terraform |
| Containers | Docker / Kubernetes |
| Monitoring | CloudWatch, Prometheus, Grafana |
| SCM | Git / GitLab |
| Security | Trivy, SonarQube, Fortify/Snyk |
| OS | Linux / RHEL |
| Scripting | Shell/Bash, Python |
| Packaging | Helm |
| Networking | TCP/IP, DNS, HTTP/HTTPS, Load Balancing |
| Observability | Logs, Metrics, Alerts, Dashboards |
| Miscelleneous | kafka, hashicorp, secretmanager, vault, clickhouse, mongoDB |
Required Experience
- 5+ years of overall IT experience, with significant experience in DevOps/Cloud/Platform Engineering.
- Strong production experience with AWS and Kubernetes/EKS.
- Strong hands-on experience with Jenkins and Ansible.
- Experience designing and supporting enterprise-scale CI/CD platforms.
- Experience working with distributed systems and production environments.
- Experience leading technical teams and mentoring engineers.
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder-management skills.
Key Competencies
- Strong technical ownership and accountability
- Architecture and solution-design capability
- Automation-first mindset
- Strong troubleshooting and problem-solving skills
- Production-focused engineering approach
- Security-conscious engineering practices
- Ability to work across Development, QA, Security, Architecture, and Operations
- Ability to mentor and technically guide engineers
- Strong documentation and communication skills
Preferred Certifications (Good To Have)
- AWS Certified Solutions Architect – Associate/Professional
- AWS Certified DevOps Engineer – Professional
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)