Lead Site Reliability Engineer (SRE)
Experience: 7+ Years
Location: Delhi NCR | Hybrid
Employment Type: Full-time
About the Role
We are looking for a hands-on Lead SRE who can own production reliability end-to-end, beyond just managing DevOps tools and infrastructure.
The ideal candidate should have strong experience with AWS, Kubernetes/EKS, CI/CD, Infrastructure as Code, observability, incident management, scalability, and distributed systems in high-traffic or high-transaction production environments.
This is a technical leadership role where you will remain hands-on while mentoring the SRE/DevOps team and driving reliability practices across engineering.
Key Responsibilities
- Lead and mentor the DevOps/SRE team and establish engineering best practices.
- Own platform reliability, availability, scalability, and operational excellence.
- Design and operate highly available cloud infrastructure.
- Build and manage production-grade Kubernetes/EKS environments.
- Drive Infrastructure as Code using Terraform or similar tools.
- Establish observability across monitoring, logging, tracing, and alerting.
- Define and monitor SLIs, SLOs, and error budgets for critical services.
- Lead production incidents, RCA, and postmortems.
- Improve resilience through automation, capacity planning, disaster recovery, and performance engineering.
- Partner with engineering teams to improve application reliability and operational readiness.
- Drive cloud cost optimisation without compromising reliability.
- Ensure infrastructure security, compliance, and operational governance.
- Champion DevSecOps practices.
- Standardise deployment, release management, and infrastructure governance.
- Evaluate and adopt modern cloud-native and platform engineering technologies.
Required Skills & Experience
- 7+ years of relevant engineering experience with strong production infrastructure ownership.
- Strong hands-on experience with AWS and Kubernetes/EKS.
- Experience with Terraform or similar IaC tools.
- Strong understanding of CI/CD and deployment automation.
- Experience with observability tools such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or New Relic.
- Strong scripting/programming skills in Python, Bash, or Go.
- Experience managing production incidents and implementing SRE practices.
- Understanding of cloud security, IAM, secrets management, and infrastructure security.
- Understanding of database reliability, backups, replication, and disaster recovery.
- Strong troubleshooting and problem-solving skills.
- Experience working with high-traffic or high-transaction production systems.
Good to Have
- Experience building or scaling an SRE function.
- Service mesh experience with Istio/Linkerd.
- Exposure to Kafka, SQS, Redis, Elasticsearch/OpenSearch and distributed messaging systems.
- Experience with Platform Engineering / Internal Developer Platforms (IDPs).
- Knowledge of FinOps and cloud cost optimisation.
- AWS, Kubernetes, or Terraform certifications.
- Experience in exchange, brokerage, FinTech, gaming, or other low-latency/high-availability environments.
Leadership Expectations
- Remain hands-on and lead by example.
- Mentor SRE/DevOps and engineering teams.
- Drive an automation-first approach.
- Work closely with Engineering, Security, QA, and Product teams.
- Build a culture of ownership, reliability, continuous improvement, and operational excellence.