For this role you must be able to work on a hybrid basis with at least 2 days per week working in the Mumbai office in Andheri.
Position
Platform Infrastructure and SRE owns the reliability, scalability, and operability of Deltatre's infrastructure and internal platform tooling - the team other engineering teams depend on to ship and run their services safely. The stack spans AWS, GCP, Kubernetes, Terraform for infrastructure as code, a full observability stack (Datadog, Prometheus, Grafana), and Deltatre's CI/CD platforms.
As Team Lead, you will both write code and lead people. This is a hybrid role: roughly half hands-on engineering work (design, code, review, and on-call participation as appropriate) and half people leadership (coaching, career growth, hiring, and connecting team output to broader engineering priorities). You will own the technical direction of a 5 to 8 person team while staying close enough to the work to make good calls under pressure.
Responsibilities
Technical direction
- Set technical direction for platform infrastructure and SRE initiatives, balancing reliability, cost, and delivery speed
- Stay hands-on: design reviews, architecture decisions, and direct contribution to high-leverage or high-risk work
- Own capacity planning, cost visibility, and infrastructure spend across our cloud providers
- Define and track operational health metrics (SLOs, error budgets, MTTR) for owned services
Reliability and incident response
- Drive incident response process and post-incident reviews, and use them to prioritise reliability investments
Team leadership
- Manage and grow a team of 5 to 8 engineers: 1:1s, career development, performance feedback, and hiring
- Partner with product and other engineering teams to translate infrastructure needs into a prioritised roadmap
- Represent the team in cross-functional planning and communicate trade-offs to stakeholders and leadership
Requirements
You have done this before - here's what we're looking for.
Required
- 6+ years' experience in infrastructure, platform, or SRE engineering, including production ownership at scale
- 1+ years' experience formally leading a team (people management or a strong tech lead track record), or clear readiness to step into that
- Deep hands-on experience with Kubernetes and at least one major public cloud (AWS, GCP, or Azure)
- Working knowledge of a second cloud provider, given our multi-cloud footprint
- Strong experience with infrastructure as code, ideally Terraform
- Experience building or operating observability stacks (metrics, logging, tracing)
- Experience owning CI/CD platforms and pipelines, not just using them
- Track record of leading incident response and driving reliability improvements from postmortems
- Clear written and verbal communication; comfortable presenting trade-offs to non-infrastructure stakeholders
Valued
- Datadog, Prometheus, and Grafana experience specifically
- Cost optimisation experience across multiple cloud providers
- Prior experience in a regulated or high-availability environment
Who Thrives Here
This is a senior-level role suited to someone who has already led infrastructure or SRE work at scale and wants to grow into leading people without giving up technical depth. You'll move between architecture decisions and 1:1s in the same afternoon, and both need your full attention.
People who do well here stay close enough to the work to make good calls under pressure, translate incident learnings into concrete reliability investments rather than blame, and are as comfortable presenting trade-offs to leadership as they are reviewing a pull request.
Process and what to expect
- Introductory conversation: a step for us to get to know each other better, and for us to answer all questions you might have around Deltatre.
- Codility assessment: a structured, self-paced technical assessment that becomes the basis for the technical conversation.
- 1-hour Technical round of Interview: we will deep dive on your experience and on the Codility assessment, and explore the technical and architectural decisions you have made in your past work.
- 1-hour Competency Based Interview: we will ask you situational questions around how you deal with real-life scenarios at work. Easier doing it, more than explaining it.
If any of these formats would be difficult for you, tell us - we adjust regularly and can usually accommodate.
Accessibility and accommodation questions are welcome at any stage. Tell us what would work for you.