About the Role
At Bajaj Broking, high availability isn't just an internal KPI—it’s the backbone of every trade, order execution, and portfolio update across millions of retail investors. During high-volatility market opens, our systems process massive spikes in concurrent transactions. Zero downtime is our baseline.
We are looking for a Lead Site Reliability Engineer (SRE) with 8–10 years of engineering experience to own the resilience, scalability, and performance of our core trading platforms. You will bridge the gap between software engineering and systems operations—building self-healing infrastructure, automating away operational toil, and maintaining ultra-low latency infrastructure.
What You’ll Do
- Architect for Scale: Design, deploy, and manage self-healing, high-concurrency cloud infrastructure and multi-tenant Kubernetes clusters.
- Eliminate Toil: Drive an "Automation-First" culture using Terraform, Ansible, and Python/Go to automate infrastructure provisioning, auto-scaling, and failovers.
- Master Observability: Build full-stack observability pipelines (Prometheus, Grafana, Datadog, ELK) to capture high-cardinality metrics, tracing, and log aggregation before issues hit production.
- Define Reliability Standards: Establish and enforce SLIs, SLOs, and Error Budgets across microservices teams to strike the right balance between rapid deployment and platform stability.
- Incident Commander & RCA: Lead high-severity incident responses, conduct blameless Root Cause Analyses (RCAs), and build long-term engineering fixes to eliminate recurring failure modes.
- CI/CD Optimization: Optimize continuous deployment pipelines (GitHub Actions, GitLab CI, ArgoCD) for zero-downtime, continuous release cycles.
- Disaster Recovery (DR) & Chaos Engineering: Conduct chaos testing and design multi-region disaster recovery protocols to guarantee business continuity under any scenario.
What We’re Looking For
- 8–10 years of hands-on experience in SRE, Platform Engineering, or Cloud Infrastructure handling mission-critical, large-scale systems.
- Containerization & Orchestration: Deep expertise with Kubernetes (EKS/GKE/Self-hosted) and Docker in production.
- Infrastructure as Code (IaC): Advanced hands-on mastery of Terraform and modular automation tools.
- Cloud Mastery: Strong expertise across AWS, Azure, or GCP core services, networking (VPCs, BGP, DNS, Load Balancers), and IAM security models.
- Coding & Scripting: Strong programming ability in Python, Go, or Bash to build custom operators, internal tools, and automation scripts.
- Observability Expert: Practical experience setting up distributed tracing, APM, and automated alerting frameworks.
Nice to Have
- Fintech Experience: Prior exposure to high-frequency trading platforms, broking systems, payment gateways, or banking backends.
- Service Mesh: Exposure to Istio or Linkerd.
- Certifications: CKA/CKAD, AWS Solutions Architect Professional, or GCP Cloud Engineer.
Equal Opportunity Employer
At Bajaj Broking, we build teams solely on technical merit, problem-solving ability, and shared commitment to technical excellence.