Job Description
We are hiring a Senior Site Reliability / Platform Engineer to own reliability, performance, and scalability of large-scale, cloud-native ecommerce platforms.
Key Responsibilities
- Own end‑to‑end reliability, availability, scalability & performance of production systems.
- Define & govern SLOs, SLIs, error budgets.
- Lead 24/7 on-call, incident response, RCA & preventive actions.
- Implement automation, self‑healing, resilient architecture patterns.
- Architect and deliver secure, scalable cloud-native platforms.
- Capacity planning & performance forecasting.
- Drive architecture reviews, chaos engineering & resilience improvements.
- Maintain runbooks, IaC diagrams, SOPs & incident playbooks.
- Lead best practices in CI/CD, release management & deployment automation.
- Ensure strong cloud security posture & compliance.
- Mentor engineers & promote operational excellence.
Tech Stack
- Cloud: Azure
- Containers: Kubernetes
- Messaging/Data: Kafka, Aerospike, MongoDB Atlas, In‑memory DBs
- CI/CD: Jenkins, Azure DevOps, GitHub Actions, Argo CD
- IaC: Terraform, ARM Templates, Ansible, Packer
- Monitoring: New Relic, Prometheus, Grafana, Azure Monitor
- Security: WAF, DDoS, Azure Front Door
- Languages/OS: Linux, Python
- Networking: DNS, NAT, Routing, Subnetting
Location & Work Mode
📍 Mohali (Onsite)
🗓 5 Days Working
🌍 UAE-Based Organization