Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.
Role: Site Reliability Engineer / DevOps Engineer
Positions: 1
Experience: 3 to 8 years
Location(s): Gurgaon
Type: On-site / Permanent
Salary: Up to INR 20 LPA (Based on fitment)
Notice Period: 15 days
About the Role
We are looking for a hands-on Site Reliability Engineer (SRE) / DevOps Engineer with strong programming and automation skills.
The role will initially involve development and automation work, helping the engineer build a strong understanding of the applications and platform. Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, observability, production reliability, incident management, and operational automation.
The ideal candidate should be comfortable working with both application code and production systems and should use engineering and automation to improve reliability and reduce manual effort.
Key Responsibilities
- Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.
- Write clean, maintainable, testable, and production-ready code.
- Participate in code reviews, debugging, testing, and technical discussions.
- Build, maintain, and improve CI/CD pipelines and automated deployment processes.
- Work with Docker and Kubernetes for application deployment and operations.
- Support on prem and cloud-based application and infrastructure deployments.
- Maintain reliable, scalable, secure, and highly available production environments.
- Implement and manage monitoring, logging, alerting, and observability solutions.
- Contribute to defining and tracking SLIs, SLOs, and error budgets.
- Troubleshoot application and production issues and perform Root Cause Analysis (RCA).
- Identify recurring operational problems and address them through automation and engineering improvements.
- Support incident response, change management, deployment governance, and disaster recovery practices.
- Maintain runbooks, SOPs, incident documentation, and technical documentation.
- Collaborate with Engineering, Product, Platform, Security, Operations, and external teams.
Technical Skills
- Strong hands-on experience with Python for development and automation.
- Experience developing scripts, APIs, integrations, utilities, or backend services.
- Good understanding of software engineering principles, debugging, logging, testing, and exception handling.
- Experience with REST APIs, JSON, Git, pull requests, and code reviews.
- Strong knowledge of Linux/Unix environments and basic Windows administration.
- Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.
- Experience with at least one cloud platform: AWS, Azure, or GCP.
- Hands-on experience with Docker and Kubernetes.
- Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.
- Familiarity with Infrastructure-as-Code tools such as Terraform is preferred.
- Experience with monitoring and observability tools such as Grafana, Prometheus, Power BI, or equivalent.
- Ability to analyze logs, metrics, alerts, and traces for troubleshooting.
- Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.
- Experience with JIRA, ServiceNow, and Confluence is desirable.
Preferred Experience
- 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, or related roles.
- Strong Python development or automation experience.
- Experience supporting applications across development, deployment, and production environments.
- Exposure to cloud-native, distributed, or production-grade systems.
- Understanding of security and compliance best practices.
- Familiarity with AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools.
Soft Skills
- Strong analytical and troubleshooting skills.
- Engineering and automation mindset.
- Good written and verbal communication skills.
- Effective cross-functional collaboration.
- Ownership-driven approach to problem solving.
- Ability to remain structured during production incidents.