Role Overview and Responsibilities
Dhruva Space is seeking experienced Senior DevOps Engineer to own infrastructure, CI/CD, deployment, monitoring, and reliability across the organization’s product portfolio. The role will also oversee baseline database operations, security, and cloud cost management in a lean engineering environment, while mentoring the DevOps team and establishing scalable platform practices.
Key responsibilities include but are not limited to:
- Design, implement, and manage CI/CD pipelines and cloud infrastructure across products.
- Own platform reliability, uptime, production deployments, and incident response.
- Manage baseline database operations, backups, disaster recovery, and performance monitoring.
- Establish and maintain production monitoring, alerting, and observability.
- Manage cloud infrastructure costs, access controls, secrets, and security posture.
- Implement infrastructure-as-code and standardized deployment practices.
- Manage containerized workloads using Kubernetes or similar platforms.
- Develop and maintain operational runbooks, documentation, and recovery procedures.
- Mentor and manage associate DevOps / platform engineers.
- Collaborate with engineering, security, and architecture teams on infrastructure and deployment requirements.
- Build scalable platform practices that support future DBA, NOC, and specialist teams.
- Support reliable infrastructure for AI/ML and agent-based workloads where required.
Candidate Requirements:
- Bachelor’s or Master’s degree in Computer Science, Engineering, IT, or a related field.
- 8–12 years of experience in DevOps, Platform Engineering, Cloud Infrastructure, or related roles.
- Strong hands-on experience with CI/CD tools and Infrastructure-as-Code, particularly Terraform or equivalent.
- Strong experience with Kubernetes/container orchestration and AWS, GCP, or Azure.
- Practical knowledge of database operations, backup/DR, and performance fundamentals.
- Strong understanding of cloud security, access control, secrets management, and compliance practices.
- Proven experience handling production incidents and reliability challenges.
- Ability to work as a cross-functional infrastructure generalist in a lean engineering environment.
- Strong leadership, mentoring, troubleshooting, and communication skills.
- Experience supporting AI/ML or agent-runtime workloads is preferred.