Job Description: Senior DevOps Engineer
Role Overview
We are building a next-generation product engineering organization where we are looking for engineers who will build and operate the engineering platform that powers rapid, reliable, and secure software delivery.
The Senior DevOps Engineer will combine strong cloud engineering, infrastructure automation, platform engineering, security, and Site Reliability Engineering (SRE) practices with the effective use of AI tools and intelligent automation to build modern cloud-native delivery platforms.
This role focuses on creating highly automated developer platforms, enabling continuous integration and continuous delivery (CI/CD), infrastructure as code, observability, security, and operational excellence. We are looking for engineers who can leverage AI as a force multiplier to automate infrastructure management, deployment pipelines, operational diagnostics, and incident response.
Experience level: 8 to 12 years
Key Responsibilities
- Platform Ownership: Own the engineering platform that enables product teams to develop, deploy, monitor, and operate software efficiently.
- SDLC partner: Design, build, and continuously improve cloud-native engineering platforms that support modern software development practices. Build self-service developer capabilities that improve engineering productivity.
- Infrastructure Automation: Design and manage Infrastructure as Code (IaC) using modern automation frameworks. Automate provisioning, configuration, scaling, and lifecycle management of cloud infrastructure.
- Reusable templates: Build reusable infrastructure templates and platform services for engineering teams.
- Continuous Integration & Continuous Delivery: Design, implement, and optimize CI/CD pipelines for rapid and reliable software delivery. Automate build, testing, deployment, rollback, and release management processes. Continuously improve deployment frequency while maintaining system stability and security.
- Cloud Platform Engineering: Build and operate cloud-native platforms on AWS, Microsoft Azure, or Google Cloud Platform. Manage Kubernetes clusters, container platforms, networking, storage, identity, and platform services. Optimize infrastructure for scalability, resilience, performance, and cost efficiency.
Additional optional responsibilities
- AI-Augmented Platform Engineering: Use AI-powered engineering tools to accelerate infrastructure automation, troubleshooting, configuration generation, documentation, and operational workflows. Validate AI-generated infrastructure definitions, deployment scripts, and operational recommendations.
- Site Reliability Engineering: Design highly available and resilient production environments. Define and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets. Lead incident management, root cause analysis, and continuous operational improvement.
Required Skills
- Strong experience with cloud platforms such as Amazon Web Services, Microsoft Azure, or Google Cloud Platform.
- Hands-on experience with Kubernetes, Docker, and container orchestration.
- Strong expertise in Infrastructure as Code using Terraform, CloudFormation, or similar technologies.
- Experience designing and managing CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent platforms.
- Experience with configuration management and automation tools such as Ansible or similar technologies.
- Strong scripting skills using Python, Bash, or PowerShell.
- Experience implementing monitoring and observability solutions such as Prometheus, Grafana, Open Telemetry, ELK, Datadog, Dynatrace, or similar platforms.
- Experience integrating security into software delivery pipelines.
- Experience using AI-powered engineering tools to improve operational efficiency and platform automation.
Desired Skills
- Platform Engineering
- Site Reliability Engineering (SRE)
- DevSecOps
- Service Mesh Technologies
- Kubernetes Operators
- GitOps, FinOps
- Chaos Engineering
- Zero Trust Security
- Cloud Networking
- AI Operations (AIOps)
- MLOps Platform Integration
- Event-Driven Architecture
What Success Looks Like in This Role
Within the first year, you will:
- Build a highly automated cloud-native engineering platform supporting multiple product teams.
- Reduce software deployment time through intelligent CI/CD automation.
- Improve platform reliability, scalability, and operational resilience.
- Establish comprehensive observability and proactive monitoring across engineering platforms.
- Implement secure-by-default DevSecOps practices across the software delivery lifecycle.
- Enable developers to provision environments and deploy applications through self-service platform capabilities.
- Reduce operational overhead by leveraging AI-assisted automation for platform management and incident resolution.
- Help define modern platform engineering practices that accelerate delivery while maintaining enterprise-grade reliability and security.