We are looking for a highly hands-on Mid–Senior DevOps Engineer with 3+ years of experience to take strong ownership of a growing DevOps ecosystem end-to-end. This is systems-heavy engineering, not standard infrastructure maintenance — the role covers architecting cloud-native distributed systems, scaling high-throughput analytics pipelines, building enterprise-grade SaaS and on-prem deployments, and driving DevOps excellence across the company. If you enjoy owning infrastructure like a product, this role will suit you well.
Key Responsibilities
- Architect and manage scalable infrastructure using AWS services; design and maintain a multi-tenant SaaS architecture handling millions of real-time events
- Own and enhance CI/CD pipelines using GitHub Actions; build automated deployments for microservices, backend systems, UI applications, and browser extensions
- Implement Infrastructure-as-Code using Terraform; build and maintain Dockerized, container-orchestrated environments
- Implement end-to-end observability (logs, metrics, traces), alerting, and incident-response systems; ensure high availability, auto-scaling, disaster recovery, and backup strategies
- Drive security best practices — network isolation, IAM policies, CVE remediation, container security
- Design and deliver on-premise versions of the SaaS platform, replicating cloud architecture in constrained enterprise environments, and automating installation, upgrades, and maintenance
- Collaborate closely with Engineering, QA, and Product teams; contribute to architecture and design decisions; act as a technical owner for DevOps initiatives
- Use AI coding tools (Claude Code, Cursor, GitHub Copilot, or similar) as a primary part of daily infrastructure and automation work
- Write and maintain specs/rules files (e.g. CLAUDE.md, .cursorrules, agent skill files) that direct AI coding tools for the DevOps and engineering teams
- Design and build AI agents and agentic workflows for infra automation, incident response, and operational tooling using LLMs (OpenAI, Claude, Gemini, or similar)
- Practice spec-driven development — translate infrastructure and platform requirements into clear specs that guide both human and AI-driven implementation
Mandatory Required Skills
- 3+ years in DevOps with hands-on ownership of production systems
- Strong AWS expertise (mandatory), including Kubernetes (EKS), VPC, EC2, RDS, ElastiCache, S3, CloudFront, Route 53, API Gateway, Lambda, and Kinesis
- Experience with ClickHouse or similar columnar databases, Docker and container orchestration, and Helm charts
- Strong GitHub Actions experience and Terraform expertise (mandatory)
- Strong scripting skills in Python, Shell, and Node.js
- Experience with high-scale distributed systems — load balancing, caching strategies, event-driven architectures, high-throughput pipelines
- Experience building SaaS-to-on-prem deployments and automating complex enterprise setups
- Hands-on experience with AI coding tools (Claude Code, Cursor, GitHub Copilot, or similar) as a core part of daily engineering workflow
Preferred
- Experience with Digital Adoption Platforms or browser-based products
- Experience with monitoring tools such as Grafana, Prometheus, or Datadog
- Experience designing and building AI agents / agentic workflows using LLMs (OpenAI, Claude, Gemini, or similar) for automation, tooling, or operational use cases
- Familiarity with spec-driven development practices and maintaining AI tool configuration/rules files (e.g. CLAUDE.md, .cursorrules, agent skill files) for a team
Nice to Have
- Exposure to SOC 2 or ISO compliance environments
- Development experience beyond scripting (full application development background)
Qualifications
Degree in Computer Science, Engineering, or a related field preferred (to be confirmed with client)
Experience Required
3+ years in DevOps with demonstrated, hands-on ownership of production infrastructure — not solely pipeline maintenance.