About the Role
We are seeking a seasoned Lead / Senior DevOps Engineer (5–8 years of experience) to lead our cloud infrastructure, automation, and continuous delivery initiatives. Evolving from traditional Linux hosting and cloud support into modern DevOps practices, you will be responsible for designing and managing resilient, scalable, and secure cloud environments on AWS, establishing Infrastructure as Code (IaC), driving container orchestration, and building robust CI/CD pipelines.
In this role, you will bridge the gap between development and operations, championing Site Reliability Engineering (SRE) principles, optimizing cloud architecture and spending, mentoring team members, and acting as the final escalation authority for mission-critical infrastructure incidents.
Key Responsibilities
- Cloud Architecture & Reliability: Architect, deploy, and maintain highly available, fault-tolerant, and secure multi-region cloud infrastructure on AWS. Enforce SRE best practices, defining and tracking SLOs, SLIs, and error budgets.
- Infrastructure as Code (IaC): Standardize and automate environment provisioning using Terraform and Ansible, eliminating manual configuration drift across development, staging, and production environments.
- Containerization & Orchestration: Manage production Docker workloads and lead the deployment, scaling, and operational management of containerized services using Amazon ECS or Kubernetes (EKS).
- CI/CD Pipeline Engineering: Design, build, and optimize automated CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, or AWS CodePipeline) to achieve zero-downtime blue/green or canary deployments.
- Stack Optimization & High Availability: Oversee and tune high-traffic web architectures (Nginx/Apache, PHP-FPM, Node.js, Python), distributed caching (Redis/Memcached), and relational/NoSQL databases (MySQL/MariaDB, PostgreSQL, AWS RDS/Aurora).
- Enterprise Observability & APM: Design and implement unified observability stacks using Prometheus, Grafana, AWS CloudWatch, and distributed logging/tracing tools (ELK/EFK, OpenTelemetry) to proactively detect regressions and system bottlenecks.
- DevSecOps & Compliance: Integrate security scans (SAST/DAST, container vulnerability scanning) into CI/CD pipelines. Implement OS and cloud hardening (CIS benchmarks, IAM least-privilege, KMS, WAF, VPC peering/transit gateways).
- Cloud FinOps & Capacity Planning: Continuously audit cloud infrastructure utilization, analyze cost drivers, and implement architectural rightsizing, auto-scaling, and purchasing strategies (Reserved Instances, Savings Plans).
- Technical Leadership & Incident Escalation: Serve as the tier-3/tier-4 technical escalation lead for high-severity platform outages, conduct blameless Root Cause Analyses (RCAs), and mentor mid-level/junior engineers.
Candidate Requirements
- 5–8 years of professional experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure Engineering, with strong roots in Linux system administration.
- Demonstrated hands-on experience designing and managing production cloud infrastructure on AWS at scale.
- Proficiency in writing production-ready Terraform configurations and modular Ansible playbooks.
- Practical production experience managing containerized applications with Docker and orchestrators like Kubernetes (EKS) or Amazon ECS.
- Proven track record of building and managing end-to-end CI/CD pipelines that support automated testing, building, and deployment.
- Strong debugging capabilities across the entire application stack: network latency, OS-level resource starvation, application runtime, and database bottlenecks.
- Excellent documentation, communication, and client/stakeholder management skills, with proven ability to lead technical discussions and RCA reviews.
- Experience in leading technical teams or mentoring junior/mid-level infrastructure engineers.
- Preferred Certifications (Bonus): AWS Certified Solutions Architect (Professional), AWS Certified DevOps Engineer (Professional), or Certified Kubernetes Administrator (CKA).