Design, implement, and maintain scalable, secure, and automated cloud infrastructure across AWS, Azure, and other cloud platforms
- Build, enhance, and optimize CI/CD pipelines to support rapid, reliable, and efficient software delivery
- Automate build, test, deployment, and release processes to improve development productivity and deployment quality
- Develop and maintain Infrastructure as Code (IaC) solutions using Terraform for infrastructure provisioning and management
- Implement DevSecOps practices by integrating security scans, compliance checks, and vulnerability management into CI/CD pipelines
- Manage and optimize containerized environments using Docker and Kubernetes
- Deploy, configure, and maintain Kubernetes clusters, applications, and Helm charts
- Ensure high availability, scalability, reliability, and security of cloud-hosted applications and infrastructure
- Implement and manage monitoring, logging, observability, and alerting solutions for applications and infrastructure
- Configure dashboards, alerts, and reporting mechanisms to proactively identify and resolve production issues
- Automate operational activities, infrastructure management, and repetitive tasks through scripting and tooling
- Monitor infrastructure performance, system utilization, deployment metrics, and service health
- Troubleshoot and resolve complex issues across cloud infrastructure, CI/CD pipelines, container platforms, and application environments
- Collaborate with Development, QA, Security, and Operations teams to streamline delivery processes and promote DevOps best practices
- Provide technical leadership, mentorship, and guidance to engineers on DevOps, cloud, automation, and infrastructure technologies
- Drive continuous improvement initiatives focused on automation, operational excellence, security, and reliability
- Participate in architecture reviews and recommend solutions for scalability, resilience, and operational efficiency
- Maintain technical documentation for infrastructure, deployment pipelines, cloud environments, monitoring solutions, and operational procedures
- Participate in Agile ceremonies such as sprint planning, stand-ups, reviews, and retrospectives
- Provide regular updates on project progress, risks, dependencies, and assigned deliverables
Mandatory Skills
- Strong experience in DevOps engineering and cloud infrastructure management
- Hands-on experience with AWS and Azure cloud platforms
- Strong experience in designing, implementing, and managing CI/CD pipelines
- Experience with Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, or similar CI/CD tools
- Strong expertise in Infrastructure as Code (IaC) using Terraform
- Hands-on experience with Docker containerization technologies
- Strong experience with Kubernetes deployment, administration, and orchestration
- Experience managing Kubernetes applications using Helm Charts
- Experience implementing DevSecOps practices, including SAST, DAST, vulnerability management, and compliance controls
- Strong experience with monitoring and logging tools such as Prometheus, Grafana, ELK Stack, CloudWatch, Azure Monitor, or similar platforms
- Proficiency in scripting and automation using Python, Bash, or PowerShell
- Experience with Git and source code management practices
- Strong understanding of cloud networking concepts including VPC/VNET, Load Balancers, VPNs, DNS, Firewalls, and Security Groups
- Strong Linux administration and troubleshooting skills
- Experience supporting production environments with high availability and reliability requirements
- Strong analytical, troubleshooting, and problem-solving skills
- Excellent communication, collaboration, and stakeholder management skills
Desirable Skills
- AWS Certified DevOps Engineer, Solutions Architect, or equivalent certifications
- Azure Administrator, Azure DevOps Engineer, or related Azure certifications
- Exposure to Google Cloud Platform (GCP)
- Experience with GitOps tools such as ArgoCD or FluxCD
- Familiarity with configuration management tools such as Ansible, Chef, or Puppet
- Experience with Service Mesh technologies such as Istio or Linkerd
- Knowledge of Site Reliability Engineering (SRE) practices
- Experience with Kafka, RabbitMQ, or other messaging platforms
- Familiarity with cloud security frameworks and governance practices
- Experience in cloud cost optimization and FinOps practices
- Experience working in Telecom, eCommerce, Financial Services, or other enterprise environments
- Strong communication and documentation skills