DevOps Engineer II – AWS & AI
Pune, Fulltime, Hybrid
Key Responsibilities • Own and maintain production AWS infrastructure with high availability, fault tolerance, and proactive monitoring
• Build and manage scalable cloud infrastructure using Terraform and Ansible on AWS
• Deploy and manage AI/LLM workloads, vector databases, and model inference pipelines
• Design and maintain secure CI/CD pipelines for microservices and AI/ML systems
• Implement security best practices including IAM policies, secrets management, and vulnerability scanning
• Set up observability and monitoring for both system and AI metrics (latency, throughput, cost, usage)
• Optimize cloud costs, particularly for GPU-backed and compute-intensive AI workloads
• Manage VPC networking, DNS, load balancing, and secure connectivity across AWS environments
• Collaborate closely with AI/ML engineers and product teams to ensure infrastructure supports fast experimentation and reliable deployment
Required Qualifications • 5-7 years of hands-on experience in DevOps, SRE, or cloud infrastructure roles
• Proven experience working at a product-based company with fast-paced engineering teams
• Strong hands-on expertise with core AWS services: EKS, ECS/Fargate, EC2, S3, RDS, SageMaker, Lambda, VPC, and CloudFront
• Experience deploying and managing AI/LLM workloads, vector databases (e.g., Pinecone, Weaviate, pgvector), and LLM APIs (OpenAI, Bedrock, etc.)
• Solid expertise in Docker, Kubernetes (EKS), Terraform, and Ansible
• Proficiency in Python scripting for automation, tooling, and infrastructure workflows — this is a mandatory requirement
• Strong understanding of Linux/Unix systems, preferably Ubuntu
• Sound knowledge of networking fundamentals, security architecture, and IAM principles on AWS • Demonstrated experience delivering AI projects in production — not just experimental or PoC environments
Preferred Skills
• Experience with LLMOps practices — prompt versioning, model monitoring, RAG pipelines, and inference optimization
• Familiarity with AI/ML frameworks (SageMaker Pipelines, MLflow, Kubeflow) and model serving patterns
• Hands-on experience with observability tools such as Prometheus, Grafana, CloudWatch, or OpenTelemetry
• Automation-first mindset with the ability to work across complex distributed systems
• Exposure to AWS Bedrock, Rekognition, Comprehend, or other managed AI/ML services
• Strong debugging, performance tuning, and root cause analysis skills