Job Description: DevOps Engineer, AI Infrastructure
Location: Bengaluru, India
Experience: 3+ years
Employment Type: Full-time
Department: Software Engineering
About the Role
We are looking for a DevOps Engineer with strong exposure to AI infrastructure to build, operate, and scale the systems that power our products.
This role goes beyond traditional CI/CD and cloud operations. You will work on Amazon EKS, GPU infrastructure, AI workloads, scalable databases, vector databases, distributed systems, observability, and production reliability. You will partner closely with application, data, and AI engineering teams to ensure that our infrastructure remains secure, performant, cost-efficient, and highly available as usage grows.
Key Responsibilities
- Design, deploy, and manage production infrastructure on AWS.
- Operate and optimize Amazon EKS clusters for microservices, data workloads, and AI/ML applications.
- Manage Kubernetes workloads, including deployments, services, ingress, autoscaling, storage, networking, and cluster upgrades.
- Build infrastructure and deployment automation using Infrastructure as Code.
- Develop and maintain reliable CI/CD pipelines for application and infrastructure releases.
- Manage GPU-enabled infrastructure for AI inference, model serving, training, and batch workloads.
- Optimize GPU utilization, scheduling, provisioning, capacity planning, and cost.
- Configure and maintain GPU-enabled Kubernetes nodes, including drivers, device plugins, node pools, taints, tolerations, and workload isolation.
- Support scaling of relational and non-relational databases for increasing traffic and data volumes.
- Improve database availability, performance, replication, read scaling, backups, failover, and disaster recovery.
- Design and operate scalable vector database infrastructure for embeddings, semantic search, and retrieval-augmented generation workloads.
- Work with distributed vector database architectures, sharding, replication, partitioning, indexing, and horizontal scaling.
- Evaluate and implement solutions for vector databases such as Milvus, Qdrant, Weaviate, OpenSearch, pgvector, Pinecone, or equivalent systems.
- Build observability across infrastructure and applications using metrics, logs, traces, dashboards, and actionable alerts.
- Establish and improve reliability practices, including SLOs, incident response, root-cause analysis, and operational runbooks.
- Identify infrastructure bottlenecks and improve system performance, availability, and cost efficiency.
- Implement security best practices for IAM, secrets management, network access, container images, and Kubernetes workloads.
- Collaborate with AI and backend engineers to productionize model-serving and data-processing pipelines.
- Automate routine operational tasks and reduce manual intervention across environments.
- Document infrastructure architecture, deployment procedures, troubleshooting guides, and recovery processes.
Required Qualifications
- 3+ years of experience in DevOps, Cloud Infrastructure, Site Reliability Engineering, Platform Engineering, or a related role.
- Strong hands-on experience with AWS services and production cloud environments.
- Practical experience managing Amazon EKS and Kubernetes in production or staging environments.
- Experience with Docker, Kubernetes networking, Helm, autoscaling, persistent volumes, and cluster troubleshooting.
- Experience with CI/CD systems such as GitHub Actions, GitLab CI, Jenkins, Argo CD, or equivalent.
- Proficiency with Infrastructure as Code tools such as Terraform, Pulumi, or CloudFormation.
- Strong Linux administration, networking, security, and troubleshooting skills.
- Experience managing GPU-based workloads or AI infrastructure.
- Understanding of GPU provisioning, utilization monitoring, scheduling, and capacity management.
- Experience scaling databases through replication, partitioning, indexing, caching, connection pooling, or sharding.
- Familiarity with distributed systems concepts such as consistency, replication, partition tolerance, fault tolerance, and horizontal scaling.
- Experience working with at least one vector database or vector search platform.
- Proficiency in scripting or programming using Python, Bash, Go, or a similar language.
- Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, CloudWatch, Datadog, or ELK.
- Strong understanding of production operations, incident management, and root-cause analysis.
Preferred Qualifications
- Experience running AI inference or model-serving workloads on Kubernetes.
- Familiarity with NVIDIA GPUs, CUDA, NVIDIA device plugins, GPU Operator, MIG, or similar technologies.
- Experience with Karpenter, Cluster Autoscaler, or custom Kubernetes autoscaling solutions.
- Experience operating Ray, KServe, Seldon, Triton Inference Server, vLLM, or similar AI-serving systems.
- Experience with distributed vector databases such as Milvus, Qdrant, Weaviate, OpenSearch, or large-scale pgvector deployments.
- Experience with PostgreSQL, Redis, Kafka, Elasticsearch/OpenSearch, or other distributed data systems.
- Knowledge of AWS services such as EC2, EKS, RDS, Aurora, ElastiCache, S3, CloudFront, IAM, VPC, and CloudWatch.
- Experience with GitOps and progressive delivery using Argo CD, Flux, Argo Rollouts, or equivalent.
- Understanding of FinOps and cloud cost optimization, particularly GPU cost management.
- Experience implementing security, compliance, backup, and disaster-recovery controls.
- AWS, Kubernetes, or relevant cloud certifications.
What Success Looks Like
In the first six months, you will be expected to:
- Improve the reliability and operational maturity of our EKS environments.
- Establish efficient deployment and rollback workflows.
- Improve GPU utilization, scheduling, monitoring, and cost visibility.
- Support scalable database and vector database architectures.
- Strengthen observability, alerting, incident response, and production documentation.
- Help AI and engineering teams deploy workloads reliably and efficiently.
- Reduce infrastructure bottlenecks and manual operational work.
Ideal Candidate
You are a hands-on infrastructure engineer who enjoys solving complex production problems. You understand that AI infrastructure requires more than deploying containers: it requires careful management of GPU capacity, workload scheduling, distributed data systems, database performance, reliability, and cost.
You are comfortable debugging issues across the application, Kubernetes, cloud, networking, and data layers. You take ownership, automate wherever possible, and can balance speed of execution with operational excellence.