Sr. DevOps Engineer
Observability Strategy (Web & AI Applications)
- Design and implement an end-to-end observability & alert stack covering both traditional web applications and AI/ML services.
- Develop and maintain tooling such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalent
- Build AI-specific observability: model latency/throughput tracking, GPU utilization, token usage, drift detection, and inference quality signals.
Infrastructure Scaling & Reliability:
- Design and manage infrastructure capable of scaling across on-premise data centers and public cloud
- Own capacity planning, load testing, auto-scaling, and cost-optimization initiatives across compute, storage, and networking.
- Implement Infrastructure as Code (Terraform, Ansible, or equivalent) to ensure environments are reproducible, version-controlled, and auditable.
- Lead disaster recovery, backup, and high-availability strategy for critical systems.
DevOps & CI/CD Delivery :
- Partner closely with Solution Architects to translate project and system designs into concrete DevOps execution plans.
- Design, build, and maintain CI/CD pipelines (Jenkins, GitHub Actions or
- ArgoCD/Flux for GitOps) across multiple projects and teams.
- Containerize and orchestrate applications using Docker and Kubernetes, including Helm chart and manifest management.
- Embed security and compliance checks (SAST/DAST, secrets scanning, image scanning) directly into the delivery pipeline (DevSecOps).
GPU Infrastructure & Model Deployment
- Provision, configure, and manage GPU infrastructure (on-prem clusters and cloud GPU instances) for model
- Deploy, scale, and monitor ML/LLM models in production using tools such as Triton Inference Server, vLLM,
- Optimize GPU utilization, cost, and throughput across multi-tenant workloads; manage CUDA/driver/toolkit
- Collaborate with data science/ML engineering teams on MLOps pipelines — model versioning, experiment
- 5–10 years of hands-on DevOps/SRE/Infrastructure engineering experience, including at least 2–3 years in a senior or lead capacity.
- Deep expertise in Git-based workflows and repository management at scale (GitHub/GitLab/Bitbucket).
- Proven experience designing observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, New Relic, OpenTelemetry).
- Strong background in cloud platforms (AWS, Azure, and/or GCP) and on-premise/hybrid infrastructure.
- Expert-level skills with Infrastructure as Code (Terraform, Ansible, CloudFormation, or Pulumi).
- Strong Kubernetes and Docker experience, including multi-cluster and multi-environment management.
- Hands-on experience building and maintaining CI/CD pipelines end to end.
- Working knowledge of GPU infrastructure (NVIDIA CUDA, drivers, NCCL) and experience deploying ML/AI models to production.
- Proficiency in scripting/automation languages: Python, Bash, and/or Go.
- Solid understanding of networking, load balancing, DNS, and security fundamentals in distributed systems.
- Experience partnering with architects and engineering leads to translate designs into infrastructure and delivery plans.
- Hands-on with SAST, DAST, and SCA tooling (e.g., SonarQube, Snyk, Checkmarx, OWASP DependencyCheck) integrated directly into CI/CD pipelines.
- Container and image security: vulnerability scanning (Trivy, Grype, Clair), minimal/hardened base images, and signed/verified image provenance (Cosign/Sigstore).
- Familiarity with Secrets management and credential hygiene using tools such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault
- Excellent communication skills and comfort operating cross-functionally with development, data science, and product teams.