Job Description: Senior DevOps & Platform Architect (MLOps)
About The Role
We are seeking a
Senior DevOps & Platform Architect with 17+ years of experience designing, building, and scaling enterprise-grade infrastructure and platform engineering solutions. The ideal candidate brings deep expertise across
DevOps, Cloud Architecture, Site Reliability Engineering, and MLOps, and can architect platforms that support both traditional software delivery and machine learning workflows at scale.
This is a senior technical leadership role — you will define platform strategy, set architectural standards, and guide engineering teams in building resilient, secure, and highly automated infrastructure for both application and ML/AI workloads.
Key Responsibilities
Platform & DevOps Architecture
- Architect and lead the design of scalable, secure, and highly available cloud infrastructure (AWS/Azure/GCP).
- Define and drive DevOps strategy, including CI/CD pipelines, Infrastructure as Code (IaC), and automation frameworks across the organization.
- Own platform architecture decisions — containerization, orchestration, service mesh, networking, and observability.
- Establish and enforce standards for security, compliance, cost optimization, and disaster recovery.
- Drive infrastructure modernization initiatives (monolith-to-microservices, on-prem to cloud, cloud-to-multi-cloud).
- Lead capacity planning, performance tuning, and reliability engineering (SRE practices — SLIs/SLOs/error budgets).
MLOps & ML Platform Engineering
- Design and build end-to-end MLOps pipelines — model training, versioning, validation, deployment, and monitoring.
- Architect scalable ML infrastructure supporting experiment tracking, feature stores, model registries, and automated retraining pipelines.
- Enable CI/CD for ML (CI/CD/CT) — integrating model testing, validation, and rollback strategies into deployment workflows.
- Collaborate with Data Science and ML Engineering teams to productionize models efficiently, ensuring reproducibility and scalability.
- Implement monitoring/observability for deployed models (data drift, model drift, performance degradation).
- Optimize GPU/compute resource utilization for training and inference workloads (including on Kubernetes).
Leadership & Collaboration
- Act as a technical leader and mentor for DevOps, SRE, and Platform engineering teams.
- Partner with engineering, data science, security, and product leadership to align platform strategy with business goals.
- Drive architectural reviews, design documents, and technical roadmaps.
- Evaluate and introduce new tools/technologies to improve platform maturity and developer experience.
- Own incident management processes and post-mortem culture for platform/production issues.
Required Skills & Qualifications
- 14+ years of experience in DevOps, Site Reliability Engineering, Cloud/Platform Architecture, or related roles.
- Proven experience architecting large-scale, production-grade infrastructure on AWS, Azure, or GCP (multi-cloud experience a plus).
- Strong hands-on expertise with Kubernetes and container orchestration at scale.
- Deep experience with Infrastructure as Code (Terraform, CloudFormation, Pulumi).
- Strong background in CI/CD tooling (Jenkins, GitLab CI, GitHub Actions, ArgoCD, Spinnaker).
- Hands-on MLOps experience with tools such as MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, or similar.
- Experience with model serving frameworks (KServe, Seldon, TorchServe, Triton Inference Server) and feature stores (Feast or similar).
- Proficiency in scripting/automation (Python, Bash, Go).
- Strong knowledge of observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, New Relic).
- Experience with service mesh and API gateway technologies (Istio, Envoy, Kong).
- Solid understanding of networking, security best practices, and compliance frameworks (SOC2, HIPAA, ISO 27001 as applicable).
- Experience with GPU infrastructure and distributed training frameworks is a strong plus.
- Excellent communication skills with a track record of technical leadership and cross-functional collaboration.
Good To Have (Preferred)
- Certifications: AWS/Azure/GCP Solutions Architect (Professional level), CKA/CKAD, or equivalent.
- Experience with data pipeline orchestration (Airflow, Dagster, Prefect).
- Familiarity with LLMOps / Generative AI deployment patterns (vector databases, RAG pipelines, model gateways).
- Experience building internal developer platforms (IDPs) or self-service infrastructure tooling.
- Background in FinOps / cloud cost optimization at scale.