Job Summary:
Senior DevOps and SRE role focused on building, scaling, and operating a self-service compute platform that enables engineering teams to provision Kubernetes clusters, workloads, virtual machines, and cloud instances on demand across public cloud and on-premises Kubernetes environments. Responsibilities include ownership of the control plane, runtime, platform reliability, uptime, deployment pipelines, monitoring, and AI-assisted operational tooling.
Location : Sunnyvale / Santa Clara / San Jose
Work Model :Hybrid – 3 days onsite
Focus : Kubernetes platform, compute self-service, AI-assisted ops
Experience : 7+yrs DevOps/SRE
Key Responsibilities
● Architect, deploy, and manage a highly available, multi-cluster Kubernetes platform enabling self-service provisioning of Kubernetes clusters, virtual machines, workloads, and cloud infrastructure.
● Develop and maintain Kubernetes operators, Custom Resource Definitions (CRDs), controllers, networking components, ingress, DNS, TLS automation, and secure platform services.
● Design and implement enterprise-grade platform capabilities, including CI/CD pipelines, Single Sign-On (SSO), Role-Based Access Control (RBAC), secrets management, encryption, observability, audit logging, usage analytics, and infrastructure automation.
● Leverage AI-powered technologies within Site Reliability Engineering (SRE) practices by implementing intelligent agents, automated runbooks, anomaly detection, predictive operations, and LLM-driven operational workflows.
● Collaborate closely with engineering teams and Client’s business units to deliver scalable, reliable infrastructure solutions supporting graphics, AI, deep learning, mobile computing, autonomous vehicles, and other next-generation platforms.
● Ensure end-to-end platform reliability by managing the Kubernetes control plane, runtime operations, release lifecycle, Helm-based deployments, phased rollouts, and disaster recovery/rollback strategies.
Required Skills
● 6+ years of hands-on experience in DevOps or Site Reliability Engineering (SRE), managing production-grade Kubernetes environments across on-premises and cloud platforms, with expertise in CRDs, Operators, Ingress, cluster networking, and multi-tenant architectures.
● Proven experience with AWS or equivalent cloud platforms, including cloud resource provisioning, OIDC/SAML-based identity federation, Single Sign-On (SSO), Role-Based Access Control (RBAC), secrets management, relational databases, caching technologies, Websocket, SSH tunneling, and messaging/queueing systems.
● Knowledge of Helm, multi-arch containers, CI/CD, Jenkins/GitLab CI, Argo/Flux, Prometheus/Grafana, Victoria Metrics, Datadog, Splunk, or Kibana.
● Strong programming skills in Python or Go, with working knowledge of TypeScript and React, and the ability to contribute across backend development, frontend interfaces, infrastructure as code (IaC), platform engineering, and operational automation.
Qualifications / Standout Experience
● B.Tech/MTech/MS/BS in Computer Science or equivalent; track record shipping internal developer platforms, Kubernetes platforms, self-service compute, and CI/CD pipelines.
● Value-Added Skills: Linux, VMs, agentic workflows, CLIs, MCPs, AI-assisted platform telemetry, strong RCA discipline, runbooks, and documentation ownership.