Role : Cloud Infrastructure Architect
Location : Hybrid-4 days a week in Redwood City, CA
Full time
We need a senior or staff-level engineer with 4+ years of experience in cloud infrastructure who is equally strong at hands-on coding and systems design. You should be comfortable working in a small, fast-moving team with high ownership and minimal red tape, and have a track record of building and operating scalable, reliable cloud infrastructure. You must have a CS or engineering degree and be
proficient in Python — this is a Python-heavy environment.
Role Requirements
- Cloud Infrastructure Engineer needed to join a growing team.
- Must handle both coding and systems design; Python, Terraform, and Kubernetes expertise required.
Candidate Requirements
- 4+ years of experience preferred; strong coding skills necessary.
- Must have a computer science or relevant engineering degree.
Team Structure and Reporting
- Reports to VP of Engineering; flat structure beyond that.
- Small engineering team of about 20, offering broad exposure across the stack.
Work Experience
- Cloud infrastructure engineering at production scale (not just backend/microservices)
- Designing and operating highly available, scalable infrastructure systems
- Managing Kubernetes-based systems for compute workloads
- Implementing infrastructure-as-code across environments
- Education
- BS in Computer Science, Electrical Engineering, or related engineering field
- Hard skills
- Strong Python coding ability (not just scripting) + Bash or Go
- Strong Kubernetes, Docker, and containerized architecture expertise
- AWS required; multi-cloud a plus (GCP, Azure, OCI, on-prem)
- Infrastructure-as-code (Terraform or Pulumi)
- CI/CD pipelines, monitoring/observability & systems-level debugging
What You'll Do
- Architect and maintain multi-cloud infrastructure (AWS primary, plus GCP, Azure, OCI, on-prem) to support customer deployments across diverse environments
- Define and implement infrastructure-as-code using Terraform or Pulumi, driving best practices across the team
- Design and manage Kubernetes-based systems for model training, inference, and data processing workloads
- Write production-quality Python code daily — this is not a scripting-only role; you'll pass rigorous coding interviews and build real systems
- Optimize CI/CD pipelines and streamline deployment of services across customer environments
- Build monitoring, alerting, and logging systems to ensure high availability and observability
- Collaborate closely with research and engineering teams to provide infrastructure support for training large-scale ML models
- Drive cost-efficiency strategies across compute and storage resources
- Respond to and resolve infrastructure incidents with ownership and urgency
Tech stack
Python, AWS, GCP, Azure, OCI, Kubernetes, Terraform, Pulumi, Docker, Helm, CI/CD, Bash, Go, CloudFormation
- Cloud infrastructure and developer tools startups (fast-moving companies with strong cloud/Kubernetes expertise)
- AI research labs and foundation model companies (teams with large-scale cloud infrastructure for training)
- Cloud-native infrastructure teams at major cloud providers and big tech