Location: Beijing/Singapore
About the Role
Operate and build the platform for our multi-cloud infrastructure and container clusters, including our self-hosted E2B (AI agent sandbox) cluster, and keep production services highly available, scalable, and cost-efficient through automation and platform engineering.
Responsibilities
- Cluster Operations & Management
- Manage and maintain container clusters (Kubernetes, Docker) and open-source middleware clusters (Kafka, Redis, Elasticsearch) across multiple business units.
- Own the self-hosted E2B (AI agent sandbox) cluster, including the Nomad / Consul orchestration layer and the Firecracker microVM runtime.
- Ensure the performance, scalability, and reliability of distributed systems.
- Linux & Cloud Infrastructure
- Handle day-to-day operations of Linux servers, including hardening, patch management, backups, and performance tuning.
- Manage cloud resources across AWS, Azure, and GCP, ensuring reliability, security, and cost efficiency.
- Infrastructure Platform Development
- Design, build, and continuously improve our infrastructure operations platform.
- Develop and maintain infrastructure management, CI/CD GitOps pipelines (Jenkins / Argo CD), monitoring and alerting (Prometheus / Grafana), and centralized logging (ELK).
- Drive platform standardization and automation.
- High Availability & Reliability
- Ensure the highest availability of production services through proactive monitoring and incident response.
- Take part in a 24/7 on-call rotation, and lead incident response, root cause analysis, and postmortems.
- Implement and maintain SLA/SLO frameworks and reliability engineering practices.
- Automation & Process Improvement
- Build self-service tools and workflows to improve team productivity.
- Establish best practices for infrastructure as code (Terraform / Ansible) and configuration management.
- Maintain thorough documentation, SOPs, and runbooks.
Requirements
- 3–4+ years of hands-on experience in Linux system operations, DevOps, or SRE (no degree requirement; ability comes first).
- Strong Linux fundamentals, with knowledge of networking basics and system security best practices.
- Deep expertise in containerization (Kubernetes, Docker), with experience operating production-grade clusters.
- Experience building and managing CI/CD pipelines (Jenkins, GitHub Actions / Argo CD).
- Familiarity with common infrastructure components: Nginx, MySQL, Redis, Kafka, and Elasticsearch.
- Proficiency in Shell and Python for scripting and automation.
- Hands-on operations experience with at least two public clouds among AWS, Azure, and GCP.
- Familiarity with infrastructure monitoring, logging, and observability tools (Prometheus, Grafana, ELK).
- Hands-on experience with Terraform.
- Ability to collaborate on technical work in Mandarin and to read, write and participate in day-to-day technical communication in English.
Preferred Qualifications
- Familiarity with or exposure to E2B / AI agent sandbox infrastructure (Firecracker microVM, Nomad, Consul).
- Experience with service mesh architecture (Istio) and eBPF.
- CKA/CKAD or AWS/Azure/GCP professional certifications.
Manus excels at various tasks in work and life, getting everything done while you rest at Manus AI.
What we offer:
- Build at the Frontier of AI Agents - Work with a fast-paced team to turn frontier AI breakthroughs into real-world impact.
- Equity & Shared success - Our share incentive plan enables employees to participate in Manus’s long-term growth and share in the value we create together.
- Unlimited Manus Tokens - Enjoy unlimited Manus tokens to experiment, build, and supercharge your productivity.
If you're passionate about cutting-edge technology and making a real impact, we’d love to hear from you!