Company Description ZySec AI builds the Sovereign Intelligence Stack, providing complete AI infrastructure that organizations fully own and control. The company focuses on mission critical organizations that must deploy AI without compromising data ownership, compliance, or security. Its flagship platform, CyberPod, is an autonomous data intelligence solution based on RAG and agentic AI workflows that turns fragmented data into real-time, context-aware insights while maintaining zero data exposure. ZySec AI emphasizes privacy-by-design, resilient architectures, and trustworthy intelligence to unlock maximum value from existing data with minimal setup and seamless integration. The team is driven by a vision of sovereign intelligence that is trusted by design and kept within the organizations that create it.
Location: Hyderabad, Telangana, India (On-site)
Role Description
- Build and maintain CI/CD pipelines for rapid application delivery.
- Provision and manage infrastructure across Azure, AWS, and private cloud.
- Deploy and manage Docker and Kubernetes workloads.
- Implement Infrastructure as Code using Terraform or equivalent tools.
- Automate operational tasks using Python, Bash, or PowerShell.
- Support deployment of AI/LLM applications, RAG applications, APIs, and agentic AI workloads.
- Implement monitoring, logging, alerting, and observability.
- Troubleshoot cloud, Kubernetes, networking, deployment, and production issues.
- Help developers create and manage development, staging, and production environments.
- Apply DevSecOps practices including secrets management, access control, and security scanning.
Qualifications
- 3–4 years of hands-on DevOps / Cloud / Platform Engineering experience.
- Strong experience with CI/CD, Git, Docker, and Kubernetes.
- Hands-on experience with Azure and/or AWS; multi-cloud experience is preferred.
- Experience with Terraform or another Infrastructure as Code platform.
- Good scripting skills in Python, Bash, or PowerShell.
- Understanding of cloud networking, IAM, compute, storage, databases, and security.
- Experience with monitoring and logging tools such as Prometheus, Grafana, Azure Monitor, CloudWatch, ELK/OpenSearch, or similar.
- Strong troubleshooting, ownership, and problem-solving skills.
- Experience with both Azure and AWS.
- Experience with private cloud or hybrid infrastructure.
- Experience supporting AI/LLM applications, RAG, AI agents, vector databases, or model-serving platforms.
- Experience with Helm, GitHub Actions, Azure DevOps, GitLab CI/CD, Jenkins, Argo CD, or Flux.
- Experience with GPU workloads and cloud/container security.