DevOps Engineer
Experience: 2–4 years Location: Bangalore (in-office) Stack: Azure (primary), AWS (parity), Bicep/Terraform, GitHub Actions, PostgreSQL, OpenTelemetry Reports to: CTO
About the Role
Zenalyst builds an enterprise agentic execution layer for the CFO's office — treasury, procure-to-pay and collections automation for large organisations.
We run a multi-tenant platform where every customer has their own physical database, deployed across Azure and AWS with full parity, and offered in three modes — managed multi-tenant, single-tenant in a customer VPC, and bring-your-own-cloud. Enterprise buyers audit us. Banks allowlist our egress IPs. Certificates expire and someone has to know before they do.
You will own that infrastructure. This is not a role where you inherit a stable estate and keep it running; you will be building it out as we scale.
Responsibilities
Infrastructure and IaC
- Own and extend our infrastructure as code — Bicep on Azure, Terraform/CDK on AWS
- Provision and manage container workloads on Azure Container Apps and ECS Fargate, including autoscaling rules driven by CPU, HTTP concurrency and queue depth
- Manage the tenant provisioning pipeline: new customer, new physical Postgres, new secrets, new connection routing
- Maintain cloud parity — every change we make on Azure has to have an AWS equivalent
CI/CD
- Own GitHub Actions pipelines across Java, Python, Node and React services
- Maintain OIDC federated deploy identity — no static cloud credentials anywhere
- Build and maintain container image pipelines: registries, scanning, signing, geo-replication
- Manage Flyway and Alembic migration execution across tenant databases safely
Networking and security
- Manage private endpoints, private DNS, NAT gateways and static egress IPs — banks allowlist these, so they cannot change casually
- Maintain egress FQDN allowlists and firewall rules; default-deny is the baseline
- Own certificate lifecycle for bank mTLS connections, including expiry alerting
- Manage secrets in Azure Key Vault and AWS Secrets Manager; keep credentials short-lived and identity-based
- Support ISO 27001 evidence collection and enterprise security reviews
Observability and reliability
- Own the OpenTelemetry pipeline: collectors, traces, metrics, structured logs into App Insights, CloudWatch and Grafana
- Build and tune alerting against our SLOs — gateway latency, error rates, model fallback rates, cross-tenant attempts, certificate expiry
- Own backup, PITR and disaster recovery posture across Postgres, Redis, document stores and object storage; run restore drills, not just backup jobs
- Participate in on-call and incident response, and write the postmortems
Cost
- Track and report per-tenant infrastructure cost; find and remove waste
- Manage environment sizing — dev, staging with scheduled shutdown, and zone-redundant production
Requirements
Essential
- 2–4 years in DevOps, SRE, platform or cloud infrastructure roles
- Hands-on with Azure or AWS in production — not just certifications
- Infrastructure as code in Terraform, Bicep, CloudFormation or CDK
- Docker, and container orchestration on Kubernetes, ECS, Container Apps or similar
- CI/CD pipeline ownership, ideally GitHub Actions
- Linux fundamentals, networking fundamentals — VPC, subnets, NAT, DNS, TLS
- Scripting in Python, Bash or Go
- Comfortable being on call and owning an incident end to end
Strongly preferred
- Managed PostgreSQL operations — HA, PITR, read replicas, connection pooling
- Secrets management (Key Vault, Secrets Manager, Vault) and workload identity
- OpenTelemetry, Prometheus, Grafana, or equivalent observability stacks
- Experience supporting a security certification — ISO 27001, SOC 2 or similar
- Multi-cloud or multi-tenant infrastructure experience
Nice to have
- mTLS and certificate lifecycle management
- Experience deploying into customer-controlled cloud environments
- Policy-as-code, image signing, supply chain security
- Cost optimisation with measurable results you can describe
Who This Suits
Someone who automates the second thing they do twice, who treats a manual production change as a bug in the pipeline, and who would rather run a restore drill than assume the backup works. If you have opinions about why default-deny egress matters, you will fit here.
What We Offer
- Ownership of the entire platform infrastructure, not a slice of it
- Real constraints worth solving: multi-cloud parity, per-tenant isolation, bank-grade network controls, enterprise audit
- Direct access to founders and the CTO; you will set the standards rather than inherit them
- Competitive fixed compensation plus ESOP
Apply
Send your CV to hello@zenalyst.ai with the subject line "DevOps Engineer — [Your Name]". If you have written IaC, pipelines or tooling you can share, include it.