Job Description
We are seeking a highly skilled Senior DevOps Engineer with deep, hands-on expertise in building and operating enterprise-grade CI/CD platforms, cloud infrastructure, and developer tooling at global scale. In this role, you will own the design, automation, and reliability of the software delivery pipeline — enabling engineering squads to ship faster, safer, and with greater confidence. You will bring a strong platform engineering mindset, infrastructure-as-code discipline, and a proven track record of reducing toil, improving system resilience, and embedding security into every stage of the software lifecycle in a regulated financial services environment.
Roles & Responsibilities
- Own the CI/CD Platform Design, build, and continuously improve enterprise CI/CD pipelines. Ensure pipelines are fast, reliable, secure, and scalable across dozens of engineering squads and technology stacks.
- Architect & Manage Cloud Infrastructure Provision, manage, and optimize cloud infrastructure across AWS, Azure, and GCP using infrastructure-as-code. Enforce cost controls, security guardrails, and operational best practices at scale.
- Drive Platform Reliability & SRE Practices Define and own SLOs, SLIs, and error budgets for platform services. Lead incident response, conduct blameless post-mortems, and implement systemic fixes to eliminate recurring failures.
- Embed Security into the Delivery Pipeline Implement DevSecOps practices — integrating SAST, DAST, SCA, container scanning, and secrets detection into every pipeline stage. Ensure compliance with enterprise security standards and regulatory requirements.
- Standardize Container & Kubernetes Operations Own the Kubernetes and OpenShift platform strategy — cluster lifecycle management, namespace governance, RBAC, networking policies, and workload autoscaling.
- Enable Developer Productivity Build internal developer platforms (IDP), self-service tooling, golden path templates, and reusable infrastructure components that reduce friction and accelerate engineering teams.
- Lead Observability & Capacity Planning Design and maintain a unified observability stack covering metrics, logging, and distributed tracing. Drive proactive capacity planning and performance optimization across all environments.
- Champion Infrastructure as Code & GitOps Establish IaC standards across the organization. Drive GitOps adoption for all infrastructure and application configuration changes, ensuring auditability and repeatability.
- Mentor & Lead Technical Direction Define DevOps standards, conduct platform design reviews, mentor junior engineers, and collaborate with architects, security teams, and engineering leads to align platform strategy with business goals.
Technical Skills Required
CI/CD & Build Systems
- Jenkins
- GitHub Actions
- GitLab CI
- ArgoCD
- Tekton
- CircleCI
- Maven / Gradle / npm build integration
- Pipeline-as-Code
- Artifact Management
- Release Automation & Versioning Strategies
Container & Kubernetes Platforms
- Docker
- Kubernetes
- OpenShift
- Helm
- Kustomize
- Operators & Custom Resource Definitions (CRDs)
- Kubernetes RBAC & Network Policies
- Pod Security Standards
- Cluster Lifecycle Management
- Service Mesh (Istio, Linkerd)
Cloud Platforms & Infrastructure
- AWS (EC2, ECS, EKS, Lambda, RDS, S3, VPC, IAM, SQS, SNS, Route53)
- GCP (GKE, Cloud Run, IAM, VPC)
- Multi-cloud & Hybrid Cloud Architecture
- Cloud Networking (VPC, peering, transit gateway, private endpoints)
- Cloud Cost Optimization & FinOps
Infrastructure as Code & GitOps
- Terraform
- Ansible
- Pulumi
- CloudFormation
- ArgoCD / Flux (GitOps)
- Crossplane
- Packer
- Terragrunt
Observability & Monitoring
- Prometheus
- Grafana
- OpenTelemetry
- ELK Stack
- Loki / Fluentd / Fluentbit
- Jaeger / Zipkin (Distributed Tracing)
- Datadog / Dynatrace / New Relic
- PagerDuty / Alertmanager
- SLO / SLI / Error Budget Management
- Google Cloud Observability
Storage
DevSecOps & Security
- SAST Tools (Checkmarx, SonarQube)
- DAST Tools (OWASP ZAP, Burp Suite)
- SCA / Dependency Scanning
- Container Image Scanning (Trivy, Aqua Security, Twistlock)
- Secrets Scanning
- HashiCorp Vault
- AWS Secrets Manager / Azure Key Vault
- mTLS & Certificate Management
- RBAC & Identity-Aware Proxy
- CIS Benchmarks & STIG Compliance
Networking & Service Delivery
- DNS & Load Balancing (Nginx, HAProxy, AWS ALB/NLB)
- API Gateway (Kong, AWS API Gateway, Apigee)
- CDN (Cloudflare, AWS CloudFront)
- Ingress Controllers (Nginx Ingress, Traefik)
- VPN & Private Connectivity (AWS Direct Connect, Azure ExpressRoute)
- Zero Trust Networking
Scripting & Automation
- Bash / Shell Scripting
- Python (automation, tooling, Lambda functions)
- Go (CLI tooling, Kubernetes controllers)
- PowerShell
- Makefile / Task runners
Reliability & SRE Practices
- Site Reliability Engineering (SRE) Principles
- Chaos Engineering (Chaos Monkey, LitmusChaos, Gremlin)
- Disaster Recovery & Business Continuity Planning
- Runbook Automation
- Incident Management (PagerDuty, OpsGenie)
- Capacity Planning & Performance Tuning
- Blue-Green & Canary Deployment Strategies
- Feature Flag Management
Developer Platform & Productivity
- Internal Developer Platforms (Backstage, Port)
- Golden Path Templates & Scaffolding
- Self-Service Infrastructure Provisioning
- Developer Experience (DX) tooling
- Pre-commit hooks & code quality gates
Salary Range: $125,000 to $140,000 per year
Qualifications: BACHELOR OF COMPUTER SCIENCE