Looking for a motivated and experienced DevOps / Site Reliability Engineer (SRE) to own and architect scalable infrastructure in a fast-paced, technologically innovative environment within a leading investment and data analytics firm. This is a key role reporting directly to executive leadership, perfect for a builder who thrives on end-to-end ownership and driving infrastructure strategy.
Role Overview
This role involves designing, deploying, and maintaining robust cloud and on-premises infrastructure, ensuring high availability, security, and scalability of the core platform supporting AI, data processing, and financial analytics. The successful candidate will serve as the sole DevOps/SRE professional, making strategic decisions that influence platform reliability and operational excellence across the organization.
Key Responsibilities
- Lead architecture design and implementation for cloud-native infrastructure, including virtual networks, container orchestration (Kubernetes, Docker), and cloud resources on AWS, Azure, or GCP.
- Develop, maintain, and optimize CI/CD pipelines, release workflows, and version control management to ensure smooth deployment processes.
- Manage container registries, secrets management (Key Vaults), storage solutions, and network configurations for secure and efficient system operation.
- Implement monitoring, observability, and alerting solutions (Prometheus, Grafana, ELK Stack, DataDog, or similar) for comprehensive system health tracking and incident response.
- Collaborate with engineering teams to support scalable, reliable systems, integrating infrastructure with AI and data workflows as needed.
- Provide strategic input into infrastructure roadmap, automation, and security best practices, ensuring compliance and operational resilience.
- Own and troubleshoot production incidents, ensuring rapid resolutions and continuous system improvements.
Core Qualifications & Requirements
- 3+ years of proven experience in DevOps, Site Reliability Engineering, or infrastructure engineering with a track record of building, deploying, and maintaining production-scale systems.
- Strong expertise in cloud infrastructure management (AWS, Azure, GCP), containerization (Docker, Kubernetes), and orchestration.
- Extensive hands-on experience with CI/CD pipelines, version control (Git), and automation tools (Jenkins, GitLab CI, CircleCI, or similar).
- Proficiency in monitoring, observability, and incident management tools (Prometheus, Grafana, DataDog, ELK, Splunk).
- Excellent communication skills with the ability to articulate technical options and system architecture to executive stakeholders.
- Ability to work independently as the sole DevOps/SRE resource, making strategic decisions and leading infrastructure initiatives.
Nice-to-Have Qualifications
- Experience in financial services, hedge funds, trading platforms, or fintech environments.
- Familiarity with infrastructure for AI and machine learning, including model deployment, MLOps pipelines, and AI workload orchestration (LangChain, MCP, MLflow, Langfuse).
- Hands-on experience with private package repositories (Artifactory, PyPI).
- Knowledge of security best practices, compliance standards, and non-compete considerations in finance.
- Prior experience at high-profile tech or quant firms such as Two Sigma, Renaissance Technologies, or similar.
Core Technical Skills
- Cloud platforms: AWS, Azure, GCP
- Containerization: Docker, Kubernetes, OpenShift
- Infrastructure as Code: Terraform, CloudFormation, ARM templates
- CI/CD tools: Jenkins, GitLab CI, CircleCI, ArgoCD
- Monitoring & observability: Prometheus, Grafana, ELK Stack, DataDog, Splunk
- Orchestration & workflow automation: Prefect, Airflow, Dagster
- Secrets management: HashiCorp Vault, AWS KMS, Azure Key Vault
- Scripting & automation: Bash, Python, Go
Career Impact
This position offers the opportunity to redefine infrastructure standards within a cutting-edge data-driven organization, impacting investment decisions and high-stakes financial systems.