Job Title: Site Reliability Engineer (SRE)
Location: The person will work a split schedule between onsite and remote.
Contract: 12 months
Only W2
Job Purpose
Job Description – Platform Engineer / Site Reliability Engineer (Hybrid)
Note: A Glider Assessment is required before the interview process.
Position Overview
We are seeking a Platform Engineer / Site Reliability Engineer (SRE) to join its Platform Engineering team. This hybrid role requires working on a split schedule between onsite and remote. The engineer will collaborate closely with product teams, enabling reliable, scalable, and high-quality platform services. The ideal candidate is a strong communicator who thrives in a paired programming environment and demonstrates initiative and ownership.
Key Responsibilities
- Design, develop, and maintain cloud-native platform services.
- Support product teams by planning work, coordinating cross-functional activities, and delivering platform capabilities.
- Deploy, manage, and optimize Kubernetes (EKS) clusters in AWS.
- Build and maintain Infrastructure as Code (IaC) using Terraform.
- Improve system reliability, scalability, observability, and performance.
- Implement monitoring, logging, tracing, and alerting solutions.
- Participate in paired programming and code reviews.
- Troubleshoot production issues and implement long-term reliability improvements.
- Collaborate with engineers to deliver secure and resilient cloud infrastructure.
- Contribute to automation and CI/CD improvements.
Required Qualifications
- 3–6 years of software development experience.
- Strong programming experience in:
- Java (preferred – approximately 70%)
- Go
- Scala
- Python
- 3–6 years of Site Reliability Engineering (SRE), observability, or distributed tracing experience.
- 3+ years of AWS cloud experience.
- 3+ years managing or deploying Kubernetes/EKS clusters.
- Experience with Terraform for infrastructure automation.
- Knowledge of OpenTelemetry for observability.
- Ability to manage multiple priorities in a fast-paced environment.
- Excellent verbal and written communication skills.
- Comfortable working in a collaborative paired programming environment.
- Self-motivated with strong initiative and problem-solving skills.
Preferred Qualifications
- Experience with Datadog and/or Grafana.
- Experience building scalable cloud-native applications.
- Familiarity with DevOps, CI/CD pipelines, and automation best practices.