IMMEDIATE INTERVIEW = Site Reliability Engineer (SRE) in Shea, AZ – HYBRID (NEED LOCAL CANDIDATE)- MUST COMPLETE ROPES ASSESSMENT
Site Reliability Engineer (SRE)
Location: Shea, AZ — Hybrid from Day 1 (3 days/week)
Assessment: MUST COMPLETE ROPES ASSESSMENT
Site Reliability Engineer (SRE) – Job Requirements
We are seeking a Site Reliability Engineer with strong experience in cloud-native operations, observability, automation, and production support for large-scale enterprise applications.
Skillset Required
- 3-5 years of experience in Site Reliability Engineering, Production Operations, or Platform Engineering supporting large-scale, high-performance applications across hybrid environments (on-premises and cloud).
- 3-5 years of experience developing automation scripts and building Application Performance Management (APM) dashboards to monitor end-to-end transaction journeys.
- Hands-on programming experience (2+ years) with one or more languages such as Go, Python, Java, or Rust.
- Working knowledge of relational and NoSQL databases including Oracle, SQL Server, PostgreSQL, MongoDB, Redis, ClickHouse, PL/SQL, or time-series databases.
- Experience with cloud migration and containerization initiatives using GCP, AWS, Azure, Rancher, OpenShift, or similar platforms.
- Experience managing containerized applications in Kubernetes environments such as GKE, RKE, or AKS.
- Strong experience implementing observability solutions using Open Telemetry (OTEL), distributed tracing, monitoring, and incident management.
- Familiarity with GraphQL frameworks such as Apollo, Prisma, or Hasura.
- Strong networking fundamentals including TCP/IP, HTTP, DNS, load balancing, and service mesh technologies.
- Experience participating in 24x7 on-call rotations and meeting incident response SLAs.
- Experience managing highly available, customer-facing platforms with a focus on reliability, automation, and operational excellence.
- Hands-on experience with monitoring and observability tools such as Splunk, Dynatrace, AppDynamics, Grafana, and Prometheus.
- Experience with CI/CD and Agile tools such as Rally, Confluence, and related DevOps platforms.
- Knowledge of in-memory caching technologies, especially Redis.
- Strong troubleshooting and debugging skills across distributed systems and API gateway architectures.
- Experience with Google Cloud services including GCS, Cloud SQL, Spanner, and BigQuery.
- Experience supporting HashiCorp Vault environments.
- Exposure to Vertex AI, Generative AI, and cloud-based analytics platforms.