We're looking for an experienced Site Reliability Engineer with strong expertise in Dynatrace and modern observability platforms. The ideal candidate will have hands-on experience improving system reliability, automating operations, and supporting cloud-native environments.
Key Skills:
✔ Dynatrace (APM, dashboards, monitoring & alerting)
✔ Observability tools – Splunk, ELK, Grafana, Prometheus (ThousandEyes is a plus)
✔ SRE, AIOps & Incident Management
✔ Terraform & Infrastructure as Code (IaC)
✔ CI/CD – Jenkins, TeamCity, Bamboo, Octopus, UDeploy
✔ AWS, Azure & GCP
✔ Production Support, Automation & Performance Optimization
✔ Jira, Confluence & Technical Documentation
What You'll Do:
- Enhance platform reliability through automation and proactive monitoring.
- Build scalable observability solutions and operational dashboards.
- Optimize cloud infrastructure and CI/CD processes.
- Improve incident response, system performance, and operational efficiency.
- Partner with engineering teams to deliver resilient, highly available platforms.