Experience required - 6 to 10 Years
- Reliability & Observability: Lead incident response, conduct RCAs, and design full-stack observability (Prometheus, Grafana, Datadog) to eliminate alert fatigue.
- AI Adoption (Mandatory): Must leverage AIOps platforms for anomaly detection and utilize LLM-powered tools to accelerate MTTR.
- AEM Management (Mandatory): Must own end-to-end reliability of AEM environments (Author, Publish, Dispatcher, and AEMaaCS) and manage OSGi configurations and replication queues.
- Infrastructure & Security: Build AWS cloud-native infrastructure using Terraform and integrate security tools (SAST/DAST) into CI/CD pipelines.
- Leadership: Mentor junior SREs, contribute to on-call rotations, and represent SRE in architecture reviews.
- CDN: Manage Cloudflare CDN and Workers.