Job Overview
Mumzworld is looking for a DevOps Engineer to join our DevOps/Platform Engineering team. You will help operate and improve the reliability, scalability, security, and observability of a high-traffic eCommerce platform. The role is hands-on and suited to someone who enjoys Linux troubleshooting, AWS, automation, Grafana, Kubernetes, and solving real production problems.
What You'll Be Doing
- Monitor and improve the reliability, availability, and performance of production and non-production environments.
- Troubleshoot Linux, application, network, database, Kubernetes, and AWS infrastructure issues.
- Participate in incident response, root cause analysis, production recovery, and post-incident improvements.
- Build automation and operational tooling using Python, JavaScript/Node.js, and Shell scripting.
- Create and maintain Grafana dashboards, alerts, metrics, logs, and distributed tracing for application and infrastructure visibility.
- Support workloads running on AWS and Kubernetes, including deployments, scaling, configuration changes, and production readiness.
- Improve system resilience through alert tuning, capacity planning, automation, runbooks, and elimination of repetitive operational toil.
- Work closely with engineering, QA, security, and platform teams to resolve issues and improve service reliability.
Experience
Core Technical Skills
- Linux - strong administration and troubleshooting across processes, CPU, memory, disk, filesystems, networking, DNS, logs, SSH, permissions, package management, and shell scripting.
- AWS - strong production experience with services such as EKS,Lambda, Cloudfront, ALB,RDS and other related services.
- Grafana & Observability - strong hands-on experience with dashboards, alerting, PromQL, Loki/log querying, metrics, tracing concepts, Kubernetes observability, and AWS monitoring.
- Programming & Automation - Python, JavaScript/Node.js, Bash/Shell, REST APIs, JSON, and scripting for automation and operational tooling.
- Kubernetes & Containers - Docker, EKS/Kubernetes, Pods, Deployments, Services, Ingress, ConfigMaps, Secrets, resource requests/limits, autoscaling, logs, events, and troubleshooting.
- Networking - TCP/IP, HTTP/HTTPS, DNS, SSL/TLS, load balancing, routing, NAT, firewalls, reverse proxies, CDN concepts, and API connectivity troubleshooting.
- DevOps & IaC - Git, CI/CD, Terraform, Helm, Infrastructure as Code, and automated deployment practices.
Qualifications
- 2-5 years of experience in DevOps, SRE, Cloud Engineering, Platform Engineering, Production Support, or a similar role.
- Understanding of incident management, SLIs/SLOs/SLAs, capacity planning, resilience, production readiness, and reducing operational toil.
- Strong troubleshooting and ownership mindset, with the ability to investigate production issues independently and communicate clearly during incidents.
- Bachelor's degree in Computer Science, Engineering, IT, or equivalent practical experience.
Good to have
Prometheus, Grafana Cloud, Loki, Tempo, OpenTelemetry, Grafana Alloy, Istio, Redis, OpenSearch/Elasticsearch, New Relic, MySQL/MariaDB, PostgreSQL, CDN/caching technologies, microservices, and experience supporting high-traffic eCommerce systems.
What Success Looks Like
- Production issues are detected early, diagnosed quickly, and resolved with clear follow-up actions.
- Grafana dashboards and alerts provide actionable visibility instead of unnecessary noise.
- Recurring operational tasks and failure patterns are reduced through automation and permanent fixes.
- AWS and Kubernetes workloads remain reliable, scalable, observable, and ready for production demand.
- You progressively take ownership of services, incidents, runbooks, and reliability improvements with less day-to-day guidance.
Why This Role Matters
Reliability directly affects the customer experience at Mumzworld. As our eCommerce platform grows, the DevOps/Platform Engineering team helps ensure that customers can browse, order, and transact without disruption. This role sits close to production and engineering: your work on observability, automation, incident response, and cloud reliability will directly improve platform stability, engineering velocity, and our ability to scale confidently