JOB DESCRIPTION
Senior DevOps Engineer
AWS & Microsoft Azure | Windows & Linux | CI/CD | Terraform | Docker | Kubernetes
Job Overview
At Beinex, customer experience is at the forefront of everything we do. We are looking for an experienced Senior DevOps Engineer to design, deploy, automate, secure, monitor and maintain application infrastructure across AWS, Microsoft Azure and on-premises environments. The role will be responsible for deploying product updates, identifying production issues, implementing integrations, improving application reliability and performance, and driving infrastructure automation and continuous improvement.
The ideal candidate should have strong hands-on experience in AWS and server administration, working knowledge of Microsoft Azure, solid Infrastructure as Code and CI/CD capabilities, Docker expertise, and fundamental-to-intermediate Kubernetes knowledge. The engineer should be capable of supporting multi-cloud deployment and operational requirements across the application lifecycle.
Key Responsibilities
1. Managing AWS Cloud Servers (Windows and Linux)
- Provision, configure, secure and manage AWS EC2 instances for Windows and Linux workloads.
- Design and manage AWS infrastructure including VPCs, subnets, routing, security groups, IAM and related networking components.
- Implement server hardening, access control, patching and security best practices.
- Monitor server health, capacity, availability and performance.
- Troubleshoot operating-system, networking, storage, application and connectivity issues.
2. Microsoft Azure Cloud Management
- Provision, configure and manage Azure Virtual Machines and cloud infrastructure.
- Work with Azure Virtual Networks (VNets), subnets, NSGs, routing, Load Balancer and Application Gateway.
- Manage Microsoft Entra ID (Azure AD), RBAC, managed identities and access controls.
- Work with Azure Storage, Azure SQL, Key Vault and Azure Monitor.
- Support Azure DevOps and Azure Pipelines for application and infrastructure deployment.
- Apply Azure security, availability, monitoring and operational best practices.
- Troubleshoot Azure infrastructure, networking, identity and application deployment issues.
3. Multi-Cloud Deployment & Operations
- Deploy and manage applications across AWS and Microsoft Azure based on project requirements.
- Understand the differences in cloud networking, identity, compute, storage, monitoring and security between AWS and Azure.
- Support consistent deployment and operational practices across multiple cloud platforms.
- Participate in designing multi-cloud and hybrid-cloud solutions where required.
- Evaluate cloud services based on reliability, security, performance, scalability and business requirements.
4. Ensuring 24x7 Production Server Availability
- Implement high-availability and fault-tolerance measures for critical production systems.
- Configure redundancy, failover and recovery mechanisms to minimise downtime.
- Monitor production environments and respond promptly to incidents and outages.
- Participate in on-call and production support activities where required.
- Analyse recurring incidents and implement permanent preventive solutions.
5. Application Performance Management
- Collaborate with developers to analyse application performance metrics and identify bottlenecks.
- Monitor CPU, memory, disk, network, application and database performance.
- Optimise server resources and application runtime environments.
- Troubleshoot latency, timeout, connection, resource and availability issues.
- Use monitoring and logging data to support data-driven performance improvements.
6. Kubernetes & Container Orchestration
- Understand fundamental Kubernetes architecture and core components.
- Work with Pods, Deployments, ReplicaSets, Services, Namespaces, ConfigMaps and Secrets.
- Understand Ingress, health probes, resource requests/limits and Persistent Volumes.
- Deploy and troubleshoot containerised applications on Kubernetes.
- Understand Kubernetes networking, service discovery and basic workload scheduling.
- Troubleshoot common issues such as CrashLoopBackOff, ImagePullBackOff, failed deployments, service connectivity and resource constraints.
- Have basic working knowledge of Amazon EKS and Azure AKS.
7. Docker & Containerisation
- Build, manage and troubleshoot Docker images and containers.
- Write and maintain Dockerfiles and container configurations.
- Manage container networking, volumes, environment variables and runtime configuration.
- Integrate Docker-based workloads with CI/CD pipelines and cloud environments.
8. Critical Thinking & Problem-Solving
- Identify potential infrastructure and application issues proactively.
- Analyse complex technical problems and design effective solutions.
- Respond promptly and effectively to production incidents, outages and security events.
- Perform root-cause analysis and implement corrective and preventive actions.
- Use logs, metrics, traces and system data to make informed technical decisions.
9. Security & Compliance
- Implement security best practices across cloud, operating-system and application environments.
- Apply least-privilege access, IAM/RBAC and secure authentication practices.
- Implement network security controls, encryption, SSL/TLS and secrets management.
- Support security monitoring, vulnerability remediation and patch management.
- Ensure infrastructure and deployment practices align with organisational compliance requirements.
12. Automation & Continuous Improvement
- Develop and maintain automation scripts using Bash, Python and PowerShell/Batch.
- Automate provisioning, configuration, deployment, monitoring, backup, patching and operational activities.
- Identify repetitive manual processes and replace them with reliable automation.
- Continuously improve infrastructure efficiency, reliability, security and deployment speed.
13. Collaboration & Documentation
- Collaborate with developers, QA, security, infrastructure and business teams.
- Support application teams during deployments, integrations and production troubleshooting.
- Document server configurations, architecture, deployment procedures and operational processes.
- Maintain troubleshooting guides, runbooks and knowledge-base documentation.
Key Skills & Requirements
1. Windows/Linux Server Administration
- Proficient in managing and maintaining Windows and Linux servers.
- Installation, configuration, patching, hardening and regular maintenance.
- Server monitoring, performance analysis and troubleshooting.
2. AWS Cloud
- Hands-on experience with EC2, IAM, VPC, Route 53, SSL/TLS, S3, CloudFormation, CloudWatch, CloudTrail, SES and RDS.
- Understanding of AWS networking, security, availability and monitoring.
- Ability to design, deploy, secure and troubleshoot AWS infrastructure.
3. Microsoft Azure
- Hands-on or working experience with Azure Virtual Machines, Entra ID, VNets, NSGs, Azure Storage, Load Balancer/Application Gateway, Azure Monitor, Key Vault, Azure SQL and Azure DevOps.
- Understanding of Azure networking, identity, RBAC and security.
- Ability to deploy and troubleshoot applications on Azure.
4. Kubernetes
- Fundamental-to-intermediate knowledge of Kubernetes.
- Understanding of Pods, Deployments, Services, Ingress, ConfigMaps, Secrets, Namespaces and Persistent Volumes.
- Basic hands-on experience with EKS or AKS preferred.
- Ability to troubleshoot common Kubernetes workload and networking problems.
5. Application Performance & Web Servers
- Application performance monitoring and troubleshooting.
- Experience deploying Python-Django applications using Gunicorn.
- Proficiency in configuring and managing Nginx and IIS.
- Understanding of application bottlenecks, CPU/memory utilisation, networking and database performance.
6. CI/CD & DevOps Tools
- Jenkins administration including master-agent architecture, users/groups/roles, plugin management, upgrades and maintenance.
- Experience developing and troubleshooting Jenkins pipelines.
- Working knowledge of Azure DevOps, Azure Repos and Azure Pipelines.
- Understanding of Git-based workflows and branching strategies.
- Ability to integrate CI/CD pipelines with Docker, Terraform, AWS and Azure.
7. Infrastructure as Code
- Strong Terraform knowledge.
- Ability to provision and manage infrastructure across AWS and Azure using Terraform.
- Understanding of Terraform modules, state management, variables, outputs and environment management.
- Experience integrating Terraform with CI/CD pipelines.
8. Scripting & Automation
- Proficiency in Bash, Python and PowerShell/Batch scripting.
- Ability to automate infrastructure and operational tasks.
9. Monitoring, Logging & Observability
- Experience with CloudWatch, Azure Monitor, Prometheus, Grafana or equivalent monitoring platforms.
- Understanding of system/application logs, metrics, alerts and incident investigation.
- Ability to establish meaningful monitoring and alerting for production workloads.
10. Security & Disaster Recovery
- Understanding of IAM/RBAC, least privilege, server hardening, SSL/TLS, encryption and secrets management.
- Knowledge of backup, restore, HA, failover, DR, RTO and RPO concepts.
capabilities may also be considered.