The Sr. DevOps Cloud Engineer resolves escalated service issues, coaches' other engineers to resolutions, engineers and implements complex projects, and maintains and oversees assigned technology pillars. This position will require 4 days in the office, candidates must be in or willing to relocate to the Nashville, TN region.
The ideal candidate will have the following minimum qualifications:
- Bachelor's or Master's degree in Computer Science, Engineering or related field or directly related year for year work experience
- 5+ years in systems, infrastructure, cloud, or platform engineering, including at least 2 years working with public cloud and infrastructure automation
Sr. DevOps Cloud Engineer Requirements:
- Demonstrated proficiency with Python, PowerShell, and Bash for automation and tooling
- Proficiency with Git, including branching strategies, pull request review, and repository and workflow management
- Demonstrated experience with CI/CD platforms such as Azure DevOps, GitHub Actions, or GitLab CI, including pipeline authoring and release automation
- Demonstrated experience with containerization and Kubernetes, including Helm, managed Kubernetes services (AKS/GKE), Traefik ingress, and container image build and registry workflows
- Hands-on experience operating production workloads in public cloud (Microsoft Azure and Google Cloud), including compute, networking, identity, storage, and cost management
- Demonstrated production experience administering software load balancers, specifically HAProxy Enterprise or equivalent, including frontend and backend configuration, ACLs, health checks, and TLS termination
- Working knowledge of Layer 4 and Layer 7 load balancing concepts, including session persistence, health checking, and HTTP routing
- Working knowledge of Linux and Windows Server administration in cloud-hosted environments, including hardening and patching
- Experience with metrics, logs, and distributed tracing, including OpenTelemetry instrumentation and SLO-based alerting
- Working knowledge of cloud security practices including least privilege, network segmentation, secrets management, hardening baselines, and vulnerability remediation
Sr. DevOps Cloud Engineer key responsibilities are:
- Troubleshoots problems related to physical and virtual infrastructure performance and resolves trouble tickets as assigned
- Manages, participates, or performs tasks related to specific projects to implement new technology as assigned
- Designs, builds, configures, integrates, deploys and tests infrastructure components, operating systems, application software and system management tools
- Designs, builds, maintains and monitors systems and platforms ensuring they consistently exceed defined goals for availability, capacity, efficiency, scalability, and performance
- Designs, builds, implements, and maintains security, redundancy, backup, and logging strategies
- Evaluates, maintains, and modifies established best practices for infrastructure maintenance and asset management
- Creates and maintains custom scripts to increase system efficiency and lower the human intervention time on manual tasks
- Creates and maintains documentation for system configurations, runbooks, and administration procedures in the team's designated knowledge base
- Produces solution designs and technical specifications, including cost estimates, resource requirements, and delivery timelines
- Sets technical goals and delivery milestones for assigned projects and drives them to completion, coordinating with project management and stakeholder teams
- Implements and utilizes configuration automation tools such as Ansible, Puppet, or Chef
- Partners with architecture and security teams to define and implement automation standards and reusable infrastructure patterns
- Implements and utilizes IaC tools such as Terraform, Bicep, and Azure Resource Manager to consistently provision resources
- Deploys and maintains monitoring and observability tooling (Prometheus, Grafana, OpenTelemetry, and enterprise platforms in use), and builds dashboards and alerting that reflect service-level objectives
- Organizes priorities, escalates problems as needed, and communicates effectively with manager, co-workers, and clients
- Participates in vendor and technology evaluations, including proof-of-concept work, technical due diligence, and total-cost analysis
- Mentors and develops other engineers through code review, pairing, runbook development, and knowledge transfer, and continuously improves team processes
- Builds and maintains CI/CD pipelines for infrastructure and platform changes, using GitOps workflows for deployment and drift detection
- Operates and upgrades managed Kubernetes platforms, including cluster lifecycle, autoscaling, ingress, and workload onboarding
- Owns vulnerability remediation and patching for assigned platforms against defined SLAs, coordinating with security on risk acceptance and exception handling
- Designs, configures, and maintains application load balancing services using HAProxy Enterprise, including frontend and backend configuration, health checks, ACL-based routing, and TLS termination
- Maintains high availability of load balancing tiers, including redundant and clustered configurations, failover testing, and zero-downtime configuration reloads and upgrades
- Onboards applications to load balancing services and implements traffic management patterns such as weighted routing, blue/green and canary releases, and rate limiting
- Manages TLS certificate lifecycle for load-balanced services, including issuance, renewal automation, and cipher and protocol standards