***This role is FTE only and is not open to C2C or any 3rd Party Assistance***
Job Title: Systems Engineer (Mid-Level)
Job Type: Direct Hire. Target range is 120-130k; 10% annual bonus, 401k match, PTO, health benefits
Job Location: Hybrid in Atlanta (Onsite Tuesday, Wednesday, Thursday)
Position Summary
- We are seeking a highly skilled Senior Systems Engineer who has owned production systems, participated in major incident response efforts, and understands how resilient systems are designed, monitored, and continuously improved. The ideal candidate combines strong infrastructure and cloud engineering experience with expertise in observability, automation, and operational excellence.
- We are looking for engineers who can develop standards, create scalable solutions, automate operational processes, and influence engineering teams across the organization.
- You will work closely with software engineers, platform teams, and technical leaders to build observability frameworks, improve incident detection and response, establish resilience strategies, and drive adoption of modern operational practices.
Key Responsibilities
- Design and implement scalable observability solutions across cloud-native and distributed environments
- Develop and champion best practices for metrics, logging, tracing, synthetic monitoring, and alerting
- Improve organizational capabilities for incident detection, triage, and response
- Drive adoption of AWS monitoring best practices across engineering teams, and help define spec-driven observability standards for cloud workloads
- Build automated remediation and response capabilities to reduce time-to-detection and time-to-resolution
- Develop automation solutions that eliminate manual operational work
- Build integrations, tooling, workflows, and self-service capabilities that improve operational efficiency
- Leverage AI and emerging technologies to enhance operational intelligence and workflow automation
- Build strong working relationships with engineering managers, architects, technical leads, and product teams
- Act as a trusted advisor on observability and reliability best practices
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field and 4+ years of relevant experience, or an equivalent combination of education and experience
- Hands-on experience designing, building, and improving AWS CloudWatch-based monitoring, metrics, and alerting across cloud-based environments
- Establish and champion cloud observability patterns as the primary standard for engineering teams operating in AWS
- Strong experience with enterprise monitoring, observability, incident management, and operational tooling
- Experience designing and implementing monitoring strategies for large-scale distributed systems
- Professional experience writing and maintaining synthetic monitoring solutions
- Proficiency in Python, JavaScript, Ruby, or similar scripting/programming languages
- Experience with source control systems and CI/CD pipelines
- Familiarity with Infrastructure-as-Code principles and automation frameworks
- Knowledge of platform integrations, APIs, and automated operational workflows
- Experience supporting production environments and participating in on-call and incident response activities
Preferred Qualifications
- Experience with modern observability platforms, APM solutions, log aggregation, and event management tools
- Background building internal platforms, shared services, or developer enablement tooling
- Experience applying AI or machine learning concepts to operational workflows
- Familiarity with Site Reliability Engineering (SRE) principles and practices
- Experience implementing automated remediation and event correlation solutions
- Exposure to large-scale cloud-native architectures and containerized environments