Responsibilities
Since H-E-B Digital Technology's inception, we've been investing heavily in our customers' digital experience, reinventing how they find inspiration from food, how they make food decisions, and how they ultimately get food into their homes. This is an exciting time to join H-E-B Digital. We're using the best available technologies to deliver modern, engaging, reliable, and scalable experiences that meet the needs of our growing audience. If you enjoy solving complex technical challenges, working in a rapidly changing environment, learning new skills, and building platforms that power enterprise-scale data and digital experiences, we want you on our team.
Our Partners thrive The H-E-B Way. In the
Senior Site Reliability Engineer, Data Platform role, that means you have a:
HEART FOR PEOPLE... you collaborate across engineering, data, and product teams, mentor others, and advocate for reliability and operational excellence.
HEAD FOR BUSINESS... you align reliability, scalability, and platform investments to business objectives while driving engineering best practices.
PASSION FOR RESULTS... you build resilient systems, automate operations, improve developer productivity, and ensure the availability of critical data platforms.
What You'll Do
As a
Senior Site Reliability Engineer supporting H-E-B's Data Platform, you will be responsible for the reliability, scalability, performance, and operational excellence of cloud-native data infrastructure and services.
Key Responsibilities
- Design, implement, and maintain highly available, resilient, and scalable data platform infrastructure.
- Develop and manage Infrastructure as Code (IaC) solutions using Terraform and other automation tools.
- Partner closely with Data Engineers, Data Scientists, Analysts, and Software Engineers to optimize platform reliability and performance.
- Develop comprehensive monitoring, alerting, observability, SLO, and capacity planning strategies aligned with business objectives.
- Monitor, troubleshoot, and optimize distributed storage, compute, and streaming systems across cloud-based environments.
- Lead root cause analysis efforts, identify systemic risks, and implement preventative solutions.
- Establish and champion best practices for platform engineering, reliability engineering, security, and operational excellence.
- Improve system resiliency through architecture reviews, fault tolerance strategies, and performance optimization initiatives.
- Build and maintain CI/CD pipelines that support rapid, reliable deployments.
- Implement security best practices and ensure compliance with enterprise and industry standards.
- Drive automation initiatives that reduce operational overhead and improve platform scalability.
- Contribute to long-term platform and reliability roadmaps.
- Stay current with emerging technologies and recommend innovative solutions that enhance platform capabilities.
Who You Are
Minimum Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent practical experience).
- 5+ years of experience in Software Engineering, Platform Engineering, Site Reliability Engineering, Cloud Engineering, or related disciplines.
- Strong experience managing and supporting large-scale cloud platforms and distributed systems.
- Extensive experience with Databricks and large-scale data processing environments.
- Deep expertise in AWS services, including EC2, S3, VPC, IAM, Lambda, CloudFormation, and related cloud technologies.
- Strong programming and automation experience with Python; experience with SQL is preferred.
- Hands-on experience with Terraform and Infrastructure as Code practices.
- Experience with distributed data and compute technologies such as Apache Spark, streaming platforms, and ETL workflows.
- Strong understanding of software engineering principles with an emphasis on reliability, scalability, observability, and performance optimization.
- Experience implementing monitoring, logging, observability, and incident response practices.
- Excellent analytical, troubleshooting, and problem-solving abilities.
- Strong communication skills with the ability to collaborate across technical and business teams.
Preferred Qualifications
- Experience with additional cloud platforms such as Google Cloud Platform (GCP) or Microsoft Azure.
- Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- Familiarity with CI/CD platforms and deployment automation tools.
- Experience supporting data lake, lakehouse, or modern data platform architectures.
- AWS, Terraform, Kubernetes, Databricks, or other relevant industry certifications.
- Experience mentoring engineers and leading cross-functional technical initiatives.
What Makes You Successful
- You take ownership of reliability and operational outcomes across complex distributed systems.
- You proactively identify risks and drive long-term solutions rather than short-term fixes.
- You can balance technical excellence with practical business needs.
- You thrive in fast-paced environments and effectively manage competing priorities.
- You influence engineering culture through collaboration, mentorship, and continuous improvement initiatives.
- You are passionate about automation, scalability, and creating exceptional platform experiences for engineering teams.
Working Conditions
- Function in a fast-paced retail and technology environment.
- Travel by car or plane, including occasional overnight stays.
- Sit for extended periods while working at a computer.
- Participate in on-call rotations and support critical production systems as needed.
- Work flexible hours when required to support business-critical initiatives.
The responsibilities listed above describe the general nature and level of work performed and are not intended to be an exhaustive list of all duties, responsibilities, or qualifications associated with this position.