Dev Ops/ SRE Engineer
Austin, TX (Hybrid)
Key Responsibilities
Deployment & Integration: Collaborate with software teams to deploy and maintain existing applications and tools, ensuring smooth integration within our infrastructure.
Monitoring & Signal Flow: Set up and manage monitoring solutions to track application health, data pipelines, and system signals. Ensure real-time visibility into performance and operational metrics.
Incident Response & Troubleshooting: Quickly assess and respond to application outages or issues, identify root causes, and coordinate resolution efforts to minimize downtime.
Automation & Optimization: Automate deployment processes, monitoring, and alerting to improve efficiency and reduce manual intervention.
Collaboration: Work closely with data engineers, security teams, and software developers to ensure applications are robust, secure, and scalable.
Documentation & Best Practices: Maintain comprehensive documentation of deployment processes, system architecture, and incident protocols.
Qualifications
Proven experience in deploying and managing enterprise software applications in a cloud or hybrid environment.
Proficient in Python
Knowledge cloud platforms AWS, GCP
Strong understanding of infrastructure automation tools (e.g., Jenkins, Terraform, Ansible, or similar).
Experience with monitoring and alerting tools (e.g., DataDog, Nagios, Prometheus, Grafana).
Familiarity with data pipelines, signal flow, and system architecture related to AI/ML applications.
Understanding of CI/CD pipelines and tools
Experience with Containerization (Docker, Kubernetes)
Skills in incident detection and response
Ability to troubleshoot complex system issues quickly and effectively.
Excellent collaboration and communication skills.
Knowledge of security best practices in deployment and system management.
Preferred Skills
Experience with AI/ML infrastructure or data platforms.
Familiarity with containerization and orchestration (Docker, Kubernetes).
Understanding of network protocols and security.