hackajob is collaborating with Verisk to connect them with exceptional professionals for this role.
Description
We are seeking a skilled DevOps Engineer to join our team as we plan for a smooth transition following an upcoming team member retirement. This role is critical for maintaining and evolving our hybrid infrastructure, which combines on-premises datacenter operations with cloud (AWS) and containerized environments. The ideal candidate will handle both hands-on physical infrastructure tasks and advanced automation/R&D initiatives to ensure reliability, security, and scalability.
This position requires proximity to our datacenter(s) in the NY/SI region for occasional on-site visits involving hardware maintenance, cabling, installations, and troubleshooting.
Responsibilities
Security and maintenance:
- Perform hardware stack maintenance on Dell servers, Cisco switches, and cable systems.
- Upgrade infrastructure components and tools to maintain performance and compatibility.
- Monitor for threats (including tracking CVEs) and execute timely remediation/mitigation of security issues.
- Maintain and configure monitoring and log collection systems (e.g., Prometheus, Grafana, ELK stack or similar).
- Lead incident management: rapidly detect, respond to, and resolve system failures in development and production environments.
- Expand capabilities of security tools such as Trivy, Dependency-Track, WAF, and others.
- Conduct regular security audits and deliver infrastructure security training to the team.
Infrastructure Development & Automation
- CI/CD: Lead migration of Jenkins from Freestyle Projects to Pipeline Projects, leveraging AI-assisted approaches where applicable.
- CI/CD: Conduct R&D to evaluate replacing Jenkins with alternative CI/CD tools and recommend/implement improvements.
- Kubernetes: Perform R&D on Service Mesh integration (e.g., Istio/Linkerd) and support future implementation.
- Kubernetes: Migrate from Nginx Ingress to Gateway API.
- AWS: Develop various integrations related to AI solutions and cloud-native services.
- Proxmox: Configure virtual machines, create new VM images, and optimize virtualization workflows.
Additional duties include collaborating with development and operations teams, automating repetitive tasks, participating in on-call rotation (as needed), and contributing to continuous improvement of our DevOps practices.
- Strong proficiency in Linux administration and troubleshooting.
- Hands-on experience with AWS services and cloud infrastructure.
- Expertise in Docker for containerization.
- Basic to intermediate knowledge of Kubernetes (including cluster management, deployments, and networking).
- Solid understanding of CI/CD tools and processes (Jenkins experience strongly preferred).
- Deep knowledge of the TCP/IP stack and networking fundamentals.
- Proficiency in at least one programming language at a middle / middle+ level (e.g., Python, Go, Java).
Qualifications
- 3–7+ years of experience in DevOps, Site Reliability Engineering (SRE), or Systems Administration roles.
- Proven track record supporting hybrid (on-prem + cloud) environments.
- Ability to travel occasionally to datacenter facilities in NY/SI region.
- Strong problem-solving skills, attention to detail, and ability to work independently or collaboratively.
- Excellent communication skills for cross-team coordination and documentation.
Nice-to-Haves / Preferred
- Experience with security tools (Trivy, Dependency-Track, WAF).
- Familiarity with HashiCorp Vault / OpenBao, Proxmox, Dell/Cisco hardware.
- Exposure to AI/ML integrations in cloud environments.
- Knowledge of Infrastructure as Code (IaC) tools (e.g., Terraform, Ansible).
What We Offer
- Opportunity to own impactful infrastructure projects and R&D initiatives.
- Collaborative environment focused on security, reliability, and innovation.
- Competitive compensation, benefits, and professional development opportunities.