Job Description – DevOps Lead
System Administrator / DevOps Lead (10+ Years)
Location: Bangalore, India
Work Mode: Hybrid
Company: The Garage Platform India Pvt. Ltd.
Website: The Garage
About The Garage
The Garage is a fast-growing technology innovation company with engineering teams in
Orlando, USA and Bangalore, India. We design, build, and operate scalable digital platforms,
enterprise solutions, and data-driven products that serve global clients across multiple
industries.
Our culture is built on ownership, collaboration, innovation, and engineering excellence. We
believe in empowering talented professionals to solve complex challenges while providing
them the autonomy to make meaningful technical decisions and drive organizational impact.
As we continue to scale, we are looking for a highly capable System Administrator / DevOps
Lead who can own our infrastructure operations, cloud ecosystem, automation strategy, and
platform reliability.
The Opportunity
This is not a traditional support role.
As the System Administrator / DevOps Lead, you will serve as the primary infrastructure
and DevOps owner for our India and US engineering teams. You will be responsible for
designing, implementing, securing, and optimizing enterprise-grade cloud environments
while ensuring high availability, performance, and security across all systems.
You will play a critical role in shaping our DevOps maturity, driving automation initiatives,
and building a scalable infrastructure foundation for future growth.
Key Responsibilities
Systems Administration
• Administer Linux and Windows server environments.
• Manage Active Directory, DNS, DHCP, firewalls, SSL certificates, identity management,
and endpoint security.
• Oversee system upgrades, patching, vulnerability remediation, and lifecycle
management.
• Ensure high availability and uninterrupted business operations.
DevOps & CI/CD Ownership
• Build, maintain, and optimize CI/CD pipelines using GitHub Actions, GitLab CI/CD,
Jenkins, or similar tools.
• Automate build, test, deployment, and release processes.
• Implement deployment strategies that enable faster and more reliable software
delivery.
• Collaborate closely with engineering teams to improve deployment efficiency.
Infrastructure as Code & Automation
• Drive infrastructure automation using Terraform, Ansible, CloudFormation, or
equivalent technologies.
• Eliminate manual processes through scripting and workflow automation.
• Develop reusable infrastructure templates and deployment frameworks.
• Promote DevOps best practices across engineering teams.
Containerization & Kubernetes
• Manage containerized applications using Docker and Kubernetes.
• Support cluster deployment, scaling, monitoring, and troubleshooting.
• Improve application reliability and operational efficiency through container
orchestration.
Monitoring, Reliability & Incident Management
• Implement and manage observability platforms such as Prometheus, Grafana, ELK
Stack, CloudWatch, Datadog, or similar tools.
• Establish proactive monitoring, alerting, and incident response processes.
• Drive root cause analysis and continuous improvement following production
incidents.
• Maintain SLAs, uptime targets, and system reliability standards.
Security & Compliance
• Implement security best practices across infrastructure and DevOps operations.
• Conduct security audits, access reviews, and vulnerability assessments.
• Ensure compliance with organizational and client security requirements.
• Partner with leadership to strengthen DevSecOps practices and governance
frameworks.
Disaster Recovery & Business Continuity
• Develop and maintain disaster recovery plans.
• Ensure backup strategies are regularly tested and validated.
• Conduct periodic recovery drills and business continuity assessments.
Leadership & Collaboration
• Serve as the technical lead for infrastructure and DevOps functions.
• Mentor junior engineers and system administrators.
• Collaborate with engineering, product, security, and leadership teams across India
and the United States.
• Drive continuous improvement initiatives and establish best-in-class operational
practices.
What We're Looking For
Required Experience
• 8-10 years of experience in System Administration, Infrastructure Engineering, Site
Reliability Engineering, or DevOps roles.
• Proven experience managing production-grade cloud environments.
• Strong hands-on expertise with AWS (preferred), Azure, or Google Cloud Platform.
Technical Skills
• Advanced Linux administration and troubleshooting.
• Strong networking fundamentals including VPCs, VPNs, routing, firewalls, DNS,
proxies, and load balancers.
• Experience with Docker, Kubernetes, and container orchestration.
• Expertise in Terraform, Ansible, CloudFormation, or other IaC tools.
• Strong scripting skills using Python, Bash, or PowerShell.
• Experience designing and maintaining CI/CD pipelines.
• Hands-on experience with monitoring and logging frameworks.
• Database administration exposure including MySQL, PostgreSQL, or MongoDB.
Professional Skills
• Strong analytical and problem-solving mindset.
• Ability to manage critical incidents calmly and effectively.
• Excellent communication and stakeholder management skills.
• Passion for automation and operational excellence.
• Ability to work independently while managing multiple priorities.
Preferred Qualifications
• Experience implementing DevSecOps practices.
• Exposure to SOC2, ISO 27001, HIPAA, or similar compliance frameworks.
• Experience supporting globally distributed engineering teams.
• Background working in high-growth startups or product companies.
What Success Looks Like
Within your first year, you will:
Establish a stable, secure, and scalable infrastructure landscape.
Improve deployment speed and reliability through automation.
Reduce manual operational effort through Infrastructure as Code.
Strengthen monitoring, incident management, and disaster recovery readiness.
Serve as a trusted technical leader for infrastructure and DevOps initiatives globally.