About Apty:
Apty is a Digital Adoption Platform (DAP) that operates directly within web applications to improve user productivity, enforce process compliance, and eliminate friction in complex enterprise workflows.
Technically, Apty sits in a unique layer: on top of third-party applications, requiring deep runtime integration, browser extensions, and a high-scale multi-tenant backend processing millions of user interactions.
This is systems-heavy engineering, not standard infrastructure maintenance.
About the role:
We are looking for a highly hands-on Mid–Senior DevOps Engineer who can take strong ownership of our DevOps ecosystem and grow into owning it end-to-end.
With a growing base of 2M+ active users, our platform requires high availability, real-time processing, and extreme scalability.
This role is not about maintaining pipelines—it’s about:
- Architecting cloud-native distributed systems
- Scaling high-throughput analytics pipelines
- Building enterprise-grade SaaS and on-prem deployments
- Driving DevOps excellence across the company
If you enjoy owning infrastructure like a product, this role will suit you well.
What you'll work on in your first 90 days:
- Take ownership of our AWS infrastructure and CI/CD ecosystem
- Optimize and scale event ingestion pipelines
- Strengthen observability, alerting, and system reliability
- Improve deployment workflows across microservices, extensions, and frontend applications
- Contribute to architecture decisions for scaling to millions of users
- Start driving on-prem deployment standardization
Key responsibilities:
Cloud infrastructure & architecture (AWS)
- Architect and manage scalable infrastructure using AWS services
- Design and maintain a multi-tenant SaaS architecture
- Scale systems handling millions of real-time events
- Optimize performance, cost, and reliability
DevOps & automation
- Own and enhance CI/CD pipelines using GitHub Actions
- Build automated deployments for:
- Microservices
- Backend systems
- UI applications
- Browser extensions
- Implement Infrastructure-as-Code using Terraform
- Build and maintain Dockerized environments
- Improve release velocity while maintaining stability
Monitoring, security & reliability
- Implement end-to-end observability across logs, metrics, and traces
- Set up alerting and incident-response systems
- Ensure: high availability, auto-scaling, disaster recovery, backup strategies
- Drive security best practices, including: network isolation, IAM policies, CVE remediation, container security
On-prem deployments
- Design and deliver on-premise versions of our SaaS platform
- Replicate cloud architecture in constrained enterprise environments
- Build automation for: installation, upgrades, maintenance
- Collaborate with customers to customize deployments
Collaboration & leadership
- Work closely with Engineering, QA, and Product teams
- Share knowledge and support fellow engineers on DevOps practices
- Help drive best practices across teams
- Contribute to architecture and design decisions
- Act as a technical owner for DevOps initiatives
Required skills & qualifications:
Core expertise
- 3+ years in DevOps, with hands-on ownership of production systems
- Strong expertise in AWS, which is mandatory
- Hands-on experience with: Kubernetes (EKS), VPC, EC2, RDS, and ElastiCache, S3, CloudFront, and Route 53, API Gateway, Lambda, and Kinesis, ClickHouse or similar columnar databases, Docker and container orchestration, Helm charts
CI/CD & infrastructure
- Strong experience with GitHub Actions
- Expertise in Terraform, which is mandatory
Programming and scripting
Strong scripting skills in:
Development experience is a strong advantage.
Scalability & distributed systems
- Experience handling high-scale distributed systems
- Strong understanding of: load balancing, caching strategies, event-driven architectures, high-throughput pipelines
On-prem & enterprise systems
- Experience building SaaS-to-on-prem deployments
- Ability to automate complex enterprise setups
Other requirements
- Strong debugging and problem-solving skills
- Ownership mindset—drives problems end-to-end
- Excellent communication and collaboration skills
Nice-to-have skills
- Experience with Digital Adoption Platforms or browser-based products
- Experience with monitoring tools such as: Grafana, Prometheus, Datadog
- Exposure to: SOC 2, ISO compliance environments
What makes you stand out
- You treat infrastructure as a product, not a support function
- You think in systems, trade-offs, and scale
- You can design from scratch, not just maintain
- You are comfortable handling high ambiguity and ownership
- You optimize for long-term reliability, not short-term fixes
Interview process
- Initial screening: 15 minutes
- Technical deep dive: 60 minutes
- DevOps/system design round: 60 minutes
- Leadership discussion: 30–45 minutes
End-to-end timeline: Approximately two weeks