About AI Logic Neural Network Pvt. Ltd.
AI Logic Neural Network Pvt. Ltd. is a technology-driven organization specializing in Artificial Intelligence, Machine Learning, Language Technology, and Enterprise Software Solutions. We build scalable, secure, and intelligent platforms that empower global businesses through innovative automation, multilingual technologies, and cloud-native applications.
We are looking for a highly skilled DevOps Engineer to join our engineering team and help build, automate, and maintain reliable cloud infrastructure, deployment pipelines, and production environments that support modern enterprise applications.
Job Summary
We are seeking an experienced DevOps Engineer with 5+ years of hands-on experience designing, implementing, and maintaining scalable CI/CD pipelines, cloud infrastructure, containerized applications, and automation frameworks. The candidate should have experience supporting both microservices-based and monolithic application architectures.
The ideal candidate will have strong expertise in AWS, Kubernetes, Docker, Infrastructure as Code (IaC), observability, deployment automation, and modern DevSecOps practices, while ensuring the security, reliability, scalability, and high availability of production environments.
Key Responsibilities
- Design and maintain CI/CD pipelines for both independently deployable microservices and monolithic applications.
- Develop deployment strategies including rolling updates, blue-green deployments, canary releases, automated rollbacks, and zero-downtime deployments.
- Manage Kubernetes resources including Deployments, StatefulSets, Services, Ingress, ConfigMaps, Secrets, Horizontal Pod Autoscalers, Pod Disruption Budgets, RBAC, and Network Policies.
- Package and deploy Kubernetes applications using Helm charts.
- Configure and manage API gateways, ingress controllers, service discovery, and secure communication between microservices.
- Implement centralized logging, metrics, alerting, and distributed tracing across applications and infrastructure.
- Establish monitoring dashboards and alerts using tools such as Prometheus, Grafana, CloudWatch, OpenSearch/ELK, OpenTelemetry, AWS X-Ray, or equivalent solutions.
- Manage application configuration and secrets using AWS Secrets Manager, Systems Manager Parameter Store, KMS, HashiCorp Vault, or similar tools.
- Manage container and artifact repositories such as Amazon ECR, GitLab Container Registry, Nexus, or Artifactory.
- Support application dependencies such as relational databases, caches, message queues, and event-streaming platforms.
- Implement autoscaling, capacity planning, performance tuning, and resource optimization for Kubernetes workloads and AWS infrastructure.
- Define and monitor Service-Level Indicators (SLIs), Service-Level Objectives (SLOs), availability targets, and production reliability metrics.
- Participate in incident response, root-cause analysis, problem management, and post-incident reviews.
- Define disaster recovery requirements, including Recovery Time Objectives (RTO), Recovery Point Objectives (RPO), backup validation, and restoration testing.
- Manage multiple environments such as Development, QA, Staging, UAT, and Production while maintaining environment consistency.
- Implement cloud cost-monitoring practices, resource tagging standards, budget alerts, and rightsizing recommendations.
- Create and maintain architecture diagrams, operational runbooks, deployment documentation, and troubleshooting procedures.
Required Skills
- Experience operating both microservices-based and monolithic applications in production environments.
- Strong understanding of Kubernetes networking, storage, ingress, autoscaling, resource limits, health probes, and workload scheduling.
- Experience with centralized logging, metrics, alerting, and application performance monitoring.
- Understanding of microservice communication patterns, API gateways, service discovery, asynchronous messaging, and distributed-system troubleshooting.
- Experience implementing rolling, blue-green, and canary deployment strategies with automated rollback capabilities.
- Hands-on experience managing application secrets, encryption keys, certificates, and configuration securely.
- Experience with container registries and artifact-management systems.
- Knowledge of production incident management, root-cause analysis, capacity planning, backup validation, and disaster recovery testing.
Preferred Skills
- Experience with GitOps tools such as Argo CD or Flux.
- Familiarity with Prometheus, Grafana, OpenSearch/ELK, OpenTelemetry, Jaeger, or AWS X-Ray.
- Experience with API gateways and ingress controllers such as AWS API Gateway, Kong, NGINX Ingress Controller, or AWS Load Balancer Controller.
- Familiarity with messaging and event-driven platforms such as Amazon SQS, SNS, MSK/Kafka, or RabbitMQ.
- Knowledge of Site Reliability Engineering (SRE) concepts including SLIs, SLOs, error budgets, incident response, and blameless postmortems.
- Experience with AWS multi-account environments and tools such as AWS Organizations or AWS Control Tower.
- Familiarity with software supply-chain security, including SBOM generation, image signing, vulnerability management, SAST, DAST, and secret scanning.
Qualifications
- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
- Minimum 5 years of professional DevOps experience supporting cloud-native and enterprise production environments.