Cloud & Data Platform Engineer (AWS)
Location: Hebbal, Bengaluru
Experience: 3–6 years
Type: Full-time
Who We Are & Why ARL
At AgResearch Labs (ARL), we're building technology and infrastructure for cleaner, more climate-resilient food production. Over more than a decade of R&D, ARL has developed indigenous aeroponics and farm-automation technologies that power Growize Farms, our commercial platform for building and operating high-performance farms at scale.
Recognised among India's Top 100 deep-tech startups by NASSCOM, ARL is now entering its next phase of growth — scaling our business, teams and operating systems.
What you will own:
You own everything from the device's connection to the cloud, through the database, and back out to alerts and notifications. Inside that boundary you are the decision-maker, not the executor of someone else's design:
- Telemetry ingest. AWS IoT Core (MQTT) → Rules → SQS → Lambda → PostgreSQL. Built idempotent and replay-safe, and proven so under duplicate, delayed and out-of-order delivery — not assumed.
- The operational data layer. Aurora PostgreSQL: partitioning strategy, indexing, retention, rollups and the maintenance jobs that keep it healthy as farms multiply.
- Imaging storage path. Bench camera imagery landing in S3 with sane layout and lifecycle tiering, laid out so batch ML inference can be added later without a migration.
- Infrastructure as code. The whole platform in Terraform or CDK, with CI/CD, environment separation and repeatable deploys. Nothing important should exist only in the console.
- Security and access. Least-privilege IAM, Cognito, encryption, audit logging, data residency in ap-south-1, and the technical controls for India's DPDP Act as its phases land.
- Backups and disaster recovery. Cross-account backups, defined RTO/RPO targets, and scheduled restore drills. The deliverable is a proven restore, not a backup job that reports green.
- Observability and self-healing. Alarms that mean something, automated remediation and rollback paths, and runbooks for the failure modes that actually occur.
- Messaging infrastructure. Transactional email and India SMS delivery, done properly — including the registration work that makes it reliable and affordable at volume.
- Cloud cost. A visible cost baseline, guardrails so the bill cannot surprise us, and the judgement to know which optimisations are worth the complexity.
- Integration runtime. The delivery side of our commerce integrations — webhooks, retries, deduplication and reconciliation jobs that survive an unreliable counterparty.
Where your boundary meets the product — payload formats, API contracts, what the data means — you and the founder define the contract together and write it down. You will not be guessing at interfaces.
What we expect from you
- Judgement over throughput. You will be the only person in the company whose full attention is on this layer. We need someone who thinks about failure modes before shipping, not after the incident.
- Evidence, not preference. You are inheriting a documented architecture that has already been independently reviewed and fact-checked. Improve it — but bring a document, a measurement or a reproduction, not a taste argument. We hold ourselves to exactly the same standard.
- Ownership through the boring parts. Restore drills, cost reviews, certificate rotation, runbook upkeep. The unglamorous work is most of what makes a platform trustworthy.
- Comfort with a small, unusual team. A meaningful share of implementation and routine operations here flows through AI agents, with humans reviewing and merging everything. You will direct and review that work as much as you type it. If that sounds interesting rather than threatening, you will enjoy this.
- Field awareness. Our devices live in polytunnels on rural connectivity, not in a data centre. Designs that assume a clean network do not survive contact with a farm.
What you bring
- 3–6 years running production workloads on AWS. Not a junior role — and not a big-team role either. You are the accountable engineer for this platform.
- Event-driven pipelines in practice. SQS, Lambda and the scars to explain idempotency, at-least-once delivery, and what happens when an ordering guarantee you assumed was there turns out not to be.
- Real PostgreSQL depth. Partitioning, index types beyond btree, vacuum behaviour, migrations on live tables. Expect detailed Postgres questions in the interview.
- Terraform or CDK as a habit, not an aspiration — and IAM you can reason about out loud.
- Python or TypeScript for pipeline and tooling code.
- Cost literacy. You have reduced a cloud bill before and can say specifically how.
Nice to Haves:
- AWS IoT Core, or any sizeable MQTT estate — provisioning, certificate rotation, topic permissions.
- Experience with India-specific delivery plumbing: transactional email at scale, DLT-registered SMS.
- Working alongside AI coding and operations agents (Claude Code, Strands or similar).
- Agri-tech, IoT hardware, or anything that had to work over rural connectivity.
What this role is not
Product frontend and backend are founder-owned. There is no data warehouse to build — we have deliberately verified we do not need one at this scale. There is no machine-learning modelling in year one; you build the storage and batch-inference plumbing, not the models. If you are looking for a pure Kubernetes platform role, this is not it.
Your first 90 days
- Days 1–30. AWS account structure, IaC baseline, CI/CD and monitoring in place. External lead-time items filed early. Interface contracts agreed and written down with the founder.
- Days 31–60. Telemetry pipeline live end-to-end for the pilot farm, with idempotency proven under deliberate fault injection. Cross-account backups running and the first restore drill completed.
- Days 61–90. Recovery objectives defined, signed off and exercised. Cost dashboard live against an agreed baseline. Imaging storage path ready for the first cameras. Runbooks written for the top failure modes.
How we work
Small, remote, IST hours, occasional farm visits. We run on primary sources and blunt verification — “the documentation says” beats “I think,” claims get checked, and being wrong early is treated as cheaper than being wrong in production. Decisions get written down with their reasoning, so nobody has to reconstruct them six months later.