AI Infrastructure Engineer — Autonomous Vehicles & Robotics
We are partnered with a high-growth autonomy company — operating at the intersection of AI, robotics, and the physical world — to hire an AI Infrastructure Engineer who will build the compute and systems backbone that turns perception models into real-time decisions at 70mph.
This is AI infrastructure where latency isn't a metric — it's a safety requirement. The models are only as good as the infrastructure that trains, validates, and deploys them into vehicles and robots operating in unpredictable, unforgiving environments.
Why Autonomy Infrastructure Is a Different Beast
Cloud-native AI infrastructure is hard. Autonomy infrastructure is harder. You're not optimizing for a chatbot's p99 — you're building systems where a training pipeline regression, a corrupted sensor log, or a deployment hiccup has consequences measured in physics, not user complaints. The data is enormous (petabytes of multimodal sensor data per week), the feedback loops span the physical and digital world, and the validation bar is orders of magnitude higher than a typical ML deployment.
What You'll Own
- Massive-scale data infrastructure — ingestion, processing, and storage pipelines for petabytes of LiDAR, camera, radar, and IMU data streaming from fleets of vehicles in the field
- Training infrastructure at scale — orchestrating distributed training across large GPU clusters for perception, prediction, and planning models with complex, multi-modal input pipelines
- Simulation and validation platforms — building the compute layer that powers millions of simulation miles, scenario replay, and regression testing before any model touches a vehicle
- Edge deployment infrastructure — packaging, delivering, and managing model updates to embedded compute platforms (NVIDIA DRIVE, custom SoCs) across a distributed fleet with strict versioning and rollback guarantees
- Data labeling and curation pipelines — scalable, quality-controlled annotation infrastructure that turns raw sensor data into training-ready datasets, including active learning loops and auto-labeling systems
- Observability and fleet telemetry — building the monitoring, alerting, and analytics layer that connects field performance back to model and infrastructure decisions
What You Bring
- 5+ years in infrastructure engineering, platform engineering, or data engineering
- Experience building and operating large-scale data pipelines for high-volume, multimodal data (video, point clouds, time-series sensor streams)
- Strong background in distributed compute orchestration — Kubernetes, Slurm, Ray, or equivalent, with production GPU cluster management experience
- Hands-on proficiency with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi)
- Familiarity with high-throughput storage architectures — object stores, data lakes, and parallel file systems tuned for ML workloads
- A strong engineering instinct for reliability, reproducibility, and auditability
Nice to Have
- Direct experience in autonomous vehicles, robotics, or aerospace infrastructure
- Familiarity with ROS/ROS2, sensor fusion pipelines, or simulation frameworks (CARLA, NVIDIA Isaac, internal replay engines)
- Background in embedded or edge deployment — model compilation, optimization (TensorRT, ONNX), and OTA update systems
- Experience with ML experiment tracking and model registry systems at scale (MLflow, Weights & Biases, custom)
- Understanding of functional safety standards (ISO 26262, SOTIF) and how they shape infrastructure requirements
Why This Space, Why Now
- The real-world AI problem — autonomy is where AI meets the hardest engineering constraints: real-time, safety-critical, physically embodied. This is infrastructure work that genuinely pushes the boundary of what's possible.
- Data moat in motion — every mile driven generates proprietary training data. The infrastructure you build compounds the company's competitive advantage daily.
- Beyond the hype cycle — autonomous systems are shipping in production today across delivery, trucking, mining, agriculture, and warehouse robotics. This isn't speculative — it's operational.
- Career-defining complexity — the intersection of fleet-scale data, real-time edge compute, and safety-critical ML deployment is one of the most challenging infrastructure problems in industry. If you can build here, you can build anywhere.