You will run the infrastructure our models are trained and served on, so that researchers spend their time on research. Our workloads combine multi-terabyte geospatial data, 3D volumes and physics simulators.
What your first year looks like
- Build and run our GPU cluster across cloud and on-premise capacity, with scheduling and cost tracking.
- Make distributed training (DDP, FSDP) fast and reliable, with checkpointing and fault tolerance.
- Build data loading for Zarr, COG and 3D volumes that keeps GPUs busy.
- Set up experiment tracking, a model registry and dataset versioning.
- Automate held-out-region evaluation and model serving.
You have:
- 4+ years in ML infrastructure or platform engineering.
- GPU workloads at scale on Kubernetes or Slurm, on AWS, GCP or Azure.
- Distributed training in PyTorch or JAX, and debugging slow jobs.
- Python, Terraform or similar, and CI/CD.
Nice to have:
- Geospatial formats such as Zarr, COG and STAC.
- On-premise GPU clusters.
- Inference serving with Triton or Ray Serve.
Interested? Write to us at gondwana@altcarbon.com with the role in the subject line.