You will build the models that learn from large volumes of Earth observation data and turn indirect measurements into calibrated predictions about the subsurface. We are hiring two engineers. One will focus on representation learning and the other on generative and probabilistic inference.
There are two problems here. The first is learning from data with very few labels. We have terabytes (and in a few regions, petabytes) of satellite imagery, hyperspectral scans, airborne geophysics, drill-core imagery and multi-element assays, but only a few hundred known deposits. Those labels are positive-only and spatially biased, and the inputs shift between sensors, regions and acquisition conditions.
The second problem is inference. The subsurface is observed only through physical forward models, so the inverse problem is ill-posed and non-unique. The task is to learn priors over 3D geology and to build fast surrogates for the physics, so that we can compute posteriors that stay calibrated.
What your first year looks like
Track A: Representation learning
- Design and train self-supervised encoders on terabyte-scale satellite data (Sentinel-2, EMIT, EnMAP, PRISMA), airborne geophysics and geological map stacks.
- Build multimodal models that align hyperspectral core scans, core imagery and assay data, trained with noisy expert labels.
- Develop positive-unlabelled prediction models that correct for sampling bias.
- Measure and improve robustness to shifts in sensor, atmosphere, vegetation and region, using physical constraints where they help.
Track B: Generative and probabilistic inference
- Train score-based and flow priors over 3D geological structure, conditioned exactly on borehole observations.
- Build neural-operator surrogates for geophysical and hydrothermal forward models, with their approximation error quantified and propagated.
- Develop amortised and simulation-based inference for inversions with millions of parameters.
- Define calibration metrics for spatial posteriors, and test them against exact methods with the applied mathematics team.
Both tracks
- With the Earth sciences and data teams, establish spatially blocked evaluation and held-out-region tests, so that reported performance reflects transfer to new regions.
- Build training infrastructure that scales across GPUs and gives reproducible results.
You have:
- 3+ years training and shipping deep models on imagery or other high-dimensional data.
- Strong programming and implementation skills, including distributed training (DDP/FSDP), high-throughput data loading and rigorous experiment tracking.
- Track A: depth in self-supervised learning (masked autoencoders, contrastive methods) and vision transformers. You also have experience learning from weak, noisy or few labels, through semi-supervised, positive-unlabelled or active learning.
- Track B: depth in at least two of the following: diffusion and flow models, variational and simulation-based inference, neural operators, model-based RL and world models.
Nice to have:
- PhD in ML, statistics, data science or applied maths.
- Publications at NeurIPS, ICML, ICLR or CVPR, widely used open-source work, or models deployed at scale.
- Diffusion posterior sampling for inverse problems (MRI, CT, deblurring, tomography).
- Neural fields, or uncertainty calibration under distribution shift.
- Experience with hyperspectral, multispectral or SAR data.
- Familiarity with geospatial tooling (COG, Zarr, STAC, xarray), or the ability to pick it up quickly.
Interested? Write to us at gondwana@altcarbon.com with the role in the subject line.