Hiring / Founding Engineer, Distributed Systems

Founding Engineer, Distributed Systems

Build the world's largest virtual GPU cluster

About C3

The next breakthroughs in cancer, fusion energy and our understanding of the universe will be found on GPUs running experiments at massive scale. The agents running these experiments need GPUs in bursts no single cloud can supply. C3 turns the world's fragmented GPU capacity into one network they can use: describe a job once, and C3 finds GPUs, reproduces the environment, positions the data, routes the work and meters only running time. C3 is building the infrastructure layer for automated research.

C3 is turning fragmented GPU compute supply into a single global network, creating the compute infrastructure layer for autoresearch at scale. If you want to work on the hardest problems in GPU computing, across hardware, distributed systems and ML research, this is an opportunity to get in early and help shape an extremely technical business. If you have 3+ years of engineering experience shipping code (exceptional open-source work counts), or can demonstrate excellence in physics, maths, computer science or a related technical field, please reach out.

Autoresearch at scale.

C3 is building the infrastructure for Karpathy-style autoresearch at scale: agents that propose, run and score experiments around the clock, each placed on the right GPU in any cloud. Our Autoresearch service is in beta.

The hard problems

C3 is writing the operating system for the world's GPUs. Its scheduler has to reason about a planet: capacity that appears and vanishes across hundreds of suppliers, prices that move like a market, datasets too large to move quickly, and agents firing thousands of experiments in bursts. It must predict where demand will land before it arrives, and keep every job running when any cloud beneath it fails. Nobody has built this. The first version runs in production today.

The role

  • Isolation at GPU speed: run untrusted agent code on GPUs inside microVMs (Firecracker, Cloud Hypervisor, KVM) with GPU passthrough, and lighter sandboxes where a VM is too heavy.
  • Instant starts: snapshot and restore whole machines (memory, filesystem, environment) so a job lands on a warm GPU in seconds, on any cloud.
  • A scheduler for a planet: place thousands of short agent jobs across clouds, regions and GPU types by capacity, price, reliability and data location, with queueing, gang scheduling, preemption and fair share.
  • Coordination that survives failure: leases, fencing, reconciliation and exactly-once metering across machines and clouds that fail independently.
  • Data at the GPU: stage images, environments and datasets ahead of the job, with content-addressed caches and lazy loading.
  • The stack today: Go agents on every machine and a TypeScript control plane on Cloudflare Workers. You help decide what comes next.
  • Shape C3: one of the first engineers, with a path to technical leadership.

Must have

  • Production experience in at least two of: job scheduling or batch systems (Slurm, PBS, LSF, or your own); orchestration (Kubernetes internals and operators, Nomad, Ray); HPC clusters; virtualisation (KVM, Firecracker, Cloud Hypervisor, QEMU). Tell us what broke and how you fixed it.
  • Systems depth: Linux internals (namespaces, cgroups, virtualisation), GPU drivers and the CUDA stack, networking and storage.
  • 3+ years in systems or infrastructure engineering. Substantial research and open-source engineering count.
  • AI leverage. Demonstrable custom agentic workflows are required. Show us systems you built and how they multiplied your output.

Your first three months

  • Ship microVM-isolated GPU execution on at least one cloud, with snapshot restore that starts a job in seconds, and make it the default for agent workloads.

Nice to have

  • Synchronisation under partial failure (distributed locks and leases, consensus or leader election, idempotency and reconciliation) from real systems; GPU passthrough (VFIO) and GPU sharing (MIG, MPS); checkpoint and restore (CRIU, VM snapshots); RDMA, InfiniBand or NCCL; systems code in Rust or Go; running production workloads on several clouds. Depth in a few is enough.

Also

  • Cambridge: mostly in person.
  • Research: optional contributions to C3's research and papers.

Interested, or know someone?

Write to sam@cthree.cloud with your CV or GitHub and a few sentences on a compute system you built and why it mattered. Excited by impossible infrastructure problems? Get in touch.