.png?1697105647)
Software Engineer (Compute Efficiency), London
On-site
Fresh
Tech
Compensation
Salary undisclosedDescription
Your impact
We are building the largest foundation models in biotech and applying them immediately to cure disease. You will play a key role and work at a grand scale to deliver the foundations that make this happen. Joining the Compute Infrastructure team, you will ensure our planet-scale accelerator fleet operates at peak health and efficiency. By partnering with in-house machine learning platforms, performance & scaling team, and AI researchers, you will design the observability systems and automated recovery mechanisms that maximize scientific throughput on every paid GPU-hour.
What you will do
- Design, deploy, and scale robust observability systems and telemetry pipelines to monitor fleetwide compute efficiency, hardware health, and workload goodput across distributed clusters
- Drive hardware efficiency and node reliability across our accelerator fleet, and integrating new hardware to leverage advancements
- Identify compute waste and efficiency bottlenecks across the fleet, partnering with ML and platform teams to actively optimize accelerator utilization and improve workload goodput
- Collaborate with teams in the AI org, and work closely with the ML Infrastructure team to identify, instrument, and improve canonical efficiency metrics across both training and inference runs
- Contribute to the efforts for consistently improving the reliability of our ML runs.
- Operate, maintain, and harden research, development, and production cloud infrastructure and cluster deployments
- Partner and collaborate with a diverse set of teams incl. science, research, product, business development and operations
- Contribute to core technical decisions (e.g. choice of tooling, infrastructure, and architectural design)
Skills and qualifications
Essential:
- Possess real world experience operating, monitoring, and debugging infrastructure for large-scale AI/ML workloads
- Have experience working in cloud compute infrastructure design, preferably GCP
- Possess strong programmings skills
- Have significant experience working and deploying in Kubernetes at scale
- Familiarity with the Nvidia GPU generations
- Proven track record of building production observability and telemetry stacks
Nice to have:
- Have a background in either ML SWE or infrastructure SRE work to build on
- Conceptual understanding of ML workload efficiency paradigms
- Have experience leading and delivering projects to multidisciplinary stakeholders
- Familiarity with Google TPU generations
- Familiarity with: workload scheduling; machine learning efficiency research; familiarity with ML-driven R&D cycles; familiarity with hardware benchmarking
Stack
GPUGCPMachine LearningFoundation ModelsKubernetes
- Posted
- Oct 6, 2026
- Last seen
- Oct 6, 2026
- First seen
- Oct 6, 2026


