Kairos
Back to jobs

Software Engineer (Compute Efficiency), London

On-site
Isomorphic LabsLondon, GB17 hours agoWebsite
Fresh
Tech

Compensation

Salary undisclosed
Apply
Share

Description

Your impact 

We are building the largest foundation models in biotech and applying them immediately to cure disease. You will play a key role and work at a grand scale to deliver the foundations that make this happen. Joining the Compute Infrastructure team, you will ensure our planet-scale accelerator fleet operates at peak health and efficiency. By partnering with in-house machine learning platforms, performance & scaling team, and AI researchers, you will design the observability systems and automated recovery mechanisms that maximize scientific throughput on every paid GPU-hour.

What you will do 

  • Design, deploy, and scale robust observability systems and telemetry pipelines to monitor fleetwide compute efficiency, hardware health, and workload goodput across distributed clusters
  • Drive hardware efficiency and node reliability across our accelerator fleet, and integrating new hardware to leverage advancements
  • Identify compute waste and efficiency bottlenecks across the fleet, partnering with ML and platform teams to actively optimize accelerator utilization and improve workload goodput
  • Collaborate with teams in the AI org, and work closely with the ML Infrastructure team to identify, instrument, and improve canonical efficiency metrics across both training and inference runs
  • Contribute to the efforts for consistently improving the reliability of our ML runs.
  • Operate, maintain, and harden research, development, and production cloud infrastructure and cluster deployments
  • Partner and collaborate with a diverse set of teams incl. science, research, product, business development and operations
  • Contribute to core technical decisions (e.g. choice of tooling, infrastructure, and architectural design)

 

Skills and qualifications 

Essential:

  • Possess real world experience operating, monitoring, and debugging infrastructure for large-scale AI/ML workloads
  • Have experience working in cloud compute infrastructure design, preferably GCP
  • Possess strong programmings skills
  • Have significant experience working and deploying in Kubernetes at scale
  • Familiarity with the Nvidia GPU generations
  • Proven track record of building production observability and telemetry stacks

Nice to have:

  • Have a background in either ML SWE or infrastructure SRE work to build on
  • Conceptual understanding of ML workload efficiency paradigms
  • Have experience leading and delivering projects to multidisciplinary stakeholders
  • Familiarity with Google TPU generations
  • Familiarity with: workload scheduling; machine learning efficiency research; familiarity with ML-driven R&D cycles; familiarity with hardware benchmarking

Stack

GPUGCPMachine LearningFoundation ModelsKubernetes
Posted
Oct 6, 2026
Last seen
Oct 6, 2026
First seen
Oct 6, 2026

Similar roles

Browse more AI jobs