Kairos
Back to jobs

Staff Software Engineer, ML Training Infrastructure

Remote
Stack AVPittsburgh, PA, US4 hours agoWebsite
Fresh
Staff / Principal
Autonomy

Compensation

Salary undisclosed
Apply
Share

Description

About the Role:

The ML Training team is dedicated to increasing Stack's AV development velocity by accelerating machine learning iterations. Our core mission is to deliver a training system that is reliable, scalable, user-friendly and observable. We are responsible for the overall AI workflow, ranging from dataset curation to training, validation, acceleration, optimization and deployment of large-scale models that power our autonomous vehicles. In addition, this team is in charge of evangelizing best practices and frameworks among Machine Learning Engineers (MLEs) across the company.

In this Staff role, you will drive the design and development of a high-performance multi-tenant AI training platform. You will balance hands-on coding with long-term technical direction by operating across ML Platform, Infrastructure, Autonomy and Safety Evaluation teams to accelerate the development of autonomous vehicles across Stack.

Responsibilities:

  • Design & evolve high-performance training platform components including orchestration, training abstractions, control plane, observability and performance tuning. 
  • Deliver end-to-end ML model pipelines across logs processing, feature extraction, dataset schema design/storage, model configuration management, model training, and profiling/acceleration workflows.
  • Analyze training infrastructure performance to identify and resolve performance bottlenecks
  • Evangelize system abstractions and tooling that enable MLEs to rapidly iterate on models.
  • Incorporate OSS tools to enable ML engineers self-sufficiently profile and optimize their workflows
  • Promote Engineering Excellence: Maintain a high bar for engineering excellence in their own work but also set a culture of engineering excellence within the team.

Qualifications: 

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 6+ years of experience with ML Platforms and building ML-based applications (modeling experience is a bonus).
  • Strong programming skills in Python, C++ or equivalent. 
  • Prior experience with Lance, PyTorch, Ray Data or equivalent technologies. 
  • Proven track record of building scalable, reliable infra in a fast-paced environment while working with MLEs across multiple modeling teams.
  • A deep understanding of design tradeoffs and ability to articulate those tradeoffs to build alignment across XFN teams.
  • Experience with model training, model optimization, or large-scale data processing pipelines.
  • Autonomous vehicles (AV) experience is a bonus.
  • Strong analytical and problem-solving skills.
  • Excellent verbal and written communication skills, with the ability to convey complex technical concepts to non-technical stakeholders.

#LI-AW1

Stack

PythonC++PyTorchAutonomous VehiclesMachine Learning
Posted
Oct 1, 2026
Last seen
Oct 1, 2026
First seen
Oct 1, 2026

Similar roles

Browse more AI jobs