Kairos
Back to jobs

DevOps Engineer, Infrastructure & Platforms

On-site
Ricursive IntelligencePalo Alto, CA, US1 day agoWebsite
Fresh
Full-time
Engineering

Compensation

Salary undisclosed
Apply
Share

Description

Ricursive Intelligence is a frontier AI lab building self-improving systems, starting with chip design. We are reinventing chip development and closing the loop between AI and the hardware that fuels it, recursively accelerating the path to artificial superintelligence. Backed by $335M from Sequoia, Lightspeed, DST, and NVIDIA Ventures, we are a growing, fast-paced team where every hire shapes the work.

The company has unmatched talent density, including IMO, IPHO, and IOAA gold medalists, pioneers who made prior breakthroughs in chip design: AlphaChip (Nature 2021), ePlace (DAC Best Paper Nominee 2014), RL-CCD (DAC Best Paper 2023), INSTA (DAC Best Paper 2025), and C3PO (ASP-DAC Best Paper 2026), chip leads for Apple Silicon (multiple generations of iPhone and iPad) and Google (TPU, OpenTitan), and top researchers and engineers from Anthropic, Google DeepMind, Stanford, and MIT.

ABOUT THE ROLE

Ricursive's research runs on infrastructure that has to keep pace with the research itself: ML training and evaluation workloads, EDA tool flows, and a fast-growing team that needs everything from cloud environments to developer workflows to just work. This role owns that foundation — the pipelines, platforms, and systems that let a small team move like a much larger one.

You will own our infrastructure-as-code (IaC), continuous integration & deployment (CI/CD), and cloud platform end-to-end that keeps the lab running day to day. Additionally, you will be collaborating closely with the team to build out the right observability stack for their needs while working around environment security limitations.

WHAT YOU WILL DO

• Design, build out and extend our self-managed cloud platform with Terraform, setting the IaC patterns the rest of the team builds on.

• Own our platform deployments spanning various environments day to day, including performance, security, reliability, scalability, and adapting to evolving business requirements.

• Design, implement, and operate CI/CD pipelines in GitHub Actions for mission-critical repositories, working within strict security and deployment restrictions.

• Architect scalable AI tooling and developer experience workflows across multiple distinct user environments, working closely with our physical design (PD) engineers.

• Build out the observability stack with the team, covering pipeline health, application-level metrics, and ML workloads.

MINIMUM QUALIFICATIONS

• BS in CS, CE, EE, or a closely related technical field, or equivalent practical experience.

• 4+ years of hands-on infrastructure/platform engineering, with ownership of systems others depend on.

• Owned production IaC architecture including maintenance and new features, not just consumed modules.

• Designed and implemented production CI/CD pipelines that build and ship reproducible artifacts, with attention to performance, scalability, and security.

• Prior experience with observability tooling for monitoring pipeline health as well as application-level metrics.

PREFERRED QUALIFICATIONS

• Hands-on experience with GCP, Kubernetes, and GitHub Actions, including custom runner setups.

• Experience running ML infrastructure for training and evaluation workloads, including GPU/TPU compute.

• Familiarity with LLM observability tooling: tracing, evaluations, and cost & latency monitoring.

• Security depth: dependency supply-chain hardening, OIDC-based auth, least-privilege secrets, and compliance work such as SOC 2 or penetration testing.

• Early-stage startup experience: built infrastructure from zero or near-zero.

Stack

LLMsGPUGCPTerraformCI/CDMachine LearningKubernetes
Posted
Sep 29, 2026
Last seen
Sep 29, 2026
First seen
Sep 29, 2026

Similar roles

Browse more AI jobs