Kairos
Back to jobs

Senior Cloud Infrastructure Engineer

On-site
Scale AISan Francisco, CA, US / New York, NY, US14 hours agoWebsite
Fresh
Senior
Horizontals EPD

Compensation

$143,000-$178,000/yr
Apply
Share

Description

Scale is powering this generative AI wave by providing the data and infrastructure for companies to build large-scale foundation models. AI is rapidly changing the world, and Scale is growing to meet that rapid demand. Our customers include OpenAI, Microsoft, Adept, Stability AI and many more major players in this space!

The Platform team is responsible for building the core abstractions and infrastructure on which the products can be built and iterated rapidly. We are looking for an Operations-focused AWS Engineer with deep experience in infrastructure management, lifecycle maintenance, and security hardening. We are looking for Ops specialists who thrive on optimizing, securing, and maintaining the health of large-scale AWS environments. You have a growth mindset and are comfortable learning new technologies.

You Will: 

  • Manage Infrastructure Lifecycle: Lead the end-to-end lifecycle of our AWS infrastructure, including routine patching, version upgrades, and system maintenance to ensure high availability.
  • Vulnerability Management: Work closely with the Security Team to identify, prioritize, and remediate vulnerabilities across our cloud footprint.
  • AWS Optimization: Continuously monitor and optimize AWS resource utilization for performance, reliability, and cost-efficiency.
  • Automate Operations: Use scripting and automation tools to streamline repetitive operational tasks such as fleet-wide patching and configuration audits.
  • Incident Response & Troubleshooting: Provide technical expertise for troubleshooting infrastructure-level issues and participate in operational health reviews.
  • Security Compliance: Build and maintain systems that adhere to strict security standards, ensuring our environment remains compliant through response and proactive mitigations.

Qualifications:

  • Extensive AWS Experience: 6+ years of experience with core AWS services (EC2, VPC, IAM, S3, RDS, EKS) and experience managing multi-account environments.
  • Patching & Upgrades: Proven track record of managing large-scale patching programs and performing major version upgrades for OS and middleware with minimal downtime.
  • Operational Tooling: Proficiency with AWS Systems Manager (SSM), Terraform, Atlantis, and Kubernetes for configuration management and automation.
  • Scripting: Familiarity with writing functional scripts (Bash, Python, etc.) to automate operational workflows - focused on system management rather than application development.
  • Security Focus: Experience with vulnerability scanning tools and a strong understanding of how to harden cloud infrastructure against common threats.
  • Monitoring & Metrics: Experience using Datadog, CloudWatch, or similar tools to monitor system health and drive optimization efforts.

Multi Cloud Experience: Experience with Azure, Google Cloud Platform is a bonus.

Stack

Generative AIPythonGCPAzureTerraformAWSFoundation ModelsKubernetes
Posted
Sep 24, 2026
Last seen
Sep 24, 2026
First seen
Sep 24, 2026

Similar roles

Browse more AI jobs