
Compensation
$320,000-$405,000/yrDescription
About the role
Anthropic's Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. At our partner-operated sites, this role manages the interface between Anthropic and the strategic site operations partner performing day-to-day data hall work.
As the site lead, you own site outcomes for your assigned sites including: deployment velocity, availability, and incident response. Rather than managing operations staff directly, you provide tactical direction, set priorities, and define the standards for the vendor's on-site teams, paired with performance oversight and ongoing operational assessment to ensure all operational commitments are met.
You will define the operational processes, quality gates, and governance rhythms for partner-operated sites. Expect to build the playbook as much as you run it, not just at a site level, but defining and developing program improvements fleet-wide.
What you’ll own
-
Operational outcomes. Own site availability, deployment milestones, and repair turnaround, verified with independent data rather than vendor self-reporting.
-
Vendor direction. Set daily and weekly priorities and lead the operating cadence, including standups and business reviews.
-
Process definition. Author and improve procedures for deployment, break-fix, change management, security, and EHS compliance. Analyze operational trends and standardize lessons across the program.
-
Performance management. Track vendor performance against SLAs and staffing commitments, driving corrective actions when necessary.
-
Incident response and on-call. Participate in the incident escalation on-call rotation. When designated Anthropic Incident Commander for a site-specific incident, direct vendor response, own communications, and close out post-incident actions.
-
Internal interface. Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.
Representative work
-
Leading weekly operations reviews and scorecards with vendor site leads.
-
Directing deployment surges to meet first-compute-online milestones.
-
Analyzing failure patterns to identify root causes and driving fixes with owners.
-
Creating break-fix ownership matrices and training vendor teams.
-
Serving as Incident Commander for facility events and producing post-mortems.
-
Establishing operational readiness for new data halls, including spares and security.
-
Identifying process gaps and codifying improvements as program standards.
You may be a good fit if you
-
Have 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role, including accountability for production availability.
-
Have managed vendors, MSPs, or contract workforces to measurable outcomes: SOWs, SLAs, operational reviews, and corrective action.
-
Carry hands-on technical depth in server, network, and rack-level infrastructure, enough to independently verify vendor claims and audit quality.
-
Have built or substantially improved operational processes, not just run them.
-
Have served in an incident command or lead-responder role and communicate clearly under ambiguity.
-
Can support non-standard hours, including an on-call rotation and availability during deployment surges and maintenance windows.
-
Bachelor's degree in relevant domain or equivalent practical experience.
Strong candidates may also have
-
Experience with third-party colocation providers or partner-operated sites, delivering IT operations outcomes inside a facility someone else runs.
-
Experience standing up operations at a new site or data hall, from commissioning handoff through first deployment.
-
Experience with GPU/accelerator or high-density liquid-cooled infrastructure.
-
Familiarity with multi-vendor sites where facilities and IT operations are performed by different partners.
-
Experience leading projects from initiation to completion across teams you didn't own.
-
Background in incident management frameworks, contract/SLA design, or EHS programs.
Stack
- Posted
- Aug 21, 2026
- Last seen
- Aug 21, 2026
- First seen
- Aug 21, 2026

