Kairos
Back to jobs

NOC Technician (Data Center and Site Ops)

On-site
xAIMemphis, TN, US2 weeks agoWebsite
Aging
Data Center

Compensation

Salary undisclosed
Apply
Share

Description

ABOUT THE ROLE:

As a NOC Technician, you are the eyes and the voice of the site — never the hands. You staff the Network / Campus Operations Center and continuously observe site health signals across xAI campuses. You detect and verify campus-impacting events, assemble the right responders, run incident communications leadership can trust, and drive every major incident to a completed report and a tracked corrective project. You work with Site Reliability Engineering, SiteOps, Facilities, Hardware Failure Analysis, SWE Platforms, and vendors — escalating correctly the first time and maintaining the institutional memory across shifts and sites.
One sentence: watch the campus, run the bridge, leave the wrench work and deep root cause to the teams that own them.

RESPONSIBILITIES:
Continuous monitoring (the watch)
• Staff the console per shift schedule to sustain 24/7 coverage (coverage posture: 2 on console per site)
• Watch the designated signal surface: cluster health dashboards, node availability, network health, facility trend panels (power/cooling), storage alarms, and threshold breaches as defined by SRE monitoring standards
• Acknowledge every page/alert within the SLA; classify it (actionable / known / noise) and log the disposition; feed noise patterns back to SRE for suppression or redesign
• Maintain a live picture of ongoing maintenance, planned work, and degraded-but-accepted states so real anomalies stand out

Detection, triage & escalation
• Detect → verify → escalate within defined time budgets; verification is signal-level (is it real, what's the blast radius), not deep diagnosis
• Operate the escalation matrix: NOC → on-call SRE → domain owners (SiteOps, Facilities, Network, Storage, HW FA, vendors); page correctly the first time
• Recommend incident declaration and severity to the on-call SRE; declare directly per runbook when thresholds are unambiguous

Incident communications & coordination
• Open and run the bridge; get the right people on within the time-to-bridge SLA
• Own stakeholder communications: first update within the SLA, then a fixed cadence until resolution
• Maintain the incident timeline in real time — timestamps, actions, decisions, engagements
• Track who owns what during the incident and call out stalls

First-pass RCA framing & closure
• Produce initial framing for major site outages: what happened, when it started, what's impacted (halls/racks/services), what changed recently, who is engaged
• Hand framing to SRE / Hardware FA for depth — the NOC does not publish root cause
• Write major-incident reports; open corrective projects in Linear with named owners and track them to closure ("filed" is not "done")

Shift operations, runbooks & improvement
• Run structured shift handoffs and keep durable shift logs; maintain cross-site awareness
• Own and continuously improve NOC runbooks: escalation matrix, comms templates, severity ladders, per-signal response procedures
• Participate in game days run by SRE; every incident where the runbook was wrong or missing produces a runbook change before the incident closes

Explicitly not this role
• Wrench work: swaps, reseats, physical recovery (SiteOps)
• Power / cooling / building plant operation (Facilities)
• Deep hardware root-cause analysis or vendor CAPA (Hardware Failure Analysis)
• Monitoring architecture, alert design, or technical SEV command (Site SRE)
• Building or operating reliability tooling such as SRT, turnback, or dashboards (SWE Platforms)

BASIC QUALIFICATIONS:
• High school diploma or equivalency certificate
• 1+ year of professional experience in a Network Operations Center (NOC), Security Operations Center (SOC), mission-control / dispatch, data center operations watch, or equivalent 24/7 monitoring and incident-communications role
• Demonstrated written and verbal communication skills under time pressure (stakeholder updates, handoffs, timelines)

PREFERRED SKILLS AND EXPERIENCE:
• Calm under pressure; excellent written and verbal communications — leadership should be able to trust your incident updates verbatim
• Pattern recognition across domains; multi-domain curiosity (compute, network, storage, power/cooling signals)
• Experience following and improving process: runbooks, escalation matrices, shift handoffs, post-incident follow-through
• Prior NOC, SOC, or critical-environment operations experience in a datacenter or hyperscale infrastructure environment
• Familiarity with reading operational dashboards, acknowledging/classifying alerts, and coordinating across on-site technicians, facilities, and engineering on-call
• Comfort with ticketing / project tracking systems (e.g. Linear, Jira) for opening and chasing corrective work to closure
• Industry certifications a plus (Network+, Security+, ITIL, or similar) — not a substitute for judgment and communications quality
• Basic familiarity with datacenter topology (racks, fabric, OOB) and how facility events affect compute availability — enough to triage and escalate correctly, not to deep-diagnose
• Bachelor's degree in IT, Computer Science, Cybersecurity, or STEM discipline preferred but not required

ADDITIONAL REQUIREMENTS:
• Must be available for on-shift rotations supporting 24/7/365 console coverage
• Shift structure (e.g. 12-hour rotations) to be confirmed; nights, weekends, and holidays are part of the role
• Must be able to work extended hours during major incidents as needed

SUCCESS LOOKS LIKE:
• Coverage attainment / shift fill rate vs plan
• Time-to-bridge for major incidents; first-update and cadence SLA attainment
• Page accuracy / escalation correctness
• % of major incidents with a complete timeline and follow-up projects tracked to done
• Not measured by: raw page counts, or heroics without a paired prevention item

Posted
Aug 13, 2026
Last seen
Aug 13, 2026
First seen
Aug 13, 2026

Similar roles

Browse more AI jobs