Infrastructure / Site Reliability Engineer (SRE)
Undisclosed employer
Compensation
$200/hr
Description
Mercor connects exceptional technical talent with leading organisations working on ambitious technology and AI initiatives. We are looking for experienced Infrastructure / Site Reliability Engineers (SREs) to join a full-time engagement focused on building and operating complex, enterprise-grade infrastructure. We are seeking engineers with strong hands-on experience building, operating, debugging, and scaling sophisticated production systems. The ideal candidate has worked extensively with Kubernetes, AWS, observability platforms such as Datadog, and modern infrastructure tooling.
This is a full-time opportunity, and candidates must be able to commit to full-time engagement.
What You'll Do
-
Build, operate, and improve highly available and scalable production infrastructure.
-
Manage and optimise Kubernetes-based production environments.
-
Design and maintain cloud infrastructure, primarily across AWS.
-
Improve system reliability, availability, scalability, and operational efficiency.
-
Build and maintain observability across infrastructure and applications using Datadog or similar platforms.
-
Investigate production incidents, perform root-cause analysis, and implement durable fixes.
-
Improve monitoring, alerting, logging, tracing, and overall production visibility.
-
Develop automation and internal tooling to reduce manual operational work.
-
Partner closely with software engineering teams on deployments, infrastructure, and production reliability.
-
Contribute to infrastructure architecture and technical decisions for complex distributed systems.
Ideal Background
-
Professional experience in Infrastructure Engineering, Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Production Engineering.
-
Hands-on experience operating complex, enterprise-grade production systems.
-
Strong production experience with Kubernetes.
-
Strong experience with AWS and cloud-native infrastructure.
-
Experience with Datadog, Prometheus, Grafana, or comparable observability platforms.
-
Experience with Infrastructure as Code using Terraform, Pulumi, or equivalent technologies.
-
Strong understanding of distributed systems, networking, containers, Linux, and cloud architecture.
-
Experience building or maintaining CI/CD and production deployment infrastructure.
-
Strong debugging, troubleshooting, and incident-response capabilities.
-
Proficiency in at least one programming or scripting language, such as Python, Go, or Bash.
Strong Signals
-
Experience operating Kubernetes and cloud infrastructure at significant production scale.
-
Experience supporting high-traffic or mission-critical applications.
-
Experience building infrastructure or platform tooling used by large engineering organisations.
-
Ownership of production reliability, on-call operations, incident response, or capacity planning.
-
Experience working within sophisticated, large-scale distributed systems.
-
Demonstrated improvements to SLOs/SLIs, observability, deployment reliability, infrastructure performance, or operational efficiency.
Why Join
-
Solve challenging reliability, scalability, and performance problems across enterprise-grade production systems.
-
Work extensively with technologies such as Kubernetes, AWS, Datadog, Terraform/Pulumi, and modern cloud-native tooling.
-
Take meaningful ownership of production reliability, observability, infrastructure architecture, and operational improvements.
-
Competitive hourly compensation reflecting your experience and technical expertise.
-
Join a network of highly skilled engineers working on ambitious projects with leading technology and AI organisations.
- Commitment
- Part-time
Skills & categories
- Posted
- Oct 7, 2026
- Slots remaining
- 10
- First seen
- Oct 7, 2026
- Last seen
- Oct 8, 2026