Site Reliability Engineer (SRE)
xAI · London, United Kingdom · Infrastructure
About this role
xAI is hiring a mid-level Site Reliability Engineer in the software engineering function based in London, United Kingdom. The posting calls out experience with Rust, Kubernetes, Terraform, Pulumi.
- Role
- Site Reliability Engineer
- Function
- software engineering
- Level
- mid
- Track
- Individual contributor
- Employment
- Full-time
- Location
- London, United Kingdom
- Department
- Infrastructure
More roles at xAI
Job description
from xAI careersABOUT xAI
xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
You will work on the team that is responsible for the backend services that power our products such as grok.com and the API. We focus on writing and maintaining highly scalable and reliable services that can efficiently process tens of thousands of queries per second. The services are hosted on a number of Kubernetes clusters (on-prem & cloud).
BASIC QUALIFICATIONS:
- Expert knowledge of Kubernetes.
- Expert knowledge of continuous deployment systems such as Buildkite and ArgoCD.
- Expert knowledge of monitoring technologies such as Prometheus, Grafana, and PagerDuty.
- Expert knowledge of infrastructure as code technologies such as Pulumi or Terraform.