
Site Reliability Engineer (SRE)
We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.
Key Responsibilities
Own reliability, availability, and performance of production services, including on-call rotation
Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
Automate deployments, scaling, and operational tasks to reduce manual toil
Manage containerized workloads on Kubernetes and cloud infrastructure
Design and maintain CI/CD pipelines for safe, frequent releases
Lead incident response and drive blameless postmortems with clear follow-ups
Perform capacity planning, performance tuning, and cost optimization
Define and track SLIs/SLOs and error budgets with product teams
Requirements
3+ years in SRE, DevOps, or production-focused engineering
Strong Linux administration and hands-on Kubernetes experience
Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
Cloud experience with AWS, GCP, or Azure
CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
Proficient scripting in Python and/or Bash
Nice to have
Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
Background in high-traffic or distributed systems
Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →