Site Reliability Engineer

remotemid

via Ashby

About this role

SITE RELIABILITY ENGINEER Platform and software · shared across customers Reports to: Director, Site Reliability Location: Remote (US) Department: Cloud Platform Engineering / SRE/Reliability POSITION SUMMARY The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution. KEY RESPONSIBILITIES - Define and operate Service Level Objectives (SLOs) aligned with customer SLAs - Build and maintain the observability stack including metrics, logs, traces, and alerting - Lead incident response and chair post-incident reviews…

Read the full description on Stninc's site →

What we'd score you on

reqspace match rubric

Five dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.

1

Skills match

For this role: python, go, datadog, grafana, prometheus…

2

Level fit

This role is mid-level. We check your trajectory against it.

3

Domain experience

Your work in the role's domain matters more than your years total. We weight recent and direct experience.

4

Recency

A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.

5

Location fit

This role is remote-eligible — we factor in your stated location and time-zone overlap.

Score yourself on this role.
Free · no card · written explanation included
See if I'm a fit →

Skills in this role

Pulled from the job description. These are the keywords we'll weight when scoring your fit.

pythongodatadoggrafanaprometheusopentelemetry

More at Stninc

See all open jobs at Stninc