S
SnappTech

Site Reliability Engineer

Tehranhybridmid

Posted 3w ago · via Recruitee

About this role

In this role, you will help scale and stabilize our systems as we grow. As part of the SRE team, you'll work on automating operations, managing incidents, and supporting the infrastructure that enables our developers and QA teams to build and release with confidence. You'll be responsible for improving system reliability, monitoring, and observability while ensuring high availability across environments. This role includes participation in a 24/7 shift or on-call rotation. Manage Incidents: Respond to incidents, perform root cause analysis, and help drive resolution and recovery. Monitor & Alert: Improve and tune monitoring systems (Grafana, Prometheus) to ensure issues are detected early.…

Read the full description on Snapp's site →

What we'd score you on

reqspace match rubric

Five dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.

1

Skills match

For this role: python, mysql, kubernetes, docker, grafana…

2

Level fit

This role is mid-level. We check your trajectory against it.

3

Domain experience

Your work in the role's domain matters more than your years total. We weight recent and direct experience.

4

Recency

A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.

5

Location fit

This role is based in Tehran. We weight your proximity and willingness to relocate.

Score yourself on this role.
Free · no card · written explanation included
See if I'm a fit →

Skills in this role

Pulled from the job description. These are the keywords we'll weight when scoring your fit.

pythonmysqlkubernetesdockergrafanaprometheuslokitempojaegerelkteams

More at Snapp

See all open jobs at Snapp