Site Reliability Engineer (Senior)

senior

via Ashby

About this role

ABOUT THE ROLE   VESSL AI의 Senior Site Reliability Engineer는 VESSL GPU 클라우드 플랫폼의 가용성과 성능을 책임지며, 개별 장애 대응을 넘어 시스템 설계와 아키텍처 결정을 리드합니다. Observability와 자동화 체계를 구축하는 데 그치지 않고, 장애가 발생하기 전에 구조적 리스크를 발견해 제거하며, 팀 전반의 안정성 관행을 수립합니다.   그만큼 안정성이 중요한 건, VESSL GPU 클라우드 플랫폼이 대규모 GPU 클러스터, InfiniBand·RoCE 기반 고성능 네트워크, Kubernetes·Slurm 기반 스케줄링 시스템 등 여러 레이어로 구성되어 있기 때문입니다. GPU 워크로드는 몇 시간에서 몇 주까지 이어지는 경우가 많고, 노드 하나의 장애나 네트워크 지연만으로도 진행 중이던 학습·추론 작업 전체가 중단될 수 있어 일반적인 웹 서비스보다 훨씬 높은 수준의 안정성이 요구됩니다. 이를 위해 GPU, NIC, 드라이버, 커널에서부터 스케줄러, 네트워크 패브릭에 이르기까지 시스템 전 레이어의 상태를 지속적으로 모니터링하고, 장애가 발생했을 때 원인을 빠르게 좁혀 복구하며, 반복되는 운영 작업을 자동화로 줄여나가는 기능 구현 역할이 중요합니다. WHAT YOU WILL DO   - Reliability & Observability: 메트릭·로그·트레이스 등 Observability 스택을 구축하고, SLI/SLO를 정의해 플랫폼의 안정성을 정량적으로 추적…

Read the full description on Vessl-ai's site →

What we'd score you on

reqspace match rubric

Five dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.

1

Skills match

For this role: python, go, kubernetes, docker, git

2

Level fit

This role is senior-level. We check your trajectory against it.

3

Domain experience

Your work in the role's domain matters more than your years total. We weight recent and direct experience.

4

Recency

A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.

5

Location fit

This role is based in a specific location. We weight your proximity and willingness to relocate.

Score yourself on this role.
Free · no card · written explanation included
See if I'm a fit →

Skills in this role

Pulled from the job description. These are the keywords we'll weight when scoring your fit.

pythongokubernetesdockergit

More at Vessl-ai

See all open jobs at Vessl-ai