Machine Learning Engineer, Evals

remotemid

via Ashby

About this role

The Role You'll work across the lab on agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together. This is a high-growth, high-ownership role on a small team, and you'll ship evaluation infrastructure that researchers depend on from day one. Responsibilities - Run the full eval pipeline end to end and reproduce known results during onboarding, pairing with a senior engineer on your first task - Build a judge calibration protocol: sample human-labeled decisions, measure agreement (κ, per-class P/R), identify drift zones, and document it so anyone can re-run it…

Read the full description on Nous Research's site →

What we'd score you on

reqspace match rubric

Five dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.

1

Skills match

For this role: r, docker, git

2

Level fit

This role is mid-level. We check your trajectory against it.

3

Domain experience

Your work in the role's domain matters more than your years total. We weight recent and direct experience.

4

Recency

A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.

5

Location fit

This role is remote-eligible — we factor in your stated location and time-zone overlap.

Score yourself on this role.
Free · no card · written explanation included
See if I'm a fit →

Skills in this role

Pulled from the job description. These are the keywords we'll weight when scoring your fit.

rdockergit

More at Nous Research

See all open jobs at Nous Research