A
Apollo ResearchScheming Research

Research Scientist/Engineer (Evaluations)

London & San Franciscoonsitemid

Posted 7mo ago · via Lever

About this role

Application deadline: We are conducting interviews actively and aim to fill this role as soon as we find someone suitable.    ABOUT THE OPPORTUNITY   We’re looking for Research Scientists/Engineers for our pre-deployment team to work on Training-Run Assessments (TRAs). You will design and build automated pipelines for assessing whether egregious misalignment or scheming are emerging at any point of frontier post-training.  This will involve evaluating and red-teaming of checkpoints at various stages of post-training as well as automated analysis of post-training data. You will get to work with frontier labs like OpenAI, Anthropic, and Google DeepMind and be among the first to interact with new models before anyone else.…

Read the full description on Apollo Research's site →

What we'd score you on

reqspace match rubric

Five dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.

1

Skills match

For this role: python, openai, anthropic

2

Level fit

This role is mid-level. We check your trajectory against it.

3

Domain experience

Your work in the role's domain matters more than your years total. We weight recent and direct experience.

4

Recency

A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.

5

Location fit

This role is based in London & San Francisco. We weight your proximity and willingness to relocate.

Score yourself on this role.
Free · no card · written explanation included
See if I'm a fit →

Skills in this role

Pulled from the job description. These are the keywords we'll weight when scoring your fit.

pythonopenaianthropic

More at Apollo Research

See all open jobs at Apollo Research