We're a small, growing team studying how machine learning systems fail, across vision, language, and medicine.
We hire in small numbers, slowly, for people who care more about finding out where a system breaks than about shipping one more benchmark win. If that sounds like you, we'd like to hear from you, even without an open role listed below.
Research
Designing and running error audits across vision, language, and medical models, the kind of work that ends in a paper, a dataset, or both.
Engineering
Building the infrastructure behind large-scale evaluation runs, reproducible benchmark generators, and the models we release publicly.
Data
Sourcing, cleaning, and stress-testing the datasets our audits depend on, often the least glamorous and most important part of the work.
Send a note, a CV, or a link to something you've built, we read everything.