Our mission
We build machine learning systems in vision, language, and medicine, then we study them: where they succeed, where they break, and how far their confidence can be trusted.
Vision
Building vision and medical-imaging models, then measuring exactly which classes or patients they still fail on.
Language
Building language systems, then measuring where they break down and the hidden costs baked into how we tokenize text.
Medicine
Building clinical-adjacent models, from CT segmentation to ECG classifiers, then studying their calibration and error structure.
Seven classifiers audited across three independent ECG datasets, errors concentrate in minority classes and specific patients, and confidence is not a reliable guide to correctness.
June 2026
A systematic tokenization cost asymmetry hides in everyday formatting choices, measured across 45 tokenizers and 2,828 meaning-preserving minimal pairs.
April 2026
A procedurally generated, self-verifying benchmark of 2,800 items across 14 models, mapping the exact point where small models collapse.
February 2026