Currently a research engineer at Extropic. Before that, I did a Math Master's at the University of Toronto.
I'm interested in applied math, training big models, and figuring out why loss doesn't go down.
Let's chat about anything, email me!
An optimizer augmentation that applies window smoothing to updates across model depth.
We empirically find that the Jacobians of transformer blocks in pre-trained LLMs have highly similar singular vectors.
A JAX/Equinox implementation of Joint-Embedding Predictive Architecture (JEPA) models and related self-supervised learning methods. Features extensible code, distributed training, and gradient checkpointing.