vision & multimodal research
I'm a researcher at Anthropic, working on vision and multimodal models.
Before that, nine years at LandingAI: building the ML team for LandingLens, co-creating Data-Centric AI with Andrew Ng, and leading VisionAgent.
The workflow I've built up over the last ~10 years of applied machine learning: eval sets, error analysis, and treating training data as a hyperparameter.
VisionAgent is a novel approach to solve complex visual reasoning tasks.
Physics of Language Models offers a deep dive into the inner workings of large language models (LLMs) and their reasoning processes.