Now
I am a Senior Research Scientist at Meta, where my work centres on large language and
multimodal models that continue to learn after they ship — agentic systems that adapt to
new tasks, retrieve and integrate new knowledge, and improve themselves rather than
remaining fixed at the moment training ends.
Doctoral research
I completed my Ph.D. at Northwestern University under Professor
Ying Wu. My thesis attacked a single
question: how do we give a learning algorithm a complete lifecycle? Real deployments face
a world that keeps moving, so a model must accumulate capability over time while retaining
what it already knows. This is the problem of continual, incremental, and lifelong
learning, and it remains the through-line of my research.
Before that I earned an M.S. in
Applied Mathematics, also at
Northwestern, specialising in analytical and computational methods for partial differential
equations, stochastic differential equations, and parallel computing.
Broader interests
Alongside the thesis work I have published across generative models (diffusion and
autoregressive), vision-language and large multimodal models, open-world and
open-vocabulary perception, model customisation and personalisation, data- and
parameter-efficient finetuning, embodied AI and robot learning, domain adaptation and
generalisation, autonomous driving, uncertainty estimation, and active, few-shot, and
semi-supervised learning.
Before the Ph.D.
Earlier work took me through image matting, semantic segmentation, and network formulation
— network pruning and optimisation, attention mechanisms, neural architecture search, and
the lottery ticket hypothesis. As an undergraduate in pure and applied mathematics
(specialising in financial mathematics and engineering) I competed steadily in data mining
and mathematical modelling contests, placing in the
Mathematical Contest in Modeling,
Kaggle, and the
SIGKDD Cup.