RSSAmplifier

Blog

Sonia Joseph

Sonia Joseph's AI research, essays, and commentary on the industry

soniajoseph.aiRSS feed ↗4 posts

Latest posts

World models and interpretability are two sides of the same coin

There has been a lot of confusion around world models. Depending on who you ask, a world model is a generative model, a 3D reconstruction model, or a latent-space prediction model. I would like to propose a fourth definition– one that goes back to the original- the Internal

Interpreting physics in video world models

Authors: Sonia Joseph, Quentin Garrido, Randall Balestriero, Matthew Kowal, Thomas Fel, Shahab Bakhtiari, Blake Richards, and Mike Rabbat. Work done at Meta Superintelligence Labs. Full preprint: https://arxiv.org/abs/2602.07050 This post walks through our main findings and intuition behind our interpretability study of physics representations in video

The logit lens can be deceptive if not used properly

The content of this short post will be already obvious to many researchers, but it's still worth writing out explicitly, especially when I see some assumptions that newer researchers are making. Tldr : The logit lens is a convenient way to investigate internal representations. But it can be misleading,

Multimodal interpretability in 2024

Multimodal interpretability, from sparse feature circuits with SAEs, to vision transformers leveraging CLIP's shared text-image space.