Michael Tschannen Β· X (formerly Twitter)

Michael Tschannen

261

posts

user avatar

@mtschannen

Research Scientist

@GoogleDeepMind

. Representation learning for multimodal understanding and generation. Personal account.

Zurich, Switzerland

Joined June 2012

  • Pinned

    user avatar

    For the past years my research focus was on unifying models and training paradigms across modalities. Today I'm excited that we're releasing our latest model aligned with this theme: Gemma 4 12B, a dense encoder-free model which processes raw text, image, and audio inputs! 1/

  • user avatar

    user avatar

    For the past years my research focus was on unifying models and training paradigms across modalities. Today I'm excited that we're releasing our latest model aligned with this theme: Gemma 4 12B, a dense encoder-free model which processes raw text, image, and audio inputs! 1/

  • user avatar

    user avatar

    Have you ever wondered how to train an autoregressive generative transformer on text and raw pixels, without a pretrained visual tokenizer (e.g. VQ-VAE)? We have been pondering this during summer and developed a new model: JetFormer πŸŒŠπŸ€– arxiv.org/abs/2411.19722 A thread πŸ‘‡ 1/

  • user avatar

    We are presenting JetFormer at ICLR this morning, poster #190. Stop by if you’re interested in unified multimodal architectures!

    user avatar

    Have you ever wondered how to train an autoregressive generative transformer on text and raw pixels, without a pretrained visual tokenizer (e.g. VQ-VAE)? We have been pondering this during summer and developed a new model: JetFormer πŸŒŠπŸ€– arxiv.org/abs/2411.19722 A thread πŸ‘‡ 1/

  • user avatar

    4o native image generation is confirmed to be some sort of autoregressive model. Maybe this is a good moment for AR skeptics to catch up on the recent literature on multimodal AR models.

Read the original on x.com β†—