sites.google.com

Research Interests

Foundation models, Transformer++ architecture, Efficient Pre-training and Knowledge Distillation.

Overview

I am a PhD student at the University of Texas at Austin, advised by Prof. Sujay Sanghavi in the Department of Electrical and Computer Engineering. For my research, I work on simple things and simple things work for me. I am currently working on understanding and improving large models (particularly language models) through training recipes. Some of my recent works have been featured in Ahead of AI magazine, Marktechpost, and the Interconnects newsletters. 

Before moving to Austin, I graduated with an M.Eng. degree in Information and Communication Engineering from Chongqing University of Posts and Telecommunications, Chongqing, China in 2019, and received a B.Tech degree in Electronics and Communication Engineering from the Maulana Abul Kalam Azad University of Technology (formerly West Bengal University of Technology), Kolkata, India. During my undergrad, I gloriously failed to scale up my startup, Tronix India, and later worked at an Indian multinational IT firm, TechMahindra. 

I am a person who stutters some info about stuttering here.

I am currently on the job market, seeking both industry and postdoctoral positions.

Internships


  • Student Researcher, Foundation Research team at Google Deepmind (May-Aug 2025) | Worked on two different projects on recurrence in language models.

    • Integrated recurrent states in Transformer blocks during test time which improved performance without any additional training.

    • Proposed a pre-training method based on iterative activation refinement.


  • Research Intern at Lightning AI (May-Aug 2024) | Topic: Efficient Fine-tuning and Continual training of LLM.


  • Applied science Intern at Amazon Science Alexa (May-Aug 2022) | Topic: Vision Language pre-training and finetuning.

Selected Blogs/Opensource


  • Looped-GPT: Looping During Pre-training improves Generalization. [Blog] [code]

  • Inheritune: A stagewise LLM training algorithm. [code]

  • Curriculum Pretraining Enables 10-Digit Addition for a 296-Parameter GPT with 99% Accuracy. [Blog]

Selected Publications 


  • [Preprint]  Mike Mozer, Shoaib Siddiqui, Danny Sawyer, Sunny Sanyal, and Rosanne Liu, "Recirculation". [paper] [tweetWork done at Google Deepmind during summer'25.

  • [ICML'25 Spotlight🏆] Sunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis and Sujay Sanghavi, "Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting". [paper]  [code]  [tweet]
    This work has been selected as a Spotlight Poster at ICML 2025, placing it among the top 2.6% of all 12,107 submissions.

  • [TMLR'26] Sunny Sanyal, Ravid Shwartz, Sujay Sanghavi and Alex Demakis, "When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models". [paper]  [code]  [tweet]

Inheritune, the training recipe introduced in this work, has enabled startups and hobbyists to pre-train small language models with extremely limited compute. Notable examples include Gumini (multi-lingual) and small LLaMA-3 variants, among others. The MAI tech report, 2026 (refer p. 22) acknowledges attention collapse, ZyphraAI (an AI startup) in their tech report uses our formulation of attention collapse to audit their model's "lazy" attention heads.

This paper was featured in a popular tech news portal Marktechpost and youtube tutorial


  • [COLM'24] Sunny Sanyal, Atula Tejaswi, Jean Kaddour, Abhishek Kumar, and Sujay Sanghavi, “Early Weight Averaging Meets High Learning Rates for LLM Pre-training”.  [paper]  [code]

Also presented at NeurIPS 2023 WANT workshop.
This paper has inspired some followup efforts in LLM pre-training with weight averaging at companies such as eBay (lilium model), Bytedance, Alibaba (Ant group- Paper1 and Paper2), Huggingface and others. It has also been featured in three widely read newsletters: [Ahead OF AI] [Interconnects] [DL Focus].


  • [Neurips'24 Dataset Track]  Jeffery Li, Alex Fang, ... Sunny Sanyal et al. "DataComp-LM: In search of the next generation of language model training sets". [paper] [code] [tweetOur work inspired  Apple's DCLM-7B model (here).

  • I have also co-authored some papers in the field of wireless networks. You can find them on my Google Scholar. 

Selected DEmos, Posters and talks


  • Gave an in-person talk at Google Deepmind's LLM investigation series: When Attention Collapses. [slides]

  • Gave a talk at Prof. Tom Goldstein's Group at UMD on Inheritune: Training Smaller Yet More Attentive Language Models. [slides]

  • Gave a talk at Lightning AI's NYC office on Training Smaller and Efficient Language models. [slides]

  • Gave a talk at ml collective on Pre-training with a little less Data and Compute. [slides] [recordings]

  • Demo at Art Gallery CVPR 2023 on Generative Masking and In-painting for Videos. [demo]

  • Poster at 6G@UT symposium 2023 on Understanding the Effectiveness of Early Weight Averaging for Training LLMs.

  • Gave a talk (in-person) at Austin Deep learning community’s main event on Do Neural Networks Overthink? [link] 

Academic Services 


Recent Updates


  • Awarded the Adaption Research Grant (Inagural cohort 2026).

  • Moving to SF this summer 2025 to join Google Deepmind.

  • LAWA accepted at COLM conference.

  • Moving to NYC this summer 2024.

  • May 2024: Gave a talk at ml collective's Deep Learning: Classics and Trends.

  • Our work featured in two newsletters.

  • Watch my CVPR 2023 demo in art gallery both at Demo hall area.

Read the original on sites.google.com ↗