Publications
See my Google Scholar for the most up-to-date list of papers.
2026
- COLMToken-Level Off-Policy Learning for Faithful Generation Under Distribution ShiftIn Conference on Language Modeling (COLM), 2026
- ICML
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right DataIn International Conference on Machine Learning (ICML), 2026*Equal Contribution - arXivValue-Aware Stochastic KV Cache Eviction for Reasoning ModelsIn arXiv, 2026*Equal Contribution
- COLM
Convergent Evolution: How Different Language Models Learn Similar Number RepresentationsIn Conference on Language Modeling (COLM), 2026 - ICML
EPSVec: Efficient and Private Synthetic Data Generation via Dataset VectorsIn International Conference on Machine Learning (ICML), 2026 - arXiv
- ICLR
Zebra-CoT: A Dataset for Interleaved Vision Language ReasoningIn International Conference on Learning Representations (ICLR), 2026*Equal Contribution - ICLR
FoNE: Precise Single-Token Number Embeddings via Fourier FeaturesIn International Conference on Learning Representations (ICLR), 2026 - COLM
Resa: Transparent Reasoning Models via SAEsIn Conference on Language Modeling (COLM), 2026 - ACL
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language ModelsIn Association of Computational Linguistics (ACL), 2026*Equal Contribution
2025
- Tech Report
- NeurIPS
VisualLens: Personalization through Visual HistoryIn Conference on Neural Information Processing Systems (NeurIPS), 2025 - ICLR
TLDR: Token-Level Detective Reward Model for Large Vision Language ModelsIn International Conference on Learning Representations (ICLR), 2025 - ICLR
Transformers Learn Low Sensitivity Functions: Investigations and ImplicationsIn International Conference on Learning Representations (ICLR), 2025*Equal Contribution - ICLR
DeLLMa: Decision Making Under Uncertainty with Large Language ModelsIn International Conference on Learning Representations (ICLR), 2025Spotlight (Top 5.1%), *Equal Contribution - NAACL
DreamSync: Aligning Text-to-Image Generation with Image Understanding FeedbackIn Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2025*Equal Contribution
2024
- NeurIPS
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear RegressionIn Conference on Neural Information Processing Systems (NeurIPS), 2024SoCalNLP Symposium 2023 Best Paper Award - NeurIPS
Pre-trained Large Language Models Use Fourier Features to Compute AdditionIn Conference on Neural Information Processing Systems (NeurIPS), 2024 - COLM
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic RepresentationsIn Conference on Language Modeling (COLM), 2024*Equal Contribution
2023
- EMNLP
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative ExamplesIn Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023
Dataset
Code