Some Random Stuff 👋
I am a researcher at Google. I received my Ph.D. in Computer Science from UIUC, where I am fortunate to be advised by Prof. Klara Nahrstedt (a member of the NAE), who leads the Multimedia Lab, and to work closely with Prof. Minjia Zhang and Prof. Chengxiang Zhai. Before starting my Ph.D., I spent wonderful time at UIUC and Shanghai Jiao Tong University.
My research focuses on agents and multimodal intelligence. I work across post-training, harnesses, real-world applications of agents, and the science of agentic systems.
I am particularly interested in multi-agent post-training and dynamic agent workflows in 2026.
When I’m not teaching large models, I prototype augmented-reality systems for fun, hoping that, one day, agentic models will power these interactive, human-centered interfaces.
I am always happy to chat about LLM/agent research, career advice, and collaborations, feel free to reach out.
Industry Experience
-
Google
-
MRS AI, Meta
-
Video Rec, Meta
-
LLM Research Team, Capital Today
-
Xin's Group, Adobe
-
IOTG, Intel
Research
-
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
-
SlideForge: Controllable Editing of Slides as Structured Artifacts
-
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
-
Spatio-Temporal LLM: Reasoning about Environments and Actions
-
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
-
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
-
Aha Moment Revisited: Are VLMs Truly Capable of Self-verification in Inference-time Scaling?
-
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
-
RecoWorld: Building Simulated Environments for Agentic Recommender Systems
-
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
-
TraceNet: Segment One Thing Efficiently
-
Anywhere Avatar: 3D Telepresence with Just a Phone and a Laptop
-
Scene Graph Driven Hybrid Interactive VR Teleconferencing
-
miVirtualSeat: A Next Generation Hybrid Telepresence System
-
Seaware: Semantic-aware View Prediction System for 360-degree Video Streaming
-
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
-
ImmerScope: Multi-view Video Aggregation at Edge towards Immersive Content Services
-
GaugeTracker: AI-Powered Cost-Effective Analog Gauge Monitoring System
-
Vesper: Learning to Manage Uncertainty in Video Streaming
-
I-Matting: Improved Trimap-Free Image Matting
-
Interactive Scene Graph Analysis for Future Intelligent Teleconferencing Systems
-
360TripleView: 360-Degree Video View Management System Driven by Convergence Value of Viewing Preferences
-
SAVG360: Saliency-aware Viewport-guidance-enabled 360-video Streaming System
-
Video 360 Content Navigation for Mobile HMD Devices
-
AnyLoc: Energy-efficient Visual Localization in Dynamic and Large-Scale Scenes without Pain
Events
- Date — Invited Talk at TBD Institute.
Acknowledgements
This simple website is built from Jiayi Pan's website and Codex.