Chelsea Finn
704
posts
Pinned
LLM post-training used to mean fine-tuning to a downstream task Robotics has been stuck in this setting, needing task-specific fine-tuning for best performance π07 changes this: It works out of the box & outperforms fine-tuned specialists Details: pi.website/pi07

00:00
One of the most important aspects of scientific discovery is deciding where to draw insights from. While LLMs are promising tools for science, we lack datasets & evaluations for this step. Help contribute to a public dataset for exactly this: tinyurl.com/45b9ykae
Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch. We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining. Paper: arxiv.org/abs/2607.27203
I'm giving a talk tomorrow at ICML on emergent physical generalization, including π0.7 🤖 3:15 pm @ SCALE workshop in Ballroom 201 scale-icml-2026.github.io
I'm giving a talk on how we can move beyond the scalar reward bottleneck for both robotics & LLMs. ICML RLxF workshop tomorrow at 1:30 pm.
RL is hitting a ceiling with human feedback. What if the world itself becomes the signal? Join us at the RLxF: RL from World Feedback 🌍 workshop at ICML 2026
@icmlconftomorrow (July 10th)! Web page: sites.google.com/view/rlxf-icml…




