Sam Park · X (formerly Twitter)

Sam Park

64

posts

user avatar

@smsampark

Researcher at GDM. Previously: Stanford Postdoc, MIT PhD, Cornell BS

San Francisco, CA

Joined July 2013

  • user avatar

    Excited about this new work led by

    @TristanThrush

    ! Obvious implications for targeted synthetic data generation and data poisoning (check out the original thread for very cool results), but wanted to highlight a few points (1/7)

    user avatar

    New paper! Want to precisely optimize synthetic training data to do practical or even wacky things? Dataset Policy Gradients get you there, letting you target any differentiable training or post-training metric. We embedded a QR code in GPT-2’s weights using only training data!

  • user avatar

    One of the coolest scaling analyses I've seen in a while!

    user avatar

    Since compute grows faster than the web, we think the future of pre-training lies in the algorithms that will best leverage ♾ compute We find simple recipes that improve the asymptote of compute scaling laws to be 5x data efficient, offering better perf w/ sufficient compute

  • user avatar

    Cool application of training data attribution / influence estimation (using TRAK) to data selection for robot imitation learning!

    user avatar

    What makes data “good” for robot learning? We argue: it’s the data that drives closed-loop policy success! Introducing CUPID 💘, a method that curates demonstrations not by "quality" or appearance, but by how they influence policy behavior, using influence functions. (1/6)

    GIF

  • user avatar

Read the original on x.com ↗