Andy Zeng
498
posts
The models are getting better everyday. This is a pretty contact-rich task, and forces us to reckon with the fact that our policies not only need to learn the quirks of different robots, but also have to do it in a way that generalizes to the fine-grained subtleties of friction
Decades of hardware development led to strong, fast, and precise robot arms. The moment we can put in general intelligence in these things, we can leverage the full spectra of capabilities that they were always meant to capture. We’re betting on a future where robot hardware
For the avid viewer -- there’s a brief moment when the robot loses it’s grip on the head of a ziptie, and so it decides to use the other hand to help readjust the grip for the pull. It’s gnarly passing by our robots everyday, and catching these random glimpses of improvisational
If anyone’s creating a benchmark for frontier physical AI models, this task is a great one to add to the roster. Sensorimotor end-to-end policies must exhibit the long-term visual memory to track and reason about where the object might be. It’s also harder to “cheat” on this
GEN-1 plays the 🐚 shell game, trained on just 1 hr of robot data. It also generalizes to unseen objects, like
@BerkayAntmen's car keys. Physical AI models should be capable of benchmark tasks like this one. It's interesting for the all the reasons
@RhodaAIcalls out --

00:00
The first time we rolled a robot into a new warehouse, it didn’t perform as well as we expected. It took us an entire day of debugging, before we realized it was something simple… the cameras were wired completely wrong. 🤦 Left camera to right gripper 🔀 right camera to left


