RobotEra has two more Olympic Medals. This time orange peeling (which PI accomplished but needed an external tool, which disqualified them) and unlocking a padlock1 25% faster than PI’s demonstration.
These (especially the key video) are approaching the speed and competence of a really skilled operator. We’re not quite as fast and dexterous as humans using their hands (a very very high bar) but it does seem like we’re approaching the speed and dexterity limits of these kinds of robots. I love what they are doing.
There is a spectrum of “generalization purity” where on the one hand you have folks who only care about performance of the models on tasks and objects completely outside the training data. To them, doing single-task fine-tuning is “fine… I guess,” if only because zero shot performance is still so bad, and you still need to eval something. But it feels like cheating to them.
On the other hand you have folks, like RobotEra, who are interested in “given what we’ve got, how can we squeeze the absolute MAX performance out of it. Let’s use every trick we can to shave a few more milliseconds off our time and see what we can get. And what we get is pretty impressive.
I love both these groups.
Videos! In real time! With audio!
There are a few interesting details the team at RobotEra shared with me. One is that they had to shrink the fingers in order to stab the orange.2 I have to imagine that thinner fingers might help with the key task because you can see what you are doing. (Thinner fingers also helped a lot with their sock-inversion win because it is easier to cram smaller fingers into a sock.3 There is an interesting tradeoff here where thin pointy fingers are good for stabbing and precise work but worse for getting good leverage on things like big objects or tools. Humans, of course, have very adaptable hands that can do both.
The other thing they really focused on was speed. To get this fast they did some interesting things. First they recorded demonstrations and trained on 15Hz samples. Then, when testing it, their model produces 15Hz output but they play it back on fast-forward (1.3x speed: about 20Hz). It is literally like watching a youtube video at 1.3x: everything happens but faster. Presumably they turned this knob up and up until things just barely worked so that is in itself an interesting finding. They also shortened the execution horizon for each rollout from 20 samples to 15 samples (1.3 seconds -> 0.75 seconds for those of you following at home). If you watch the lock video closely you can see the 0.75 second period as little hiccups in the motion.
I’m also just going to quote their email to me directly:
For Key Unlocking
Since the robot lacks tactile feedback, it struggles to learn fine manipulation patterns directly from standard teleoperation. Inspired by experiments where grasp aperture increases under tactile deprivation, we increased motion redundancy for the left hand holding the lock — using a larger range of rotation and wider gripper opening — which greatly improved success rates.
For accurate key insertion, we slowed down the key-insertion phase during data collection, similar to how a person carefully aims at a keyhole. This gives the wrist camera higher visual resolution on the keyhole, allowing the model to predict smaller, more precise actions and significantly improve insertion success.
For Orange Peeling
A major challenge is handling various out-of-distribution (OOD) states.
We collected recovery-focused data by continuing to demonstrate on partially peeled oranges after failed inference attempts, teaching the robot how to recover from mistakes.
We also added orange pose adjustment actions, so the robot can actively reposition the orange when it is in a hard-to-peel orientation.
A bunch of interesting things here about how folks are squeezing performance. If you have a black box that is data-in behavior-out how do you optimize? You optimize the data. This really feels like a black art, which is both neat and scary. Implied here is that the limit to speeding up replay was the actual key insertion which they compensated for by demonstrating that part slowed down so when they sped it up it was a more normal speed.
One of the tricks with talented human demonstrations is that they demonstrate how to do it when everything goes well but robots will always screw up and then have no idea what should happen when they screw up. The RobotEra team had to intentionally capture orange peeling data showing what to do if you are in a bad state (recover and reposition the orange).
Totally interesting stuff!
Thanks RobotEra for the submissions & peek into your process!
Unlocking a padlock isn't even one of my original tasks but it's a good one so we'll all pretend.
It feels funny to disqualify PI for using a sharp tool but allow RobotEra for changing the fingers into a sharp tool, but PI got like a bajillion medals so hopefully they don’t begrudge this one.
This is the kind of Robotics Professional Insider hot-take you subscribe to get.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.