We’re releasing WorldDiT, a pareto frontier model architecture bringing world and robot action modeling into one compact diffusion transformer, without relying on a large pretrained VLM action backbone.
WorldDiT in action
At 399 million total parameters, WorldDiT lies on the reported LIBERO Pareto frontier for model size and mean success.
Among publicly released methods in the comparison that do not use a large pretrained VLM action backbone, it reports the highest mean success.
WorldDiT unifies future visual world state prediction with continuous robot action generation in a single diffusion transformer.
If you use WorldDiT in your research, feel free to cite the paper.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.