Zizhang Li · X (formerly Twitter)

  1. X
  2. Zizhang Li

Zizhang Li

86

posts

user avatar

@zizhang_li

Joined April 2022

  • user avatar

    Video world model as game engine is really cool when it supports multi-player and playable! I remember

    @Po_lhr

    just asked us to play it for ~10 mins to shoot a demo video, but the model is so interesting while with consistent performance and we kept playing for more than 1 hour.

    user avatar

    🎮 Real-time multiplayer world model 👥 Arbitrary number of players 🧠 Generated entirely by a neural network MultiGen is a real-time multiplayer diffusion game engine that supports an arbitrary number of players through a shared memory-based world model, rather than limiting

    00:00

  • user avatar

    RealWonder is a *real-time* video model that directly operates on 3D simulation, enabling single image 3D action-based interaction! Check out Koven’s🧵 for more details👇

    user avatar

    Replying to @Koven_Yu

    RealWonder thus unlocks a new capability: Real-time future video prediction on a single image and 3D physical actions. Checkout our website for more: liuwei283.github.io/RealWonder/

    00:00

  • user avatar

    user avatar

    #ICCV2025 🤩3D world generation is cool, but it is cooler to play with the worlds using 3D actions 👆💨, and see what happens! — Introducing *WonderPlay*: Now you can create dynamic 3D scenes that respond to your 3D actions from a single image! Web: kyleleey.github.io/WonderPlay/ 🧵1/7

    00:00

  • user avatar

    #CVPR2026 🤩 PerpetualWonder: interactive 4D scene generation with long-horizon actions. From a single image, build a world you can interact with a sequence of ✨3D actions (point force / wind / gravity). Project: johnzhan2023.github.io/PerpetualWonde… More details in 🧵: 1/5

    00:00

  • user avatar

    user avatar

    Our new work, coupled diffusion sampling, allows fast and diverse multi-view editing! Not by training — but by having one diffusion model guide another. 🤝 As a bonus: it can construct video-editing datasets and can make video models generate longer videos. 🧵👇

    00:00

Read the original on x.com ↗