(Visualization of RenderOcc's prediction, which is supervised only with 2D labels.)
RenderOcc is a novel paradigm for training vision-centric 3D occupancy models only with 2D labels. Specifically, we extract a NeRF-style 3D volume representation from multi-view images, and employ volume rendering techniques to establish 2D renderings, thus enabling direct 3D supervision from 2D semantics and depth labels.
@inproceedings{pan2024renderocc,
title={Renderocc: Vision-centric 3d occupancy prediction with 2d rendering supervision},
author={Pan, Mingjie and Liu, Jiaming and Zhang, Renrui and Huang, Peixiang and Li, Xiaoqi and Xie, Hongwei and Wang, Bing and Liu, Li and Zhang, Shanghang},
booktitle={2024 IEEE International Conference on Robotics and Automation (ICRA)},
pages={12404--12411},
year={2024},
organization={IEEE}
}