01 · Introduction

Camera control · Video diffusion

SCoPESightline-Coordinate Positional Encoding for Video Diffusion Transformers

Give every video token its camera ray as a second positional coordinate—so precise camera control becomes part of the coordinate system, not an added control module.

Minghao Yin1* Jiahao Lu2* Wenbo Hu3† Wang Zhao3 Ying Shan3 Kai Han1‡
1The University of Hong Kong 2The Hong Kong University of Science and Technology 3ARC Lab, Tencent

* Equal contribution  ·  † Project lead  ·  ‡ Corresponding author

Each result is generated from a single input frame and a target camera trajectory. The inset visualizes the commanded motion.

Abstract

Camera control as a property of the coordinate system. SCoPE gives each token its camera ray as a second address. Ray geometry inside attention resolves appearance ambiguity and follows the target trajectory across diverse scenes.

A lightweight Normalize–Gate–Inject path supports both metric and up-to-scale poses while preserving pretrained RoPE bit-exactly. The same geometry improves controllability and video fidelity at both 5B and 14B scales.

<0.1%new parameters −29%rotation error −43%FVD at 14B
03

Video Overview

02:00

SCoPE in two minutes.A concise introduction to the motivation, ray-based positional encoding, qualitative results, and evaluation.

Key Ideas


Same scene · many cameras

Scenes

One generated scene browsed under six different camera trajectories. The inset visualizes the commanded camera path.

WASD · third-person control

Motions

WASD inputs drive the camera trajectory that moves the third-person subject. Overlay keys reflect the per-frame motion.


Revisit result · 01 / 03

07 / Emergent property

Revisit consistency,by construction.

Turn 90°, pass through a gateway, then return: previously seen regions remain anchored instead of drifting.

The outbound and return views reuse the same ray coordinates, addressing the scene from the same geometric positions.

Encoding, not memorization.Revisit consistency follows from the coordinate system itself—not from an added memory mechanism.

More results

Real-World Scenes

Diverse real captures driven along varied camera paths from a single frame.

More results

In-the-Wild Videos

Everyday footage re-rendered under controllable camera motion.

More results

Stylized & Artistic Scenes

Illustrated and painterly inputs animated with 3D-consistent camera control.

Citation

@article{yin2026scope,
  title   = {SCoPE: Sightline-Coordinate Positional Encoding for Video Diffusion Transformers},
  author  = {Yin, Minghao and Lu, Jiahao and Hu, Wenbo and Zhao, Wang and Shan, Ying and Han, Kai},
  journal = {arXiv preprint arXiv:2606.27345},
  year    = {2026}
}