Supplementary Video Results
Interaction guide: Hover to play the video. Click to show the control bar.
Dataset source: We use self-collected videos to present the results.
Each video shows the full analogy canvas: demonstration pair (top) and query input → generated output (bottom).
Tasks using structured visual signals (depth, flow, events, points) as conditions for generation.
Held-out manipulation tasks requiring temporal coherence and transformation reasoning.
Camera motion and view transfer between demonstration and query scenes.