AnyCamVLA-LVSM: Novel View Synthesis for Camera Adaptation

Synthesize a novel camera viewpoint from two source images using the LVSM model finetuned on LIBERO-Plus scenes by AnyCamVLA.

LVSM takes two genuinely different camera views of the same scene as context (here, the LIBERO third-person agentview as View 1 and the wrist camera as View 2) and renders a novel viewpoint. The examples below are real two-view pairs with their true camera poses. Sharpest results come from viewpoints near View 1: keep View mix at 0 and use the small yaw / pitch / translation controls to move the synthesized camera.

Paper: AnyCamVLA (IROS 2026) | Model: heo0224/AnyCamVLA-LVSM | Code: GitHub

0 1
-2 2
-2 2
-2 2
-45 45
-45 45
-45 45
Examples