Decoupling physical grasp synthesis from task-dependent reasoning allows robotic systems to leverage foundation model improvements without retraining, and enables cross-hand generalization through explicit kinematic and stability constraints.
AdaRoboVLG combines vision-language models with robotic grasping by separating physical grasp synthesis from task understanding. A base policy generates and evaluates grasp candidates using kinematics and stability checks, while foundation models provide task-specific context.