Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
Paper • 2512.08333 • Published
RETAIN baseline (arXiv:2512.08333, Robust Finetuning of VLA Policies via Parameter Merging) for the absolute-joint DROID pi0 policy.
The language-model weights of the finetuned specialist are interpolated back toward the pretrained generalist MolmoBot-Pi0-DROID, while the vision encoder and action expert are kept fully finetuned:
theta_llm = (1 - alpha) * theta_base + alpha * theta_finetuned (alpha = 0.8)
vision_tower + gemma_expert + action heads : fully finetuned
Flat pi0 checkpoint: model.safetensors + metadata.pt + assets/droid_equad/norm_stats.json.
Actions are absolute future joint positions q[g+i+1].