Sparse pretraining learns reusable subtask skills, but fails OOD because the instruction interface isn't calibrated. One demonstration per distinct instruction recovers large OOD success.
OOD = tasks never seen during pretraining | +1 demo/instruction = one demonstration for each of the 64 task instructions
L5 Pick & Place (64 tasks) — pretrained on B sparse diagonal tasks, then finetuned with 1 demo per instruction
Takeaway: Even with as few as 4 pretraining tasks, one demonstration per instruction is enough to jump from near-zero to >50% OOD success — the skills were already there.
* B=8 baseline uses a single-seed evaluation (no 3-seed pooled baseline available for that checkpoint).