Skills live in the parameters — only the instruction needs aligning

Sparse pretraining learns reusable subtask skills, but fails OOD because the instruction interface isn't calibrated. One demonstration per distinct instruction recovers large OOD success.

OOD = tasks never seen during pretraining  |  +1 demo/instruction = one demonstration for each of the 64 task instructions

Pretrained only (OOD)
+ 1 demo / instruction (OOD)

OOD Success: Before vs. After 1-Demo Finetuning

L5 Pick & Place (64 tasks) — pretrained on B sparse diagonal tasks, then finetuned with 1 demo per instruction

Takeaway: Even with as few as 4 pretraining tasks, one demonstration per instruction is enough to jump from near-zero to >50% OOD success — the skills were already there.

* B=8 baseline uses a single-seed evaluation (no 3-seed pooled baseline available for that checkpoint).