Dependent Two-Stage Pick-and-Place · fixed budget of 12 training tasks
Instruction gives o₁, o₂, and c₂ — but not c₁. Stage 2 reserves the mug, so stage 1 must avoid it.
4 × 3 = 12 possible (o₁, c₂) cases. Budget stays at 12 tasks; only how many relation cases appear changes.
Filled cells = an illustrative spread of 12 observed pairs for this Cov level (schematic order, not the true experimental identity). Ringed cell = the (milk, mug) example from the task above.
Broader dependent-pair coverage raises OOD success and lowers same-container violations.
Caveat: with a fixed budget, concentrating demos on fewer pairs can favor seen (ID) combinations. Intermediate ID points are noisy — the clean signal is the OOD relational gain above.
36-task Dependent 2S-PP (l4_2s_no_cont0). Means ± SE across seeds. Cov4 / Cov6 / Cov9 / Cov12 keep 12 training tasks while varying observed (o₁, c₂) pairs.