Behavioral cloning on successes gives you a policy with no notion of what a bad grasp looks like and how to avoid it, and nothing to learn from once it drifts off-distribution. To learn those, you need failure and suboptimal data. (3/13)