Anything that reasons about the quality of behavior — RL, reward modeling, failure detection, world models — benefits from seeing behavior that isn't optimal. However, there is no good existing mixed-quality real-world dataset for researchers to use. (4/13)