i think the only thing i'm lacking now is RL to give the model the goal of winning, not just pure imitation of player habits. also more data, there is really no limit to the amount of replay data i would like to have, it's clearly still overfocusing on a few weird rounds (e.g. the backdash habit)