switched to training on whole rounds instead of small context windows, i'm pretty shocked by how big the improvement is compared to what i had before. this also enables performance optimizations which make it way easier to run larger models in realtime. here's a Johnny vs Faust round as an example