switched to training on whole rounds instead of small context windows, i'm pretty shocked by how big the improvement is compared to what i had before. this also enables performance optimizations which make it way easier to run larger models in realtime. here's a Johnny vs Faust round as an example
Публикация
Обсуждение
0 прямых ответов
Показаны не все ответы
Публичный счётчик показывает 3, а источник передал 0. Остальные ответы могли быть удалены, скрыты или не выданы публично.
Ответов пока нет
Обсуждение ещё не началось.