As with Deepseek R1, there is a lot of value in going back and just reading some of Moonshot's technical papers. They are out there: arxiv.org/abs/2510.26692

Kimi Linear: An Expressive, Efficient Attention ArchitectureWe introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-c...arxiv.org